CryoThrust commented on issue #12097:
URL: https://github.com/apache/seatunnel/issues/12097#issuecomment-5551076901

   A few correctness/recovery constraints seem important before implementation, 
especially if the design combines arithmetic ranges with boundary probing:
   
   - The split predicate must define `NULL` handling and inclusive/exclusive 
boundaries explicitly. Adjacent splits should be disjoint while covering every 
row, including nullable split keys and duplicate boundary values.
   - `MIN/MAX` and probe results need a consistency contract. If planning runs 
outside a transaction or the source changes during enumeration, document 
whether snapshot isolation is required, or how the reader prevents 
gaps/duplicates caused by moving boundaries.
   - A split cap should be a cap on planned work, not an implicit row-count 
guarantee. For highly skewed keys, expose when a bounded split plan may produce 
uneven reader load and keep checkpoint/retry identity stable for each generated 
range.
   - For database-side sampling, specify the fallback when sampling is 
unsupported or returns too few distinct keys. The fallback must not silently 
revert to a full client-side column scan for very large tables.
   - Please include acceptance tests for nullable keys, duplicate boundary 
values, empty ranges, skewed distributions, concurrent inserts/deletes during 
planning, and enumerator restart from a checkpoint.
   
   These constraints would make the optimization measurable without changing 
the source's read-completeness or recoverability semantics.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to