CryoThrust commented on issue #12097: URL: https://github.com/apache/seatunnel/issues/12097#issuecomment-5551076901
A few correctness/recovery constraints seem important before implementation, especially if the design combines arithmetic ranges with boundary probing: - The split predicate must define `NULL` handling and inclusive/exclusive boundaries explicitly. Adjacent splits should be disjoint while covering every row, including nullable split keys and duplicate boundary values. - `MIN/MAX` and probe results need a consistency contract. If planning runs outside a transaction or the source changes during enumeration, document whether snapshot isolation is required, or how the reader prevents gaps/duplicates caused by moving boundaries. - A split cap should be a cap on planned work, not an implicit row-count guarantee. For highly skewed keys, expose when a bounded split plan may produce uneven reader load and keep checkpoint/retry identity stable for each generated range. - For database-side sampling, specify the fallback when sampling is unsupported or returns too few distinct keys. The fallback must not silently revert to a full client-side column scan for very large tables. - Please include acceptance tests for nullable keys, duplicate boundary values, empty ranges, skewed distributions, concurrent inserts/deletes during planning, and enumerator restart from a checkpoint. These constraints would make the optimization measurable without changing the source's read-completeness or recoverability semantics. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
