CryoThrust commented on issue #12097:
URL: https://github.com/apache/seatunnel/issues/12097#issuecomment-5551288601

   The proposed direction is promising, but I would separate the work into two 
independently verifiable phases so the optimization does not change read 
semantics and split planning at the same time.\n\n1. First establish a 
bounded-dispatch contract: keep the existing `ChunkRange` generation algorithm, 
but introduce a reader/enumerator handoff that limits in-flight/pending split 
metadata and preserves deterministic split IDs across checkpoint/restart. This 
gives a measurable memory bound without changing predicates.\n2. Then add 
database-side sampling or boundary probing per dialect behind an explicit 
capability, with a conservative fallback that never silently downloads the full 
split-key column for a large table.\n\nFor the boundary algorithm, the 
acceptance contract should explicitly cover: nullable keys (a dedicated NULL 
split or documented exclusion), duplicate probe boundaries (guaranteed 
progress), inclusive/exclusive predicates with no overlap/gap, planning 
snapshot consistenc
 y, empty ranges, and restart identity. I would also report plan metrics such 
as rows/bytes scanned during planning, generated split count, peak pending 
metadata, and fallback reason.\n\nA small first PR could therefore add 
deterministic split-plan/property tests and a bounded pending-state 
abstraction, before any dialect-specific SQL. That keeps the performance claim 
measurable and makes it easier to review correctness independently from 
database-specific sampling syntax.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to