CryoThrust commented on issue #12097: URL: https://github.com/apache/seatunnel/issues/12097#issuecomment-5551288601
The proposed direction is promising, but I would separate the work into two independently verifiable phases so the optimization does not change read semantics and split planning at the same time.\n\n1. First establish a bounded-dispatch contract: keep the existing `ChunkRange` generation algorithm, but introduce a reader/enumerator handoff that limits in-flight/pending split metadata and preserves deterministic split IDs across checkpoint/restart. This gives a measurable memory bound without changing predicates.\n2. Then add database-side sampling or boundary probing per dialect behind an explicit capability, with a conservative fallback that never silently downloads the full split-key column for a large table.\n\nFor the boundary algorithm, the acceptance contract should explicitly cover: nullable keys (a dedicated NULL split or documented exclusion), duplicate probe boundaries (guaranteed progress), inclusive/exclusive predicates with no overlap/gap, planning snapshot consistenc y, empty ranges, and restart identity. I would also report plan metrics such as rows/bytes scanned during planning, generated split count, peak pending metadata, and fallback reason.\n\nA small first PR could therefore add deterministic split-plan/property tests and a bounded pending-state abstraction, before any dialect-specific SQL. That keeps the performance claim measurable and makes it easier to review correctness independently from database-specific sampling syntax. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
