SEZ9 commented on issue #10203: URL: https://github.com/apache/seatunnel/issues/10203#issuecomment-6104908429
@ryanmeowy, thanks for digging into the runtime path — this is the right boundary to work on. Your framing matches my understanding: the enumerator fixes its table set once (`IncrementalSourceEnumerator.run()` only calls `assignSplits()`, and `IncrementalSplitAssigner` builds incremental splits from the startup tables), and a reader already running `IncrementalSourceStreamFetcher` never picks up a new split because `IncrementalSourceSplitReader.checkSplitOrStartNext()` returns early. Reusing the existing snapshot unit (`MySqlSnapshotSplitReadTask` low/high read plus the bounded `[low, high]` backfill in `MySqlSnapshotFetchTask`) as the admission mechanism is a reasonable starting point. On the two constraints: - Scoping v1 to exactly-once jobs because the backfill only exists under `isExactlyOnce()` is acceptable as an explicit, documented limitation. Please also spell out what happens if the option is enabled on a non-exactly-once job (reject at config validation, or warn and disable?) — silently missing changes is the outcome to avoid. - The `database.server.id` collision needs more work before choosing. Stopping the live client for the admission read is simpler, but it pauses every already-captured table, and "restart from the last emitted position" needs a precise answer for failure mid-way: if the job dies after the snapshot read but before the backfill completes, what is persisted so that recovery neither re-admits the table nor loses the `[low, high]` range? Please evaluate both options (reserved second server-id slot vs. stop/restart) against checkpoint failure and recovery, not only the happy path. On cadence, the point that `AbstractSchemaChangeResolver.support()` only handles `ALTER TABLE` is a good reason to poll. Please also state the interval, who owns the timer, and how it behaves across restore. To make this a reviewable design, the remaining items are the ones listed earlier in this thread: enumerator ownership of discovery, persisted state for admitted and in-flight tables, how a new split is actually delivered to a reader given the early return in `checkSplitOrStartNext()`, duplicate suppression when discovery and restore both see the same table, and the exact MySQL vs. OceanBase scope. Existing `table_pattern` restore semantics must stay backward compatible — the current `capturedTables - checkpointCapturedTables` behavior in `restore()` should not change as a side effect. Please keep this as a design write-up in this issue for now, and hold off on opening a PR until the pending per-table restore watermark work mentioned above is resolved, since both touch the same state and exactly-once path. <!-- streview-comment:1669 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
