SEZ9 commented on issue #10203:
URL: https://github.com/apache/seatunnel/issues/10203#issuecomment-6104908429

   @ryanmeowy, thanks for digging into the runtime path — this is the right 
boundary to work on. Your framing matches my understanding: the enumerator 
fixes its table set once (`IncrementalSourceEnumerator.run()` only calls 
`assignSplits()`, and `IncrementalSplitAssigner` builds incremental splits from 
the startup tables), and a reader already running 
`IncrementalSourceStreamFetcher` never picks up a new split because 
`IncrementalSourceSplitReader.checkSplitOrStartNext()` returns early. Reusing 
the existing snapshot unit (`MySqlSnapshotSplitReadTask` low/high read plus the 
bounded `[low, high]` backfill in `MySqlSnapshotFetchTask`) as the admission 
mechanism is a reasonable starting point.
   
   On the two constraints:
   
   - Scoping v1 to exactly-once jobs because the backfill only exists under 
`isExactlyOnce()` is acceptable as an explicit, documented limitation. Please 
also spell out what happens if the option is enabled on a non-exactly-once job 
(reject at config validation, or warn and disable?) — silently missing changes 
is the outcome to avoid.
   - The `database.server.id` collision needs more work before choosing. 
Stopping the live client for the admission read is simpler, but it pauses every 
already-captured table, and "restart from the last emitted position" needs a 
precise answer for failure mid-way: if the job dies after the snapshot read but 
before the backfill completes, what is persisted so that recovery neither 
re-admits the table nor loses the `[low, high]` range? Please evaluate both 
options (reserved second server-id slot vs. stop/restart) against checkpoint 
failure and recovery, not only the happy path.
   
   On cadence, the point that `AbstractSchemaChangeResolver.support()` only 
handles `ALTER TABLE` is a good reason to poll. Please also state the interval, 
who owns the timer, and how it behaves across restore.
   
   To make this a reviewable design, the remaining items are the ones listed 
earlier in this thread: enumerator ownership of discovery, persisted state for 
admitted and in-flight tables, how a new split is actually delivered to a 
reader given the early return in `checkSplitOrStartNext()`, duplicate 
suppression when discovery and restore both see the same table, and the exact 
MySQL vs. OceanBase scope. Existing `table_pattern` restore semantics must stay 
backward compatible — the current `capturedTables - checkpointCapturedTables` 
behavior in `restore()` should not change as a side effect.
   
   Please keep this as a design write-up in this issue for now, and hold off on 
opening a PR until the pending per-table restore watermark work mentioned above 
is resolved, since both touch the same state and exactly-once path.
   
   <!-- streview-comment:1669 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to