DanielLeens commented on PR #12258:
URL: https://github.com/apache/seatunnel/pull/12258#issuecomment-5661833626

   Pushed `16e377bb7b` (test-only, `MariaDbCDCCheckpointRestoreIT.java`): since 
the `order by id` fix in `2367cfc7` still hits the identical `index [6][0], 
expected: <21> but was: <12>` failure on its own CI run, sorting has already 
ruled out physical-scan-order as the cause — a sorted comparison can't diverge 
at a fixed index/value pair from row ordering alone. That means the two result 
sets actually differ in content, and the single differing-index message can't 
distinguish "sink has an extra duplicate `12`" from "sink hasn't caught up to 
`21`/`22` yet within the 2-minute window."
   
   Working backwards from the reported index/value: for `21` to land at 
position 6 in a correctly-converged, `ORDER BY id`-sorted 8-row set 
(`1,2,3,11,12,12,21,22`), the sink would have to contain *more* rows with 
`id=12` than the two the test expects (one from before the checkpoint, one 
replayed after `dropPrimaryKey`) — consistent with the restored job also 
redelivering the `11`/`12` rows that were already part of the checkpoint being 
restored from, on top of the correctly-replayed post-checkpoint duplicate. I 
want to flag this as a reasoned hypothesis from the numbers, not an observed 
fact — I don't have a live dump of the sink table at failure time, and I did 
not find the redelivery decision itself inside this connector's own code 
(`MariaDbBinlogFetchTask`/`MariaDbSourceFetchTaskContext` delegate binlog 
resume entirely to the reused 
`io.debezium.connector.mysql.MySqlStreamingChangeEventSource`); it would come 
from whatever offset/GTID state this connector's own `MariaDbSource
 FetchTaskContext`/`MariaDbGtidUtils` hands to that shared reader on restore, 
which I have not traced to a definitive line yet.
   
   Rather than guess further at production code without ground truth, I added a 
diagnostic-only change: `awaitSourceAndSinkConsistent` now logs the full source 
and sink row sets when the polling window is exhausted, so the next CI run 
shows the actual divergence directly (extra duplicate vs. missing/lagging row) 
instead of leaving it to be inferred. No production code touched.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to