DanielLeens commented on PR #11780:
URL: https://github.com/apache/seatunnel/pull/11780#issuecomment-5322696028

   Correction to my CI note in the review above (2026-08-13): I wrote that 
`all-connectors-it-1` (`PostgresCDCIT#testAddFieldWithRestore`) and 
`all-connectors-it-6` (`OpengaussCDCIT`) were "entirely untouched by this PR's 
diff" and looked like pre-existing flakiness. That was wrong — I re-checked the 
actual job logs and both failures are a deterministic 
`java.lang.ArrayIndexOutOfBoundsException: 4`, same 3-frame stack on every 
retry:
   
   ```
   at 
org.apache.seatunnel.api.table.type.SeaTunnelRowType.getFieldType(SeaTunnelRowType.java:73)
   at 
org.apache.seatunnel.api.table.type.SeaTunnelRow.getBytesSize(SeaTunnelRow.java:148)
   at 
org.apache.seatunnel.engine.server.task.SeaTunnelSourceCollector.collect(SeaTunnelSourceCollector.java:253)
   ```
   
   In both jobs this fires immediately after the new `Restored source collector 
schema from checkpoint for tables: [...]` log line — i.e. inside 
`SeaTunnelSourceCollector#restoreSchema`, reached via the 
`restoreCollectorSchema()` call this PR adds to 
`IncrementalSourceReader#pollNext()`. `Collector#restoreSchema` and 
`SeaTunnelSourceCollector#restoreSchema` are both new code introduced by this 
PR, so this isn't unrelated CDC IT flakiness: it reproduces on every retry, in 
both the Postgres and openGauss CDC jobs, which share `IncrementalSourceReader` 
with the MySQL path this PR targets (`PostgresCDCIT#testAddFieldWithRestore` 
itself predates this PR).
   
   From the stack, it looks like after restore, `collect()`'s `getBytesSize()` 
call is being made with a `SeaTunnelRowType` that has fewer fields than the row 
actually being collected (index 4 out of bounds) — on the 
add-column-then-restore-from-savepoint scenario. Worth checking whether 
`restoreSchema()` (restoring the pre-evolution checkpointed schema) is racing 
with the CDC add-column `SchemaChangeEvent` handling on the Postgres/openGauss 
WAL-replay path before the first post-restore row is collected — that path may 
behave differently from MySQL's binlog/DDL flow.
   
   Since this reproduces deterministically and sits directly in this PR's new 
schema-restore code, syncing with `dev` and rerunning would not be expected to 
help here. This looks like a real regression that needs to be fixed (or the 
restore mechanism scoped away from Postgres/openGauss if not yet supported 
there) before merge.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to