201811510411lw commented on issue #12361:
URL: https://github.com/apache/seatunnel/issues/12361#issuecomment-5709711355
Additional source references from SeaTunnel 2.3.13:
1. **The final snapshot split has no upper key bound.**
The splitter appends `ChunkRange.of(chunkStart, null)`:
[AbstractJdbcSourceChunkSplitter.java, lines
338–339](https://github.com/apache/seatunnel/blob/8c9d47f5ea7e634d2d27952c7ee7a55269218950/seatunnel-connectors-v2/connector-cdc/connector-cdc-base/src/main/java/org/apache/seatunnel/connectors/cdc/base/source/enumerator/splitter/AbstractJdbcSourceChunkSplitter.java#L338-L339).
2. **Exactly-once processing buffers the entire split in memory.**
`pollSplitRecordsIfExactlyOnce()` accumulates records in a
`LinkedHashMap` until snapshot and binlog backfill processing finish, then
materializes the output list:
[IncrementalSourceScanFetcher.java, lines
148–202](https://github.com/apache/seatunnel/blob/8c9d47f5ea7e634d2d27952c7ee7a55269218950/seatunnel-connectors-v2/connector-cdc/connector-cdc-base/src/main/java/org/apache/seatunnel/connectors/cdc/base/source/reader/external/IncrementalSourceScanFetcher.java#L148-L202).
In our case, inserts continued for approximately 17 hours between split
planning and processing the final split. Because its range is `[L, +infinity)`,
newly inserted rows above the original maximum key also fall into this split. A
later count found **812,737 rows**, although `snapshot.split.size` was
**81,920**. The configured split-size target does not cap the actual row count
of this already-planned final split.
The suspected failure mechanism is therefore **growth of the open-ended
final split combined with whole-split in-memory buffering**. After a failure,
checkpoint recovery retries the same unfinished range, rebuilding the buffer
and encountering memory pressure again.
We understand that the open-ended range may be needed for CDC coverage. Is
there a supported way to handle this growing final split with bounded memory
while preserving snapshot/binlog and recovery consistency?
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]