201811510411lw opened a new issue, #12361:
URL: https://github.com/apache/seatunnel/issues/12361

   ### Search before asking
   
   - [x] I had searched in the 
[issues](https://github.com/apache/seatunnel/issues?q=is%3Aissue+label%3A%22bug%22)
 and found no similar issues.
   
   
   ### What happened
   
   A MySQL CDC initial snapshot repeatedly fails on the last split of a large, 
actively written table when exactly_once=true.
   
   Observed behavior:
   - snapshot.split.size=81920, Source parallelism=2.
   - The table was divided into 658 snapshot splits. The first 657 completed, 
but the final split repeatedly failed.
   - The final split has a lower bound but no upper bound: splitStart=[L], 
splitEnd=null. Its query is equivalent to SELECT * FROM source_db.large_table 
WHERE ID >= L.
   - This split was reached approximately 17 hours after planning, while 
application inserts continued.
   - A read-only count later found 812,737 rows in this range, approximately 
9.9 times the configured split-size target.
   
   With 8 GiB TaskManager process memory and 5 GiB task heap, the reader 
stalled around 379,108 rows with heavy Full GC and a TaskManager heartbeat 
timeout.
   
   We increased process memory to 10 GiB and task heap to 8 GiB, then resumed 
from a completed savepoint. A subsequent attempt reached 566,206 rows before 
failing with:
   java.lang.OutOfMemoryError: GC overhead limit exceeded
   
   Another Reader was processing a different split concurrently, so these are 
whole-TaskManager memory observations.
   
   Checkpoints completed before the failures, but recovery retried the same 
unfinished final split and rebuilt its buffer.
   
   The suspected mechanism is the combination of an open-ended final split and 
whole-split in-memory buffering in pollSplitRecordsIfExactlyOnce(). Continued 
inserts can make the final split much larger than it was when initially planned.
   
   Expected behavior:
   The snapshot should have a supported way to complete within a bounded memory 
budget while preserving concurrent insert/update/delete handling and 
checkpoint/savepoint consistency.
   
   Is there an existing safe configuration or recommended mitigation for this 
case?
   
   This describes an observed workload; a minimal automated reproducer is not 
yet available.
   
   ### SeaTunnel Version
   
   2.3.13
   
   ### SeaTunnel Config
   
   ```conf
   Relevant configuration excerpt. Connection details, the full table list, 
transforms and sink settings are omitted; this is not a standalone reproducer.
   
   
   env {
     job.mode = "STREAMING"
     parallelism = 2
     checkpoint.interval = 300000
   }
   
   source {
     MySQL-CDC {
       startup.mode = "initial"
       exactly_once = true
       snapshot.split.size = 81920
     }
   }
   
   
   The source uses an increasing `BIGINT` primary key. The topology is MySQL 
CDC -> transforms -> Paimon multi-table sink, with Source and writer 
parallelism 2.
   
   Flink settings for the second attempt:
   
   
   taskmanager.memory.process.size: 10240m
   taskmanager.memory.task.heap.size: 8g
   taskmanager.memory.managed.size: 256m
   taskmanager.numberOfTaskSlots: 2
   state.backend: rocksdb
   ```
   
   ### Running Command
   
   ```shell
   Submitted through Flink Kubernetes Operator in application mode:
   
   
   Entry class: org.apache.seatunnel.core.starter.flink.SeaTunnelFlink
   Arguments: --config /opt/seatunnel.streaming.conf
   ```
   
   ### Error Exception
   
   ```log
   java.lang.OutOfMemoryError: GC overhead limit exceeded
   
   
   The TaskManager fatal exception handler stopped the process. JobManager 
subsequently reported a TaskManager heartbeat timeout and restored the job from 
a completed checkpoint. The text above is the observed error message; no 
synthetic stack trace is included.
   ```
   
   ### Zeta or Flink or Spark Version
   
   Flink 1.18.1
   
   ### Java or Scala Version
   
   Java 8
   
   ### Screenshots
   
   _No response_
   
   ### Are you willing to submit PR?
   
   - [ ] Yes I am willing to submit a PR!
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to