dybyte commented on issue #10329:
URL: https://github.com/apache/seatunnel/issues/10329#issuecomment-5480631374

   > Adding a production data point, because this issue is currently framed as 
"storage grows" — in our case it made the cluster **unbootable**, and the 
failure looks like something else entirely (a "poisoned"/corrupt IMap store), 
which cost us two days of diagnosis.
   > 
   > **Setup:** SeaTunnel 2.3.13, Zeta single node (`master_and_worker`, 
Hazelcast 5.1), 40 streaming SQL Server-CDC → Doris jobs, `parallelism = 1`, 
`checkpoint.interval = 5000`. `map-store` enabled for `engine*` with 
`initial-mode: EAGER` and `FileMapStoreFactory` (`type: hdfs`, `fs.defaultFS: 
file:///`, local volume). Heap `-Xms1g -Xmx5g`, container memory limit 7 GiB.
   > 
   > ### Growth
   > One `wal.txt` segment per JVM run, `engine_runningJobMetrics`:
   > 
   > segment written    duration        size
   > 9 jobs running     1 h 16 min      18 MB
   > 40 jobs running    23 h 25 min     1.52 GB
   > 40 jobs running    6 h 9 min       402 MB
   > 40 jobs running    **7 d 14 h**    **11.95 GB**
   > ≈ **65 MB/h ≈ 1.5 GB/day**, stable across segments. Whole map ≈ 13 GB. 
Next biggest map, `engine_checkpoint-id-map`, 492 MB; every other `engine*` map 
combined < 6 MB. The live content of the metrics map is ~40 entries.
   > 
   > The size is invisible while the cluster runs, because the WAL is only ever 
read back at startup. Ours had not been restarted for 7.5 days when a routine 
`dnf update` + reboot happened.
   > 
   > ### What happens at the next start
   > `CoordinatorService` becomes master, restores the running jobs, and that 
triggers the EAGER map-store load of ~12 GB into a 5 GB heap:
   > 
   > ```
   > java.lang.OutOfMemoryError: Java heap space
   > Dumping heap to /tmp/seatunnel/dump/zeta-server/java_pid13.hprof ...
   > /opt/seatunnel/bin/seatunnel-cluster.sh: line 201: 13 Killed   java 
${JAVA_OPTS} ...
   > ```
   > 
   > (first JVM killed by the cgroup OOM-killer while writing the heap dump; 
the restarts after it die inside the JVM instead:)
   > 
   > ```
   > at 
org.apache.seatunnel.engine.server.CoordinatorService.restoreAllRunningJobFromMasterNodeSwitch(CoordinatorService.java:499)
   > at 
org.apache.seatunnel.engine.server.CoordinatorService.restoreJobFromMasterActiveSwitch(CoordinatorService.java:528)
   > at 
com.hazelcast.map.impl.proxy.MapProxySupport.initializeMapStoreLoad(MapProxySupport.java:330)
   > at 
com.hazelcast.instance.impl.OutOfMemoryErrorDispatcher.onOutOfMemory(OutOfMemoryErrorDispatcher.java:187)
   > com.hazelcast.core.HazelcastInstanceNotActiveException: Hazelcast instance 
is not active!
   > WARN [h.m.i.r.BasicRecordStoreLoader] - Could not load keys from map store 
  (repeated, once per partition record store)
   > ```
   > 
   > 49 `OutOfMemoryError`s in ~12 minutes. `OutOfMemoryErrorDispatcher` shuts 
the Hazelcast instance down, the starter's main thread exits — but non-daemon 
threads keep PID 1 alive, so **the container stays `Up` and the Docker health 
check can even flip back to healthy while REST 8080 is never opened**. That is 
the part that makes this hard to recognise: it does not crash-loop, it zombies. 
Pipelines were frozen for two days.
   > 
   > A plain restart replays the same file and dies the same way. 
`initial-mode: LAZY` is not a workaround either — restores then race the load 
(`RetryableHazelcastException: Map engine_ownedSlotProfilesIMap is still 
loading data from external store` → `Job init failed` → `UNKNOWABLE`).
   > 
   > Recovery that works: stop the node, move `state/imap` aside (**keep 
`state/checkpoint`**), start, re-submit each job with 
`isStartWithSavePoint=true`. All 40 jobs came back from their checkpoints with 
no data loss.
   > 
   > ### Suggestions
   > 1. Since [[Feature][Zeta] Add compaction support to IMAP external storage 
#10399](https://github.com/apache/seatunnel/pull/10399) is a large change and 
still in review, a **bounded interim guard** would already remove the 
"unbootable cluster" class of failure: either a max total size per map for the 
file store (rotate/reset instead of growing forever), or at least a WARN when a 
map's store passes some fraction of `-Xmx` at load time. Today nothing in the 
logs hints at the size before it is too late.
   > 2. Not persisting the metrics maps at all would remove ~95% of the write 
volume for free — filed separately as [[Improve][Zeta] Do not persist 
engine_runningJobMetrics / engine_finishedJobMetrics via map-store 
#12022](https://github.com/apache/seatunnel/issues/12022), since it is a 
config-level default that could ship well before [[Feature][Zeta] Add 
compaction support to IMAP external storage 
#10399](https://github.com/apache/seatunnel/pull/10399) lands.
   > 
   > We now run with the metrics maps excluded from `map-store` and a watchdog 
that rebuilds the store when it passes 1 GB: the store went from 13 GB to a few 
MB, and restarts restore all 40 jobs in ~90 s.
   
   @albgen Thanks for the detailed production data.
   
   `engine_runningJobMetrics` has already been excluded from persistence by 
#11244, so I'm now trying to understand whether the remaining append-only 
growth is significant enough to justify adding a general compaction mechanism.
   
   You measured `engine_checkpoint-id-map` growing at around 60 MB/day. In your 
production environment, do you consider that growth a real operational problem 
over long-running deployments, or would it be acceptable after excluding the 
metrics map?
   
   I'm working on #10399, but since it introduces a relatively large 
compaction/storage change, I'd like to understand whether the remaining growth 
actually justifies that complexity.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to