albgen commented on issue #10329:
URL: https://github.com/apache/seatunnel/issues/10329#issuecomment-5479456762
Adding a production data point, because this issue is currently framed as
"storage grows" — in our case it made the cluster **unbootable**, and the
failure looks like something else entirely (a "poisoned"/corrupt IMap store),
which cost us two days of diagnosis.
**Setup:** SeaTunnel 2.3.13, Zeta single node (`master_and_worker`,
Hazelcast 5.1), 40 streaming SQL Server-CDC → Doris jobs, `parallelism = 1`,
`checkpoint.interval = 5000`. `map-store` enabled for `engine*` with
`initial-mode: EAGER` and `FileMapStoreFactory` (`type: hdfs`, `fs.defaultFS:
file:///`, local volume). Heap `-Xms1g -Xmx5g`, container memory limit 7 GiB.
### Growth
One `wal.txt` segment per JVM run, `engine_runningJobMetrics`:
| segment written | duration | size |
|---|---|---|
| 9 jobs running | 1 h 16 min | 18 MB |
| 40 jobs running | 23 h 25 min | 1.52 GB |
| 40 jobs running | 6 h 9 min | 402 MB |
| 40 jobs running | **7 d 14 h** | **11.95 GB** |
≈ **65 MB/h ≈ 1.5 GB/day**, stable across segments. Whole map ≈ 13 GB. Next
biggest map, `engine_checkpoint-id-map`, 492 MB; every other `engine*` map
combined < 6 MB. The live content of the metrics map is ~40 entries.
The size is invisible while the cluster runs, because the WAL is only ever
read back at startup. Ours had not been restarted for 7.5 days when a routine
`dnf update` + reboot happened.
### What happens at the next start
`CoordinatorService` becomes master, restores the running jobs, and that
triggers the EAGER map-store load of ~12 GB into a 5 GB heap:
```
java.lang.OutOfMemoryError: Java heap space
Dumping heap to /tmp/seatunnel/dump/zeta-server/java_pid13.hprof ...
/opt/seatunnel/bin/seatunnel-cluster.sh: line 201: 13 Killed java
${JAVA_OPTS} ...
```
(first JVM killed by the cgroup OOM-killer while writing the heap dump; the
restarts after it die inside the JVM instead:)
```
at
org.apache.seatunnel.engine.server.CoordinatorService.restoreAllRunningJobFromMasterNodeSwitch(CoordinatorService.java:499)
at
org.apache.seatunnel.engine.server.CoordinatorService.restoreJobFromMasterActiveSwitch(CoordinatorService.java:528)
at
com.hazelcast.map.impl.proxy.MapProxySupport.initializeMapStoreLoad(MapProxySupport.java:330)
at
com.hazelcast.instance.impl.OutOfMemoryErrorDispatcher.onOutOfMemory(OutOfMemoryErrorDispatcher.java:187)
com.hazelcast.core.HazelcastInstanceNotActiveException: Hazelcast instance
is not active!
WARN [h.m.i.r.BasicRecordStoreLoader] - Could not load keys from map store
(repeated, once per partition record store)
```
49 `OutOfMemoryError`s in ~12 minutes. `OutOfMemoryErrorDispatcher` shuts
the Hazelcast instance down, the starter's main thread exits — but non-daemon
threads keep PID 1 alive, so **the container stays `Up` and the Docker health
check can even flip back to healthy while REST 8080 is never opened**. That is
the part that makes this hard to recognise: it does not crash-loop, it zombies.
Pipelines were frozen for two days.
A plain restart replays the same file and dies the same way. `initial-mode:
LAZY` is not a workaround either — restores then race the load
(`RetryableHazelcastException: Map engine_ownedSlotProfilesIMap is still
loading data from external store` → `Job init failed` → `UNKNOWABLE`).
Recovery that works: stop the node, move `state/imap` aside (**keep
`state/checkpoint`**), start, re-submit each job with
`isStartWithSavePoint=true`. All 40 jobs came back from their checkpoints with
no data loss.
### Suggestions
1. Since #10399 is a large change and still in review, a **bounded interim
guard** would already remove the "unbootable cluster" class of failure: either
a max total size per map for the file store (rotate/reset instead of growing
forever), or at least a WARN when a map's store passes some fraction of `-Xmx`
at load time. Today nothing in the logs hints at the size before it is too late.
2. Not persisting the metrics maps at all would remove ~95% of the write
volume for free — filed separately as #12022, since it is a config-level
default that could ship well before #10399 lands.
We now run with the metrics maps excluded from `map-store` and a watchdog
that rebuilds the store when it passes 1 GB: the store went from 13 GB to a few
MB, and restarts restore all 40 jobs in ~90 s.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]