albgen commented on issue #10329:
URL: https://github.com/apache/seatunnel/issues/10329#issuecomment-5680676242

   @dybyte Yes, but because of the record count, not the MB: 
`FileMapStore.loadAll()` re-reads and deserializes the whole WAL on every call.
   
   Measured on 2.3.13 (40 jobs, checkpoint every 5 s, metrics maps excluded): 
`engine_checkpoint-id-map` grows ~690 k records/day. After ~3 days we had 2.45 
M records, and a start with ~565 k already peaked at 5.1 GB of an 8 GB heap, so 
a restart after about a day of uptime is no longer safe for us.
   
   But this map may not need persistence at all: on restore `CheckpointManager` 
overwrites the counter with `latestCheckpointId + 1` from checkpoint storage. 
If that's right, it could be excluded like `engine_runningJobMetrics` (#11244), 
and the remaining growth is small.
   
   For #12022 I narrowed the scope to `engine_finishedJobMetrics` and opened a 
docs PR: #12328.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to