SEZ9 commented on issue #12266: URL: https://github.com/apache/seatunnel/issues/12266#issuecomment-5862190429
Thanks — the storage-first split (Phase A keyed load with an explicitly selective file implementation, then Phase B widening the durability sample outside the measured SingleShot window) looks right, and holding both PRs until #12173 is healthy sounds good. Two things would help make the "no OOM under `initialStoredJobCount=1000`" criterion verifiable: could you note the heap budget used for the normal Benchmarks / Diagnostics forks, and roughly how much heap the earlier full-WAL reload consumed when it OOM'd? If those numbers are not readily available, a constrained-heap diagnostic run alongside Phase A's deterministic tests would be fine as supporting evidence, as Daniel suggested. When Phase B lands, it would also be good to include a test or assertion that the durability sample only runs in the iteration/trial tear-down hooks, so the measured growth scores cannot silently absorb the reload cost. <!-- streview-comment:1364 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
