SEZ9 commented on issue #12060: URL: https://github.com/apache/seatunnel/issues/12060#issuecomment-5564633681
@goutamadwant thanks for the calibration and the follow-up reproduction. On the A/A calibration: agreed that identical-code post-readiness ratios of 0.951–1.216 on Java 8 and 0.921–1.046 on Java 11 do not support a production speedup claim or a small regression threshold yet. Separating startup-inclusive timing from writes after coordinator activation and read-back was the right change, and keeping production storage and cleanup untouched is correct. I also agree with your caveat that activation plus read-back does not mean background startup or JIT activity has settled, so I would not lock in a readiness boundary from those runs alone. Your local reproduction on the f6f18be7 artifact (pipelineCount=1, storedDagCount=100, 3 forks, 3 warmups, 5 measurements) shows declining samples within each fork with no overlapping GC pauses in those small-DAG batches, and the teardown behavior you describe (three sampled values reloaded through full-WAL scans, tombstone appends, allocation across those reads growing from about 70 MB to 140 MB with 100 pipelines) means later samples run under different conditions than earlier ones. That fixture-history effect needs to be removed before the readiness boundary or confirmation budget can be decided, otherwise the A/A spread is partly measuring the fixture. Suggested order, in line with the previous comment: 1. Fix the fixture independently: bound the teardown verification so it no longer does full-WAL scans and does not grow work across iterations. Keep this benchmark-only and separate from the framework additions, which stay on hold as you proposed. 2. Re-run the same parameter combination with identical settings (forks, warmups, measurements, writes per batch, JVM flags) on the corrected fixture and post the per-fork samples in the same format for comparison. 3. Repeat the Java 8 and Java 11 A/A calibration on the corrected fixture. If the identical-code ratios tighten, that spread can inform the readiness boundary and confirmation budget; if not, the remaining variance is not fixture-driven. 4. Once that new baseline exists, continue with the remaining non-GC waits and the large-DAG stalls that overlap G1 pauses. A benchmark-only PR for the fixture fix is welcome whenever it is ready, and reproduction results, ruled-out hypotheses and open questions are welcome here as you go. <!-- streview-comment:867 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
