SEZ9 commented on issue #12060:
URL: https://github.com/apache/seatunnel/issues/12060#issuecomment-5564633681

   @goutamadwant thanks for the calibration and the follow-up reproduction.
   
   On the A/A calibration: agreed that identical-code post-readiness ratios of 
0.951–1.216 on Java 8 and 0.921–1.046 on Java 11 do not support a production 
speedup claim or a small regression threshold yet. Separating startup-inclusive 
timing from writes after coordinator activation and read-back was the right 
change, and keeping production storage and cleanup untouched is correct. I also 
agree with your caveat that activation plus read-back does not mean background 
startup or JIT activity has settled, so I would not lock in a readiness 
boundary from those runs alone.
   
   Your local reproduction on the f6f18be7 artifact (pipelineCount=1, 
storedDagCount=100, 3 forks, 3 warmups, 5 measurements) shows declining samples 
within each fork with no overlapping GC pauses in those small-DAG batches, and 
the teardown behavior you describe (three sampled values reloaded through 
full-WAL scans, tombstone appends, allocation across those reads growing from 
about 70 MB to 140 MB with 100 pipelines) means later samples run under 
different conditions than earlier ones. That fixture-history effect needs to be 
removed before the readiness boundary or confirmation budget can be decided, 
otherwise the A/A spread is partly measuring the fixture.
   
   Suggested order, in line with the previous comment:
   
   1. Fix the fixture independently: bound the teardown verification so it no 
longer does full-WAL scans and does not grow work across iterations. Keep this 
benchmark-only and separate from the framework additions, which stay on hold as 
you proposed.
   2. Re-run the same parameter combination with identical settings (forks, 
warmups, measurements, writes per batch, JVM flags) on the corrected fixture 
and post the per-fork samples in the same format for comparison.
   3. Repeat the Java 8 and Java 11 A/A calibration on the corrected fixture. 
If the identical-code ratios tighten, that spread can inform the readiness 
boundary and confirmation budget; if not, the remaining variance is not 
fixture-driven.
   4. Once that new baseline exists, continue with the remaining non-GC waits 
and the large-DAG stalls that overlap G1 pauses.
   
   A benchmark-only PR for the fixture fix is welcome whenever it is ready, and 
reproduction results, ruled-out hypotheses and open questions are welcome here 
as you go.
   
   <!-- streview-comment:867 -->


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to