goutamadwant commented on issue #12060:
URL: https://github.com/apache/seatunnel/issues/12060#issuecomment-5557301528

   @nzw921rx thanks for clarifying. I went back to the original 
finishedJobDagStore and reproduced the variation locally before using separate 
diagnostic experiments.
   
   Setup: original artifact from f6f18be7, macOS ARM, Temurin 11.0.32.1+1, 
local file storage, 4 GiB heap, G1, AlwaysPreTouch, DisableExplicitGC, 
ActiveProcessorCount=4 and preferIPv4Stack=true. ActiveProcessorCount controls 
JVM ergonomics, not CPU affinity.
   
   For pipelineCount=1 and storedDagCount=100, I retained 3 forks, 3 
single-shot warmups and 5 measurements, with 100 writes per batch.
   
   All measured samples below are us/DAG, in iteration order:
   - Fork 1: 211.88750, 174.54000, 177.11208, 150.49709, 123.12500
   - Fork 2: 181.67375, 160.18583, 155.25875, 163.31625, 127.75667
   - Fork 3: 189.42084, 159.59458, 132.79375, 147.09708, 108.95291
   
   Separate batch-level recordings still show declining measurements with the 
coordinator active and no overlapping GC pauses in those small-DAG batches. 
That weakens coordinator readiness as a sufficient explanation, but does not 
rule out all background activity.
   
   Further diagnostic runs found:
   - serialization CPU cost falls during measurement, and exercising the actual 
write path changes the pattern beyond simply waiting
   - teardown reloads three sampled values through full-WAL scans; deletes 
append tombstones, so verification work grows between iterations
   - with 100 pipelines, allocation across those three reads grows from about 
70 MB after the first measured batch to 140 MB after the last
   - some large-DAG stalls overlap recorded G1 pauses, while other non-GC waits 
remain unresolved
   
   The verification allocations are outside the write timer. They change 
conditions for later samples rather than directly contributing to the current 
Score.
   
   I also tested isolated WAL state. It bounds verification work but did not 
consistently improve latency, so I am not treating it as a variance fix. 
Diagnostic scores remain separate from the unchanged reproduction.
   
   These findings do not establish a production defect or explain every 
historical CI outlier. I will keep the proposed framework additions on hold. 
Does investigating the remaining non-GC waits look like the most useful next 
step, or should we first discuss the fixture-history effect?


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to