goutamadwant commented on issue #12060: URL: https://github.com/apache/seatunnel/issues/12060#issuecomment-5557301528
@nzw921rx thanks for clarifying. I went back to the original finishedJobDagStore and reproduced the variation locally before using separate diagnostic experiments. Setup: original artifact from f6f18be7, macOS ARM, Temurin 11.0.32.1+1, local file storage, 4 GiB heap, G1, AlwaysPreTouch, DisableExplicitGC, ActiveProcessorCount=4 and preferIPv4Stack=true. ActiveProcessorCount controls JVM ergonomics, not CPU affinity. For pipelineCount=1 and storedDagCount=100, I retained 3 forks, 3 single-shot warmups and 5 measurements, with 100 writes per batch. All measured samples below are us/DAG, in iteration order: - Fork 1: 211.88750, 174.54000, 177.11208, 150.49709, 123.12500 - Fork 2: 181.67375, 160.18583, 155.25875, 163.31625, 127.75667 - Fork 3: 189.42084, 159.59458, 132.79375, 147.09708, 108.95291 Separate batch-level recordings still show declining measurements with the coordinator active and no overlapping GC pauses in those small-DAG batches. That weakens coordinator readiness as a sufficient explanation, but does not rule out all background activity. Further diagnostic runs found: - serialization CPU cost falls during measurement, and exercising the actual write path changes the pattern beyond simply waiting - teardown reloads three sampled values through full-WAL scans; deletes append tombstones, so verification work grows between iterations - with 100 pipelines, allocation across those three reads grows from about 70 MB after the first measured batch to 140 MB after the last - some large-DAG stalls overlap recorded G1 pauses, while other non-GC waits remain unresolved The verification allocations are outside the write timer. They change conditions for later samples rather than directly contributing to the current Score. I also tested isolated WAL state. It bounds verification work but did not consistently improve latency, so I am not treating it as a variance fix. Diagnostic scores remain separate from the unchanged reproduction. These findings do not establish a production defect or explain every historical CI outlier. I will keep the proposed framework additions on hold. Does investigating the remaining non-GC waits look like the most useful next step, or should we first discuss the fixture-history effect? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
