SEZ9 commented on issue #12058: URL: https://github.com/apache/seatunnel/issues/12058#issuecomment-5628347156
@Rangsh thanks — this is what I was asking for: wall/CPU flame graphs and the GC summary taken on unchanged `origin/dev` rather than on modified code. The screenshots plus the attached `flame-wall-forward.html` / `flame-wall-reverse.html` / `flame-cpu-forward.html` / `flame-cpu-reverse.html` and `gc-profile-summary.md` give us a baseline to read. Two things to close the root-cause step before we go back to any code changes: 1. **Name the root cause from the baseline.** The CPU zoom shows `WALWorkHandler` → `HdfsWriter.write/flush` → `FSDataOutputStream.hsync` → `RawLocalFileSystem$LocalFSFileOutputStream.write`, but `hsync` only matches `0.82%` of CPU samples. Could you state, from `flame-wall-reverse.html`, which frame(s) dominate wall time for `CheckpointStorageBenchmark.checkpointOverviewIncrementalUpdate$`? Is the time spent blocked in the sync path (wall, not CPU), or somewhere else? One sentence with the dominating stack and its share is enough. If the `205.64` / `115.29` samples do not line up with GC pauses, that also points back to this path. 2. **Keep the loop as agreed.** Once we have a single stated cause on `origin/dev`, apply one change at a time and re-run `profile_benchmarks.sh` (`wall` / `cpu` / `gc`) to show the original hotspot is gone from the flame graph / GC summary — not just that the A/B number improved. That is the verification I want to see before we discuss the fix itself. The environment (Corretto 11.0.26 / Apple M1, `tools/benchmarks/profile_benchmarks.sh`) is good; please keep it identical for the follow-up run so the before/after flames are comparable. <!-- streview-comment:948 --> -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
