nzw921rx opened a new pull request, #12021:
URL: https://github.com/apache/seatunnel/pull/12021

   ### Purpose of this pull request
   
   The existing benchmark workflow can identify an unexpected JMH `Score`, 
`Error`, or `CV`, and its normal PR mode can compare a candidate with its 
baseline on the same worker. It does not currently retain enough runtime 
evidence to explain whether an unstable or slower result comes from CPU 
hotspots, waiting, lock contention, allocation pressure, or JVM activity.
   
   This PR adds an opt-in diagnostic path for one exact benchmark method:
   
   - adds `cpu`, `wall`, `lock`, `gc`, and `all` profiling choices to the 
manual `Benchmarks` workflow;
   - adds an independent `capture_jfr` choice for offline JVM analysis;
   - requires one exact benchmark method, one selected Java version, and 
exactly one JMH fork, preventing accidental all-benchmark diagnostic runs and 
profiler files being overwritten by later forks;
   - uses JMH's async-profiler integration for CPU, wall-clock, and lock 
diagnostics, retaining the raw JFR and text summary and converting the 
recording to forward and reverse HTML flame graphs with async-profiler's 
bundled `jfrconv`;
   - uses JMH's built-in GC profiler for allocation and collection metrics, and 
its built-in JFR profiler for JVM recordings;
   - uses async-profiler's `ctimer` event on the hosted Linux runner so CPU 
diagnostics do not depend on hardware `perf_event` permissions;
   - runs the `all` choice as separate CPU, wall-clock, lock, and GC steps so 
each profiler has an independent result and artifact directory;
   - generates a standalone JSON and Markdown diagnostic report. Profiled 
scores are explicitly marked as diagnostic and are not mixed with normal 
benchmark results;
   - skips empty lock flame graphs when no contention event was sampled;
   - installs async-profiler 4.5 with download retries and SHA-256 verification;
   - adds focused tests for the new diagnostic runner and report, and adds 
coverage for the existing normalized-result and regression-report tools;
   - runs the Python benchmark-tool tests in a dedicated backend CI job only 
when `tools/benchmarks/**` changes;
   - documents local installation, commands, result interpretation, and 
workflow behavior in English and Chinese.
   
   The two benchmark paths remain intentionally separate:
   
   - normal runs continue to execute the Java 8/11 matrix and support the 
existing `baseline -> PR -> PR -> baseline` same-worker comparison;
   - diagnostic runs analyze one target. With an empty `pr_number`, that target 
is `seatunnel_ref` (default: `dev`); with `pr_number`, it is the selected PR. 
Profiling the baseline is not automatically duplicated because profiler 
overhead makes its score unsuitable for the normal regression comparison.
   
   Scheduled benchmark behavior is unchanged. No new SeaTunnel runtime 
dependency or binary is added.
   
   ### Does this PR introduce _any_ user-facing change?
   
   Yes, for benchmark maintainers and contributors. The manual `Benchmarks` 
workflow gains inputs for an exact diagnostic benchmark, Java 8 or 11, profiler 
mode, optional JFR capture, and optional JMH duration arguments. The same 
diagnostics can also be run locally with 
`tools/benchmarks/profile_benchmarks.sh`.
   
   Normal scheduled benchmark runs and released SeaTunnel runtime behavior are 
unchanged.
   
   ### How was this patch tested?
   
   Formatting:
   
   ```bash
   ./mvnw spotless:apply
   ```
   
   Benchmark-tool unit tests, including selector validation, one-fork 
enforcement, non-empty output rejection, async-profiler conversion, zero-sample 
lock handling, GC report rendering, normalized JMH results, and 
baseline/candidate regression calculations:
   
   ```bash
   python3 -m unittest discover -s tools/benchmarks -p 'test_*.py'
   ```
   
   Result: 21 tests passed.
   
   Workflow and shell validation:
   
   ```bash
   bash -n tools/benchmarks/profile_benchmarks.sh
   actionlint .github/workflows/backend.yml .github/workflows/benchmarks.yml
   git diff --check
   ```
   
   Build the benchmark runner used by the local diagnostic tests:
   
   ```bash
   ./mvnw -Pbenchmark -pl seatunnel-benchmarks -am -DskipTests package
   ```
   
   CPU and lock profiling were run locally against 
`IntermediateQueueBenchmark.disruptorRecordHandoff`. Both produced raw JFR/text 
artifacts and valid forward/reverse HTML flame graphs:
   
   ```bash
   export ASYNC_PROFILER_HOME="$(brew --prefix async-profiler)"
   
   bash tools/benchmarks/profile_benchmarks.sh profile cpu \
     --benchmark 'IntermediateQueueBenchmark.disruptorRecordHandoff$' \
     --output 
seatunnel-benchmarks/target/profiles/intermediate-queue-disruptor-fixed-cpu \
     -- -wi 0 -i 1 -r 1s
   
   bash tools/benchmarks/profile_benchmarks.sh profile lock \
     --benchmark 'IntermediateQueueBenchmark.disruptorRecordHandoff$' \
     --output 
seatunnel-benchmarks/target/profiles/intermediate-queue-disruptor-fixed-lock \
     -- -wi 0 -i 1 -r 1s
   ```
   
   The JMH GC and JFR profilers were also exercised through the new runner. The 
GC run generated the standalone allocation/collection table, and the JFR run 
generated a readable `.jfr` artifact:
   
   ```bash
   bash tools/benchmarks/profile_benchmarks.sh profile gc \
     --benchmark 'IntermediateQueueBenchmark.disruptorRecordHandoff$' \
     --output seatunnel-benchmarks/target/profiles/pr-validation-gc \
     -- -wi 0 -i 1 -r 1s
   
   bash tools/benchmarks/profile_benchmarks.sh capture jfr \
     --benchmark 'IntermediateQueueBenchmark.disruptorRecordHandoff$' \
     --output seatunnel-benchmarks/target/profiles/pr-validation-jfr \
     -- -wi 0 -i 1 -r 1s
   ```
   
   ### Check list
   
   * [x] No new Jar binary package is added.
   * [x] English and Chinese benchmark documentation is updated.
   * [x] No incompatible SeaTunnel API, SPI, configuration, or runtime behavior 
is introduced.
   * [x] This PR does not modify connector code.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to