nzw921rx opened a new issue, #12060:
URL: https://github.com/apache/seatunnel/issues/12060

   ## Description
   
   This is a focused investigation and improvement task for the latency 
variance observed when storing finished JobDAG information.
   
   It is not a performance regression report and does not attempt to compare 
the performance of Java 8 with Java 11.
   
   The `finishedJobDagStore` benchmark shows high sample-to-sample variance 
across every tested parameter combination:
   
   - Java 8 CV: approximately `15.97%–60.53%`
   - Java 11 CV: approximately `10.88%–35.77%`
   
   | Parameters | Java 8 Error / CV | Java 11 Error / CV |
   | --- | ---: | ---: |
   | `pipelineCount=1, storedDagCount=0` | 37.02% / 34.63% | 38.24% / 35.77% |
   | `pipelineCount=1, storedDagCount=100` | 64.71% / 60.53% | 28.26% / 26.44% |
   | `pipelineCount=10, storedDagCount=0` | 25.88% / 24.20% | 34.14% / 31.94% |
   | `pipelineCount=10, storedDagCount=100` | 27.86% / 26.06% | 26.54% / 24.83% 
|
   | `pipelineCount=100, storedDagCount=0` | 19.40% / 18.15% | 11.63% / 10.88% |
   | `pipelineCount=100, storedDagCount=100` | 17.08% / 15.97% | 20.70% / 
19.37% |
   
   Benchmark run:
   
   https://github.com/apache/seatunnel/actions/runs/33723940255
   
   Other benchmarks executed by the same workflow were substantially more 
stable. The repeated variance across the complete parameter matrix is therefore 
worth investigating.
   
   The goals of this issue are to:
   
   1. Identify whether the variance comes from the benchmark fixture, JobDAG 
serialization, Hazelcast IMap operations, WAL/FileMapStore persistence, or the 
execution environment.
   2. Improve the benchmark fixture or methodology if it causes the instability.
   3. Optimize the production storage path if profiling confirms an 
implementation bottleneck.
   4. Preserve finished-job state correctness and durability.
   
   ## Relevant benchmark
   
   ```text
   IMapDagStorageBenchmark.finishedJobDagStore
   ```
   
   ## Profiling tools
   
   For an overview of the Zeta benchmark suite and usage, see the [SeaTunnel 
Zeta Benchmark Guide](https://seatunnel.apache.org/docs/engines/zeta/benchmark).
   
   SeaTunnel provides the `Benchmarks Diagnostics` workflow for running one 
exact benchmark method with CPU, wall-clock, lock, GC, or JFR profiling:
   
   
https://github.com/apache/seatunnel/actions/workflows/benchmarks_diagnostics.yml
   
   The benchmark can also be run directly from the GitHub Actions page by 
selecting **Benchmarks Diagnostics**, clicking **Run workflow**, and providing 
the target branch, tag, commit SHA, or trusted PR number together with the 
exact benchmark method.
   
   ## Local profiling
   
   Build the benchmark module:
   
   ```bash
   ./mvnw -Pbenchmark -pl seatunnel-benchmarks -am -DskipTests package
   ```
   
   Run a profiler locally:
   
   ```bash
   bash tools/benchmarks/profile_benchmarks.sh profile cpu \
     --repository . \
     --benchmark 'IMapDagStorageBenchmark.finishedJobDagStore$'
   ```
   
   Replace `cpu` with `wall`, `lock`, or `gc` as needed.
   
   Capture a JFR recording:
   
   ```bash
   bash tools/benchmarks/profile_benchmarks.sh capture jfr \
     --repository . \
     --benchmark 'IMapDagStorageBenchmark.finishedJobDagStore$'
   ```
   
   CPU, wall-clock, and lock profiling require `ASYNC_PROFILER_HOME`. The 
GitHub Actions workflow installs the required profiler automatically.
   
   ## Expected outcome
   
   This issue should result in both root-cause analysis and a focused 
improvement:
   
   - Identify and explain the primary source of the latency variance.
   - Improve the benchmark when the fixture or methodology is responsible.
   - Optimize the production storage path when an implementation problem is 
confirmed.
   - Add or update benchmark coverage to validate the improvement.
   - Provide before-and-after results collected with the same JDK, runner, 
benchmark arguments, and storage configuration.
   - Compare both operation latency and variance before and after the change.
   - Add appropriate tests and confirm that finished-job state correctness and 
durability remain unchanged.
   
   The implementation should focus on the confirmed source and avoid unrelated, 
broad state-store refactoring.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to