goutamadwant opened a new issue, #12189:
URL: https://github.com/apache/seatunnel/issues/12189

   ### Search before asking
   
   - [x] Searched existing issues and PRs. This is a focused follow-up to 
#11616, not a replacement for the umbrella. Existing JMH comparison work 
(#12089, #12021) measures a different benchmark; #11553 supplies the existing 
CLI single-run framework.
   
   ### Description
   
   Add an offline comparison of two saved CLI benchmark runs so an aggregate 
improvement cannot hide tasks that regressed.
   
   The discussion in #11616 asks for task-level pass-to-fail and fail-to-pass 
transitions alongside headline deltas. Current single-run reports do not pair 
those identities across revisions. This issue covers only that reporting slice, 
not a held-out dataset, prompt changes, or an automatic CI acceptance gate.
   
   ### Usage Scenario
   
   A contributor changes connector knowledge or config-generation behavior. 
They have baseline and candidate `results.json` files and need to see which 
tasks improved, regressed, or could not be compared, without another model call.
   
   ### Reproduction and current behavior
   
   Baseline: `dev` at `8bea8c681cacccbd99fa642ca0853718980b9234`.
   
   An offline fixture uses one model, one trial per task, and three tasks with 
identical recorded settings:
   
   - Baseline: `lost` passes; `gained-a` and `gained-b` fail: 1/3 passed.
   - Candidate: `lost` fails; `gained-a` and `gained-b` pass: 2/3 passed.
   - Existing single-run Markdown reports show 33.3% and 66.7%, but neither 
identifies the paired regression of `lost`.
   
   This is a reporting-gap reproduction with synthetic saved results, not a 
measured model-accuracy improvement. The existing 17 benchmark tests passed 
before implementation. Importing the proposed comparison module failed on the 
baseline because that module does not exist.
   
   The proposed branch includes 
`seatunnel-cli/tests/test_benchmark_comparison.py`. Its 
`test_aggregate_gain_does_not_hide_task_regression` exercises the existing 
single-run renderer and the new comparator against the same inputs:
   
   ```shell
   cd seatunnel-cli
   python -m pytest -q tests/test_benchmark_comparison.py
   ```
   
   Use the CLI's documented development dependencies for this test command. The 
comparison command itself is standard-library-only.
   
   ### Proposed behavior
   
   ```shell
   python -m benchmark.compare baseline/results.json candidate/results.json 
--out comparison.md
   ```
   
   - Pair model/task/trial identities and show first-attempt and 
within-repair-budget transitions plus aggregate deltas.
   - In the fixture, report +33.3 percentage points while explicitly showing 
`lost` as pass-to-fail and the other two tasks as fail-to-pass.
   - Require matching recorded model settings, task fingerprints, requested 
gates, trial counts, and repair budgets.
   - List missing, skipped, incomplete, or incompatible evidence and exclude it 
from both denominators. Do not turn missing evidence into an apparent failure 
or improvement.
   - Never overwrite an existing output file or modify input files. Do not 
print raw model configuration or error payloads.
   
   ### Compatibility and limits
   
   New runs add a task fingerprint; existing single-run Markdown and CSV output 
remain byte-for-byte unchanged. Old results without fingerprints need fresh 
runs for a supported comparison; do not infer their task definitions from 
today's files.
   
   Recorded metadata cannot prove equivalent external model versions, validator 
behavior, or execution environments. The report describes observations, not 
causation, statistical significance, or a CI pass/fail policy. No provider 
calls or SeaTunnel jobs are executed by the comparator, and no new dependency 
is introduced.
   
   ### Validation
   
   - Python 3.10.20 and 3.11.15: complete CLI suite passed on each, 129 tests 
plus 3 subtests, including 46 comparison cases.
   - Coverage includes real runner serialization, legitimate generation/gate 
failures, inconsistent records, skipped gates, deterministic output, unchanged 
single-run reports, and input/output preservation.
   - A separate Python `-S` smoke check confirms standard-library-only 
operation.
   - Black, Ruff, repository Spotless, and full-repository `./mvnw -q 
-DskipTests verify` passed. The full build used Java 11 and skips tests; Java 8 
runtime testing is not applicable to this Python feature.
   
   ### Related issues and implementation
   
   Related to #11616. This does not claim ownership of its other roadmap items 
or prior maintainer approval of this exact implementation.
   
   [Reviewable compare 
branch](https://github.com/apache/seatunnel/compare/dev...goutamadwant:feature/benchmark-revision-comparison).
 A PR will be linked after review.
   
   ### Are you willing to submit a PR?
   
   - [x] Yes, I am willing to submit a PR.
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct).
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to