goutamadwant opened a new issue, #12189: URL: https://github.com/apache/seatunnel/issues/12189
### Search before asking - [x] Searched existing issues and PRs. This is a focused follow-up to #11616, not a replacement for the umbrella. Existing JMH comparison work (#12089, #12021) measures a different benchmark; #11553 supplies the existing CLI single-run framework. ### Description Add an offline comparison of two saved CLI benchmark runs so an aggregate improvement cannot hide tasks that regressed. The discussion in #11616 asks for task-level pass-to-fail and fail-to-pass transitions alongside headline deltas. Current single-run reports do not pair those identities across revisions. This issue covers only that reporting slice, not a held-out dataset, prompt changes, or an automatic CI acceptance gate. ### Usage Scenario A contributor changes connector knowledge or config-generation behavior. They have baseline and candidate `results.json` files and need to see which tasks improved, regressed, or could not be compared, without another model call. ### Reproduction and current behavior Baseline: `dev` at `8bea8c681cacccbd99fa642ca0853718980b9234`. An offline fixture uses one model, one trial per task, and three tasks with identical recorded settings: - Baseline: `lost` passes; `gained-a` and `gained-b` fail: 1/3 passed. - Candidate: `lost` fails; `gained-a` and `gained-b` pass: 2/3 passed. - Existing single-run Markdown reports show 33.3% and 66.7%, but neither identifies the paired regression of `lost`. This is a reporting-gap reproduction with synthetic saved results, not a measured model-accuracy improvement. The existing 17 benchmark tests passed before implementation. Importing the proposed comparison module failed on the baseline because that module does not exist. The proposed branch includes `seatunnel-cli/tests/test_benchmark_comparison.py`. Its `test_aggregate_gain_does_not_hide_task_regression` exercises the existing single-run renderer and the new comparator against the same inputs: ```shell cd seatunnel-cli python -m pytest -q tests/test_benchmark_comparison.py ``` Use the CLI's documented development dependencies for this test command. The comparison command itself is standard-library-only. ### Proposed behavior ```shell python -m benchmark.compare baseline/results.json candidate/results.json --out comparison.md ``` - Pair model/task/trial identities and show first-attempt and within-repair-budget transitions plus aggregate deltas. - In the fixture, report +33.3 percentage points while explicitly showing `lost` as pass-to-fail and the other two tasks as fail-to-pass. - Require matching recorded model settings, task fingerprints, requested gates, trial counts, and repair budgets. - List missing, skipped, incomplete, or incompatible evidence and exclude it from both denominators. Do not turn missing evidence into an apparent failure or improvement. - Never overwrite an existing output file or modify input files. Do not print raw model configuration or error payloads. ### Compatibility and limits New runs add a task fingerprint; existing single-run Markdown and CSV output remain byte-for-byte unchanged. Old results without fingerprints need fresh runs for a supported comparison; do not infer their task definitions from today's files. Recorded metadata cannot prove equivalent external model versions, validator behavior, or execution environments. The report describes observations, not causation, statistical significance, or a CI pass/fail policy. No provider calls or SeaTunnel jobs are executed by the comparator, and no new dependency is introduced. ### Validation - Python 3.10.20 and 3.11.15: complete CLI suite passed on each, 129 tests plus 3 subtests, including 46 comparison cases. - Coverage includes real runner serialization, legitimate generation/gate failures, inconsistent records, skipped gates, deterministic output, unchanged single-run reports, and input/output preservation. - A separate Python `-S` smoke check confirms standard-library-only operation. - Black, Ruff, repository Spotless, and full-repository `./mvnw -q -DskipTests verify` passed. The full build used Java 11 and skips tests; Java 8 runtime testing is not applicable to this Python feature. ### Related issues and implementation Related to #11616. This does not claim ownership of its other roadmap items or prior maintainer approval of this exact implementation. [Reviewable compare branch](https://github.com/apache/seatunnel/compare/dev...goutamadwant:feature/benchmark-revision-comparison). A PR will be linked after review. ### Are you willing to submit a PR? - [x] Yes, I am willing to submit a PR. ### Code of Conduct - [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
