goutamadwant opened a new issue, #12240: URL: https://github.com/apache/seatunnel/issues/12240
### Search before asking - [x] I searched existing issues and pull requests. This is a focused follow-up to the benchmark direction in #11616, not a duplicate of the existing comparison implementation. ### Description Add an opt-in public paraphrase regression suite to the CLI benchmark while keeping the existing 100-task benchmark unchanged. The current baseline uses fixed prompt wording. A separate set of equivalent requests would let contributors check whether a knowledge-pack or generation change remains stable when the same task is expressed differently. Each alternative prompt should reuse its original task's assertions and execution fixtures, so wording changes do not silently change the evaluation contract. This proposal adds evaluation infrastructure. It does not claim that a particular model currently fails these prompts or that model accuracy has improved. #### Reproduction of the missing capability Baseline: `apache/seatunnel` revision `75fd4ed4b2e63579a57b461285997951490c4fe7`. From `seatunnel-cli`, run: ```bash python -m benchmark.runner --suite paraphrase --level l1 ``` On the original baseline, this was confirmed on Python 3.10 with exit status 2: ```text runner.py: error: unrecognized arguments: --suite paraphrase ``` Argument parsing stopped before any provider call. The original loader returned the existing 100 tasks; there was no selectable paraphrase suite. This is a reproducible capability gap, not a model-quality failure. #### Proposed behavior and acceptance criteria - Add `--suite paraphrase`; retain `--suite baseline` as the default. - Start with 12 reviewed alternative prompts: four routing tasks, four CDC tasks and four connector-option/mode tasks, including one Chinese prompt. Selecting this suite runs the variants only, not a combined 112-task baseline. - Define variants using only parent task ID, full parent-contract fingerprint and alternative prompt text. Copy assertions, schemas, execution fixtures and other evaluation fields from the canonical task without allowing overrides. - Give variants distinct IDs and retain parent provenance in saved results. A changed parent contract requires explicit review and repinning. - Reject missing/duplicate parents, invalid definitions, unchanged/empty wording and invalid variant selections before provider setup. - Reuse existing generation, scoring, repair, reporting and revision comparison. Preserve baseline task order, definitions, fingerprints and report formats. - Document the suite and its limitations in English, Chinese and the benchmark README. #### Before and after Before: contributors can compare results for fixed baseline prompts, but cannot select a reviewed alternative-wording suite through the benchmark runner. After: contributors can run a separate public regression suite and inspect task-level transitions using the existing comparison tools, without replacing or expanding the default benchmark. #### Prepared implementation and validation Prepared commit: `8e364acfb2638ce7ca3c7c51ae035176d9e9e3ae`. From `seatunnel-cli`: ```bash python -m pytest tests -q python -m black --check benchmark/paraphrases.py tests/test_benchmark_paraphrases.py python -m ruff check --isolated --select E4,E7,E9,F benchmark/paraphrases.py benchmark/runner.py tests/test_benchmark_paraphrases.py ``` - Complete CLI suite: 181 tests and 3 subtests passed on each of Python 3.10.20 and Python 3.11.15, compared with 129 tests and 3 subtests on the original baseline. - The 52 added cases cover inheritance, mutation isolation, parent drift, malformed definitions, suite/tier/task selection, scoring parity, saved provenance, skipped execution gates and cross-revision comparison. - All 12 prompts were independently checked against their canonical requirements and fixtures. The original 100-task definitions remain identical. - Root formatting and the full Java 11 no-tests build passed for the branch and unchanged baseline. [Fork Build](https://github.com/goutamadwant/seatunnel/actions/runs/34320639314) also passed for the prepared commit. This is Python benchmark tooling, so Python 3.10/3.11 are the directly tested runtimes. No Java connector/runtime code is changed; Java 8 benchmark execution is not applicable and is not claimed. The full Maven compilation/package check used Java 11. The corpus is public, not an unseen holdout. Running the actual benchmark still calls the configured model, including L1-only runs. The reported offline tests made no model calls and ran no engine jobs; they do not establish accuracy gains, generalization, output-data equivalence or downstream production impact. ### Usage Scenario Regression review for CLI knowledge packs and generation changes: evaluate alternative wording against the same assertions, report pass-to-fail and fail-to-pass transitions, and preserve a reproducible default baseline for historical comparisons. ### Related issues - Benchmark discipline and Stage 1 direction: #11616. - Existing saved-result comparison: #12189 / #12192. This proposal reuses it rather than creating another comparator. - [Prepared compare branch](https://github.com/apache/seatunnel/compare/dev...goutamadwant:test/benchmark-paraphrase-suite). ### Are you willing to submit a PR? - [x] Yes, I am willing to submit a PR. The implementation is prepared on the linked compare branch. ### Code of Conduct - [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
