andygrove opened a new issue, #5870:
URL: https://github.com/apache/datafusion-comet/issues/5870

   ### What is the problem the feature request solves?
   
   #5843 split CI into a PR tier and a merge-queue tier. This issue proposes 
the next set of tier changes, based on the last 30 days of CI data (every `CI` 
run on a pull request from Aug 12 to Sep 11, 1,288 runs that actually executed, 
plus 168 push-to-main runs of the full pipeline used as the flake baseline).
   
   "Unique catches" below means PR runs where the job failed while the Linux 
build, lint and the Spark 4.1 Linux test profile were all green, i.e. no 
cheaper job would have flagged the PR.
   
   **Cost and signal of the current PR tier**
   
   | Job | Runner hrs/day | PR fail | Main fail | Unique catches / month | Wall 
clock |
   |---|---|---|---|---|---|
   | Spark SQL 4.1, `sql_core-1/2/3` (+ build) | ~96 | 8.5% | 4.5% | 33 | 100 
min |
   | Spark SQL 4.1, `sql_hive-1/2/3` | ~65 | 1.4% | 1.3% | 7 | 110 min |
   | Iceberg 1.11 | 84 | 6.4% | 2.6% | 17 | 100 min |
   | Linux tests, Spark 3.4 profile | 46 | 11.9% | 6.0% | 25 | 60 min |
   | Linux tests, Spark 3.5 profile | 44 | 11.3% | 6.0% | 20 | 60 min |
   | Linux tests, Spark 4.0 profile | 48 | 10.5% | 5.4% | 15 | 60 min |
   | Linux tests, Spark 4.2 profile | 45 | 10.0% | 4.8% | 19 | 60 min |
   | Linux tests, Spark 4.1 profile | 48 | 9.9% | 4.8% | 58 | 60 min |
   | Linux build, lint, rust-test, TPC verify | 58 | 15.3% | 5.8% | 120 | 55 
min |
   | Delta gate, PyArrow UDF, CodeQL, preflight | ~9 | low | 0% | 33 | under 20 
min |
   
   The PR tier costs roughly 540 runner-hours per day. The five Linux test 
profiles co-fail with each other 70 to 88% of the time. 60% of Iceberg 1.11 
failures die at its "Build Comet" step, which the Linux build already catches.
   
   **The queue tier is not healthy enough to take more jobs yet**
   
   On known-good main, the jobs that are now queue-only went red 32% of the 
time over the month and 40% in September. With `max_entries_to_build: 2` and a 
2.5 hour pipeline that eviction rate will stall the queue. Two culprits:
   
   - macOS: red on main 11% of the time (8% in the week of Sep 8). 79 of its 89 
unique PR failures are the `[scans]` suite crashing with SIGSEGV in 
`hdfsThreadDestructor` (#5023, proposed fix #5036, not yet merged). A Sep 9 
comment on #5023 reports the same crash on `ubuntu-24.04`.
   - Iceberg 1.8/1.9/1.10: red on main 24% of the time in September. About half 
of that was #5758 (fixed Sep 8); the rest is artifact-store 403/409 errors, 
Maven connection resets, and a Scala 2.13 Iceberg test flake.
   
   Without macOS and the old Iceberg versions, the queue tier's red rate on 
main is 5 to 9%.
   
   ### Describe the potential solution
   
   In order:
   
   1. **Stabilize the queue tier first.**
      - Land #5036 (libhdfs thread destructor use-after-free). Until the crash 
rate on main is near zero, make macOS `[scans]` label-only rather than 
queue-tier.
      - Track the Iceberg 1.8/1.9/1.10 red rate on main after #5758 and the 
artifact retry changes (#5842). If it stays above a few percent, move those 
three to label-only as well.
   
   2. **Move `Spark SQL Tests (Spark 4.1)` `sql_hive-1/2/3` to queue-only.** 65 
runner-hours per day for 7 unique catches a month, and it is the 110 minute 
long pole. Keeps `sql_core-1/2/3` and `catalyst` on PRs, which catch 33 real 
failures a month across 33 different branches. PR wall clock drops to about 100 
minutes. This needs a `hive` on/off input on `spark_sql_test_reusable.yml` and 
a corresponding POLICY entry in `dev/ci/compute-changes.py`.
   
   3. **Move `Iceberg Spark SQL Tests (Iceberg 1.11)` to queue-only** (opt in 
on PRs with `run-iceberg-tests`). 84 runner-hours per day for 17 unique 
catches, which become roughly one queue eviction every two days.
   
   4. **Optionally move the Spark 4.0 and 4.2 Linux test profiles to 
queue-only.** 93 runner-hours per day for 34 unique catches. Keep 3.4 (Scala 
2.12 and the oldest Spark, the most unique catches of the non-4.1 profiles) and 
4.1 on PRs. Same mechanism as #5843 for the macOS build: a profile filter input 
on `pr_build_linux.yml` driven by POLICY.
   
   Steps 2 and 3 save about 150 runner-hours per day (28% of PR spend) for 
about 24 escapes a month. Adding step 4 brings the saving to about 240 
runner-hours per day (45%).
   
   Not recommended: dropping the Spark SQL `sql_core` modules from PRs. They 
are the best-value heavy job after the Linux build.
   
   ### Additional context
   
   - The merge queue ruleset is active but there were zero `merge_group` runs 
at the time of the analysis (#5843 and #5865 merged just before it took 
effect), so queue-tier behavior above is inferred from push-to-main runs.
   - Unique-catch counts include some flakes; the "Main fail" column is the 
calibration for that. On main, a flaky failure hits all Linux profiles 80 to 
100% of the time, so the per-profile unique catches are mostly real 
version-specific failures.
   - The `.github/workflows/README.md` tier table and diagram need updating 
with each step.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to