andygrove opened a new pull request, #6393: URL: https://github.com/apache/datafusion-comet/pull/6393
## Which issue does this PR close? Closes #6388. ## Rationale for this change The Spark SQL suite sets the length of every merge-queue run: its build job plus its slowest shard. Two of the seven rows ran far longer than the rest (medians over 129 queue runs, 2026-09-11 to 2026-09-29): `sql_hive-1` at 66.7 minutes and `sql_core-1` at 59.6, against 21.1 for `sql_hive-2` and 28.1 for `sql_core-3`. Per-suite timings from the job logs show where that time goes; the table is in the issue. ## What changes are included in this PR? Two new rows in `dev/ci/spark-sql-modules.py`. Each moved suite is excluded from the row it came from with sbt's `-<glob>` testOnly syntax (supported by every sbt the Spark builds use, 1.8.2 through 1.11.7), and its new row keeps the old row's tag filters, so every test still runs exactly once: - `sql_hive-4` runs `org.apache.spark.sql.hive.client.HivePartitionFilteringSuites`, which runs its suite against every Hive client version and took 21 to 35 minutes of `sql_hive-1`. It is a single class, so no name filter splits it further. - `sql_core-4` runs the untagged and Extended tests under `org.apache.spark.sql.execution.datasources.*` and `org.apache.spark.sql.connector.*` (13 to 19 minutes of `sql_core-1`, plus 11 of `sql_core-2`). Their Slow tests stay in `sql_core-3` with every other Slow test. The globs are package prefixes, so suites added to those packages later land in the new rows without an edit. The comments and docs that counted seven rows now say nine. In the two queue runs I sampled, the slowest row would drop from 72 and 51 minutes to about 48 and 40, which takes roughly 20 minutes off a merge-queue run. The cost is two more runners per Spark SQL run, each with about 9 minutes of setup. ## How are these changes tested? `python3 dev/ci/check-ci-config.py` (including its module-group partition check), `dev/local-ci.sh --print-config`, `actionlint` and prettier pass locally. I'm applying `run-spark-4.1-tests` so the new rows run before this is queued. The check is that the suites listed in `sql_core-4` and `sql_hive-4` no longer appear in `sql_core-1`, `sql_core-2` or `sql_hive-1`, and that the per-module test totals match a recent queue run. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
