adriangb commented on PR #24086: URL: https://github.com/apache/datafusion/pull/24086#issuecomment-5900228209
# Benchmark summary: `02429ca` (arrow-rs pin with apache/arrow-rs#11241 in place of apache/arrow-rs#11223) This revision is rebased on `main` after #25809, which draws the simulated object store latency at random instead of round-robin. The arrow-rs pin (`pydantic/arrow-rs` `6bea7c8969`) is now 60.0.0 + https://github.com/apache/arrow-rs/pull/11241 + https://github.com/apache/arrow-rs/pull/10555 (current tips). The DataFusion code did not change. `read_ahead_bytes` = 100 MB on the branch side only. GKE runs: adriangbot `c4a-highmem-16`, compared with merge-base `953518c`. "Speedup" is the ratio of the per-query time totals. The previous column is the speedup of `4a9f276` from the [previous summary](https://github.com/apache/datafusion/pull/24086#issuecomment-5857231608). Because #25809 changed the latency model, the two columns are not exactly comparable. ## With `SIMULATE_LATENCY`, `pushdown_filters=false` | suite | query total base → branch | speedup | previous | faster / slower / same | peak memory base → branch | |---|---|---|---|---|---| | tpch_sf1 | 17.02 s → 9.27 s | **1.84x** | 1.75x | 21 / 0 / 1 | 890 MiB → 1.0 GiB | | tpch_sf10 | 113.68 s → 15.71 s | **7.24x** | 7.03x | 22 / 0 / 0 | 3.6 GiB → 4.4 GiB | | tpcds_sf1 | 55.62 s → 56.43 s | 0.99x | 1.00x | 37 / 46 / 16 | 1.1 GiB → 1.0 GiB | | clickbench_partitioned | 84.45 s → 42.64 s | **1.98x** | 1.96x | 38 / 1 / 4 | 16.4 GiB → 17.0 GiB | ## With `SIMULATE_LATENCY`, `pushdown_filters=true` | suite | query total base → branch | speedup | previous | faster / slower / same | peak memory base → branch | |---|---|---|---|---|---| | tpch_sf1 | 38.58 s → 9.52 s | **4.05x** | 3.73x | 22 / 0 / 0 | 677 MiB → 713 MiB | | tpcds_sf1 | 107.74 s → 62.91 s | **1.71x** | 1.72x | 94 / 5 / 0 | 1.2 GiB → 1.1 GiB | | clickbench_partitioned | 105.25 s → 39.09 s | **2.69x** | 2.75x | 41 / 0 / 2 | 9.5 GiB → 11.0 GiB | | tpch_sf1, `DATAFUSION_RUNTIME_MEMORY_LIMIT=1G` | 39.53 s → 9.63 s | **4.10x** | 3.82x | 22 / 0 / 0 | 674 MiB → 847 MiB | ## Notes - The new arrow-rs pin does not change the results. All speedups are within a few percent of `4a9f276`. The tpch pushdown runs are a little faster (4.05x and 4.10x, previously 3.73x and 3.82x). - tpcds without pushdown is flat, as before (0.99x). The query counts changed from 9 / 12 / 78 to 37 / 46 / 16 because the random latency draws make each query more variable: the standard deviation on the base side is 30–60% of the mean for many queries. The mean-based total is also flat (68.4 s → 69.3 s). - tpcds Q27 is no longer slower in the no-pushdown run: its mean is 1.12x faster. It is 1.15x slower with pushdown (189 → 219 ms), previously 1.55x. - tpcds Q1 and Q44 are slower in both tpcds runs: Q1 is 1.29x slower with pushdown and 1.56x without, and Q44 is 1.31x slower with pushdown and 1.33x (mean) without. These are small queries (100–250 ms). I will profile them. - The 1 GB pool run has no failures. Runs: [pushdown + latency](https://github.com/apache/datafusion/pull/24086#issuecomment-5897086822), [latency](https://github.com/apache/datafusion/pull/24086#issuecomment-5897087136), [1 GB pool](https://github.com/apache/datafusion/pull/24086#issuecomment-5897087498). -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
