NoahKusaba commented on PR #2518: URL: https://github.com/apache/datafusion-ballista/pull/2518#issuecomment-6019412339
Closing — benchmarking showed this doesn't help any plan Ballista currently produces. TPC-H SF10 (main @ d6d8bd91 vs this branch, 22 queries): identical results, no time difference beyond noise. Passthrough shuffle files were already well-sized (main: 2,098 batches / 6,631 rows per batch; branch: 2,102 / 6,618), so there was nothing to coalesce. The motivating case — `UnorderedRangeRepartitionExec` feeding the passthrough writer — isn't planned by any rule today. The range scatter that is planned (parallel windows) writes through `RangeShuffleWriterExec`, and its streaming merge already emits full batches (1M-row window query: 7,936 rows per batch, none under 1,024). I'll be more careful about any other performance PR's in the future, seems like AI isn't ready to 1 shot them. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
