NoahKusaba commented on PR #2518:
URL: 
https://github.com/apache/datafusion-ballista/pull/2518#issuecomment-6019412339

   Closing — benchmarking showed this doesn't help any plan Ballista currently 
produces.
   
   TPC-H SF10 (main @ d6d8bd91 vs this branch, 22 queries): identical results, 
no time difference beyond noise. Passthrough shuffle files were already 
well-sized (main: 2,098 batches / 6,631 rows per batch; branch: 2,102 / 6,618), 
so there was nothing to coalesce.
   
   The motivating case — `UnorderedRangeRepartitionExec` feeding the 
passthrough writer — isn't planned by any rule today. The range scatter that is 
planned (parallel windows) writes through `RangeShuffleWriterExec`, and its 
streaming merge already emits full batches (1M-row window query: 7,936 rows per 
batch, none under 1,024).
   
   
   I'll be more careful about any other performance PR's in the future, seems 
like AI isn't ready to 1 shot them.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to