Rachelint opened a new pull request, #26116:
URL: https://github.com/apache/datafusion/pull/26116

   ## Which issue does this PR close?
   
   Experimental follow-up to #26100; no issue closed.
   
   ## Rationale for this change
   
   Test a bounded group-count policy for Partial aggregation: consume complete 
input batches, flush once the table reaches the target batch size, and reserve 
space for twice that many groups. With one grouping set and input batches no 
larger than the target, a table stays below twice the target before flushing.
   
   This branch starts directly at `634a59fff53e19c630657d728c4db3a57204727e`, 
before the StringView and general group-key reserve changes.
   
   ## What changes are included in this PR?
   
   - Add experimental `datafusion.execution.partial_aggregation_flush_batch` 
(default false). Enabled mode overrides the byte and row thresholds, with the 
existing soft-limit/nested-state exclusions and a single grouping set 
requirement.
   - Process each input batch once. Reserve `2 * batch_size` group slots 
initially and after threshold flushes; the capacity is a hint, and larger input 
batches remain valid.
   - Preallocate primitive and StringView group keys, supported multi-column 
builders, hash/scratch storage, and count/primitive/average accumulator value 
vectors. Other layouts retain their existing allocation behavior.
   - Retain hash allocations across threshold flushes. Keep ordinary string 
payload growth and 2 MiB view payload blocks. Memory-pressure flushes use the 
original memory-releasing path.
   
   ## What is the testing strategy for this PR?
   
   Passed `cargo check --locked`, `cargo fmt --all -- --check`, `cargo clippy 
--all-targets --all-features -- -D warnings`, the 10 existing Partial stream 
tests, and `aggregate_partial_flush.slt` / `information_schema.slt`. Extended 
the SQL test with integer, StringView, mixed, and nullable keys and conflicting 
thresholds. Generated and formatted the configuration docs.
   
   ECS comparison uses the same release binary for 2 MiB flushing, 8192-group 
flushing without reservation, and this mode (8192 groups / 16384 reserved). 
Neoverse N2, 12 partitions, input batch 8192, 99,997,497 source rows, three 
interleaved rounds with two warm samples per round after discarding the first 
run.
   
   ## Are there any user-facing changes?
   
   An opt-in experimental setting and defaulted reservation hooks on 
aggregation traits. Disabled mode preserves the original threshold behavior. 
This is a draft performance experiment.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to