sesteves opened a new issue, #24072:
URL: https://github.com/apache/datafusion/issues/24072

   ## Bug
   
   Grouped approx_distinct can panic while DataFusion materializes 
intermediate HyperLogLog states. In DataFusion 54.1.0, each HLL state contains 
16,384 one-byte registers but is represented as ScalarValue::Binary and 
declared as Arrow DataType::Binary.
   
   GroupsAccumulatorAdapter::state collects one scalar state per group and 
calls ScalarValue::iter_to_array. Arrow BinaryArray uses signed 32-bit 
offsets, so this deterministic boundary is unrepresentable:
   
   text
   131,072 groups * 16,384 bytes = 2,147,483,648 bytes
   
   
   The resulting panic is:
   
   text
   byte array offset overflow
   
   
   The stack reaches GroupedHashAggregateStream::emit -> 
GroupsAccumulatorAdapter::state -> ScalarValue::iter_to_array.
   
   ## Spill does not prevent it
   
   The branch-54 grouped hash aggregate spill path first calls 
emit(EmitTo::All, true), so all current group states are materialized in one 
record batch before spill can write them. Terminal output similarly 
materializes EmitTo::All and only slices the resulting batch afterward.
   
   This means the panic can occur while trying to spill, and memory pressure is 
not applied before a potentially multi-gigabyte temporary state array is built.
   
   ## Suggested direction
   
   1. Represent dense fixed-width HLL register state without cumulative 
variable-width offsets, such as FixedSizeBinary(16384) for the branch-54 
implementation.
   2. Materialize terminal and spill aggregate state in bounded chunks rather 
than building EmitTo::All first.
   3. Propagate memory or spill exhaustion as 
DataFusionError::ResourcesExhausted rather than unwinding.
   
   A direct 131,072-group regression requires about 2 GiB per HLL expression, 
so practical CI coverage can verify the fixed-size state schema and use a small 
configured emission-byte threshold or synthetic fixed-size accumulator to prove 
bounded draining.
   
   ## Related work
   
   - #18907 describes unbounded terminal aggregate emission.
   - #23178 documents performance concerns around repeated EmitTo::First.
   - #24061 is related ongoing spill work in newer aggregate streams.
   
   Newer DataFusion has evolved its grouped HLL implementation, including 
sparse state, so the exact state representation on main may need a different 
design. The invariant remains that dense HLL state and grouped spill/output 
must not depend on one variable-width Arrow array with 32-bit cumulative 
offsets.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to