sunchao commented on PR #25584:
URL: https://github.com/apache/datafusion/pull/25584#issuecomment-5879767242

   @jayzhan211 @comphead @viirya Updated in 
548a66268ab1928ed47b2cf778656b4386542815.
   
   Reshaped this around the narrow buffered design: one cross-side column 
comparison, optionally negated, using shared min/max accumulators and 
vectorized comparison. The expression compiler, configuration option, 
const-generic state-machine copy, separate consumer, manual yields, and five 
metrics are removed. The helper is now named `SemiAntiComparison`; compound 
predicates and mark joins keep ordinary evaluation. A planner-level min/max 
rewrite that also reaches hash joins remains separate work. The production 
change is still larger than the proposed sketch because it keeps admission 
refusal, early-witness cost gates, and cache reuse across outer batches. A 
synchronous helper keeps reduction and probing outside the async traversal body.
   
   For memory pressure, existing buffering/spilling runs first. Refused summary 
admission falls back without adding a resource-exhaustion error; successful 
reduction replaces the buffered group with owned extrema and shrinks the same 
reservation before further outer input. Long-string tests cover admission 
refusal and forced spilling, including probes across input batches. 
`peak_mem_used` measures this reservation including admitted overlap. There is 
still one optional grow/shrink per summarized group, but no additional consumer 
or per-clause/per-row pool traffic. Groups below seven rows avoid additional 
summary reservation traffic; the description records the concurrent shared-pool 
and separate-pool benchmark settings.
   
   On timestamps, the timezone Arc remains shared. The conservative allowance 
follows the existing `ScalarValue::size()` convention, including its 
timezone-length charge, and the initial bound matches it. This does not assert 
an allocation for each timezone clone.
   
   The final source passed 12,211 Rust tests (including all 44 
default-configuration join-fuzz tests), all 525 SQL logic files, 
all-target/all-feature Clippy, formatting, and the full lint suite.
   
   Performance review remains open. The unfiltered semi timing difference is 
strongly sensitive to benchmark binary layout: it reproduced with identical 
production code and identical timed work, while native profiles locate the 
difference in unchanged Arrow key-comparison code. The adverse single-row range 
result did not reproduce under native sampling, which does not establish that 
it is fixed. SQL Q13 also has an adverse mean with substantial process 
variation. I have retained all raw results and diagnostic provenance, and am 
not claiming blanket non-regression or performance readiness. Six 
structural/test threads are resolved by this revision. The traversal-cost and 
benchmark threads remain open. The updated description records the exact 
base/head provenance, validation scope, and measurement limits; its obsolete 
performance tables have been removed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to