sunchao commented on PR #25584: URL: https://github.com/apache/datafusion/pull/25584#issuecomment-5879767242
@jayzhan211 @comphead @viirya Updated in 548a66268ab1928ed47b2cf778656b4386542815. Reshaped this around the narrow buffered design: one cross-side column comparison, optionally negated, using shared min/max accumulators and vectorized comparison. The expression compiler, configuration option, const-generic state-machine copy, separate consumer, manual yields, and five metrics are removed. The helper is now named `SemiAntiComparison`; compound predicates and mark joins keep ordinary evaluation. A planner-level min/max rewrite that also reaches hash joins remains separate work. The production change is still larger than the proposed sketch because it keeps admission refusal, early-witness cost gates, and cache reuse across outer batches. A synchronous helper keeps reduction and probing outside the async traversal body. For memory pressure, existing buffering/spilling runs first. Refused summary admission falls back without adding a resource-exhaustion error; successful reduction replaces the buffered group with owned extrema and shrinks the same reservation before further outer input. Long-string tests cover admission refusal and forced spilling, including probes across input batches. `peak_mem_used` measures this reservation including admitted overlap. There is still one optional grow/shrink per summarized group, but no additional consumer or per-clause/per-row pool traffic. Groups below seven rows avoid additional summary reservation traffic; the description records the concurrent shared-pool and separate-pool benchmark settings. On timestamps, the timezone Arc remains shared. The conservative allowance follows the existing `ScalarValue::size()` convention, including its timezone-length charge, and the initial bound matches it. This does not assert an allocation for each timezone clone. The final source passed 12,211 Rust tests (including all 44 default-configuration join-fuzz tests), all 525 SQL logic files, all-target/all-feature Clippy, formatting, and the full lint suite. Performance review remains open. The unfiltered semi timing difference is strongly sensitive to benchmark binary layout: it reproduced with identical production code and identical timed work, while native profiles locate the difference in unchanged Arrow key-comparison code. The adverse single-row range result did not reproduce under native sampling, which does not establish that it is fixed. SQL Q13 also has an adverse mean with substantial process variation. I have retained all raw results and diagnostic provenance, and am not claiming blanket non-regression or performance readiness. Six structural/test threads are resolved by this revision. The traversal-cost and benchmark threads remain open. The updated description records the exact base/head provenance, validation scope, and measurement limits; its obsolete performance tables have been removed. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
