sunchao commented on PR #11279:
URL: https://github.com/apache/arrow-rs/pull/11279#issuecomment-5963463304

   @cetra3 @viirya Thanks for pushing on both the duplicate-range case and the 
comparison with `interleave + gc()`. I have kept this as a draft and revised 
the implementation and rationale.
   
   The proposed API is now `Interleaver::new().with_compact_byte_views(true)`, 
with separate opt-in `with_preserve_byte_view_sharing(true)`. The latter keys 
by effective source address plus length, copies each distinct range once into 
fresh output storage, and sizes from those unique bytes. It handles aliases 
across sliced inputs without content hashing. It does not merge equal strings 
in separate allocations, and it never shares copied buffers across separate 
output calls.
   
   `interleave + gc()` already gives independent payload storage and exact 
payload capacity in the ordinary cases. The additional justification here is 
lower temporary allocation in the direct compact path when gathering selected 
rows, plus optional preservation of repeated source ranges during copying. A 
spilling join needs to account for the gather before retaining independently 
releasable partitions; the lookup table must also be charged when range sharing 
is enabled.
   
   The exact-base comparison is now kernel base `705e6405b` (benchmark-only 
companion `363eb96df`) versus final head `d8726603c`, with identical 
fixture/lockfile and two balanced rounds. In the random 13-20-byte/8192-row 
case, direct compact is 101.214 us versus 101.258 us for `interleave + gc()`. 
The 20-60-byte case remains 3.8% slower (113.667 versus 109.458 us), so I am 
not claiming a uniform speedup.
   
   The stronger memory evidence is selecting 3 rows from 32,768 source buffers: 
peak additional live requested allocation is 424 bytes versus 131,664 bytes for 
`interleave + gc()`, with the same 48 bytes of output payload capacity. For 
8,192 selections of 64 repeated 100-400-byte ranges, opt-in sharing reduces 
output payload capacity from 1,979,008 to 15,461 bytes and time from 136.333 to 
96.144 us. Views and other result metadata are additional.
   
   The adverse case matters too: on 8,192 unique 13-20-byte values, sharing 
takes 423.686 us versus 72.661 us for `interleave + gc()`, and peak extra 
allocation grows from 398,080 to 745,536 bytes. Direct compact is 67.524 us and 
266,584 bytes. This is why sharing is independently opt-in. Plain-interleave 
control timings and both rounds are also included in the description, including 
the remaining slower controls. Allocation counters measure requested live 
bytes, not process RSS.
   
   The copy loop now rewrites gathered source views in place, uses a 
single-buffer fast path, and reuses existing interleave validity handling. The 
signed-size issue is fixed: production blocks are capped at `i32::MAX`, larger 
totals split, and unsupported individual lengths or buffer indices return an 
error. Tests are in the existing `mod tests`, including ownership after source 
drop, aliases, null-hidden payload, and small-limit buffer splitting. The final 
source passes 459 library tests, 20 doc tests, forced array validation, Clippy, 
and all 168 benchmark smoke cases.
   
   Benchmark-only coverage is in #11346. The PR description has the 
exact-base/final-head results and reproduction details. API placement relative 
to the coalescer and shared copy helpers remains open; these kernel 
measurements are not an end-to-end join or spill-throughput claim.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to