sunchao commented on PR #11279: URL: https://github.com/apache/arrow-rs/pull/11279#issuecomment-5963463304
@cetra3 @viirya Thanks for pushing on both the duplicate-range case and the comparison with `interleave + gc()`. I have kept this as a draft and revised the implementation and rationale. The proposed API is now `Interleaver::new().with_compact_byte_views(true)`, with separate opt-in `with_preserve_byte_view_sharing(true)`. The latter keys by effective source address plus length, copies each distinct range once into fresh output storage, and sizes from those unique bytes. It handles aliases across sliced inputs without content hashing. It does not merge equal strings in separate allocations, and it never shares copied buffers across separate output calls. `interleave + gc()` already gives independent payload storage and exact payload capacity in the ordinary cases. The additional justification here is lower temporary allocation in the direct compact path when gathering selected rows, plus optional preservation of repeated source ranges during copying. A spilling join needs to account for the gather before retaining independently releasable partitions; the lookup table must also be charged when range sharing is enabled. The exact-base comparison is now kernel base `705e6405b` (benchmark-only companion `363eb96df`) versus final head `d8726603c`, with identical fixture/lockfile and two balanced rounds. In the random 13-20-byte/8192-row case, direct compact is 101.214 us versus 101.258 us for `interleave + gc()`. The 20-60-byte case remains 3.8% slower (113.667 versus 109.458 us), so I am not claiming a uniform speedup. The stronger memory evidence is selecting 3 rows from 32,768 source buffers: peak additional live requested allocation is 424 bytes versus 131,664 bytes for `interleave + gc()`, with the same 48 bytes of output payload capacity. For 8,192 selections of 64 repeated 100-400-byte ranges, opt-in sharing reduces output payload capacity from 1,979,008 to 15,461 bytes and time from 136.333 to 96.144 us. Views and other result metadata are additional. The adverse case matters too: on 8,192 unique 13-20-byte values, sharing takes 423.686 us versus 72.661 us for `interleave + gc()`, and peak extra allocation grows from 398,080 to 745,536 bytes. Direct compact is 67.524 us and 266,584 bytes. This is why sharing is independently opt-in. Plain-interleave control timings and both rounds are also included in the description, including the remaining slower controls. Allocation counters measure requested live bytes, not process RSS. The copy loop now rewrites gathered source views in place, uses a single-buffer fast path, and reuses existing interleave validity handling. The signed-size issue is fixed: production blocks are capped at `i32::MAX`, larger totals split, and unsupported individual lengths or buffer indices return an error. Tests are in the existing `mod tests`, including ownership after source drop, aliases, null-hidden payload, and small-limit buffer splitting. The final source passes 459 library tests, 20 doc tests, forced array validation, Clippy, and all 168 benchmark smoke cases. Benchmark-only coverage is in #11346. The PR description has the exact-base/final-head results and reproduction details. API placement relative to the coalescer and shared copy helpers remains open; these kernel measurements are not an end-to-end join or spill-throughput claim. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
