XiaoHongbo-Hope commented on PR #10320:
URL: https://github.com/apache/paimon/pull/10320#issuecomment-5890993064

   Follow-up benchmark for eadc4b968: on one fixed, read-only 2,128-row Arrow 
batch (about 5,362 top-level Variant fields per row), batch variant_get 
selected 152 exact-typed numeric fields. Interleaved A/B on the same in-memory 
Arrow data measured 15.97/16.03 s before vs. 2.1906/2.1912 s after this change 
(about 7.3x faster for the decoding stage). All selected Arrow arrays were 
equal. Planning, Parquet I/O, and application work are excluded. The reader 
already accessed Arrow buffers directly; this commit replaces per-field Python 
offset parsing/sorting with NumPy arrays and caches field slots by the exact 
object layout. Additional regression tests cover three-byte offsets, changing 
layouts under identical metadata, and malformed duplicate IDs/offsets; 219 
local Variant tests pass.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to