XiaoHongbo-Hope commented on PR #10320: URL: https://github.com/apache/paimon/pull/10320#issuecomment-5890993064
Follow-up benchmark for eadc4b968: on one fixed, read-only 2,128-row Arrow batch (about 5,362 top-level Variant fields per row), batch variant_get selected 152 exact-typed numeric fields. Interleaved A/B on the same in-memory Arrow data measured 15.97/16.03 s before vs. 2.1906/2.1912 s after this change (about 7.3x faster for the decoding stage). All selected Arrow arrays were equal. Planning, Parquet I/O, and application work are excluded. The reader already accessed Arrow buffers directly; this commit replaces per-field Python offset parsing/sorting with NumPy arrays and caches field slots by the exact object layout. Additional regression tests cover three-byte offsets, changing layouts under identical metadata, and malformed duplicate IDs/offsets; 219 local Variant tests pass. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
