JingsongLi commented on PR #10148:
URL: https://github.com/apache/paimon/pull/10148#issuecomment-5805508451
Requirement fit: SUPPORTED. Full shared-shredding MAP reads are a real
reader path, and the measured end-to-end gains on several layouts are
substantial. The change is scoped to replacing per-element Arrow-to-Python
conversion for qualifying batches.
Implementation: CLEAN in the changed path. I checked LIST/LARGE_LIST slices,
null/empty mappings, duplicate and unknown field IDs, value ordering, and the
small-batch fallback. The fast path's null-count guard falls back safely if the
count is unknown. I found no actionable regression.
Verification on an isolated checkout: all 4 new focused tests passed.
Related shared-shredding reader and selected-key tests passed with ORC tests
excluded (30 passed, 23 subtests). ORC tests hit a local PyArrow
sysctlbyname('hw.l1dcachesize') environment error, so I cannot claim local ORC
integration coverage. The posted performance numbers predate the exact final
cleanup; a current-head benchmark would firm up the performance claim before
release.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]