Zouxxyy opened a new pull request, #10002:
URL: https://github.com/apache/paimon/pull/10002

   ### Purpose
   
   PyPaimon rejects Arrow `large_string` inputs even though they represent the 
same Paimon `STRING` type as Arrow `string`. Some adapters work around this by 
narrowing strings independently, which makes schema handling inconsistent and 
imposes 32-bit offset limits before writing.
   
   - Map both Arrow string layouts to Paimon `STRING`, including nested fields. 
Share schema compatibility and safe input preparation across core and adapters 
while retaining their column policies.
   - Preserve input string layouts and widen differing string fields when 
buffering mixed batches. Remove the Daft and Ray string-narrowing workarounds, 
and preserve Arrow-backed pandas columns and generated index fields.
   - Keep the existing `string` read contract with checked layout conversion 
and a fixed Arrow schema within each Parquet file. Align map keys during schema 
evolution reads and validate projected struct fields by name.
   
   BYTES/BLOB distinctions and existing numeric schema-evolution rules are 
retained. Dependency constraints are unchanged; `string_view` support is 
outside this change. New regression coverage is in core and adapter tests, with 
no new LeRobot-specific tests.
   
   ### Tests
   
   Validated after rebasing onto `apache/master` at `3b889d239f`:
   
   - PyArrow 19.0.1: 686 passed, 1 skipped across core, write/read schema 
evolution, Ray/Daft adapters, and existing multimodal suites. The existing 
video-decoding test skipped because libtorchcodec could not load the local 
FFmpeg libraries.
   - PyArrow 16.1.0: 338 passed across core, write/read schema evolution, 
row-id updates, and authorization masking.
   - Flake8 on changed Python files, license-header validation, and `git diff 
--check` passed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to