TheR1sing3un opened a new pull request, #10058:
URL: https://github.com/apache/paimon/pull/10058
### Purpose
Writing Python rows such as `[{"id": 1}, {"id": 2, "caption": "keep-me"}]`
to an existing multimodal table silently stores NULL for the second row's
caption. Untyped Arrow row conversion discovers columns from the first row,
before alignment to the table schema. It also loses the row count for inputs
containing only empty dictionaries.
Build Python row columns using the target schema before Arrow inference.
Construct MAP columns with their declared type so dictionary and key-value-pair
inputs, including MAP BLOB values, can be written. Keep the existing inference
and safe casts for other columns, including string-to-number conversion and
rejection of fractional values for integer fields.
### Tests
- Multimodal tests selected with `-k 'not ray'`: 86 passed; ARRAY BLOB
read/stream regression: 1 passed.
- Three new regressions failed before the fix and pass on PyArrow 19 and 16.
They cover persisted later-row scalar/BLOB values, MAP inputs in row and column
form, null/empty values, empty rows, and safe casts.
- Flake8, changed-file license headers, Python 3.6 grammar checks, and `git
diff --check` passed.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]