JingsongLi opened a new issue, #1026: URL: https://github.com/apache/paimon-rust/issues/1026
The current shared-shredding MAP writer still follows Java's older plain allocation / first-row width inference. Current Java supports plain, sequential and LRU placement (LRU by default), and chooses the next file's width from completed-file statistics. There are also two integration gaps: - Primary-key Parquet writes bypass the shared shredding format wrapper, so configured MAP layouts are not applied to data or input changelog files. - Shredded Parquet files carry two ARROW:schema footer entries. PyArrow reads the first schema and misses the MAP reconstruction metadata from the second entry. These prevent PyPaimon from enabling native MAP shared-shredding writes and updates reliably. Align the allocator/context lifecycle and schema validation with Java, route PK output through the shared format writer while preserving value-only statistics, and publish one final Arrow schema. Cover rolling, changelogs, duplicate keys, metadata interoperability and native/Python cross-reads. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
