JingsongLi opened a new pull request, #1027:
URL: https://github.com/apache/paimon-rust/pull/1027

   ### Purpose
   
   Closes #1026.
   
   Complete the Rust core support needed to enable PyPaimon Native MAP 
shared-shredding writes and updates. The baseline ignores placement policies, 
infers width from the first row, bypasses shredding for PK files, and emits 
duplicate Arrow schemas that PyArrow resolves differently from arrow-rs.
   
   ### Brief change log
   
   - Implement Java's plain, sequential and LRU allocators, with LRU as the 
default. Values for duplicate keys come from the last input occurrence; 
overflow duplicates follow Java's map semantics.
   - Select physical column counts using Java's completed-file window (20 
files, p90, slack), starting from max-columns. Publish widths only after a 
successful close. Append writers await pending MAP file closes before choosing 
the next width; ordinary rolling remains asynchronous.
   - Share format context across rolled files. PK data and input changelog use 
independent contexts scoped to each buffer flush, matching Java's 
writer-factory lifetime.
   - Route PK output through the shared format writer and keep statistics 
scoped to logical value fields.
   - Emit exactly one final ARROW:schema footer entry containing shredding 
metadata, readable consistently by PyArrow and arrow-rs.
   - Validate shared-shredding schema/options at schema creation/evolution and 
writer opening. Keep Rust's existing 16384-column reader bound consistent with 
writer configuration by rejecting oversized configurations before writing.
   
   ### Tests
   
   - `cargo test --locked -p paimon --lib`: 3608 passed, 6 ignored.
   - `cargo clippy --locked -p paimon --all-targets --features fulltext,vortex 
-- -D warnings`: passed.
   - `cargo fmt --all -- --check`: passed.
   - Regression coverage includes Java allocator examples, LRU eviction and 
metadata, p90/slack/window boundaries, invalid options, append/PK rolling, 
independent data/changelog histories, per-flush reset, value statistics, and a 
single final Arrow schema.
   - A companion PyPaimon PR exercises native and Python cross-reads, selected 
MAP keys, nested values, duplicate keys, data-evolution updates and upserts 
against a locally rebuilt native package.
   
   ### API and Format
   
   No public Rust API changes or new file layout version. Uses Java's existing 
PIP-43 physical MAP layout and metadata. Configured shared-shredding PK files 
now use the requested physical layout. Default placement follows current Java 
(LRU).
   
   ### Documentation
   
   The companion PyPaimon PR documents the newly enabled Native MAP 
write/update path. Implementation comments identify the Java reference classes 
and context lifetime.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to