jianguotian opened a new pull request, #89:
URL: https://github.com/apache/paimon-mosaic/pull/89

   ## Purpose
   
   Reduce schema BPE construction overhead by avoiding repeated name encoding 
and using a lower-overhead pair counter for large schemas, while preserving 
deterministic serialized output.
   
   ## Changes
   
   - use an adaptive dense 65,536-slot pair counter for large BPE inputs while 
retaining the `HashMap` path for smaller inputs
   - store BPE tokens as `u8`
   - return and reuse the encoded names produced during vocabulary construction
   - preserve deterministic pair tie-breaking, BPE rule ordering, and 
serialized schema bytes
   - add compatibility and boundary tests for the full legacy token space and 
the production decoder
   
   ## Performance evidence
   
   The motivating end-to-end profile showed schema serialization decreasing 
from 3.757% to 0.486% of whole-process samples across the complete optimization 
set. This PR targets that path, but this measurement is not presented as an 
isolated benchmark of this PR alone.
   
   ## Compatibility
   
   - no Mosaic file-format change
   - no schema-encoding or BPE rule-order change
   - no FFI or public API change
   
   ## Tests
   
   - `cargo fmt --all -- --check`
   - `cargo clippy --locked --offline --all-targets --workspace -- -D warnings`
   - `cargo test --locked --offline --workspace`
   - `cargo build --locked --offline --all-targets`
   - `cargo deny check licenses`
   - `python3 tools/dependencies.py check`
   - `git diff --check HEAD^`
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to