jianguotian opened a new pull request, #89: URL: https://github.com/apache/paimon-mosaic/pull/89
## Purpose Reduce schema BPE construction overhead by avoiding repeated name encoding and using a lower-overhead pair counter for large schemas, while preserving deterministic serialized output. ## Changes - use an adaptive dense 65,536-slot pair counter for large BPE inputs while retaining the `HashMap` path for smaller inputs - store BPE tokens as `u8` - return and reuse the encoded names produced during vocabulary construction - preserve deterministic pair tie-breaking, BPE rule ordering, and serialized schema bytes - add compatibility and boundary tests for the full legacy token space and the production decoder ## Performance evidence The motivating end-to-end profile showed schema serialization decreasing from 3.757% to 0.486% of whole-process samples across the complete optimization set. This PR targets that path, but this measurement is not presented as an isolated benchmark of this PR alone. ## Compatibility - no Mosaic file-format change - no schema-encoding or BPE rule-order change - no FFI or public API change ## Tests - `cargo fmt --all -- --check` - `cargo clippy --locked --offline --all-targets --workspace -- -D warnings` - `cargo test --locked --offline --workspace` - `cargo build --locked --offline --all-targets` - `cargo deny check licenses` - `python3 tools/dependencies.py check` - `git diff --check HEAD^` -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
