discivigour opened a new pull request, #9928: URL: https://github.com/apache/paimon/pull/9928
### Purpose Use Avro's `BufferedBinaryEncoder` for data-file records so small field encodings are batched before they enter the Avro block buffer. This improves nested MAP/ARRAY writes and avoids a regression for common scalar schemas. Add coverage for closing with a partially filled encoder buffer and for mixing buffered records with raw block copies across the supported codecs. ### Tests - `mvn -pl paimon-format -am clean test -Dmaven.repo.local=/opt/homebrew/opt/maven/repository -DskipITs -Dcheckstyle.skip=true -Dtest=AvroFileFormatTest -Dsurefire.failIfNoSpecifiedTests=false` (37 tests passed) A/B benchmarks ran on the same 4-core ECS/JDK 11 process against `fx_perf_rest.codex_db`; each mode used identical inputs and output settings, and all rows/fields, file layouts, block distributions, and decompressed payload CRCs matched. | Workload | Data per sample | Buffered mean | Direct mean | Result | | --- | ---: | ---: | ---: | --- | | 20 `MAP<STRING, ARRAY<INT>>` fields; 20 keys/map; 300 ints/key | 5 GiB | 63.362 s | 73.479 s | Buffered 13.77% faster, 3/3 | | 20 `BYTES` fields; 512 B/field | 3 GiB | 33.873 s | 33.515 s | Direct 1.06% faster in write stage, 2/3; direction was not stable | | 36 scalar fields; 6 each of `INT`, `FLOAT`, `BOOLEAN`, `STRING`, `DECIMAL(18,4)`, `DATE` | 5 GiB | 77.142 s | 78.710 s | Buffered 1.99% faster, 3/3 | For the BYTES workload, including commit time was 34.036 s for Buffered and 34.360 s for Direct, so it did not demonstrate a stable end-to-end Direct advantage. Results are limited to these three-run workloads and this ECS/OSS environment. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
