JingsongLi opened a new pull request, #959:
URL: https://github.com/apache/paimon-rust/pull/959

   ## Summary
   
   - Support Format Table reads and writes for CSV, JSON, and TEXT, and enable 
ORC and Mosaic Format Table writes.
   - Add Java-compatible text compression names and file suffixes for gzip, 
bzip2, deflate, snappy, lz4, and zstd. Snappy and LZ4 use Hadoop's block 
framing; readers detect the codec from the file suffix.
   - Add Format Table round-trip tests covering the supported formats, all six 
text codecs, external compressed files, partitioning, projections, and options.
   
   ## Validation
   
   - `cargo fmt --all --check`
   - `cargo clippy --locked -p paimon --all-targets --features fulltext,vortex 
--offline -- -D warnings`
   - `cargo test -p paimon --all-targets --features fulltext,vortex --locked 
--offline` (3,417 library tests passed; 6 ignored; all integration targets 
passed)
   - `cargo test -p paimon-rest-server --all-targets --locked --offline`
   
   ## Limitations and assumptions
   
   - `orc-rust` 0.8's writer currently emits uncompressed ORC and supports a 
limited set of primitive Arrow types. This implementation rejects unsupported 
types but does not apply the requested ORC compression codec.
   - Text readers currently materialize each data file before parsing, so 
memory use scales with file size.
   - The Hadoop Snappy/LZ4 framing follows Hadoop's implementation; direct 
Java-process interoperability was not run.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to