JingsongLi opened a new pull request, #989:
URL: https://github.com/apache/paimon-rust/pull/989

   ## Purpose
   
   Close the core gaps exposed by enabling scalar BLOB Arrow writes in PyPaimon 
Native. This covers file grouping, rolling, metadata, descriptors and URI I/O 
as one write/read path.
   
   ## Changes
   
   - Finish each normal + dedicated file group together and preserve group 
order in commit messages. This keeps Blob row IDs aligned when normal files 
roll within one writer.
   - Check Blob payload size after each row, including descriptor-backed 
payloads. Share the data-file UUID and counter across the normal and dedicated 
writers, like Java's DataFilePathFactory.
   - Set data-evolution physical file sequence ranges to `0..row_count-1`, 
honor `data-evolution.write-cols-optimization.enabled`, and generate indexes 
for the normal columns.
   - Keep completed groups owned until prepare_commit. Abort and late 
close/write failures remove earlier groups, external files and index sidecars.
   - Validate inline descriptor/view values before writing. Use Java's prefix 
parsing for schema-declared descriptors, including v1 and trailing bytes, while 
preserving raw-byte detection for ordinary Blob payloads.
   - Resolve known descriptors consistently in data-evolution and primary-key 
reads.
   - Add a shared URI reader for HTTP(S) Blob references. It uses decoded GET 
streams, supports gzip/deflate, verifies short reads/statuses, and reuses one 
response for a multi-chunk Blob copy. Ordinary paths continue using the table 
FileIO and its provider/credentials.
   
   The HTTP reader follows Java's decoded-stream offsets rather than HTTP wire 
ranges. Unknown-length descriptors may need an additional size pass; that pass 
does not retain the payload in memory.
   
   ## Verification
   
   - Reproduced the original rolling, grouping, sequence and inline-validation 
failures before the fixes.
   - Rust table unit tests: 1,679 passed, 3 ignored.
   - Dedicated Blob integration tests cover independent two-column rolls, 
null/empty payloads, external paths, indexes, optimized write columns, v1/v2 
inline descriptors and primary-key reads.
   - URI tests cover redirects, chunked bodies, decoded gzip/deflate offsets, 
one-response sequential reads, short reads, unsupported encodings and HTTP 
errors.
   - PyPaimon end-to-end tests use the rebuilt binding and exercise both 
Python/Native planning, reading and committing, stream reuse, abort/failure 
cleanup and an 18 MiB HTTP Blob copied with one GET.
   
   Final local results:
   
   - `cargo test --locked -p paimon --lib table::`: 1,679 passed, 3 ignored.
   - `cargo test --locked -p paimon --lib arrow::format::blob::tests`: 46 
passed.
   - URI reader unit test: passed (including gzip/deflate decoded offsets).
   - Dedicated Blob and existing external-data integration suites: 12 passed.
   - `cargo clippy --locked --all-targets --workspace --features 
fulltext,vortex -- -D warnings`: passed.
   - `cargo fmt --all -- --check` and `git diff --check`: passed.
   - Expanded PyPaimon regression with all five Native options: 1,238 passed, 3 
skipped, 67 subtests passed. Three Vortex cases were deselected because the 
optional Python Vortex dependency is absent; no Vortex behavior is changed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to