SteNicholas opened a new pull request, #392: URL: https://github.com/apache/paimon-cpp/pull/392
### Purpose Linked issue: close #389 #388 taught the C++ reader to read top-level `ARRAY<BLOB>` fields, but tables with such a column were still rejected on creation and write, so an embedding engine had to route these writes back to the Java SDK. This makes the C++ writer store `ARRAY<BLOB>` fields in dedicated blob files, aligned with Java Paimon (apache/paimon#8181): - `ARRAY<BLOB>` fields are routed to blob files, matching Java's `BlobType#isBlobFileField`, and are no longer rejected on table creation and write. `MAP<..., BLOB>` remains read-only. - Each `ARRAY<BLOB>` row is encoded as the nested payload of the BLOB file spec: array magic, version, element count, element data, delta-varint element lengths and the element index length. A null array is a `-1` entry in the file index, an empty array has an element count of zero, and a null element has length `-1`. The bytes match Java's `BlobFormatWriterTest#testArrayBlobGoldenBytes`. - Elements may be raw bytes or serialized `BlobDescriptor`s, whose referenced data is copied on write. `blob-write-null-on-missing-file` and `blob-write-null-on-fetch-failure` apply per element, so only the unreachable element becomes NULL, not the array. - In a data-evolution partial update, a one-element array holding the placeholder sentinel is written as a placeholder entry (`-2`). - `ARRAY<BLOB>` rows are left to `Flush()` and `Finish()` to persist instead of flushing the blob file after every row, as Java does unless its `BlobConsumer` asks for a flush. - Copying blob data, for `BLOB` as well as `ARRAY<BLOB>`, continues after short reads; a read that fails or ends before the reported length fails the write. - The write fails instead of truncating when a compressed element length index or file index exceeds its signed 32-bit length field. - NULL warnings and fetch and copy errors name the field, its row in the blob file and, for an `ARRAY<BLOB>`, the element index; output write and flush errors name the blob file. - `AppendCompactCoordinator` keeps rejecting `ARRAY<BLOB>` tables: its rewrite reads and writes plain append files, so it can neither merge data-evolution blob layers nor write blob files. Differences from Java, also documented in `write.rst`: - A placeholder is identified by the reserved bytes `_PAIMON_BLOB_PLACEHOLDER`; Java uses a dedicated placeholder object, which cannot collide with a user value. - A missing file is detected with `FileSystem::Exists`. Java detects a missing file for an `ARRAY<BLOB>` element only from an HTTP 404, and does not write NULL for a 404 under `blob-write-null-on-fetch-failure` alone. ### Tests - `BlobFormatWriterArrayBlobTest.TestArrayBlobGoldenBytes`: the written bytes match Java's golden bytes for an empty array, inline, null, empty and descriptor elements, a null array and a placeholder. - `BlobFormatWriterArrayBlobTest.TestArrayBlobRoundTrip`: null and empty arrays, null and empty elements, a descriptor slice, an element larger than the copy buffer and a sentinel element outside placeholder mode round trip, read back as data and as descriptors. - `BlobFormatWriterArrayBlobTest.TestArrayBlobPlaceholder`: only an array holding exactly the sentinel becomes a placeholder; a strict reader rejects it and a placeholder-aware reader emits the sentinel. - `BlobFormatWriterArrayBlobTest.TestArrayBlobWriteNullOnUnreachableElements`: missing files, invalid offsets and truncated descriptors under every combination of the write-null options, including the per-cause metrics and the element named in errors. - `BlobFormatWriterArrayBlobTest.TestArrayBlobCopyWithShortReads` and `BlobFormatWriterWriteNullTest.TestCopyWithShortReads`: short reads are continued, and a failed read or an early end of data fails the write, for `ARRAY<BLOB>` and `BLOB`. - `BlobFormatWriterArrayBlobTest.TestArrayBlobOutputFailureFailsWrite`: write failures and short writes at each stage of an entry, and flush failures deferred to `Finish()`. - `BlobFormatWriterTest.TestCreateWithInvalidParameters` and `BlobFormatWriterWriteNullTest.TestWriteNullOnMissingFile`: non-BLOB lists are rejected, and a failed `BLOB` row is named in its error. - `BlobUtilsTest.ValidateContainerBlobWriteSchema`, `BlobUtilsTest.SeparateArrayBlobFields`, `BlobFileContextTest.ContainerBlobFields`, `TableSchemaTest.CreatingArrayBlobSchemaIsAllowed` and `AppendCompactCoordinatorTest.TestValidateFailsOnArrayBlobTable`: schema validation, blob file routing and the compaction rejection. - `BlobTableInteTest.TestArrayBlobWriteAndPartialUpdate`: writes, a partial update that fails to commit without first row ids and succeeds with them, placeholder rows, null arrays, and a missing descriptor written as a NULL element. - `BlobTableInteTest.TestArrayBlobWriteDescriptorElements`: descriptor elements are copied into the blob file, so the table stays readable after the source files are deleted, as data and as descriptors. - `BlobTableInteTest.TestArrayBlobWithBlobFieldAcrossMultipleBlobFiles`: `BLOB` and `ARRAY<BLOB>` columns rolling across several blob files, read with row ranges and projections, then partially updated across files. ### API and Format No public symbol is added or changed. `include/paimon/defs.h` documents that the blob write-null options apply per `ARRAY<BLOB>` element and which failures `blob-write-null-on-fetch-failure` converts, and `include/paimon/append/append_compact_coordinator.h` documents the `ARRAY<BLOB>` and `MAP<..., BLOB>` restriction. The storage format is unchanged: the writer produces the `ARRAY<BLOB>` payload of the existing Paimon BLOB file spec, which the C++ reader already reads since #388. ### Documentation - `docs/source/user_guide/write.rst` gains a "Writing BLOB Columns" section: how to build `BLOB` and `ARRAY<BLOB>` fields, accepted element forms, the write-null options, which writes are partial updates, the first row id they must carry, the placeholder sentinel and the differences from Java. - `docs/source/user_guide/data_types.rst` documents the `BLOB` and `ARRAY<BLOB>` types and their table requirements. - `docs/source/user_guide/compaction.rst` clarifies automatic compaction and that `AppendCompactCoordinator` rejects `ARRAY<BLOB>` tables. ### Generative AI tooling Generated-by: Claude Code (Claude Opus 5.5) 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
