JingsongLi opened a new pull request, #10338:
URL: https://github.com/apache/paimon/pull/10338
### Purpose
Part 1 of the composite BTree index series, split from #10327. Add the
common-layer storage primitives for ordered, multi-column BTree keys and
complete-key point lookups.
- Encode typed tuple components with length delimiters and NULL markers;
compare keys lexicographically in declared field order.
- Construct composite BTree indexers through the existing
primary/extra-fields factory API while preserving scalar index creation.
- Copy tuple keys retained by the writer so reused input rows cannot mutate
posting-list grouping or min/max metadata.
- Add `GlobalIndexReader.visitCompositeEqual` with a default unsupported
result and a BTree implementation using the existing point lookup, Bloom filter
and row-ID filtering.
- Require the exact tuple arity when serializing keys or querying complete
tuples. SQL equality with a NULL literal returns an empty result.
- Retain composite files conservatively in factory-level scalar metadata
selection until composite predicate planning is introduced.
The existing BTree file versions and posting encodings are reused. This PR
is independently buildable and covers common storage and reader APIs. Table
query planning and SQL construction follow in the next parts.
### Series
1. **This PR:** common tuple storage and complete-key point lookup.
2. Common equality matcher/pruning/evaluated coverage; core multi-column
build, equality query/coverage; Python manifest and scalar-reader compatibility.
3. Spark/Flink SQL build integration, end-to-end equality tests and
documentation.
4. Composite prefix/range/IN/NULL queries, key-only filtering, scan budgets
and index selection.
### Tests
JDK 8 verification, without `fast-build`:
```sh
mvn -pl paimon-common -am -DwildcardSuites=none -DfailIfNoTests=false \
-Dtest=CompositeBTreeIndexTest,BTreeIndexReaderTest,LazyFilteredBTreeIndexReaderTest,BTreeIndexWriterCloseTest,SortedFileMetaSelectorTest
test
```
- 339 tests passed, including 4 focused composite cases and existing scalar
BTree regressions.
- Cover BTree versions 1/2, Bloom filters, LZ4 compression, reused mutable
rows, duplicate-key postings, missing keys, split-local row ranges, typed
ordering, NULLs, negative integers, embedded delimiters, malformed tuple arity
and physical ROW-type rejection.
- Checkstyle, Spotless, RAT and Enforcer passed; `git diff --check` passed.
### Split verification
The original implementation was captured at `c738a41321`; its branch was
preserved. The first patch was reconstructed in an isolated worktree and
rebased onto current master. Its patch is byte-for-byte identical across the
rebase.
An independent dependency review identified the tuple-arity validation as an
explicit correctness addition, included with regression assertions. The
full-series tree-equivalence check remains for the final part, with this
addition recorded as an intended correction.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]