JingsongLi opened a new pull request, #10338:
URL: https://github.com/apache/paimon/pull/10338

   ### Purpose
   
   Part 1 of the composite BTree index series, split from #10327. Add the 
common-layer storage primitives for ordered, multi-column BTree keys and 
complete-key point lookups.
   
   - Encode typed tuple components with length delimiters and NULL markers; 
compare keys lexicographically in declared field order.
   - Construct composite BTree indexers through the existing 
primary/extra-fields factory API while preserving scalar index creation.
   - Copy tuple keys retained by the writer so reused input rows cannot mutate 
posting-list grouping or min/max metadata.
   - Add `GlobalIndexReader.visitCompositeEqual` with a default unsupported 
result and a BTree implementation using the existing point lookup, Bloom filter 
and row-ID filtering.
   - Require the exact tuple arity when serializing keys or querying complete 
tuples. SQL equality with a NULL literal returns an empty result.
   - Retain composite files conservatively in factory-level scalar metadata 
selection until composite predicate planning is introduced.
   
   The existing BTree file versions and posting encodings are reused. This PR 
is independently buildable and covers common storage and reader APIs. Table 
query planning and SQL construction follow in the next parts.
   
   ### Series
   
   1. **This PR:** common tuple storage and complete-key point lookup.
   2. Common equality matcher/pruning/evaluated coverage; core multi-column 
build, equality query/coverage; Python manifest and scalar-reader compatibility.
   3. Spark/Flink SQL build integration, end-to-end equality tests and 
documentation.
   4. Composite prefix/range/IN/NULL queries, key-only filtering, scan budgets 
and index selection.
   
   ### Tests
   
   JDK 8 verification, without `fast-build`:
   
   ```sh
   mvn -pl paimon-common -am -DwildcardSuites=none -DfailIfNoTests=false \
     
-Dtest=CompositeBTreeIndexTest,BTreeIndexReaderTest,LazyFilteredBTreeIndexReaderTest,BTreeIndexWriterCloseTest,SortedFileMetaSelectorTest
 test
   ```
   
   - 339 tests passed, including 4 focused composite cases and existing scalar 
BTree regressions.
   - Cover BTree versions 1/2, Bloom filters, LZ4 compression, reused mutable 
rows, duplicate-key postings, missing keys, split-local row ranges, typed 
ordering, NULLs, negative integers, embedded delimiters, malformed tuple arity 
and physical ROW-type rejection.
   - Checkstyle, Spotless, RAT and Enforcer passed; `git diff --check` passed.
   
   ### Split verification
   
   The original implementation was captured at `c738a41321`; its branch was 
preserved. The first patch was reconstructed in an isolated worktree and 
rebased onto current master. Its patch is byte-for-byte identical across the 
rebase.
   
   An independent dependency review identified the tuple-arity validation as an 
explicit correctness addition, included with regression assertions. The 
full-series tree-equivalence check remains for the final part, with this 
addition recorded as an intended correction.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to