JingsongLi opened a new pull request, #10327:
URL: https://github.com/apache/paimon/pull/10327

   ### Purpose
   
   Support multi-column BTree global indexes for data evolution tables. Queries 
such as `materialization_id = ? AND episode_index = ?` currently read each 
single-column posting list before intersecting them, which is expensive when 
one condition matches many row IDs. A composite index reads the posting list 
for the complete tuple directly.
   
   - Add typed, length-delimited tuple keys with lexicographic comparison, null 
handling, full-key bloom filtering and min/max pruning. Reuse the existing 
BTree file versions and retain single-column behavior.
   - Build composite indexes in Spark and Flink with `index_column => 
'materialization_id,episode_index'`. Carry all key columns through sorting, 
metadata, incremental builds and refresh after component updates.
   - Prefer the longest complete composite equality match in planning-time and 
reader-side execution, independently of predicate order. Allow composite and 
single-column indexes to coexist.
   - Track coverage for the selected index path so full/detail searches scan 
uncovered rows without borrowing coverage from a different index. Preserve 
coverage when key metadata prunes every index file.
   - Skip composite index files in PyPaimon's scalar reader and document the 
supported behavior.
   
   The initial query path supports equality predicates on every indexed column. 
Prefix and range queries use existing single-column indexes or table scans. 
PyPaimon does not build or read composite indexes. No production-scale 
benchmark has been run.
   
   ### Tests
   
   All commands below passed without the `fast-build` profile:
   
   ```sh
   mvn -pl paimon-core -am -DfailIfNoTests=false -DwildcardSuites=none \
     
-Dtest=CompositeBTreeIndexTest,CompositeBTreeTableTest,BtreeGlobalIndexTableTest,GlobalIndexQueryTest,GlobalIndexEvaluatorTest,SortedGlobalIndexScannerTest,SortedGlobalIndexWriterTest,IndexManifestFileHandlerTest,DataEvolutionBatchScanTest,BitmapGlobalIndexTableTest,MultiValueGlobalIndexTableTest,BTreeIndexWriterCloseTest,BTreePostingListTest
 test
   
   mvn -pl paimon-flink/paimon-flink-common,paimon-spark/paimon-spark-ut -am \
     -Pflink1 -Pspark3 -DfailIfNoTests=false \
     
-Dtest=SortedGlobalIndexITCase#testCompositeBTreeIndex,SortedIndexTopoBuilderTest,CreateGlobalIndexProcedureTest
 \
     
-DwildcardSuites=org.apache.paimon.spark.procedure.CompositeBTreeIndexProcedureTest
 test
   
   PYTHONPATH=paimon-python python3 -m unittest 
pypaimon.tests.global_index_scalar_search_mode_test
   ```
   
   - 167 common/core tests, 12 Flink tests, 15 Spark Java tests, one Spark 
Scala integration test and seven Python tests passed.
   - New regressions cover two- and three-column lookups, reversed predicate 
order, OR branches, contradictions, nulls, negative integers, embedded 
delimiters, mutable input rows, BTree versions 1/2, partial coverage in both 
directions, serialized reader-side plans, incremental builds and refresh after 
either component changes.
   - The single-column index files are removed in a joint-query regression to 
prove that composite lookups do not read their posting lists.
   - Maven Checkstyle, Spotless, RAT and Enforcer checks passed; `git diff 
--check` passed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to