JingsongLi opened a new pull request, #10327:
URL: https://github.com/apache/paimon/pull/10327
### Purpose
Support multi-column BTree global indexes for data evolution tables. Queries
such as `materialization_id = ? AND episode_index = ?` currently read each
single-column posting list before intersecting them, which is expensive when
one condition matches many row IDs. A composite index reads the posting list
for the complete tuple directly.
- Add typed, length-delimited tuple keys with lexicographic comparison, null
handling, full-key bloom filtering and min/max pruning. Reuse the existing
BTree file versions and retain single-column behavior.
- Build composite indexes in Spark and Flink with `index_column =>
'materialization_id,episode_index'`. Carry all key columns through sorting,
metadata, incremental builds and refresh after component updates.
- Prefer the longest complete composite equality match in planning-time and
reader-side execution, independently of predicate order. Allow composite and
single-column indexes to coexist.
- Track coverage for the selected index path so full/detail searches scan
uncovered rows without borrowing coverage from a different index. Preserve
coverage when key metadata prunes every index file.
- Skip composite index files in PyPaimon's scalar reader and document the
supported behavior.
The initial query path supports equality predicates on every indexed column.
Prefix and range queries use existing single-column indexes or table scans.
PyPaimon does not build or read composite indexes. No production-scale
benchmark has been run.
### Tests
All commands below passed without the `fast-build` profile:
```sh
mvn -pl paimon-core -am -DfailIfNoTests=false -DwildcardSuites=none \
-Dtest=CompositeBTreeIndexTest,CompositeBTreeTableTest,BtreeGlobalIndexTableTest,GlobalIndexQueryTest,GlobalIndexEvaluatorTest,SortedGlobalIndexScannerTest,SortedGlobalIndexWriterTest,IndexManifestFileHandlerTest,DataEvolutionBatchScanTest,BitmapGlobalIndexTableTest,MultiValueGlobalIndexTableTest,BTreeIndexWriterCloseTest,BTreePostingListTest
test
mvn -pl paimon-flink/paimon-flink-common,paimon-spark/paimon-spark-ut -am \
-Pflink1 -Pspark3 -DfailIfNoTests=false \
-Dtest=SortedGlobalIndexITCase#testCompositeBTreeIndex,SortedIndexTopoBuilderTest,CreateGlobalIndexProcedureTest
\
-DwildcardSuites=org.apache.paimon.spark.procedure.CompositeBTreeIndexProcedureTest
test
PYTHONPATH=paimon-python python3 -m unittest
pypaimon.tests.global_index_scalar_search_mode_test
```
- 167 common/core tests, 12 Flink tests, 15 Spark Java tests, one Spark
Scala integration test and seven Python tests passed.
- New regressions cover two- and three-column lookups, reversed predicate
order, OR branches, contradictions, nulls, negative integers, embedded
delimiters, mutable input rows, BTree versions 1/2, partial coverage in both
directions, serialized reader-side plans, incremental builds and refresh after
either component changes.
- The single-column index files are removed in a joint-query regression to
prove that composite lookups do not read their posting lists.
- Maven Checkstyle, Spotless, RAT and Enforcer checks passed; `git diff
--check` passed.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]