JingsongLi opened a new pull request, #10339:
URL: https://github.com/apache/paimon/pull/10339
### Purpose
Part 2 of the composite BTree series, following merged #10338 and split from
#10327. Connect complete-key equality lookups to data evolution index
construction, query planning and coverage accounting.
- Match a conjunction of equalities against every indexed field, independent
of predicate order, and construct the point key in declared index order. Prefer
the longest fully matched composite definition before reading single-column
postings.
- Prune tuple files using typed min/max metadata. NULL or contradictory
equalities yield no candidates; incomplete keys and missing metadata remain
conservative.
- Build ordered tuple keys through the core sorted-index writer/scanner,
preserve ordered definitions in manifests, and refresh when either indexed
component changes.
- Track the row ranges covered by the index paths that actually contributed,
including ranges whose files were safely pruned by metadata. Preserve
unsupported-reader fallback and scan uncovered rows in full/detail modes.
- Keep composite files out of scalar TopN, vector/full-text prefilters and
Python scalar readers. Allow scalar and distinct composite BTree definitions to
coexist in Java and Python manifests.
- Reuse the ordered `List<DataField>` factory APIs and tuple-arity
validation merged in #10338.
SQL topology integration in Spark/Flink follows in part 3. Composite
prefix/range/IN/IS NULL bounds and scan-budget/index-selection extensions
follow in part 4. This PR introduces no new index file format or table options.
### Tests
- **283 Java tests passed on JDK 8**, with Checkstyle, Spotless and Enforcer
enabled (`fast-build` was not used for final verification).
- **92 Python tests passed** across scalar search modes, manifest writes,
and vector prefilters. Flake8 passed for the changed Python files using the
repository configuration. The root RAT licensing check and `git diff --check`
also passed.
- Cover two/three-column keys, reversed predicate order, nested OR branches,
contradictory equalities, NULL components, typed tuple pruning, metadata gaps,
serialized reader-side plans, partial coverage in both directions,
full/detail/fast modes, dropped columns, incremental builds and refresh of
either component.
- Remove scalar index data files while retaining their manifest entries to
prove joint equality queries avoid those postings. A separate test builds
matching two- and three-column indexes, removes the shorter index's data files,
and proves eager/deferred queries choose the longest complete key.
- An isolated semantic mutation reversing the longest-key preference failed
that regression with the expected missing-shorter-index-file error. Production
source was restored and the original regression passed again.
```sh
# Prepare the bundled codegen plugin in a fresh checkout.
mvn -pl paimon-codegen-loader -am -DskipTests package
mvn -pl paimon-core -am -DwildcardSuites=none -DfailIfNoTests=false \
-Dtest=CompositeBTreePredicateTest,CompositeBTreeIndexTest,CompositeBTreeTableTest,GlobalIndexQueryTest,GlobalIndexEvaluatorTest,SortedFileMetaSelectorTest,SortedGlobalIndexScannerTest,SortedGlobalIndexWriterTest,IndexManifestFileHandlerTest,DataEvolutionBatchScanTest,BtreeGlobalIndexTableTest,BitmapGlobalIndexTableTest,MultiValueGlobalIndexTableTest,VectorSearchRowFilterExactnessTest,FullTextSearchBuilderTest,IndexQuerySplitTest
test
PYTHONPATH=paimon-python python3 -m unittest \
pypaimon.tests.global_index_scalar_search_mode_test \
pypaimon.tests.index_manifest_write_test \
pypaimon.tests.vector_search_filter_test
```
### Split verification
Reconstructed from the reviewed equality snapshot `4e34273c09` in a new
worktree based on master `df43423696`. The complete source branch at
`c738a41321` was preserved. The merged field-list API and arity checks remain
in place; no composite range-planning or engine topology changes are included.
Independent dependency, query/coverage, and storage/Python reviews found no
remaining P0–P2 issues. The longest-key selection regression is recorded as an
explicit test addition for the final-series tree-equivalence check.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]