JingsongLi opened a new pull request, #10339:
URL: https://github.com/apache/paimon/pull/10339

   ### Purpose
   
   Part 2 of the composite BTree series, following merged #10338 and split from 
#10327. Connect complete-key equality lookups to data evolution index 
construction, query planning and coverage accounting.
   
   - Match a conjunction of equalities against every indexed field, independent 
of predicate order, and construct the point key in declared index order. Prefer 
the longest fully matched composite definition before reading single-column 
postings.
   - Prune tuple files using typed min/max metadata. NULL or contradictory 
equalities yield no candidates; incomplete keys and missing metadata remain 
conservative.
   - Build ordered tuple keys through the core sorted-index writer/scanner, 
preserve ordered definitions in manifests, and refresh when either indexed 
component changes.
   - Track the row ranges covered by the index paths that actually contributed, 
including ranges whose files were safely pruned by metadata. Preserve 
unsupported-reader fallback and scan uncovered rows in full/detail modes.
   - Keep composite files out of scalar TopN, vector/full-text prefilters and 
Python scalar readers. Allow scalar and distinct composite BTree definitions to 
coexist in Java and Python manifests.
   - Reuse the ordered `List<DataField>` factory APIs and tuple-arity 
validation merged in #10338.
   
   SQL topology integration in Spark/Flink follows in part 3. Composite 
prefix/range/IN/IS NULL bounds and scan-budget/index-selection extensions 
follow in part 4. This PR introduces no new index file format or table options.
   
   ### Tests
   
   - **283 Java tests passed on JDK 8**, with Checkstyle, Spotless and Enforcer 
enabled (`fast-build` was not used for final verification).
   - **92 Python tests passed** across scalar search modes, manifest writes, 
and vector prefilters. Flake8 passed for the changed Python files using the 
repository configuration. The root RAT licensing check and `git diff --check` 
also passed.
   - Cover two/three-column keys, reversed predicate order, nested OR branches, 
contradictory equalities, NULL components, typed tuple pruning, metadata gaps, 
serialized reader-side plans, partial coverage in both directions, 
full/detail/fast modes, dropped columns, incremental builds and refresh of 
either component.
   - Remove scalar index data files while retaining their manifest entries to 
prove joint equality queries avoid those postings. A separate test builds 
matching two- and three-column indexes, removes the shorter index's data files, 
and proves eager/deferred queries choose the longest complete key.
   - An isolated semantic mutation reversing the longest-key preference failed 
that regression with the expected missing-shorter-index-file error. Production 
source was restored and the original regression passed again.
   
   ```sh
   # Prepare the bundled codegen plugin in a fresh checkout.
   mvn -pl paimon-codegen-loader -am -DskipTests package
   
   mvn -pl paimon-core -am -DwildcardSuites=none -DfailIfNoTests=false \
     
-Dtest=CompositeBTreePredicateTest,CompositeBTreeIndexTest,CompositeBTreeTableTest,GlobalIndexQueryTest,GlobalIndexEvaluatorTest,SortedFileMetaSelectorTest,SortedGlobalIndexScannerTest,SortedGlobalIndexWriterTest,IndexManifestFileHandlerTest,DataEvolutionBatchScanTest,BtreeGlobalIndexTableTest,BitmapGlobalIndexTableTest,MultiValueGlobalIndexTableTest,VectorSearchRowFilterExactnessTest,FullTextSearchBuilderTest,IndexQuerySplitTest
 test
   
   PYTHONPATH=paimon-python python3 -m unittest \
     pypaimon.tests.global_index_scalar_search_mode_test \
     pypaimon.tests.index_manifest_write_test \
     pypaimon.tests.vector_search_filter_test
   ```
   
   ### Split verification
   
   Reconstructed from the reviewed equality snapshot `4e34273c09` in a new 
worktree based on master `df43423696`. The complete source branch at 
`c738a41321` was preserved. The merged field-list API and arity checks remain 
in place; no composite range-planning or engine topology changes are included.
   
   Independent dependency, query/coverage, and storage/Python reviews found no 
remaining P0–P2 issues. The longest-key selection regression is recorded as an 
explicit test addition for the final-series tree-equivalence check.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to