TheR1sing3un opened a new pull request, #10006:
URL: https://github.com/apache/paimon/pull/10006

   ### Purpose
   
   Full-text search currently restricts `pre_filter` to partition columns, and 
hybrid queries reject data predicates when a full-text route is present. 
Support ordinary-column prefilters on data-evolution tables, applying the 
predicate before full-text Top-K selection and forwarding it to both hybrid 
routes.
   
   Reuse the vector filter's snapshot-pinned data-read helper to collect exact 
matching row IDs from the filter columns, then pass that bitmap to native 
full-text search. This read uses full scalar-index coverage so partial scalar 
indexes cannot exclude rows covered by the full-text plan. Partition-only 
queries retain their pruning path, and vector candidate-refinement policy is 
unchanged. Raw fallback preserves the unfiltered BM25 corpus; full mode also 
covers selected partitions with no full-text index files. Document the 
additional filter-column read cost.
   
   ### Tests
   
   - Native full-text indexes with no scalar index, BTree and Bitmap; partial 
scalar coverage, mixed indexed/raw data and raw-only partitions.
   - Non-partitioned tables, filtered Top-K, stable BM25 scores, AND/OR 
partition semantics, hybrid routing, empty results, historical 
snapshots/deletions and concurrent commits.
   - Relevant multimodal, vector filtering, exact-filter resource cleanup, Ray 
filter exactness, primary-key golden data and global-index coverage regressions 
passed.
   - Repository-configured Flake8, Apache license checks and Python 3.6 grammar 
checks passed.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to