SteNicholas opened a new issue, #403: URL: https://github.com/apache/paimon-cpp/issues/403
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Paimon-cpp version `main` at commit `d7c5f607900ffd63b2e067800dc4b9097c24d888`. ### Minimal reproduce step 1. Create a data-evolution table with row tracking and `deletion-vectors.enabled = true`. Write 10 rows whose text column matches `apple` with different relevance, then build a full-text global index on that column. 2. Delete the three rows that match best, using deletion vectors. 3. Search for `apple` with `limit = 3` and `with_score = true` through `GlobalIndexScan::CreateReader` → `VisitFullTextSearch`. 4. Pass the result to `ScanContextBuilder::SetGlobalIndexResult`, then scan and read. ### What doesn't meet your expectations? The index still ranks the deleted rows, so they take all three top-k slots. The read then drops them, because deletion vectors are applied only at read time (`src/paimon/core/operation/data_evolution_split_read.cpp`). The query returns 0 rows instead of 3. Nothing in the global index scan path builds an include set from deletion vectors. Java excludes deleted rows before ranking (apache/paimon#8459, apache/paimon#9031): - `GlobalIndexLiveRowFilter.liveRows(table, snapshot, partitionFilter, rowRanges)` returns `null` unless `deletion-vectors.enabled` is set. - Otherwise it: 1. takes the union of every data file's `nonNullRowIdRange()` from a `ScanMode.ALL` snapshot read; 2. subtracts each deletion vector's positions, shifted by the file's `firstRowId`. This happens in a second pass so a deleted row is never added back by another file (apache/paimon#8939). - The computation is pinned to the plan snapshot (apache/paimon#8946, apache/paimon#9031). - The live set is intersected into each shard's include row ids. Only live rows are ranked, and BM25 statistics still come from the full index. **Expected:** `limit = k` returns k live rows whenever at least k live rows match. ### Anything else? Sub-issue of #399 (step 3: search correctness). Suggestions: - Compute live rows next to `GlobalIndexScanImpl`, pinned to the same snapshot as the index scan, and intersect them into the full-text include set. - The same helper can be reused by the table-level read ({{S5}}) and by vector search, which Java also filters (apache/paimon#8459). This builds on the include-row-id semantics and shard clipping in {{S2}} and {{S3}}. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
