SteNicholas opened a new issue, #403:
URL: https://github.com/apache/paimon-cpp/issues/403

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   ### Paimon-cpp version
   
   `main` at commit `d7c5f607900ffd63b2e067800dc4b9097c24d888`.
   
   ### Minimal reproduce step
   
   1. Create a data-evolution table with row tracking and 
`deletion-vectors.enabled = true`. Write 10 rows whose text column matches 
`apple` with different relevance, then build a full-text global index on that 
column.
   2. Delete the three rows that match best, using deletion vectors.
   3. Search for `apple` with `limit = 3` and `with_score = true` through 
`GlobalIndexScan::CreateReader` → `VisitFullTextSearch`.
   4. Pass the result to `ScanContextBuilder::SetGlobalIndexResult`, then scan 
and read.
   
   ### What doesn't meet your expectations?
   
   The index still ranks the deleted rows, so they take all three top-k slots. 
The read then drops them, because deletion vectors are applied only at read 
time (`src/paimon/core/operation/data_evolution_split_read.cpp`). The query 
returns 0 rows instead of 3. Nothing in the global index scan path builds an 
include set from deletion vectors.
   
   Java excludes deleted rows before ranking (apache/paimon#8459, 
apache/paimon#9031):
   
   - `GlobalIndexLiveRowFilter.liveRows(table, snapshot, partitionFilter, 
rowRanges)` returns `null` unless `deletion-vectors.enabled` is set.
   - Otherwise it:
     1. takes the union of every data file's `nonNullRowIdRange()` from a 
`ScanMode.ALL` snapshot read;
     2. subtracts each deletion vector's positions, shifted by the file's 
`firstRowId`. This happens in a second pass so a deleted row is never added 
back by another file (apache/paimon#8939).
   - The computation is pinned to the plan snapshot (apache/paimon#8946, 
apache/paimon#9031).
   - The live set is intersected into each shard's include row ids. Only live 
rows are ranked, and BM25 statistics still come from the full index.
   
   **Expected:** `limit = k` returns k live rows whenever at least k live rows 
match.
   
   ### Anything else?
   
   Sub-issue of #399 (step 3: search correctness).
   
   Suggestions:
   - Compute live rows next to `GlobalIndexScanImpl`, pinned to the same 
snapshot as the index scan, and intersect them into the full-text include set.
   - The same helper can be reused by the table-level read ({{S5}}) and by 
vector search, which Java also filters (apache/paimon#8459).
   
   This builds on the include-row-id semantics and shard clipping in {{S2}} and 
{{S3}}.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to