SteNicholas opened a new issue, #402:
URL: https://github.com/apache/paimon-cpp/issues/402

   ### Search before asking
   
   - [x] I searched in the 
[issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar.
   
   ### Paimon-cpp version
   
   `main` at commit `d7c5f607900ffd63b2e067800dc4b9097c24d888`.
   
   ### Minimal reproduce step
   
   1. Create a data-evolution table (row tracking, `bucket = -1`). Build a 
`tantivy-fulltext` or `lucene-fts` global index on a text column over two row 
ranges, for example `[0, 99]` and `[100, 199]`. Each range should have at least 
5 rows matching `apple`.
   2. Get the reader with `GlobalIndexScan::CreateReader(field, 
"tantivy-fulltext")`. `GlobalIndexScanImpl::CreateReaders` wraps each range in 
`OffsetGlobalIndexReader(reader, range.from)` and combines them in one 
`UnionGlobalIndexReader` 
(`src/paimon/core/global_index/global_index_scan_impl.cpp:198-209`).
   3. Call `VisitFullTextSearch` for `apple` with `limit = 5` and `with_score = 
true`. Repeat with `with_score = false`.
   4. Call it again with `pre_filter = {5, 150}`.
   
   ### What doesn't meet your expectations?
   
   **Union reader (step 3)**
   
   - Step 3 returns 10 rows. `UnionGlobalIndexReader::VisitFullTextSearch` 
(`src/paimon/common/global_index/union_global_index_reader.cpp:159-180`) ORs 
the per-shard results without a final top-k. A search with `limit = k` can 
therefore return up to shards × k rows, and they are not necessarily the global 
top k by score.
   - If two shards return the same row id, for example from overlapping index 
ranges, `BitmapScoredGlobalIndexResult::Or` fails with `not support two 
BitmapScoredGlobalIndexResult or with same row id` 
(`src/paimon/common/global_index/bitmap_scored_global_index_result.cpp:88-109`).
 ORing a scored result with an unscored one drops the scores.
   - Java does this differently (apache/paimon#8301, apache/paimon#8809, 
apache/paimon#9955):
     - each shard is asked for `limit` candidates;
     - the results are combined with `ScoredGlobalIndexResult#merge`, where the 
first score wins for duplicate row ids;
     - `ScoredGlobalIndexResult#topK(limit)` is applied last, ordered by score 
descending, then row id ascending.
   
   **Offset reader (step 4)**
   
   - `OffsetGlobalIndexReader::VisitFullTextSearch` 
(`src/paimon/common/global_index/offset_global_index_reader.cpp:148-164`) 
copies the whole global `pre_filter` into every shard. It subtracts the offset 
without clipping to the shard range:
     - For the second shard, `5 - 100` is added as a negative `int64_t`, which 
is a huge unsigned id in the bitmap.
     - For the first shard, `150` is kept even though it is outside that shard.
     - A shard with no filtered rows still runs a search.
   - The reader only knows `offset_`, not the end of the shard range.
   - In Java (apache/paimon#8459), `OffsetGlobalIndexReader(wrapped, offset, 
to)` calls `FullTextSearch#offsetRange(offset, to)`. That intersects the 
include set with `[offset, to]` before shifting it. The native reader then 
returns an empty result for an empty include set.
   
   **Expected**
   
   - `limit = k` returns at most k rows: the top k by score across all shards, 
with ties broken by row id.
   - Each shard receives only its own filtered row ids, already converted to 
local ids.
   - A shard whose include set is empty is skipped.
   
   ### Anything else?
   
   Sub-issue of #399 (step 3: search correctness).
   
   Suggested fix:
   - Pass `range.to` to `OffsetGlobalIndexReader`, clip the include set before 
shifting it, and skip shards whose include set is empty.
   - Add a scored merge that tolerates duplicate row ids and a `TopK(limit)` on 
scored results, and apply both in 
`UnionGlobalIndexReader::VisitFullTextSearch`. For the unscored path, truncate 
the union to `limit`.
   - Add regression tests to `offset_global_index_reader_test.cpp` and 
`union_global_index_reader_test.cpp`.
   
   This fix can use the include-row-id semantics from {{S2}}, but it does not 
depend on them.
   
   ### Are you willing to submit a PR?
   
   - [x] I'm willing to submit a PR!
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to