SteNicholas opened a new issue, #402: URL: https://github.com/apache/paimon-cpp/issues/402
### Search before asking - [x] I searched in the [issues](https://github.com/apache/paimon-cpp/issues) and found nothing similar. ### Paimon-cpp version `main` at commit `d7c5f607900ffd63b2e067800dc4b9097c24d888`. ### Minimal reproduce step 1. Create a data-evolution table (row tracking, `bucket = -1`). Build a `tantivy-fulltext` or `lucene-fts` global index on a text column over two row ranges, for example `[0, 99]` and `[100, 199]`. Each range should have at least 5 rows matching `apple`. 2. Get the reader with `GlobalIndexScan::CreateReader(field, "tantivy-fulltext")`. `GlobalIndexScanImpl::CreateReaders` wraps each range in `OffsetGlobalIndexReader(reader, range.from)` and combines them in one `UnionGlobalIndexReader` (`src/paimon/core/global_index/global_index_scan_impl.cpp:198-209`). 3. Call `VisitFullTextSearch` for `apple` with `limit = 5` and `with_score = true`. Repeat with `with_score = false`. 4. Call it again with `pre_filter = {5, 150}`. ### What doesn't meet your expectations? **Union reader (step 3)** - Step 3 returns 10 rows. `UnionGlobalIndexReader::VisitFullTextSearch` (`src/paimon/common/global_index/union_global_index_reader.cpp:159-180`) ORs the per-shard results without a final top-k. A search with `limit = k` can therefore return up to shards × k rows, and they are not necessarily the global top k by score. - If two shards return the same row id, for example from overlapping index ranges, `BitmapScoredGlobalIndexResult::Or` fails with `not support two BitmapScoredGlobalIndexResult or with same row id` (`src/paimon/common/global_index/bitmap_scored_global_index_result.cpp:88-109`). ORing a scored result with an unscored one drops the scores. - Java does this differently (apache/paimon#8301, apache/paimon#8809, apache/paimon#9955): - each shard is asked for `limit` candidates; - the results are combined with `ScoredGlobalIndexResult#merge`, where the first score wins for duplicate row ids; - `ScoredGlobalIndexResult#topK(limit)` is applied last, ordered by score descending, then row id ascending. **Offset reader (step 4)** - `OffsetGlobalIndexReader::VisitFullTextSearch` (`src/paimon/common/global_index/offset_global_index_reader.cpp:148-164`) copies the whole global `pre_filter` into every shard. It subtracts the offset without clipping to the shard range: - For the second shard, `5 - 100` is added as a negative `int64_t`, which is a huge unsigned id in the bitmap. - For the first shard, `150` is kept even though it is outside that shard. - A shard with no filtered rows still runs a search. - The reader only knows `offset_`, not the end of the shard range. - In Java (apache/paimon#8459), `OffsetGlobalIndexReader(wrapped, offset, to)` calls `FullTextSearch#offsetRange(offset, to)`. That intersects the include set with `[offset, to]` before shifting it. The native reader then returns an empty result for an empty include set. **Expected** - `limit = k` returns at most k rows: the top k by score across all shards, with ties broken by row id. - Each shard receives only its own filtered row ids, already converted to local ids. - A shard whose include set is empty is skipped. ### Anything else? Sub-issue of #399 (step 3: search correctness). Suggested fix: - Pass `range.to` to `OffsetGlobalIndexReader`, clip the include set before shifting it, and skip shards whose include set is empty. - Add a scored merge that tolerates duplicate row ids and a `TopK(limit)` on scored results, and apply both in `UnionGlobalIndexReader::VisitFullTextSearch`. For the unscored path, truncate the union to `limit`. - Add regression tests to `offset_global_index_reader_test.cpp` and `union_global_index_reader_test.cpp`. This fix can use the include-row-id semantics from {{S2}}, but it does not depend on them. ### Are you willing to submit a PR? - [x] I'm willing to submit a PR! -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
