goankur commented on code in PR #16705:
URL: https://github.com/apache/lucene/pull/16705#discussion_r4150449362
##########
lucene/core/src/java/org/apache/lucene/search/RescoreTopNQuery.java:
##########
@@ -70,25 +92,71 @@ public Query rewrite(IndexSearcher indexSearcher) throws
IOException {
}
DoubleValues rescores = rewrittenValueSource.getValues(leaf,
getDoubleValues(innerScorer));
DocIdSetIterator iterator = innerScorer.iterator();
+ while (iterator.nextDoc() != DocIdSetIterator.NO_MORE_DOCS) {
+ rescoreInto(queue, rescores, leaf.docBase, iterator.docID());
+ originalCount++;
+ }
+ }
+ return originalCount;
+ }
+
+ /**
+ * Starts the loads for every candidate before scoring any of them, so that
more than one read is
+ * in flight when the values live on slow storage. A rerank shortlist is
spread over every
+ * segment, so the prefetches for all segments are issued before any scoring
rather than a segment
+ * at a time. Only doc ids are buffered, never values.
+ */
+ private int rescoreWithPrefetch(
+ IndexReader reader, Weight weight, DoubleValuesSource
rewrittenValueSource, HitQueue queue)
+ throws IOException {
+ final List<LeafReaderContext> leaves = reader.leaves();
+ final DoubleValues[] leafValues = new DoubleValues[leaves.size()];
+ final int[][] leafDocs = new int[leaves.size()][];
Review Comment:
Ok, I used the structure from your #16656 sketch [in
comment](https://github.com/apache/lucene/pull/16656#issuecomment-5748849380) :
the ring holds PREFETCH_WINDOW = 128 (leaf, doc) entries, issues the prefetch
for the incoming candidate before scoring, and once full scores the oldest
entry.
Measured on a
[CohereLabs/msmarco-v2.1-embed-english-v3](https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3):
- 100M-vector, 4KB-aligned
- 1-bit BBQ index with float32 rescoring (405 GB index, 60 GB RAM)
- open-loop Poisson at 280 QPS
- 16 searcher threads
- Binary codes, HNSW graph and vector metadata preloaded
- Recall from a **preceding** 50k-query warm-up pass
- Same 50K queries replayed to collect the latency and I/O metrics
baseline | candidate (PrefetchRing) | delta
-- | -- | --
recall | 0.9601 | 0.9601 | identical
achieved QPS | 281.4 | 281.4 | 0
shed | 0.000% | 0.000% | 0
mean | 28.91 ms | 29.36 ms | +1.6%
p99 | 48.63 ms | 49.27 ms | +1.3%
p999 | 62.69 ms | 58.21 ms | −7.1%
max | 78.21 ms | 71.95 ms | −8.0%
aqu-sz | 13.86 | 11.83 | −14.6%
rareq-sz | 4.01 kB | 4.01 kB | identical
MB/s | 379.3 | 358.6 | −5.5%
IOPS | 96,970 | 91,674 | −5.5%
r_await | 0.143 ms | 0.129 ms | −9.8%
#### IMPORTANT NOTE
1. For anyone reproducing these numbers: they require the
`MemorySegmentIndexInput#prefetch` power-of-two backoff to be **bypassed**. At
defaults the shared counter samples away roughly `94%` of prefetch requests for
a `500-candidate` shortlist, so this PR's benefit largely doesn't materialize
on current main. Once #16145, is merged with an acceptable solution, these
numbers should be revaluate. I am happy to do so if it lands in the next week
or so beyond which I won't be able to retain the 100M vector index.
2. Alignment to 4K is a small one line change that is also required, without
it a vector is split across page boundary leading to unnecessary doubling of
read bandwidth (8K per vector). I intend to create a fast-follow PR with that
change once this is merged.
3.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]