goankur commented on code in PR #16705:
URL: https://github.com/apache/lucene/pull/16705#discussion_r4150449362


##########
lucene/core/src/java/org/apache/lucene/search/RescoreTopNQuery.java:
##########
@@ -70,25 +92,71 @@ public Query rewrite(IndexSearcher indexSearcher) throws 
IOException {
       }
       DoubleValues rescores = rewrittenValueSource.getValues(leaf, 
getDoubleValues(innerScorer));
       DocIdSetIterator iterator = innerScorer.iterator();
+      while (iterator.nextDoc() != DocIdSetIterator.NO_MORE_DOCS) {
+        rescoreInto(queue, rescores, leaf.docBase, iterator.docID());
+        originalCount++;
+      }
+    }
+    return originalCount;
+  }
+
+  /**
+   * Starts the loads for every candidate before scoring any of them, so that 
more than one read is
+   * in flight when the values live on slow storage. A rerank shortlist is 
spread over every
+   * segment, so the prefetches for all segments are issued before any scoring 
rather than a segment
+   * at a time. Only doc ids are buffered, never values.
+   */
+  private int rescoreWithPrefetch(
+      IndexReader reader, Weight weight, DoubleValuesSource 
rewrittenValueSource, HitQueue queue)
+      throws IOException {
+    final List<LeafReaderContext> leaves = reader.leaves();
+    final DoubleValues[] leafValues = new DoubleValues[leaves.size()];
+    final int[][] leafDocs = new int[leaves.size()][];

Review Comment:
   Ok, I used the structure from your #16656 sketch [in 
comment](https://github.com/apache/lucene/pull/16656#issuecomment-5748849380) : 
the ring holds PREFETCH_WINDOW = 128 (leaf, doc) entries, issues the prefetch 
for the incoming candidate before scoring, and once full scores the oldest 
entry. 
   
   Measured on a 
[CohereLabs/msmarco-v2.1-embed-english-v3](https://huggingface.co/datasets/CohereLabs/msmarco-v2.1-embed-english-v3):
 
   - AWS `g6.4xl` instance: 64 GB RAM, 16 vCPU, 1x600 GB NVMe SSD 
   - 100M-vector, 4KB-aligned 
   - 1-bit BBQ index with float32 rescoring (405 GB index, 60 GB RAM, 550 TB 
SSD)
   - open-loop Poisson at 280 QPS
   - 16 searcher threads
   - Binary codes, HNSW graph and vector metadata preloaded
   - Recall from a **preceding** 50k-query warm-up pass
   - Same 50K queries replayed to collect the latency and I/O metrics
   
   
   baseline | candidate  (PrefetchRing) | delta
   -- | -- | --
   recall | 0.9601 | 0.9601 | identical
   achieved QPS | 281.4 | 281.4 | 0
   shed | 0.000% | 0.000% | 0
   mean | 28.91 ms | 29.36 ms | +1.6%
   p99 | 48.63 ms | 49.27 ms | +1.3%
   p999 | 62.69 ms | 58.21 ms | −7.1%
   max | 78.21 ms | 71.95 ms | −8.0%
   aqu-sz | 13.86 | 11.83 | −14.6%
   rareq-sz | 4.01 kB | 4.01 kB | identical
   MB/s | 379.3 | 358.6 | −5.5%
   IOPS | 96,970 | 91,674 | −5.5%
   r_await | 0.143 ms | 0.129 ms | −9.8%
   
   #### IMPORTANT NOTE
   
   1. For anyone reproducing these numbers: they require the 
`MemorySegmentIndexInput#prefetch` power-of-two backoff to be **bypassed**. At 
defaults the shared counter samples away roughly `94%` of prefetch requests for 
a `500-candidate` shortlist, so this PR's benefit largely doesn't materialize 
on current main. Once #16145, is merged with an acceptable solution, these 
numbers should be revaluate. I am happy to do so if it lands in the next week 
or so beyond which I won't be able to retain the 100M vector index.
   2. Alignment to 4K is a small one line change that is also required, without 
it a vector is split across page boundary leading to unnecessary doubling of 
read bandwidth (8K per vector). I intend to create a fast-follow PR with that 
change once this is merged. 
   



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to