goankur commented on code in PR #16705:
URL: https://github.com/apache/lucene/pull/16705#discussion_r4099817791
##########
lucene/core/src/java/org/apache/lucene/search/RescoreTopNQuery.java:
##########
@@ -71,24 +95,82 @@ public Query rewrite(IndexSearcher indexSearcher) throws
IOException {
DoubleValues rescores = rewrittenValueSource.getValues(leaf,
getDoubleValues(innerScorer));
DocIdSetIterator iterator = innerScorer.iterator();
while (iterator.nextDoc() != DocIdSetIterator.NO_MORE_DOCS) {
- int docId = iterator.docID();
- if (rescores.advanceExact(docId)) {
- double v = rescores.doubleValue();
- queue.insertWithOverflow(new ScoreDoc(leaf.docBase + docId, (float)
v));
- } else {
- queue.insertWithOverflow(new ScoreDoc(leaf.docBase + docId, 0f));
- }
+ rescoreInto(queue, rescores, leaf.docBase, iterator.docID());
originalCount++;
}
}
- int i = 0;
- ScoreDoc[] scoreDocs = new ScoreDoc[queue.size()];
- for (ScoreDoc topDoc : queue) {
- scoreDocs[i++] = topDoc;
+ return originalCount;
+ }
+
+ /**
+ * Starts the loads for every candidate before scoring any of them, so that
more than one read is
+ * in flight when the values live on slow storage. Each segment's prefetches
are issued on the
+ * searcher's executor: prefetching costs real CPU per candidate, so issuing
a whole shortlist
+ * from a single thread caps how many reads can be outstanding. Only doc ids
are buffered, never
+ * values.
+ */
+ private int rescoreWithPrefetch(
+ IndexSearcher indexSearcher,
+ IndexReader reader,
+ Weight weight,
+ DoubleValuesSource rewrittenValueSource,
+ HitQueue queue)
+ throws IOException {
+ final List<LeafReaderContext> leaves = reader.leaves();
+ final DoubleValues[] leafValues = new DoubleValues[leaves.size()];
+ final int[][] leafDocs = new int[leaves.size()][];
+ final List<Callable<Void>> tasks = new ArrayList<>(leaves.size());
+ for (int i = 0; i < leaves.size(); i++) {
+ final int idx = i;
+ final LeafReaderContext leaf = leaves.get(i);
+ tasks.add(
+ () -> {
+ Scorer innerScorer = weight.scorer(leaf);
+ if (innerScorer == null) {
+ return null;
+ }
+ DoubleValues rescores = rewrittenValueSource.getValues(leaf, null);
+ DocIdSetIterator iterator = innerScorer.iterator();
+ int[] docs = new int[16];
+ int count = 0;
+ while (iterator.nextDoc() != DocIdSetIterator.NO_MORE_DOCS) {
+ int docId = iterator.docID();
+ rescores.prefetch(docId);
+ if (count == docs.length) {
+ docs = ArrayUtil.grow(docs, count + 1);
+ }
+ docs[count++] = docId;
+ }
+ leafValues[idx] = rescores;
+ leafDocs[idx] = ArrayUtil.copyOfSubArray(docs, 0, count);
+ return null;
+ });
+ }
+ indexSearcher.getTaskExecutor().invokeAll(tasks);
Review Comment:
Task executor is out in the next commit. It was in the diff while the
description claimed otherwise, my bad. Agreed on the split too: per-leaf gives
one task most of the shortlist. The follow-up PR will pool
candidates across leaves into equal ranges, with a `DoubleValues` per task
per leaf.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]