jimczi opened a new pull request, #16749: URL: https://github.com/apache/lucene/pull/16749
Read advice belongs to a mapping, not to a file. A merge reads the raw vectors front to back and once, a graph search reads them at random, and both can happen on the same file at the same time. A merge used to change the advice on the mapping searches were using and put it back at the end. Now `Lucene99FlatVectorsReader.getMergeInstance()` maps the data file again and asks for sequential reads that are not worth keeping in the page cache. The mapping is created when a merge first asks for it, shared by the merge instances that follow, and released by `finishMerge()` when the last one is done. Searches keep their own mapping. The reopen says a merge is reading, with `IOContext.merge()`. That's a merge context with no `MergeInfo`, since the reader knows a merge is reading the file but not how large the merge is. A `Directory` can route it like any other merge read, and `DirectIODirectory` falls back to the file length when the merge size is unknown. The writers had nothing to go on either, every vectors data file was created with a bare context. Each one now says what its own file is, and the quantized format tells the raw delegate that its vectors are only read back to rescore. So the raw vectors say no-reuse when they are written and when a merge reads them, matching what they already say for search. Same idea as #16684 for stored fields. With both in, `IndexInput#updateIOContext` has no callers left in Lucene. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
