[
https://issues.apache.org/jira/browse/SOLR-16560?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18118858#comment-18118858
]
Andrzej Bialecki commented on SOLR-16560:
-----------------------------------------
I spent some Claude tokens and this is what Claude Opus thinks about this
report:
{noformat}
Real churn site — incremental path replaces the writer anyway:
- IndexFetcher.java:666 closeIndexWriter(solrCore, true) before download
- IndexFetcher.java:731 openIndexWriter(solrCore) in finally
- → DefaultSolrCoreState.changeWriter :216 createMainIndexWriter(...) → fresh
SolrIndexWriter, fresh empty ReaderPool
- → SolrCore.java:2330 DirectoryReader.openIfChanged(currentReader,
writer.get(), true) hits Lucene's writer != this.writer
branch → writer.getReader(...), no oldReaders passed → every segment rebuilt.
Side effect nobody noted on the issue: per-segment IndexFingerprint cache
(SolrCore.java:2143) fully invalidated each fetch.
Microbenchmark: 0.38–0.79 ms/segment full rebuild vs 7–13 ms incremental reopen
→ ~3–5×, ~40–60 ms at 100 segments. At _default poll cadence (7.5 s via
ReplicateFromLeader.java:165) ≈ 0.4% of one core per replica.
10.0 tag is code-justified — every line in the chain predates 8.8 and is
unchanged today. But magnitude is likely smaller on 10.x (blocktree no longer
loads FST on open) and unmeasured.
Fix: open PULL-replica searchers non-NRT from the Directory at core open
(SolrCore.initSearcher:1218-1233, plus :2325/:2402), discriminator
Replica.Type.PULL.{noformat}
For now I'm demoting this from Critical to Major.
> Pull replication appears to consume excessive CPU as it does not use Segment
> Reader pooling
> -------------------------------------------------------------------------------------------
>
> Key: SOLR-16560
> URL: https://issues.apache.org/jira/browse/SOLR-16560
> Project: Solr
> Issue Type: Bug
> Components: SolrCloud
> Affects Versions: 8.8, 10.0
> Reporter: Patson Luk
> Priority: Critical
>
> While we are experimenting with adding PULL replica to our solr cluster, it's
> found from profiling that IndexFetcher seems to impose much more CPU overhead
> than anticipated , and most of those CPU time are spent in
> `SegmentReader.init` which is called by `SolrCore.openNewSearch` from the
> `IndexFetcher`.
> With some debugging, it's found that for every replication on an updated
> collection, a new `SegmentReader` is created for every segment for such
> collection (not only the ones that are pulled down), compared to
> `SolrCore.openNewSearch` triggered from regular commit, which ONLY creates
> `SegmentReader` for new segments, which old `SegmentReader`s are obtained
> from `ReaderPool`.
> Unfortunately, such pool does not work for `IndexFetcher` as it opens a new
> `IndexWriter` on every run at
> [here|https://github.com/apache/solr/blob/main/solr/core/src/java/org/apache/solr/handler/IndexFetcher.java#L749],
> which creates a new `ReaderPool`.
> I am not familiar enough to tell whether such new `IndexWriter` is always
> needed, but opening `SegmentReader` for every segment of a collection seems
> excessive for pull replication.
> Any thoughts on this please? Many thanks!! :)
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]