[ 
https://issues.apache.org/jira/browse/SOLR-18316?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18098422#comment-18098422
 ] 

Justin Pealing commented on SOLR-18316:
---------------------------------------

I believe thisĀ _might_ be related to SOLR-18301, but I'm not familiar enough to 
understand if this is the same root cause or not.

> Slow Solrcloud recovery following restart of Zookeeper leader
> -------------------------------------------------------------
>
>                 Key: SOLR-18316
>                 URL: https://issues.apache.org/jira/browse/SOLR-18316
>             Project: Solr
>          Issue Type: Bug
>          Components: SolrCloud
>    Affects Versions: 10.0
>         Environment: Docker Desktop (Windows)
>            Reporter: Justin Pealing
>            Priority: Major
>         Attachments: docker-compose.yml
>
>
> Since updating to Solrcloud 10 we have been experiencing periods of downtime 
> when the Zookeeper leader restarts. I have a docker-compose file which 
> reproduces this issue, and I believe this is a regression from 9.10.1.
> h2. Steps to reproduce
> The attached docker-compose.yml creates a standard 2 node Solrcloud cluster 
> with a 3 node Zookeeper, as well as a short script to create some empty 
> collections. (This doesn't exactly match our production environment but I 
> wanted to make sure none of our configuration was causing this issue).
> {{docker compose up -d}}
> {{Wait for Solr to come up}}
> {{docker compose run --rm collections 10}}
> Once the collections are up, identify and restart the Zookeeper leader (e.g. 
> `docker compose restart zoo1`). This results in 1 or more of the Solr nodes 
> showing as "(DEAD)" in the dashboard, and a number of replicas showing as 
> "Down" for a period of time (which may or may not result in collections being 
> unavailable, depending on which nodes are affected), after which the 
> Solrcloud cluster self-recovers.
> The amount of downtime appears to be proportional to the number of 
> collections. With 1 collection recovery takes a few seconds, with 10 recovery 
> takes a few minutes. In live we have hundreds of collections and we have seen 
> Solr remain down for hours in this scenario (presumably it would eventually 
> recover). Restarting the down Solr nodes also works for recovery.
> This is reproducible with 100% reliability when restarting the Zookeeper 
> leader with the attached docker-compose.yml, but when testing with 
> `solr:9.10.1` (replace the image tags in the docker-compose.yml) there is no 
> perceptible downtime.
> h2. Logs
> I've not included too many logs here as the issue is reproducible in 
> docker-compose, but I see a large volume of these logs on the affected node:
> {{2026-07-23 10:49:44.838 | 2026-07-23 09:49:44.838 INFO 
> (zkCallback-14-thread-5) [] o.a.s.c.Overseer Overseer 
> (id=216172940400852994-172.21.0.6:8983_solr-n_0000000000) closing }}
> {{2026-07-23 10:49:44.838 | 2026-07-23 09:49:44.838 INFO 
> (zkCallback-14-thread-5) [] o.a.s.c.Overseer Overseer 
> (id=216172940400852994-172.21.0.6:8983_solr-n_0000000000) closing }}
> {{2026-07-23 10:49:44.838 | 2026-07-23 09:49:44.838 INFO 
> (zkCallback-14-thread-5) [] o.a.s.c.Overseer Overseer 
> (id=216172940400852994-172.21.0.6:8983_solr-n_0000000000) closing }}
> {{2026-07-23 10:49:44.840 | 2026-07-23 09:49:44.839 INFO 
> (zkCallback-14-thread-5) [] o.a.s.c.Overseer Overseer 
> (id=216172940400852994-172.21.0.6:8983_solr-n_0000000000) closing }}
> {{2026-07-23 10:49:44.840 | 2026-07-23 09:49:44.840 INFO 
> (zkCallback-14-thread-5) [] o.a.s.c.Overseer Overseer 
> (id=216172940400852994-172.21.0.6:8983_solr-n_0000000000) closing }}
> {{2026-07-23 10:49:44.841 | 2026-07-23 09:49:44.841 INFO 
> (zkCallback-14-thread-5) [] o.a.s.c.OverseerElectionContext I am going to be 
> the leader 172.21.0.6:8983_solr }}
> {{2026-07-23 10:49:44.843 | 2026-07-23 09:49:44.843 ERROR 
> (zkCallback-14-thread-5) [] o.a.s.c.LeaderElector node exists 
> =>org.apache.zookeeper.KeeperException$NodeExistsException: KeeperErrorCode = 
> NodeExists for /overseer_elect/leader }}
> {{2026-07-23 10:49:44.843 | at 
> org.apache.zookeeper.KeeperException.create(KeeperException.java:125)}}



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to