[ 
https://issues.apache.org/jira/browse/SOLR-18334?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

David Smiley updated SOLR-18334:
--------------------------------
    Priority: Minor  (was: Major)

> PRS collection creation returns too early; can cause restore failure
> --------------------------------------------------------------------
>
>                 Key: SOLR-18334
>                 URL: https://issues.apache.org/jira/browse/SOLR-18334
>             Project: Solr
>          Issue Type: Bug
>          Components: SolrCloud
>            Reporter: David Smiley
>            Assignee: David Smiley
>            Priority: Minor
>
> When a PRS (per-replica states) collection is created, 
> {{CreateCollectionCmd}} returns before the local {{ZkStateReader}} can see 
> the new collection. The PRS branch writes the collection's {{state.json}} 
> directly to ZooKeeper and then submits an asynchronous 
> {{{}RefreshCollectionMessage{}}}, but unlike the non-PRS branch it never 
> calls {{zkStateReader.waitForState(...)}} before returning. Any caller that 
> reads the collection from cluster state immediately after creating it can 
> therefore fail with "Could not find collection".
> {{RestoreCmd$RestoreOnANewCollection.process}} is such a caller — it even 
> documents the assumption:
> {code:java}
> createCoreLessCollection(...);
> // note: when createCollection() returns, the collection exists (no race)
> ...
> DocCollection restoreCollection =
>     
> rc.zkStateReader.getClusterState().getCollection(rc.restoreCollectionName);  
> // throws here
> {code}
> When the race is lost, the restore operation fails on the Overseer with:
> {code:java}
> o.a.s.c.a.c.CollectionHandlingUtils Collection 
> backuprestore_testskipconfigset-restore, operation restore failed
> org.apache.solr.common.SolrException: Could not find collection : 
> backuprestore_testskipconfigset-restore
>     at 
> org.apache.solr.common.cloud.ClusterState.getCollection(ClusterState.java:123)
>     at 
> org.apache.solr.cloud.api.collections.RestoreCmd$RestoreOnANewCollection.process(RestoreCmd.java:267)
> {code}
> The restored collection is left core-less, and the test subsequently fails 
> querying it:
> {code:java}
> AbstractIncrementalBackupTest > testSkipConfigset FAILED
> org.apache.solr.common.SolrException: Could not find a healthy node to handle 
> the request.
>     at 
> org.apache.solr.client.solrj.impl.CloudSolrClient.sendRequest(CloudSolrClient.java:1081)
>     ...
>     at 
> org.apache.solr.cloud.api.collections.AbstractIncrementalBackupTest.testSkipConfigset(AbstractIncrementalBackupTest.java:421)
> {code}
> The test masks the true failure: the async restore status from 
> {{processAndWait}} isn't checked and {{waitForRecoveriesToFinish}} is called 
> with {{{}failOnTimeout=false{}}}, so the failure only surfaces at the query.
> *CI evidence:* this is a long-standing low-rate flake (\{{~}}0.5–2% per 
> week), affecting all three subclasses of {{AbstractIncrementalBackupTest}} — 
> consistent with the test-framework randomization only sometimes enabling PRS. 
> Weekly failures/runs from fucit.org's Jenkins reports:
> ||Week||LocalFSCloudIncrementalBackupTest||GCSIncrementalBackupTest||S3IncrementalBackupTest||
> |2026-20|6/396|3/362|1/395|
> |2026-22|7/329|6/302|1/323|
> |2026-24|6/492|5/446| |
> |2026-26|6/455|3/420| |
> |2026-28..31|1|1| |
> *Fix:* in the PRS branch of {{{}CreateCollectionCmd{}}}, after submitting 
> {{{}RefreshCollectionMessage{}}}, wait for the collection to be visible in 
> the local {{ZkStateReader}} ({{{}waitForState(collectionName, 30s, 
> Objects::nonNull){}}}) — the same guarantee the non-PRS branch already 
> provides. This restores the invariant that {{RestoreCmd}} (and potentially 
> other callers) rely on.
> _(Written by Claude Fable 5.  Investigating low-rate test failures relating 
> to backups.)_



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to