[ 
https://issues.apache.org/jira/browse/SOLR-18497?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

ASF GitHub Bot updated SOLR-18497:
----------------------------------
    Labels: pull-request-available  (was: )

> Concurrent RELOAD of a core that failed to load creates the core once per 
> request
> ---------------------------------------------------------------------------------
>
>                 Key: SOLR-18497
>                 URL: https://issues.apache.org/jira/browse/SOLR-18497
>             Project: Solr
>          Issue Type: Bug
>            Reporter: Serhiy Bzhezytskyy
>            Priority: Major
>              Labels: pull-request-available
>          Time Spent: 10m
>  Remaining Estimate: 0h
>
> A core that failed to load stays in {{coreInitFailures}}, and 
> {{CoreContainer.reload}} then takes the branch that creates it from the 
> stored descriptor. That branch waits for the per-core reservation and calls 
> {{createFromDescriptor}} without checking again whether the core was loaded 
> in the meantime. When several RELOAD requests arrive together after the cause 
> of the failure was fixed, every one of them creates the core: the first 
> succeeds, the others fail on the index lock with HTTP 500, and each failure 
> is recorded as a new init failure. The core then serves requests while 
> CoreAdmin STATUS still lists it under {{initFailures}}.
> To reproduce (user-managed, {{<home>/configsets/_default}} copied from the 
> distribution):
> {code:bash}
> bin/solr start --user-managed --solr-home <home>
> curl '.../solr/admin/cores?action=CREATE&name=f1&configSet=_default'
> bin/solr stop
> mv <home>/configsets/_default <home>/configsets/_hidden      # f1 now fails 
> to load at startup
> bin/solr start --user-managed --solr-home <home>
> mv <home>/configsets/_hidden <home>/configsets/_default      # the cause is 
> fixed
> for i in 1 2 3 4 5 6 7 8; do curl 
> '.../solr/admin/cores?action=RELOAD&core=f1' & done; wait
> {code}
> One request returns 200 and seven return 500, "Unable to create core [f1]", 
> caused by
> {noformat}
> LockObtainFailedException: Index dir '.../f1/data/index/' of core 'f1' is 
> already locked.
> {noformat}
> The core answers queries afterwards, but {{STATUS}} still shows f1 under 
> {{initFailures}}. A single RELOAD creates the core once and leaves 
> {{initFailures}} empty.
> The reload should not create the core again when another request loaded it 
> while this one waited for the reservation.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to