Marc Byrd created SOLR-18492:
--------------------------------

             Summary: V2HttpCall never retries on stale cluster state - RETRY 
action silently overwritten to ADMIN
                 Key: SOLR-18492
                 URL: https://issues.apache.org/jira/browse/SOLR-18492
             Project: Solr
          Issue Type: Bug
            Reporter: Marc Byrd



Found as a side discovery while working SOLR-18487 (see that ticket and 
{{apache/solr#4973}}) - a separate, pre-existing issue, but with the same 
exposure pattern as that bug.

{{HttpSolrCall.extractRemotePath()}} sets {{action = RETRY}} (when it can't 
resolve a remote core URL for a collection, e.g. due to stale local ZK state) 
without returning:

{code:java}
cores.getZkController().zkStateReader.forceUpdateCollection(collectionName);
action = RETRY;
// falls through here, no return
{code}

{{V2HttpCall.init()}} only checks for {{action == REMOTEPROXY}} afterward - no 
equivalent check for {{RETRY}}:

{code:java}
if (core == null) {
  extractRemotePath(collectionName);
  if (action == REMOTEPROXY) {
    action = ADMIN_OR_REMOTEPROXY;
    ...
    return;
  }
  // no check for RETRY here
}
...
if (core == null) {
  initAdminRequest(path);  // unconditionally sets action = ADMIN
  return;
}
{code}

So when {{action == RETRY}}, execution falls through to the unconditional 
{{core == null}} check below, overwriting {{action}} to {{ADMIN}}. 
{{SolrServlet.dispatch()}}'s retry handling ({{case RETRY -> dispatch(request, 
response, true)}}) can therefore never fire for a V2 request - a 
stale-cluster-state V2 request gets silently misclassified as a plain admin 
request instead of transparently retrying after the forced state refresh.

V1 already handles this correctly - {{HttpSolrCall.init()}}'s own (non-V2) 
dispatch logic has {{if (action != null) return;}} immediately after its 
equivalent {{REMOTEPROXY}} check, letting {{RETRY}} reach 
{{SolrServlet.dispatch()}} properly. {{V2HttpCall}}'s override is simply 
missing the equivalent line.

Age: confirmed present in {{releases/solr/10.0.0}} and current 
{{branch_10_1}}/{{main}} - this predates 10.1 and is not a new regression.

Why this matters more in 10.1: the defect is old, but exposure to it has grown, 
for the same reason as SOLR-18487/SOLR-18324. Before 10.1, an admin action 
hitting this exact staleness condition from the Admin UI went through V1 (which 
already handles {{RETRY}} correctly). SOLR-15752 made the Admin UI 
V2-exclusive, so the same action now routes through the broken path 
unconditionally. Separately, SOLR-18072 and related V2 work keep adding admin 
operations with no V1 equivalent at all, so some operations have no fallback to 
mask this gap even for users who haven't deliberately adopted the V2 UI. Same 
shape as SOLR-18487: a latent gap in V2's cross-node dispatch plumbing, freshly 
exposed by 10.1 removing the V1 safety net that was accidentally covering for 
it.

Suggested fix: add {{if (action == RETRY) return;}} right after the existing 
{{REMOTEPROXY}} check in {{V2HttpCall.init()}}, mirroring {{HttpSolrCall}}'s 
own correct pattern.

Reproduction: confirmed only by direct code reading. Four separate attempts at 
an automated repro (immediate cross-node query, async collection creation 
exploiting the window before a replica is marked active, querying a node 
immediately after it joins an already-populated cluster, and investigating 
direct {{ZkStateReader}} cache manipulation) did not succeed without resorting 
to reflection into private internals or hand-crafted raw ZK state, which seemed 
too fragile to be worth it. This suggests the staleness window may be narrow in 
practice even though the dead-code path itself is unambiguous.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to