Marc Byrd created SOLR-18492:
--------------------------------
Summary: V2HttpCall never retries on stale cluster state - RETRY
action silently overwritten to ADMIN
Key: SOLR-18492
URL: https://issues.apache.org/jira/browse/SOLR-18492
Project: Solr
Issue Type: Bug
Reporter: Marc Byrd
Found as a side discovery while working SOLR-18487 (see that ticket and
{{apache/solr#4973}}) - a separate, pre-existing issue, but with the same
exposure pattern as that bug.
{{HttpSolrCall.extractRemotePath()}} sets {{action = RETRY}} (when it can't
resolve a remote core URL for a collection, e.g. due to stale local ZK state)
without returning:
{code:java}
cores.getZkController().zkStateReader.forceUpdateCollection(collectionName);
action = RETRY;
// falls through here, no return
{code}
{{V2HttpCall.init()}} only checks for {{action == REMOTEPROXY}} afterward - no
equivalent check for {{RETRY}}:
{code:java}
if (core == null) {
extractRemotePath(collectionName);
if (action == REMOTEPROXY) {
action = ADMIN_OR_REMOTEPROXY;
...
return;
}
// no check for RETRY here
}
...
if (core == null) {
initAdminRequest(path); // unconditionally sets action = ADMIN
return;
}
{code}
So when {{action == RETRY}}, execution falls through to the unconditional
{{core == null}} check below, overwriting {{action}} to {{ADMIN}}.
{{SolrServlet.dispatch()}}'s retry handling ({{case RETRY -> dispatch(request,
response, true)}}) can therefore never fire for a V2 request - a
stale-cluster-state V2 request gets silently misclassified as a plain admin
request instead of transparently retrying after the forced state refresh.
V1 already handles this correctly - {{HttpSolrCall.init()}}'s own (non-V2)
dispatch logic has {{if (action != null) return;}} immediately after its
equivalent {{REMOTEPROXY}} check, letting {{RETRY}} reach
{{SolrServlet.dispatch()}} properly. {{V2HttpCall}}'s override is simply
missing the equivalent line.
Age: confirmed present in {{releases/solr/10.0.0}} and current
{{branch_10_1}}/{{main}} - this predates 10.1 and is not a new regression.
Why this matters more in 10.1: the defect is old, but exposure to it has grown,
for the same reason as SOLR-18487/SOLR-18324. Before 10.1, an admin action
hitting this exact staleness condition from the Admin UI went through V1 (which
already handles {{RETRY}} correctly). SOLR-15752 made the Admin UI
V2-exclusive, so the same action now routes through the broken path
unconditionally. Separately, SOLR-18072 and related V2 work keep adding admin
operations with no V1 equivalent at all, so some operations have no fallback to
mask this gap even for users who haven't deliberately adopted the V2 UI. Same
shape as SOLR-18487: a latent gap in V2's cross-node dispatch plumbing, freshly
exposed by 10.1 removing the V1 safety net that was accidentally covering for
it.
Suggested fix: add {{if (action == RETRY) return;}} right after the existing
{{REMOTEPROXY}} check in {{V2HttpCall.init()}}, mirroring {{HttpSolrCall}}'s
own correct pattern.
Reproduction: confirmed only by direct code reading. Four separate attempts at
an automated repro (immediate cross-node query, async collection creation
exploiting the window before a replica is marked active, querying a node
immediately after it joins an already-populated cluster, and investigating
direct {{ZkStateReader}} cache manipulation) did not succeed without resorting
to reflection into private internals or hand-crafted raw ZK state, which seemed
too fragile to be worth it. This suggests the staleness window may be narrow in
practice even though the dead-code path itself is unambiguous.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]