dsmiley opened a new pull request, #4936:
URL: https://github.com/apache/solr/pull/4936

   # Tests
   
   Full local `./gradlew test` run: BUILD SUCCESSFUL, `:solr:core:test` 5824 
tests / 0 failures (272 skipped), all other modules passed too.
   
   Targeted before landing: 16 highest-risk test classes covering every 
split-shard test (`SplitShardTest`, `SplitShardWithNodeRoleTest`, 
`SplitByPrefixTest`, `CollectionsAPISolrJTest`, 
`CollectionsAPIAsyncDistributedZkTest`, `TestCoordinatorRole`, 
`DeleteShardTest`, `InactiveShardRemoverTest`, `TestSolrCloudSnapshots`, 
`AbstractIncrementalBackupTest`, `TestCollectionsAPIViaSolrCloudCluster`, 
`OverseerTest`) plus backup/restore-adjacent tests (`SnapshotExportToolTest`, 
`LocalFSCloudIncrementalBackupTest`, `TestLocalFSCloudBackupRestore`, 
`IndexSizeEstimatorTest`, `PlacementPluginIntegrationTest`) — all passing 
before the full run confirmed it.
   
   I manually audited every call site that combines `waitForActiveCollection` 
with split/restore/slice-state-changing code (~12 files) to confirm none of 
them call it *after* a state-changing operation without an intervening stricter 
wait — each already re-waits via `activeClusterShape`/a custom predicate before 
proceeding, so none relied on the old permissive (replica-only) behavior.
   
   Also reproduced the originating flake's absence: 
`SnapshotExportToolTest.testExportUsesSnapshotStateNotLiveIndex`, which 
intermittently failed with `HTTP 510: Could not find a healthy node` under load 
(observed repeatedly on crave.io's CI, a 96-core remote runner that widens the 
race window — untracked by Jenkins/Develocity/fucit.org since that CI doesn't 
publish scans), passed 10/10 runs with varied seeds against this fix.
   
   # Description
   
   `MiniSolrCloudCluster.waitForActiveCollection()`'s two predicates 
(`expectedShardsAndActiveReplicas`, `expectedActive()`) only checked 
`Replica.isActive()`, never the owning `Slice`'s own state. A slice mid-restore 
(`CONSTRUCTION`) or a split's transient sub-shard state can already report its 
replica as active before the slice itself transitions to `ACTIVE`, so 
`waitForActiveCollection()` could return success while the collection was still 
unqueryable — `CloudSolrClient` only routes to active slices.
   
   This caused an intermittent failure in `SnapshotExportToolTest` 
(SOLR-18403): `RestoreCmd` returns before its shard leaves `CONSTRUCTION`, the 
Overseer's state update lands asynchronously, and a query issued right after 
`waitForActiveCollection()` returned could 510.
   
   # Solution
   
   - `expectedShardsAndActiveReplicas` now counts only 
`DocCollection.getActiveSlices()` and their replicas — matching the pattern 
`SolrCloudTestCase.activeClusterShape` already used elsewhere.
   - `expectedActive()` now rejects any slice outside `ACTIVE`/`INACTIVE` 
state. A split's inactive parent slice (and its replicas, which stay "active" 
indefinitely post-split) is intentionally still excluded from the count.
   - `activeClusterShape` is now a thin delegate to 
`expectedShardsAndActiveReplicas`, since the two implementations were identical 
duplicate logic.
   
   Test-framework-only change, no production code or user-facing behavior 
touched, so no changelog entry.
   
   *Developed with assistance from Claude (Anthropic); investigation, design 
tradeoffs, and verification reviewed by me.*


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to