dsmiley opened a new pull request, #4936: URL: https://github.com/apache/solr/pull/4936
# Tests Full local `./gradlew test` run: BUILD SUCCESSFUL, `:solr:core:test` 5824 tests / 0 failures (272 skipped), all other modules passed too. Targeted before landing: 16 highest-risk test classes covering every split-shard test (`SplitShardTest`, `SplitShardWithNodeRoleTest`, `SplitByPrefixTest`, `CollectionsAPISolrJTest`, `CollectionsAPIAsyncDistributedZkTest`, `TestCoordinatorRole`, `DeleteShardTest`, `InactiveShardRemoverTest`, `TestSolrCloudSnapshots`, `AbstractIncrementalBackupTest`, `TestCollectionsAPIViaSolrCloudCluster`, `OverseerTest`) plus backup/restore-adjacent tests (`SnapshotExportToolTest`, `LocalFSCloudIncrementalBackupTest`, `TestLocalFSCloudBackupRestore`, `IndexSizeEstimatorTest`, `PlacementPluginIntegrationTest`) — all passing before the full run confirmed it. I manually audited every call site that combines `waitForActiveCollection` with split/restore/slice-state-changing code (~12 files) to confirm none of them call it *after* a state-changing operation without an intervening stricter wait — each already re-waits via `activeClusterShape`/a custom predicate before proceeding, so none relied on the old permissive (replica-only) behavior. Also reproduced the originating flake's absence: `SnapshotExportToolTest.testExportUsesSnapshotStateNotLiveIndex`, which intermittently failed with `HTTP 510: Could not find a healthy node` under load (observed repeatedly on crave.io's CI, a 96-core remote runner that widens the race window — untracked by Jenkins/Develocity/fucit.org since that CI doesn't publish scans), passed 10/10 runs with varied seeds against this fix. # Description `MiniSolrCloudCluster.waitForActiveCollection()`'s two predicates (`expectedShardsAndActiveReplicas`, `expectedActive()`) only checked `Replica.isActive()`, never the owning `Slice`'s own state. A slice mid-restore (`CONSTRUCTION`) or a split's transient sub-shard state can already report its replica as active before the slice itself transitions to `ACTIVE`, so `waitForActiveCollection()` could return success while the collection was still unqueryable — `CloudSolrClient` only routes to active slices. This caused an intermittent failure in `SnapshotExportToolTest` (SOLR-18403): `RestoreCmd` returns before its shard leaves `CONSTRUCTION`, the Overseer's state update lands asynchronously, and a query issued right after `waitForActiveCollection()` returned could 510. # Solution - `expectedShardsAndActiveReplicas` now counts only `DocCollection.getActiveSlices()` and their replicas — matching the pattern `SolrCloudTestCase.activeClusterShape` already used elsewhere. - `expectedActive()` now rejects any slice outside `ACTIVE`/`INACTIVE` state. A split's inactive parent slice (and its replicas, which stay "active" indefinitely post-split) is intentionally still excluded from the count. - `activeClusterShape` is now a thin delegate to `expectedShardsAndActiveReplicas`, since the two implementations were identical duplicate logic. Test-framework-only change, no production code or user-facing behavior touched, so no changelog entry. *Developed with assistance from Claude (Anthropic); investigation, design tradeoffs, and verification reviewed by me.* -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
