Nick Shanin created SOLR-18505:
----------------------------------
Summary: DirectUpdateHandlerTest.testExpungeDeletes is flaky: a
deletion-driven merge races the post-commit searcher sample
Key: SOLR-18505
URL: https://issues.apache.org/jira/browse/SOLR-18505
Project: Solr
Issue Type: Bug
Reporter: Nick Shanin
AI-generated, human-approved text below.
DirectUpdateHandlerTest fails intermittently in CI (seen on a Solr Tests via
Crave run for an unrelated PR). With the run's seed, 71E7F210A62A9B8C, it fails
deterministically in local forced re-runs: 6 of 6 across current main and the
parent of the SOLR-18317 merge, so it predates those changes and is not caused
by them.
Failure signature
- testExpungeDeletes: AssertionError "maxDoc !> numDocs ... expected some
deletions" at DirectUpdateHandlerTest.java:494 (one run showed the sibling
variant "expected:<5> but was:<4>" at line 501).
- testDeleteRollback: AssertionError in teardown (knock-on, see below).
- classMethod: ObjectTracker reports unreleased objects (knock-on).
Root cause
testExpungeDeletes adds a duplicate document, leaving one segment 50 percent
deleted. The test class pins TieredMergePolicy, and that policy treats any
segment whose deleted percentage exceeds deletesPctAllowed (20 percent by
default in Lucene 10.4) as eligible for a deletion-driven merge, even in a
two-segment index where tier merging never fires. The merge runs on the
ConcurrentMergeScheduler thread during the second commit. If it lands before
Solr opens the post-commit searcher, the sample at line 494 reads the
already-merged index and finds no deletions left to observe, so the assertion
fails even though nothing is wrong with the update handler. In failing runs the
surviving segment's diagnostics record source=merge and there is no live-docs
file, confirming the deletions were expunged by the background merge before the
sample.
The seed pins everything except merge-thread scheduling, which is why the
failure is near-deterministic on a quiet machine with this seed but flips on
loaded CI runners.
Knock-on failures
The failed assertion skips closing the sample request, the leaked request pins
the core's directory, and the next test's deleteCore() then fails in
CachingDirectoryFactory.close; that is the whole of the testDeleteRollback
teardown failure and the ObjectTracker report. testDeleteRollback's own body
does not fail.
Proposed fix
In testExpungeDeletes only: set deletesPctAllowed to 100.0 on the live writer's
TieredMergePolicy for the duration of the test, saving and restoring the
previous value in a finally block. That removes the deletion-merge trigger
while forceMergeDeletes still performs a real expunge, so the test keeps
verifying what it verifies today. Also put the two sample requests in
try-with-resources so a failed assertion can never leak a searcher again.
(NoMergePolicy is not a substitute: its forced-deletes merge lookup returns
null, which would turn the expunge into a no-op and gut the test.)
A PR with this fix is prepared and will be linked here.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]