[
https://issues.apache.org/jira/browse/SOLR-18505?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18123530#comment-18123530
]
Nick Shanin commented on SOLR-18505:
------------------------------------
AI-generated, human-approved text below.
Root cause found and a fix is up as PR #5028:
[https://github.com/apache/solr/pull/5028]
The duplicate add in testExpungeDeletes leaves a segment 50 percent deleted.
The TieredMergePolicy the class pins admits any segment over deletesPctAllowed
(20 percent by default) to a deletion-driven merge, and that merge runs on the
ConcurrentMergeScheduler thread during the second commit. When it lands before
the post-commit searcher opens, the sample at line 494 reads the already-merged
index and finds no deletions. In failing runs the surviving segment's
diagnostics show source=merge with no live-docs file. The testDeleteRollback
teardown failure and the ObjectTracker report are knock-ons: the failed
assertion skipped closing the sample request, which pinned the directory for
the next test's core delete.
The PR wraps the core's live merge policy for the duration of the test in a
FilterMergePolicy that returns null from findMerges while still delegating
findForcedDeletesMerges, so the expunge the test verifies stays a genuine
forced merge, and puts both sample requests in try-with-resources. Proof at the
PR head: five forced re-runs with the seed 71E7F210A62A9B8C pass 7 of 7, plus
two random-seed runs. Noted honestly in the PR: the race is timing dependent
and stopped reproducing on the current base later the same day (it failed 6 of
6 earlier), so the fix rests on removing the mechanism by construction rather
than on a same-day before/after pair.
> DirectUpdateHandlerTest.testExpungeDeletes is flaky: a deletion-driven merge
> races the post-commit searcher sample
> ------------------------------------------------------------------------------------------------------------------
>
> Key: SOLR-18505
> URL: https://issues.apache.org/jira/browse/SOLR-18505
> Project: Solr
> Issue Type: Bug
> Reporter: Nick Shanin
> Priority: Minor
> Labels: pull-request-available
> Time Spent: 10m
> Remaining Estimate: 0h
>
> AI-generated, human-approved text below.
> DirectUpdateHandlerTest fails intermittently in CI (seen on a Solr Tests via
> Crave run for an unrelated PR). With the run's seed, 71E7F210A62A9B8C, it
> fails deterministically in local forced re-runs: 6 of 6 across current main
> and the parent of the SOLR-18317 merge, so it predates those changes and is
> not caused by them.
> Failure signature
> - testExpungeDeletes: AssertionError "maxDoc !> numDocs ... expected some
> deletions" at DirectUpdateHandlerTest.java:494 (one run showed the sibling
> variant "expected:<5> but was:<4>" at line 501).
> - testDeleteRollback: AssertionError in teardown (knock-on, see below).
> - classMethod: ObjectTracker reports unreleased objects (knock-on).
> Root cause
> testExpungeDeletes adds a duplicate document, leaving one segment 50 percent
> deleted. The test class pins TieredMergePolicy, and that policy treats any
> segment whose deleted percentage exceeds deletesPctAllowed (20 percent by
> default in Lucene 10.4) as eligible for a deletion-driven merge, even in a
> two-segment index where tier merging never fires. The merge runs on the
> ConcurrentMergeScheduler thread during the second commit. If it lands before
> Solr opens the post-commit searcher, the sample at line 494 reads the
> already-merged index and finds no deletions left to observe, so the assertion
> fails even though nothing is wrong with the update handler. In failing runs
> the surviving segment's diagnostics record source=merge and there is no
> live-docs file, confirming the deletions were expunged by the background
> merge before the sample.
> The seed pins everything except merge-thread scheduling, which is why the
> failure is near-deterministic on a quiet machine with this seed but flips on
> loaded CI runners.
> Knock-on failures
> The failed assertion skips closing the sample request, the leaked request
> pins the core's directory, and the next test's deleteCore() then fails in
> CachingDirectoryFactory.close; that is the whole of the testDeleteRollback
> teardown failure and the ObjectTracker report. testDeleteRollback's own body
> does not fail.
> Proposed fix
> In testExpungeDeletes only: set deletesPctAllowed to 100.0 on the live
> writer's TieredMergePolicy for the duration of the test, saving and restoring
> the previous value in a finally block. That removes the deletion-merge
> trigger while forceMergeDeletes still performs a real expunge, so the test
> keeps verifying what it verifies today. Also put the two sample requests in
> try-with-resources so a failed assertion can never leak a searcher again.
> (NoMergePolicy is not a substitute: its forced-deletes merge lookup returns
> null, which would turn the expunge into a no-op and gut the test.)
> A PR with this fix is prepared and will be linked here.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]