[
https://issues.apache.org/jira/browse/SOLR-18505?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated SOLR-18505:
----------------------------------
Labels: pull-request-available (was: )
> DirectUpdateHandlerTest.testExpungeDeletes is flaky: a deletion-driven merge
> races the post-commit searcher sample
> ------------------------------------------------------------------------------------------------------------------
>
> Key: SOLR-18505
> URL: https://issues.apache.org/jira/browse/SOLR-18505
> Project: Solr
> Issue Type: Bug
> Reporter: Nick Shanin
> Priority: Minor
> Labels: pull-request-available
> Time Spent: 10m
> Remaining Estimate: 0h
>
> AI-generated, human-approved text below.
> DirectUpdateHandlerTest fails intermittently in CI (seen on a Solr Tests via
> Crave run for an unrelated PR). With the run's seed, 71E7F210A62A9B8C, it
> fails deterministically in local forced re-runs: 6 of 6 across current main
> and the parent of the SOLR-18317 merge, so it predates those changes and is
> not caused by them.
> Failure signature
> - testExpungeDeletes: AssertionError "maxDoc !> numDocs ... expected some
> deletions" at DirectUpdateHandlerTest.java:494 (one run showed the sibling
> variant "expected:<5> but was:<4>" at line 501).
> - testDeleteRollback: AssertionError in teardown (knock-on, see below).
> - classMethod: ObjectTracker reports unreleased objects (knock-on).
> Root cause
> testExpungeDeletes adds a duplicate document, leaving one segment 50 percent
> deleted. The test class pins TieredMergePolicy, and that policy treats any
> segment whose deleted percentage exceeds deletesPctAllowed (20 percent by
> default in Lucene 10.4) as eligible for a deletion-driven merge, even in a
> two-segment index where tier merging never fires. The merge runs on the
> ConcurrentMergeScheduler thread during the second commit. If it lands before
> Solr opens the post-commit searcher, the sample at line 494 reads the
> already-merged index and finds no deletions left to observe, so the assertion
> fails even though nothing is wrong with the update handler. In failing runs
> the surviving segment's diagnostics record source=merge and there is no
> live-docs file, confirming the deletions were expunged by the background
> merge before the sample.
> The seed pins everything except merge-thread scheduling, which is why the
> failure is near-deterministic on a quiet machine with this seed but flips on
> loaded CI runners.
> Knock-on failures
> The failed assertion skips closing the sample request, the leaked request
> pins the core's directory, and the next test's deleteCore() then fails in
> CachingDirectoryFactory.close; that is the whole of the testDeleteRollback
> teardown failure and the ObjectTracker report. testDeleteRollback's own body
> does not fail.
> Proposed fix
> In testExpungeDeletes only: set deletesPctAllowed to 100.0 on the live
> writer's TieredMergePolicy for the duration of the test, saving and restoring
> the previous value in a finally block. That removes the deletion-merge
> trigger while forceMergeDeletes still performs a real expunge, so the test
> keeps verifying what it verifies today. Also put the two sample requests in
> try-with-resources so a failed assertion can never leak a searcher again.
> (NoMergePolicy is not a substitute: its forced-deletes merge lookup returns
> null, which would turn the expunge into a no-op and gut the test.)
> A PR with this fix is prepared and will be linked here.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]