[
https://issues.apache.org/jira/browse/HDDS-16437?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Siyao Meng updated HDDS-16437:
------------------------------
Description:
h3. Finding
optimizeDirDeletesAndSubmitRequest discards the batching result and increments
deletion counters and operator metrics (numDirsSentForPurge, INFO log, AOS/snap
last-run metrics) for the full pre-batch list even when a later batch's Ratis
submit returns non-success, so only committed batches are durable while metrics
over-report; the drift is cumulative and never reconciled. The deletion itself
self-corrects via retry, no data corruption.
h3. Classification
* Verdict: REPRODUCED
* Severity: Medium
* Source: Specula TLA+ model checking and confirmation debate, finding CR-7
h3. Reproduce
{noformat}
Ozone commit: 9fbf9ee0cb1bd2f5f5d437b6719ebbe5309351fb
Specula: v1.1.0 (commit c6aa3dfa)
Target: om-fso-recursive-deletion
Guidance:
campaigns/ozone-9fbf9ee/targets/005-om-fso-recursive-deletion/.prompt-extra.md
{noformat}
{code:none}
specula run --agent=claude-code --effort=high --keep-original --max-parallel=2 \
--enable-reviews --confirm-debate --tlc-memory-limit=28G --tlc-worker-limit=8
\
"om-fso-recursive-deletion|apache/ozone|Java|Use the target-specific
.prompt-extra.md"
{code}
Discovered under HDDS-16429 (Specula TLA+ verification effort). The TLA+
specification, counterexample, and confirmation debate live in the Specula run
artifacts.
Generated with Specula (Claude Opus 4.8).
h3. Steps to reproduce
# The DirectoryDeletingService background task builds a purge-path list large
enough to span more than one Ratis batch.
# optimizeDirDeletesAndSubmitRequest submits the list via
submitPurgePathsWithBatching, which flushes the requests in Ratis-byte-limited
batches, in order.
# One batch's submit response is not success (Ratis rejects it, leader change,
or state-machine error). submitPurgePathsWithBatching stops and abandons the
remaining batches; earlier batches are already committed durably.
# The caller ignores the return value and increments deletedDirsCount /
movedDirsCount / movedFilesCount and the DirectoryDeletion metrics for the full
pre-batch list.
# The counters and OM metrics over-report the number of purged
directories/files versus what was durably committed. The un-purged directories
remain in the deleted-directory table and are reprocessed and counted again on
the next run.
Reproduced in a unit test that spies the service, commits the first batch and
fails the rest, and asserts deletedDirsCount reflects only the committed batch.
Without the fix the counter increments by the full list (20) while only the
first batch (6) committed.
h3. Root cause
optimizeDirDeletesAndSubmitRequest tallies dirNum, subDirNum, subFileNum and
subdirDelNum across the entire purge-path list, then increments the progress
counters and metrics from those tallies unconditionally.
submitPurgePathsWithBatching returns after the first failed batch (abandoning
the rest), but its return value is discarded, so the counters are advanced for
directories that were never purged. The counters are monotonic and nothing
reconciles them, and the abandoned directories are re-counted on the next run.
h3. Suggested fix
Have submitPurgePathsWithBatching return the sublist of requests that were
durably committed (the committed prefix; on a failed or null submit it returns
only what committed so far), and have optimizeDirDeletesAndSubmitRequest
recompute dirNum, subDirNum, subFileNum and subdirDelNum from that committed
list before updating the counters and metrics. This makes the counters reflect
exactly the purges that were applied, on both full and partial commits, and
avoids double-counting on retry. See [^HDDS-16437.001.patch]. The patch adds
testPurgeDirectoriesPartialCommitCountsOnlyCommitted, which fails without the
fix (counts the full list) and passes with it.
Patch is an AI-drafted proposal (Specula + Claude); pending human build, full
test, and review before merge.
was:
h3. Finding
optimizeDirDeletesAndSubmitRequest discards the batching result and increments
deletion counters and operator metrics (numDirsSentForPurge, INFO log, AOS/snap
last-run metrics) for the full pre-batch list even when a later batch's Ratis
submit returns non-success, so only committed batches are durable while metrics
over-report; the drift is cumulative and never reconciled. The deletion itself
self-corrects via retry, no data corruption.
h3. Classification
* Verdict: REPRODUCED
* Severity: Medium
* Source: Specula TLA+ model checking and confirmation debate, finding CR-7
h3. Reproduce
{noformat}
Ozone commit: 9fbf9ee0cb1bd2f5f5d437b6719ebbe5309351fb
Specula: v1.1.0 (commit c6aa3dfa)
Target: om-fso-recursive-deletion
Guidance:
campaigns/ozone-9fbf9ee/targets/005-om-fso-recursive-deletion/.prompt-extra.md
{noformat}
{code:none}
specula run --agent=claude-code --effort=high --keep-original --max-parallel=2 \
--enable-reviews --confirm-debate --tlc-memory-limit=28G --tlc-worker-limit=8
\
"om-fso-recursive-deletion|apache/ozone|Java|Use the target-specific
.prompt-extra.md"
{code}
Discovered under HDDS-16429 (Specula TLA+ verification effort). The TLA+
specification, counterexample, and confirmation debate live in the Specula run
artifacts.
Generated with Specula (Claude Opus 4.8).
> FSO directory deletion over reports purge metrics when a batch Ratis submit
> fails
> ---------------------------------------------------------------------------------
>
> Key: HDDS-16437
> URL: https://issues.apache.org/jira/browse/HDDS-16437
> Project: Apache Ozone
> Issue Type: Bug
> Reporter: Siyao Meng
> Priority: Minor
> Attachments: HDDS-16437.001.patch
>
>
> h3. Finding
> optimizeDirDeletesAndSubmitRequest discards the batching result and
> increments deletion counters and operator metrics (numDirsSentForPurge, INFO
> log, AOS/snap last-run metrics) for the full pre-batch list even when a later
> batch's Ratis submit returns non-success, so only committed batches are
> durable while metrics over-report; the drift is cumulative and never
> reconciled. The deletion itself self-corrects via retry, no data corruption.
> h3. Classification
> * Verdict: REPRODUCED
> * Severity: Medium
> * Source: Specula TLA+ model checking and confirmation debate, finding CR-7
> h3. Reproduce
> {noformat}
> Ozone commit: 9fbf9ee0cb1bd2f5f5d437b6719ebbe5309351fb
> Specula: v1.1.0 (commit c6aa3dfa)
> Target: om-fso-recursive-deletion
> Guidance:
> campaigns/ozone-9fbf9ee/targets/005-om-fso-recursive-deletion/.prompt-extra.md
> {noformat}
> {code:none}
> specula run --agent=claude-code --effort=high --keep-original
> --max-parallel=2 \
> --enable-reviews --confirm-debate --tlc-memory-limit=28G
> --tlc-worker-limit=8 \
> "om-fso-recursive-deletion|apache/ozone|Java|Use the target-specific
> .prompt-extra.md"
> {code}
> Discovered under HDDS-16429 (Specula TLA+ verification effort). The TLA+
> specification, counterexample, and confirmation debate live in the Specula
> run artifacts.
> Generated with Specula (Claude Opus 4.8).
> h3. Steps to reproduce
> # The DirectoryDeletingService background task builds a purge-path list large
> enough to span more than one Ratis batch.
> # optimizeDirDeletesAndSubmitRequest submits the list via
> submitPurgePathsWithBatching, which flushes the requests in
> Ratis-byte-limited batches, in order.
> # One batch's submit response is not success (Ratis rejects it, leader
> change, or state-machine error). submitPurgePathsWithBatching stops and
> abandons the remaining batches; earlier batches are already committed durably.
> # The caller ignores the return value and increments deletedDirsCount /
> movedDirsCount / movedFilesCount and the DirectoryDeletion metrics for the
> full pre-batch list.
> # The counters and OM metrics over-report the number of purged
> directories/files versus what was durably committed. The un-purged
> directories remain in the deleted-directory table and are reprocessed and
> counted again on the next run.
> Reproduced in a unit test that spies the service, commits the first batch and
> fails the rest, and asserts deletedDirsCount reflects only the committed
> batch. Without the fix the counter increments by the full list (20) while
> only the first batch (6) committed.
> h3. Root cause
> optimizeDirDeletesAndSubmitRequest tallies dirNum, subDirNum, subFileNum and
> subdirDelNum across the entire purge-path list, then increments the progress
> counters and metrics from those tallies unconditionally.
> submitPurgePathsWithBatching returns after the first failed batch (abandoning
> the rest), but its return value is discarded, so the counters are advanced
> for directories that were never purged. The counters are monotonic and
> nothing reconciles them, and the abandoned directories are re-counted on the
> next run.
> h3. Suggested fix
> Have submitPurgePathsWithBatching return the sublist of requests that were
> durably committed (the committed prefix; on a failed or null submit it
> returns only what committed so far), and have
> optimizeDirDeletesAndSubmitRequest recompute dirNum, subDirNum, subFileNum
> and subdirDelNum from that committed list before updating the counters and
> metrics. This makes the counters reflect exactly the purges that were
> applied, on both full and partial commits, and avoids double-counting on
> retry. See [^HDDS-16437.001.patch]. The patch adds
> testPurgeDirectoriesPartialCommitCountsOnlyCommitted, which fails without the
> fix (counts the full list) and passes with it.
> Patch is an AI-drafted proposal (Specula + Claude); pending human build, full
> test, and review before merge.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]