[
https://issues.apache.org/jira/browse/HDDS-16297?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Andrey Yarovoy reassigned HDDS-16297:
-------------------------------------
Assignee: Andrey Yarovoy
> Bound per-directory purge size so a single dense directory doesn't hold the
> bucket write lock for the full per-task budget
> --------------------------------------------------------------------------------------------------------------------------
>
> Key: HDDS-16297
> URL: https://issues.apache.org/jira/browse/HDDS-16297
> Project: Apache Ozone
> Issue Type: Improvement
> Reporter: Andrey Yarovoy
> Assignee: Andrey Yarovoy
> Priority: Major
>
> *Problem*
> HDDS-16220 reduced OM apply-thread bucket-write-lock contention for the
> _multi-bucket_ directory-purge regime by grouping {{{}PurgePathRequest{}}}s
> per bucket ({{{}VolumeBucketId{}}}) so a single purge transaction no longer
> acquires several buckets' write locks together. It does not address the
> orthogonal _single dense directory_ case: one directory with a large fan-out
> of sub-files/sub-dirs.
> In {{{}DirectoryDeletingService.prepareDeleteDirRequest{}}}, a directory's
> sub-entries are gathered up to the remaining per-task budget
> ({{{}remainNum{}}}, seeded from {{{}ozone.path.deleting.limit.per.task{}}})
> and {{wrapPurgeRequest}} bundles them into {*}one indivisible
> {{PurgePathRequest}}{*}. On the apply side,
> {{OMDirectoriesPurgeRequestWithFSO.validateAndUpdateCache}} acquires the
> bucket write lock once ({{{}acquireWriteLocks(BUCKET_LOCK, …){}}}) and
> applies every prepared entry under that single hold. So a single dense
> directory can consume the whole per-task budget worth of entries under one
> uninterrupted write-lock acquisition — freezing readers on that hot bucket
> for the duration.
> *Why not split on the apply side*
> {{validateAndUpdateCache}} applies *one already-committed Ratis transaction*
> and produces *one {{OMClientResponse}}* flushed atomically by the double
> buffer at {{{}context.getIndex(){}}}. The apply path receives a committed log
> entry and cannot re-split it into multiple transactions. Cycling the bucket
> write lock between chunks _within_ a single apply would:
> * not reduce the single apply thread's total serial work (the thing that
> bounds commit throughput), and
> * expose a partially-tombstoned subtree under one transaction index,
> departing from how every other OM request handler holds its lock once and
> releases in {{{}finally{}}}.
> *Proposed approach (DDS-side)*
> Introduce a per-{{{}PurgePathRequest{}}} (per-directory) entry cap that is
> distinct from the existing per-{_}round{_}
> {{{}ozone.path.deleting.limit.per.task{}}}. When a directory's pending
> sub-entries exceed the cap, emit them across *multiple* purge transactions
> instead of one. The normal per-transaction apply loop then releases and
> re-acquires the bucket write lock between chunks for free, bounding each
> individual lock hold — without any bespoke intra-apply lock cycling.
> *Correctness surface to cover*
> * *Per-transaction atomicity of a directory purged in pieces* — a
> half-purged directory becomes observable between transactions. Note this
> already happens across DDS rounds today (a recursive delete is drained over
> many transactions), so the new behavior is a finer-grained instance of an
> existing property, not a new one — but it should be stated and tested.
> * *{{deletedTable}} move ordering / idempotency on apply-retry* — chunked
> emission must remain replay-safe if a transaction is re-applied.
> * *Snapshot chain ownership* —
> {{fromSnapshotInfo.setLastTransactionInfo(...)}} records the purge point on
> the snapshot; when a directory spans multiple transactions, define which
> transaction owns that update so the snapshot's {{lastTransactionInfo}} stays
> consistent.
> * *Progress/termination* — ensure the per-directory cap can't stall drain
> progress for a pathological single directory.
> *Value axis*
> This is a *read-availability* improvement on hot buckets (bounding the
> write-lock hold seen by readers), not a write-throughput one — the single
> apply thread still does the same total work. Same distinction called out on
> the write-lock-scope work; state it on the JIRA so the two aren't read as
> contradictory.
> *Testing*
> * DDS unit test: a directory with sub-entries above the cap emits N
> {{{}PurgePathRequest{}}}s across N transactions rather than one.
> * Apply-side re-validation / idempotency test for chunked emission.
> * Snapshot-deep-clean coverage (extend
> {{{}TestSnapshotDirectoryCleaningService{}}}) for a dense directory spanning
> transactions.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]