symious commented on code in PR #10822:
URL: https://github.com/apache/ozone/pull/10822#discussion_r3764507111


##########
hadoop-hdds/docs/content/design/s3-versioning.md:
##########
@@ -0,0 +1,433 @@
+---
+title: S3-compatible Object Versioning
+summary: Bucket-level, S3-compatible object versioning with O(1) version 
writes and built-in reclamation
+date: 2026-07-21
+jira: HDDS-15728
+status: accepted
+author: Symious
+---
+<!--
+  Licensed under the Apache License, Version 2.0 (the "License");
+  you may not use this file except in compliance with the License.
+  You may obtain a copy of the License at
+   http://www.apache.org/licenses/LICENSE-2.0
+  Unless required by applicable law or agreed to in writing, software
+  distributed under the License is distributed on an "AS IS" BASIS,
+  WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
+  See the License for the specific language governing permissions and
+  limitations under the License. See accompanying LICENSE file.
+-->
+
+# Summary
+
+Add S3-compatible object versioning to Ozone: the full three-state bucket state
+machine (Unversioned / Enabled / Suspended), per-key version chains with delete
+markers and null versions, the S3 versioning APIs on the S3 Gateway, and 
built-in
+version reclamation (`maxVersions`, background expiration) — with O(1) metadata
+cost per version operation and zero regression on non-versioned paths.
+
+# Status
+
+Defined in the markdown header.
+
+# Problem statement (Motivation / Abstract)
+
+Amazon S3 provides bucket-level object versioning: a single key can retain 
multiple
+versions, so users can recover objects that were accidentally overwritten or
+deleted. A large part of the S3 ecosystem (backup software, data lake 
components,
+DR tooling) depends on the versioning APIs (`PutBucketVersioning`,
+`ListObjectVersions`, object operations with a `versionId`). Ozone exposes an
+S3-compatible API through the S3 Gateway but does not support object versioning
+today: the bucket-level `isVersionEnabled` boolean cannot express the Suspended
+state, `OmKeyInfo.keyLocationVersions` tracks block locations within one record
+rather than object versions, and the gateway has no versioning endpoints.
+
+This proposal implements versioning with S3-compatible semantics, usable by
+standard S3 clients (AWS CLI / SDKs) without modification. The metadata cost 
of a
+version operation is decoupled from the number of versions (one extra small KV
+write per operation), and reclamation controls are built into the feature 
itself
+to avoid the unbounded version accumulation problems commonly seen on S3 (the 
S3
+troubleshooting guide documents list degradation and throttling on keys with
+millions of versions, and leaves the fix to user-configured Lifecycle rules 
that
+are often forgotten).
+
+# Non-goals
+
+- **MFA delete** — depends on the AWS IAM/MFA device ecosystem; Ozone has no
+  counterpart infrastructure. The `MfaDelete` field of `PutBucketVersioning`
+  returns NotImplemented.
+- **A full S3 Lifecycle rule engine** — only the minimal reclamation 
capabilities
+  that versioning itself requires are included.
+- **Version-aware cross-cluster replication.**
+- **Versioning for FSO / LEGACY bucket layouts** — the first version supports
+  OBJECT_STORE buckets only; enabling versioning on other layouts returns
+  NotImplemented. Combining FSO's directory/rename semantics with per-key 
version
+  chains is disproportionately complex, and S3 tooling scenarios essentially 
use
+  the OBS layout. FSO support can be evaluated as an independent follow-up.
+- **Coexistence with Ozone snapshots on the same bucket** — the version-aware
+  reclamation the combination needs is deferred past the first version, so OM
+  rejects the combination outright rather than leaving it to convention. The
+  snapshot section below states the enforcement and what lifts it.
+
+# Technical Description (Architecture and implementation details)
+
+## Bucket state machine
+
+```
+Unversioned (default) ──enable──▶ Enabled ◀──enable── Suspended
+                                     │                     ▲
+                                     └──────suspend────────┘
+        (Once Enabled, a bucket can never return to Unversioned)
+```
+
+A `BucketVersioningStatusProto` enum (`UNVERSIONED` / `VERSIONING_ENABLED` /
+`VERSIONING_SUSPENDED`) is added as an optional field on `BucketInfo` and
+`BucketArgs`. The legacy `isVersionEnabled` boolean is kept and maintained in
+two-way sync (`ENABLED → true`, otherwise `false`; records without the enum are
+interpreted via the boolean), so old and new clients/OMs coexist during rolling
+upgrades. OM enforces the state machine on `SetBucketProperty`: transitions 
back
+to `UNVERSIONED` are rejected with `INVALID_REQUEST`, preserving S3's
+data-protection promise that no single state change can silently destroy
+historical versions.
+
+## Metadata layout: keyTable (current) + versionedKeyTable (noncurrent)
+
+A new column family, **versionedKeyTable**, splits responsibilities with 
keyTable:
+
+- **keyTable** (existing, semantics unchanged) always holds each key's 
**current
+  version** — a regular object or a delete marker. Plain GET / HEAD / 
ListObjects
+  read paths are unchanged.
+- **versionedKeyTable** (new) holds all **noncurrent** versions (including
+  noncurrent delete markers), each as a complete `OmKeyInfo`. The RocksDB key 
is
+
+  ```
+  /{volume}/{bucket}/{keyName}\x00{Long.MAX_VALUE - versionId}
+  ```
+
+  (fixed-width hex suffix; the separator is `0x00` rather than `/` because OBS 
key
+  names contain `/` verbatim, which would interleave a key's versions with 
those of
+  keys nested under it), so all versions of a key are physically adjacent and
+  ordered newest to oldest: `ListObjectVersions` and version promotion are a
+  single seek plus a sequential read. The table is registered in
+  `OmMetadataManagerImpl.getTableBucketPrefix` alongside the other key tables, 
so
+  bucket-prefixed iteration and SST filtering resolve its prefix.
+
+`OmKeyInfo` gains three optional proto fields (old records deserialize
+compatibly): `versionId` (int64, assigned once at version creation, then 
frozen),
+`isDeleteMarker` (a marker is a record with this flag and no data blocks — no
+datanode storage), and `isNullVersion` (the single overwritable "null version"
+slot per key). Keys written before versioning was enabled are interpreted as 
null
+versions on read — **zero migration**, matching S3's "enabling versioning does
+not change existing objects".
+
+Every keyTable ↔ versionedKeyTable update rides OM's existing atomic
+multi-table WriteBatch commit (the same double-buffer pattern used today for
+keyTable + deletedTable on overwrite): no new transaction mechanism and no
+cross-table consistency problem. The new column family is introduced under the 
OM
+layout feature / finalization framework (`OMLayoutFeature.OBJECT_VERSIONING`):
+before finalization, requests carrying a versioning status are rejected.
+
+## VersionId: a pluggable generator
+
+`versionId` generation is abstracted behind a `VersionIdGenerator` interface,
+chosen per cluster by class name (`ozone.om.versioning.version-id-generator`), 
so
+a deployment can plug in its own. The generator is cluster-wide and may be 
changed
+on a running cluster; it is not recorded in bucket metadata. Every generator 
must
+satisfy, **for itself**: strictly increasing within a key (a later version's 
id is
+always greater than every earlier version's id of that key); frozen once 
assigned;
+`0` and `1` never handed out (`1` is the first-version sentinel).
+
+That guarantee binds one generator, not a sequence of them, so the write path
+enforces it at commit: a commit whose id does not come after the key's current
+version is rejected (`INVALID_REQUEST`), and the operator deletes the key's
+versions before writing under the new generator. The check costs no read in the
+steady state — the current version holds the key's largest id, so an id above 
it
+cannot be taken; only a record predating versioning, which carries no id to 
order
+against, falls back to a versionedKeyTable lookup.
+
+- **`TransactionIndexVersionIdGenerator` (default)** — the OM Ratis transaction
+  index of the committing transaction, used directly. No allocator state of any
+  kind. Externally encoded as an opaque URL-safe string. objectID's epoch bits 
are
+  deliberately not applied: dropping them keeps versionIds small positive 
longs, so
+  the versionedKeyTable ordering stays plain signed arithmetic, and uniqueness
+  rests on the monotonicity of the Ratis log index — a log that was reset or 
rolled
+  back is caught by the commit-time check rather than silently accepted.
+- **`PinnedFirstVersionIdGenerator` (opt-in)** — only the first version is
+  special: it takes the reserved sentinel `FIRST_VERSION_ID = 1`, below any 
usable
+  transaction index, so it sorts oldest and can be referenced without listing 
the
+  key's versions first. All later versions are the commit transaction's index; 
no
+  persistent allocator. First versions are detected by "no current version in
+  keyTable", reusing the lookup the write path performs anyway. Known 
trade-off: if
+  every version of a key is permanently deleted and the key is recreated, the 
new
+  first version takes the sentinel again; only deployments that accept this 
should
+  configure this generator.
+
+How a versionId is rendered on the wire — the opaque encoding, and whether the
+pinned first version is presented as a fixed literal or derived from the 
keyName —
+is part of the `?versionId=` read path (T4), not of generation.
+
+The null version is not a special ID value but the `isNullVersion` attribute — 
it
+carries a normally generated id like any other version, since a null created
+between two versioned writes is not the oldest one;
+`versionId=null` requests resolve to "locate this key's null slot".
+
+## Request handling
+
+- **PUT (Enabled)**, atomic in one WriteBatch: move the current version (if 
any)
+  into versionedKeyTable; write the new record into keyTable as current.
+- **PUT (Suspended)**: the new record takes the null slot and becomes the 
current
+  version — if the null slot is already the current version, overwrite it in 
place;
+  otherwise move the current version into versionedKeyTable, write the new 
record
+  into keyTable as current, and delete the old null record (if any) from
+  versionedKeyTable, blocks to deletedTable in both cases. Versions accumulated
+  while Enabled are unaffected. **PUT (Unversioned)**: unchanged from today.
+- **GET/HEAD**: without versionId, read keyTable; a current delete marker 
returns
+  404 with `x-amz-delete-marker: true`. With versionId, check the current 
version
+  first, then point-look-up versionedKeyTable; if the addressed version is a 
delete
+  marker, current or not, the gateway returns 405 with `x-amz-delete-marker: 
true`
+  and `Allow: DELETE`.
+- **DELETE without versionId (Enabled)**: move current into versionedKeyTable 
and
+  write a delete marker as the new current. (Suspended: the marker takes the 
null
+  slot exactly as a suspended PUT does.) **DELETE ?versionId=x**: permanently
+  delete that version (blocks to deletedTable); if it was current, trigger 
version
+  promotion.
+- **Version promotion** — the invariant is that keyTable always holds a key's
+  current version. When a permanent delete removes the current version, one
+  `seek` on the key's versionedKeyTable prefix yields the newest noncurrent
+  version (reverse ordering), which is moved back into keyTable unchanged — a
+  pure positional move; the record's content stays frozen. If no noncurrent
+  version remains, the key disappears entirely. Executed in the same WriteBatch
+  as the delete, under the bucket write lock. Deleting a delete marker this way
+  is exactly the S3 "restore an object" flow.
+- **ListObjects**: scans keyTable only, skipping keys whose current version is 
a
+  marker — list scan volume is not amplified by versions (an advantage over 
S3's
+  single-namespace implementation). **ListObjectVersions**: merges keyTable and
+  versionedKeyTable in key order (both are key-name prefixed, so the merge is
+  naturally ordered), with `key-marker` / `version-id-marker` pagination.
+
+## Reclamation, quota, observability
+
+Built-in bucket-level controls: `maxVersions` (default 100, 0 = unlimited; 
markers
+count toward the limit), `noncurrentVersionExpiration` (opt-in), and
+expired-delete-marker cleanup (enabled by default: when only a marker remains,
+the whole key is removed) — all enforced by a new **VersionCleanupService**
+following the `KeyDeletingService` pattern. Versions beyond `maxVersions` are

Review Comment:
   Or we can decoupling "lifecycle" and "XxxTableScanService", but seems need 
many changes for this option.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to