Devesh Kumar Singh created HDDS-16652:
-----------------------------------------

             Summary: Avoid persisting redundant StorageType in datanode block 
metadata
                 Key: HDDS-16652
                 URL: https://issues.apache.org/jira/browse/HDDS-16652
             Project: Apache Ozone
          Issue Type: Sub-task
            Reporter: Devesh Kumar Singh
            Assignee: Devesh Kumar Singh


PR #11223 (HDDS-16388. Datanode putBlock Support StorageType) carries 
storageTypeID in ContainerProtos.DatanodeBlockID.

  The value is required on client-to-datanode requests to:

  - create a missing container on a volume of the requested StorageType; and
  - validate block operations against the existing container’s 
ContainerData.storageType before mutation.

  However, the same value is also persisted in every RocksDB BlockData record 
because the database codec and network path use the same serialization method:

  PutBlock request
    -> BlockData.getFromProtoBuf()
    -> Java BlockID.storageType
    -> DatanodeStore.putBlockByID(...)
    -> BlockData.CODEC
    -> BlockData.getProtoBufMessage()
    -> BlockID.getDatanodeBlockIDProtobuf()
    -> storageTypeID persisted in RocksDB

  Storage type is a container-level physical property. All local blocks in a 
container reside on the container’s selected volume, and the authoritative 
value is already persisted in ContainerData/
  container YAML and available from HddsVolume.

  Current production validation consumes the request value before the block is 
persisted. No production consumer has been identified that requires storage 
type to survive a per-block RocksDB round trip.
  The known readers are test assertions in 
TestContainerPersistence#testPutBlockWithStorageType and the planned patch 16 
TestOzoneStoragePolicy integration test.

  For current enum values, the protobuf field normally adds approximately two 
raw bytes per block record. At one billion blocks, that is roughly 2 GB of 
uncompressed protobuf payload, excluding RocksDB
  WAL, compaction and write-amplification effects. More importantly, it creates 
a duplicate source of truth that could become stale if a container replica is 
later imported or moved without rewriting
  every block record.

  The implementation should separate wire serialization from RocksDB 
serialization:

  Wire request serialization -> include storageTypeID
  RocksDB serialization      -> omit storageTypeID
  GetBlock response          -> omit it or derive it from ContainerData

  Existing records containing or omitting the optional field must remain 
readable without an eager database migration.

  Origin: PR #11223 review discussion 
(https://github.com/apache/ozone/pull/11223#pullrequestreview-5354128214).

  ### Acceptance criteria

  - New RocksDB BlockData and lastChunkInfoTable records do not persist 
storageTypeID.
  - Client-to-datanode requests continue carrying storage type for container 
creation and validation.
  - Mismatched valid and invalid storage types are rejected before any chunk or 
block mutation.
  - ContainerData.storageType and the selected HddsVolume remain authoritative.
  - Existing block records with or without the optional field remain readable 
across restart and mixed-version operation.
  - GetBlock behavior is documented; if it must return storage type, the value 
is derived from the owning container.
  - Tests cover DISK, SSD and ARCHIVE requests, mismatch rejection, incremental 
chunk lists, restart compatibility and the chosen GetBlock contract.
  - Patch 16 tests are updated so they do not require redundant per-block 
persistence.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to