devmadhuu commented on code in PR #11223:
URL: https://github.com/apache/ozone/pull/11223#discussion_r4134719956
##########
hadoop-hdds/common/src/main/java/org/apache/hadoop/hdds/client/BlockID.java:
##########
@@ -103,6 +122,9 @@ public ContainerProtos.DatanodeBlockID
getDatanodeBlockIDProtobuf() {
if (replicaIndex != null) {
blockID.setReplicaIndex(replicaIndex);
}
+ if (storageType != null) {
+ blockID.setStorageTypeID(StorageTypeUtils.getID(storageType));
Review Comment:
@xichen01 Do we need to persist `storageTypeID` in every block record?
BlockData uses `getDatanodeBlockIDProtobuf() `for its RocksDB serialization, so
this field will be stored in every block entry. Since storage type is a
container-level property already persisted in ContainerData/container YAML, and
the request value is validated against it before PutBlock persists the block,
the per-block persisted copy appears redundant. Could we keep storageTypeID
only in the client-to-datanode request and omit it from the RocksDB BlockData
codec? If GetBlock needs to return it, it can be populated from the owning
container when building the response. This would avoid per-block metadata
overhead and the possibility of stale block metadata after container movement.
Based on math: For 1 billion blocks, it will consume roughly 2GB of extra
space.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]