Rajan Dhabalia created HDFS-17974:
-------------------------------------
Summary: HDFS DataNode Affinity for Tenant Isolation
Key: HDFS-17974
URL: https://issues.apache.org/jira/browse/HDFS-17974
Project: Hadoop HDFS
Issue Type: Improvement
Components: hdfs
Reporter: Rajan Dhabalia
h2. Summary
Introduce a pluggable, regex-based DataNode affinity mechanism that maps HDFS
paths to dedicated DataNode pools.
This enables *tenant/dataset-level storage and I/O isolation* without requiring
separate HDFS clusters or rack-based workarounds.
h2. Motivation
The current HDFS block placement policy has no native path-based mechanism to
restrict data to a specific DataNode pool. This can cause noisy-neighbor
interference for isolation-sensitive workloads.
The feature provides:
* Path-to-DataNode-pool mapping
* Tenant/dataset I/O isolation
* More predictable performance
* Independent capacity planning
* Runtime configuration updates without NameNode restart
h2. Design
Introduce a pluggable `DatanodeAffinityManager` configured through:
{code:java}
dfs.datanode.affinity.manager.classname
{code}
An affinity rule maps:
{code:java}
HDFS path regex -> DataNode hostname regex
{code}
Example:
{code:java}
/data/tenantA/.* -> dn-tenantA-.*
/data/tenantB/.* -> dn-tenantB-.*
{code}
The manager resolves the hostname regex against registered DataNodes and builds
a restricted `NetworkTopology` containing only eligible DataNodes for each
affinity group.
h2. NameNode Integration
*DatanodeManager*
* Identifies DataNodes belonging to affinity pools during registration.
* Removes affinity-only DataNodes from the default placement topology.
* Prevents non-affinity workloads from using dedicated DataNodes.
* Refreshes affinity state through `hdfs dfsadmin -refreshNodes`.
*BlockManager*
For each block placement request:
# Match the source path against configured affinity groups.
# Use the group's `BlockPlacementPolicy` if matched.
# Select targets from the group's restricted topology.
# Fall back to the default placement policy when no group matches.
This avoids large exclusion lists on the placement hot path.
h2. Pluggable Implementation
`DatanodeAffinityManager` is an abstraction that allows different affinity
sources without changing block-placement logic.
Built-in implementation:
{code:java}
FileDatanodeAffinityManager
{code}
It loads affinity rules from a JSON configuration file and reloads them through
`dfsadmin -refreshNodes`.
h2. Configuration
||Property||Default||Description||
|`dfs.datanode.affinity.manager.classname`|Empty|`DatanodeAffinityManager`
implementation. Empty disables the feature.|
|`dfs.datanode.affinity.file.path`|Empty|JSON affinity configuration used by
`FileDatanodeAffinityManager`.|
Example:
{code:java}
/data/tenantA/.* -> dn-tenantA-.*
/data/tenantB/.* -> dn-tenantB-.*
{code}
h2. Operational Visibility
Add:
{code:java}
hdfs fsck <path> -favored-nodes
{code}
to display the DataNodes resolved for a path.
Example:
{code:java}
hdfs fsck /data/tenantA -favored-nodes
{code}
This provides a dry-run mechanism to validate affinity configuration before
enabling it.
h2. Runtime Refresh
Affinity configuration and DataNode membership can be updated using:
{code:java}
hdfs dfsadmin -refreshNodes
{code}
No NameNode restart is required.
Supported changes include:
* Adding/removing DataNodes from an affinity pool
* Updating path-to-pool mappings
* Changing hostname matching rules
h2. Benefits
* *Tenant Isolation:* Affinity-enabled data is placed only on its dedicated
DataNode pool.
* *Predictable Performance:* Reduces noisy-neighbor impact and isolates I/O
capacity.
* *Default Pool Protection:* Dedicated DataNodes are excluded from normal
placement.
* *Efficient Placement:* Restricted topology limits placement to the relevant
pool.
* *Runtime Configuration:* Changes take effect without NameNode restart.
h2. Expected Impact
||Dimension||Expected Impact||
|Cross-tenant interference|Eliminated within the isolated DataNode pool|
|Tail latency|Reduced when contention exists|
|Placement scope|Limited to affinity pool|
|Default-pool throughput|No expected regression|
|Configuration changes|Runtime refresh|
The primary goal is {*}isolation and performance predictability{*}, not raw
throughput improvement.
h2. Implementation
Key changes:
* `DatanodeAffinityManager` abstraction
* `FileDatanodeAffinityManager`
* Path-regex to DataNode-hostname-regex mapping
* Per-group restricted `NetworkTopology`
* `DatanodeManager` topology integration
* Per-group `BlockPlacementPolicy` in `BlockManager`
* Affinity-aware `chooseTarget4NewBlock`
* Runtime refresh via `dfsadmin -refreshNodes`
* `hdfs fsck -favored-nodes` validation
h2. Backward Compatibility
* Disabled by default.
* Existing block placement behavior is unchanged when affinity is not
configured.
* Non-matching paths continue to use the default placement policy.
* No HDFS client changes are required.
* Can be enabled selectively for specific tenants or datasets.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]