[
https://issues.apache.org/jira/browse/HDDS-16609?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119995#comment-18119995
]
Sumit Agrawal commented on HDDS-16609:
--------------------------------------
[~aryangupta1998]
Another thing, we can delay FCR till last HB is success with ICR, this will
reduce the number of ICR in this scenario avoiding FCR till last HB is success
> Limit DN report queue and send small batches to avoid SCM re-registration
> failures
> ----------------------------------------------------------------------------------
>
> Key: HDDS-16609
> URL: https://issues.apache.org/jira/browse/HDDS-16609
> Project: Apache Ozone
> Issue Type: Improvement
> Reporter: Aryan Gupta
> Assignee: Aryan Gupta
> Priority: Major
>
> When an SCM is down or asks DN to re-register, the DN keeps collecting
> incremental reports (ICR) in memory.
> If this queue becomes too large, heartbeat/register payload can become too
> big and cross RPC message size limits.
> Then DN fails to register again, and the problem repeats.
> We should make DN smarter in this case.
> Proposed change:
> * Send queued incremental reports in multiple small batches (not one very
> large request).
> * Add a max payload cap per heartbeat request.
> * If queue becomes too old/too large, drop or compact old ICR entries.
> * Keep important command status handling safe (do not blindly drop critical
> status reports).
> * Add metrics/logs for queue size, dropped reports, and batch sends.
> Expected result:
> * DN can re-register reliably after SCM outage.
> * No oversized heartbeat/register RPC due to huge queued reports.
> * Better stability when one SCM node is slow/down.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]