jojochuang commented on code in PR #550:
URL: https://github.com/apache/ozone-site/pull/550#discussion_r4067026195


##########
docs/06-troubleshooting/19-datanode-out-of-space.md:
##########
@@ -0,0 +1,40 @@
+---
+sidebar_label: Datanode space
+---
+
+# Datanode running out of space
+
+Insufficient storage space can cause writes to slow down or fail. When too few 
Datanodes have enough space for the required replication, SCM cannot allocate 
new containers or pipelines and clients may retry their writes.
+
+## Check storage usage
+
+1. Open the SCM web UI and examine the Datanode table. Sort by **Used Space 
Percent** to find nodes approaching capacity.
+2. Follow a Datanode's hostname link to its web UI. In **Volume Information**, 
compare usage across individual disks. A node's overall utilization can hide a 
full disk.
+3. Compare **Ozone Available**, **Filesystem Available**, **Reserved**, and 
**Non-Ozone Used**. Other files on the same filesystem can reduce the space 
available to Ozone.
+4. Check the Datanode logs for storage errors, such as `No volumes have enough 
space for a new container`. Check the filesystems holding Ratis transaction 
logs as well as those holding container data.

Review Comment:
   TODO: We should display this error in the datanode web UI or Recon UI.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to