[ 
https://issues.apache.org/jira/browse/HBASE-30359?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18112056#comment-18112056
 ] 

mazhengxuan commented on HBASE-30359:
-------------------------------------

Hi [~asolomon] [~andor] [~vinayakhegde], could you take a look at the proposed 
metrics?

I am thinking of starting with checkpoint age, queued WAL count/bytes, and 
archive/checkpoint failures and latency. These would be aggregated at the 
RegionServer level, without table or WAL dimensions.

I am less sure about PITR restore/WALPlayer metrics. Since restore is an 
on-demand CLI/MapReduce job, it may be better to handle those separately under 
HBASE-16565.

Does this scope make sense? I would especially like to confirm that checkpoint 
age should only advance after the checkpoint is persisted, and that 
pendingWALBytes should be based on persisted replication offsets.

> Add RegionServer-level metrics for continuous backup health
> -----------------------------------------------------------
>
>                 Key: HBASE-30359
>                 URL: https://issues.apache.org/jira/browse/HBASE-30359
>             Project: HBase
>          Issue Type: Improvement
>          Components: backup&restore
>            Reporter: mazhengxuan
>            Assignee: mazhengxuan
>            Priority: Major
>
> HBASE-28957 adds continuous WAL archival for PITR through 
> ContinuousBackupReplicationEndpoint. At the moment, it is difficult to tell 
> from metrics whether the backup is making durable progress or building up a 
> backlog.
> The existing replication metrics are not quite enough here. For example, 
> ageOfLastShippedOp is updated before the backup WAL writer is closed and the 
> checkpoint is persisted. Also, ContinuousBackupReplicationEndpoint reports an 
> archive failure by returning false, while failedBatches is only incremented 
> when an exception escapes from the endpoint.
> I am thinking of adding the following RegionServer metrics:
>  * checkpointAge: Time since the oldest successfully persisted checkpoint 
> handled by this RegionServer. This is the most direct indication of the 
> current PITR recovery point.
>  * queuedWALs and pendingWALBytes: The amount of WAL data waiting to be 
> covered by persisted replication offsets. WAL count alone is not enough, as 
> WAL sizes can vary significantly.
>  * walArchiveFailures and walArchiveTime: Number of failed archive attempts 
> and time spent in ContinuousBackupReplicationEndpoint.replicate(). This 
> should include failures reported by returning false.
>  * checkpointFailures and checkpointTime: Failures and time spent closing 
> backup WAL writers and storing the checkpoint before the replication offset 
> is advanced.
> These metrics would be aggregated at RegionServer process level across the 
> continuous-backup replication sources and WAL groups hosted by the process. I 
> do not plan to add table or WAL dimensions, as they would create unnecessary 
> cardinality. Existing shippedBatches, shippedBytes, and shippedOps metrics 
> can continue to provide throughput information.
> A few related metrics are still open for discussion:
>  * PITR restore failure count and duration. PITR restore runs as an on-demand 
> CLI/MapReduce job, so I am not sure a RegionServer metric is the right place 
> for it.
>  * Separate bulk-load upload metrics. For now, these failures could be 
> included in walArchiveFailures.
>  * Lower-level object-store request and cleanup metrics. These may be useful 
> later, but the checkpoint, backlog, failure, and latency metrics should first 
> tell us whether that extra detail is needed.
> For checkpointAge, I propose updating it only after the backup WAL writers 
> have been closed and the checkpoint has been stored successfully. 
> pendingWALBytes should be calculated from persisted offsets rather than only 
> the reader's in-memory position.
> Does this look like the minimum useful set? In particular, feedback on the 
> checkpoint and pendingWALBytes definitions would be helpful.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to