[
https://issues.apache.org/jira/browse/HBASE-4038?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13056701#comment-13056701
]
Nicolas Spiegelberg commented on HBASE-4038:
--------------------------------------------
@Todd: Correct. We also have a separate internal task to look at a
runtime-enabled sampling approach for hot read diagnosis. However, right now,
our main applications with non-uniform distribution are heavily write-dominant
so write analysis is more important for us.
@Jason: Tracking Block Cache usage would give us hot read analysis. You would
have the same problem where there is not a 1:1 Block:Row mapping, so you would
need further investigation either way. Really, you want general read/write
request stats so you know which servers to drill down into. Note that the
metrics necessary for this approach could also be used by the load balancer.
> Hot Region Diagnosis
> --------------------
>
> Key: HBASE-4038
> URL: https://issues.apache.org/jira/browse/HBASE-4038
> Project: HBase
> Issue Type: Improvement
> Components: client, regionserver
> Affects Versions: 0.92.0
> Reporter: Nicolas Spiegelberg
> Assignee: Nicolas Spiegelberg
> Priority: Minor
>
> We should provide a basic way for end users to operationally diagnose hot row
> problems. Thinking about a 2-phase approach:
> 1. Diagnose hot regions
> 2. Inspect those regions/servers to find the hot rows.
> To diagnose hot regions, we could query the master or regionservers for these
> regions + sort. To inspect the regions for hot rows, we could write another
> script to analyze the HLogs on a server and basically do: sort log|uniq
> -n|sort -n|top
--
This message is automatically generated by JIRA.
For more information on JIRA, see: http://www.atlassian.com/software/jira