[ 
https://issues.apache.org/jira/browse/HBASE-30394?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

huginn updated HBASE-30394:
---------------------------
          Component/s: BlockCache
    Affects Version/s: 2.4.11
          Description: 
## What happens

TinyLfuBlockCache reports aggregate cache size and block count through 
getCurrentDataSize() and getDataBlockCount(), even though these APIs are 
intended to report data-block-only statistics. As a result, data-block metrics 
are indistinguishable from aggregate cache metrics when index or metadata 
blocks are cached.

## When it happens

This occurs when TinyLfuBlockCache is used and the cache contains a mixture of 
data, index, or metadata blocks. The current implementation derives both values 
from the aggregate Caffeine cache size and entry count.

## Impact

Operators and monitoring cannot accurately determine the amount and number of 
data blocks held by TinyLfuBlockCache, which can make cache composition and 
capacity analysis misleading.

## Root cause

On master, TinyLfuBlockCache.getCurrentDataSize() returns getCurrentSize() and 
getDataBlockCount() returns getBlockCount(). The insertion and eviction paths 
do not maintain separate data-block counters, so the implementation cannot 
satisfy the data-block-specific BlockCache contract.

<!-- File: 
hbase-server/src/main/java/org/apache/hadoop/hbase/io/hfile/TinyLfuBlockCache.java,
 upstream master -->
<!-- Lines: 185-197, 226-232, 296-317, 412-419 -->

## Proposed fix

Maintain the current data-block size and count while blocks are inserted and 
removed. Update the counters only for blocks whose BlockType is data, and 
return those counters from getCurrentDataSize() and getDataBlockCount().

## Reproduction

Testing evidence will be added by the reporter.

  was:
What happens

When ReplicationSourceShipper.clearWALEntryBatch is interrupted while waiting 
for the shipper and reader threads to stop, the warning log can leave its final 
placeholder unresolved instead of including the interruption value in the 
message.

When it happens

During replication source shutdown or termination, if the wait in 
clearWALEntryBatch is interrupted before both threads stop.

Impact

The warning is incomplete and makes it harder to identify the interruption that 
prevented cleanup from completing.

Root cause

In 
hbase-server/src/main/java/org/apache/hadoop/hbase/replication/regionserver/ReplicationSourceShipper.java,
 the InterruptedException branch passes the exception as the last argument to a 
message with three placeholders. SLF4J treats a final Throwable specially, so 
the last placeholder is not populated as intended.

Proposed fix

Format the interrupted exception explicitly with e.toString() for the final 
placeholder, preserving the peer ID and thread name in the warning message.

Reproduction

Testing evidence will be added by the reporter.


> Report data block size and count separately in TinyLfuBlockCache
> ----------------------------------------------------------------
>
>                 Key: HBASE-30394
>                 URL: https://issues.apache.org/jira/browse/HBASE-30394
>             Project: HBase
>          Issue Type: Bug
>          Components: BlockCache
>    Affects Versions: 2.4.11
>            Reporter: huginn
>            Priority: Major
>              Labels: pull-request-available
>
> ## What happens
> TinyLfuBlockCache reports aggregate cache size and block count through 
> getCurrentDataSize() and getDataBlockCount(), even though these APIs are 
> intended to report data-block-only statistics. As a result, data-block 
> metrics are indistinguishable from aggregate cache metrics when index or 
> metadata blocks are cached.
> ## When it happens
> This occurs when TinyLfuBlockCache is used and the cache contains a mixture 
> of data, index, or metadata blocks. The current implementation derives both 
> values from the aggregate Caffeine cache size and entry count.
> ## Impact
> Operators and monitoring cannot accurately determine the amount and number of 
> data blocks held by TinyLfuBlockCache, which can make cache composition and 
> capacity analysis misleading.
> ## Root cause
> On master, TinyLfuBlockCache.getCurrentDataSize() returns getCurrentSize() 
> and getDataBlockCount() returns getBlockCount(). The insertion and eviction 
> paths do not maintain separate data-block counters, so the implementation 
> cannot satisfy the data-block-specific BlockCache contract.
> <!-- File: 
> hbase-server/src/main/java/org/apache/hadoop/hbase/io/hfile/TinyLfuBlockCache.java,
>  upstream master -->
> <!-- Lines: 185-197, 226-232, 296-317, 412-419 -->
> ## Proposed fix
> Maintain the current data-block size and count while blocks are inserted and 
> removed. Update the counters only for blocks whose BlockType is data, and 
> return those counters from getCurrentDataSize() and getDataBlockCount().
> ## Reproduction
> Testing evidence will be added by the reporter.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to