[ 
https://issues.apache.org/jira/browse/CASSANDRA-21701?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Michael Semb Wever updated CASSANDRA-21701:
-------------------------------------------
    Test and Documentation Plan: 
CompactionIteratorTest (13 tests, 1 of them new) and ValidatorTest (8 tests) 
pass on JDK 17.

The new {{testBytesReadFollowsTheScannersOnEveryCall}} drives a compaction of 
three unfiltereds through a scanner that reports ten bytes per partition, and 
then compares {{getBytesRead()}} with the scanner. Without the production 
change it fails with "expected:<10> but was:<0>", because the count is below 
the refresh interval.
                         Status: Patch Available  (was: Open)

> repair_validations reports a bytes_read figure that trails the work done
> ------------------------------------------------------------------------
>
>                 Key: CASSANDRA-21701
>                 URL: https://issues.apache.org/jira/browse/CASSANDRA-21701
>             Project: Apache Cassandra
>          Issue Type: Bug
>          Components: Observability/Logging
>            Reporter: Michael Semb Wever
>            Priority: Normal
>             Fix For: 5.0.x, 6.0.x, 7.x
>
>
> {{CompactionIterator}} refreshes its {{bytesRead}} field once every 
> {{UNFILTERED_TO_UPDATE_PROGRESS}} unfiltereds:
> {code:java}
> if ((++compactedUnfiltered) % UNFILTERED_TO_UPDATE_PROGRESS == 0)
>     updateBytesRead();
> {code}
> and {{getBytesRead()}} returned that field. Two things follow. The field lags 
> behind the scanners by up to a hundred unfiltereds, and it is still short of 
> the total once the iteration is over, because the last refresh lands on the 
> last multiple of a hundred.
> {{ValidationManager}} assigns the value to {{ValidationState.bytesRead}} once 
> per partition:
> {code:java}
> state.bytesRead = vi.getBytesRead();
> {code}
> and {{LocalRepairTables}} reports it as the {{bytes_read}} column of 
> {{system_views.repair_validations}}. An operator watching a repair therefore 
> sees a figure that trails the work done, and a validation of fewer than a 
> hundred unfiltereds reports zero throughout.
> {{getBytesRead()}} now sums the scanners on every call. The field remains for 
> {{getCompactionInfo}}, which reports from a background thread, so the refresh 
> interval still keeps that path off the scanners. The validation path asks 
> once per partition rather than once per unfiltered, so the summation is not 
> on a hot loop.
> Patch: 
> [mck/upstream/stale-bytes-read/5.0|https://github.com/thelastpickle/cassandra/tree/mck/upstream/stale-bytes-read/5.0]
> Provenance: 
> [fa9e504c11|https://github.com/datastax/cassandra/commit/fa9e504c1178927d12d40b340048818714f969c7]
>  by [~blambov]. That commit reaches the same result, on a class that has 
> since diverged: it folds a second accessor, {{getTotalBytesScanned}}, into 
> {{bytesRead()}}, and that second accessor does not exist here. That commit 
> carries tests for classes that do not exist here; this patch adds one that 
> fits the upstream class.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to