[
https://issues.apache.org/jira/browse/CASSANDRA-21701?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Michael Semb Wever updated CASSANDRA-21701:
-------------------------------------------
Test and Documentation Plan:
CompactionIteratorTest (13 tests, 1 of them new) and ValidatorTest (8 tests)
pass on JDK 17.
The new {{testBytesReadFollowsTheScannersOnEveryCall}} drives a compaction of
three unfiltereds through a scanner that reports ten bytes per partition, and
then compares {{getBytesRead()}} with the scanner. Without the production
change it fails with "expected:<10> but was:<0>", because the count is below
the refresh interval.
Status: Patch Available (was: Open)
> repair_validations reports a bytes_read figure that trails the work done
> ------------------------------------------------------------------------
>
> Key: CASSANDRA-21701
> URL: https://issues.apache.org/jira/browse/CASSANDRA-21701
> Project: Apache Cassandra
> Issue Type: Bug
> Components: Observability/Logging
> Reporter: Michael Semb Wever
> Priority: Normal
> Fix For: 5.0.x, 6.0.x, 7.x
>
>
> {{CompactionIterator}} refreshes its {{bytesRead}} field once every
> {{UNFILTERED_TO_UPDATE_PROGRESS}} unfiltereds:
> {code:java}
> if ((++compactedUnfiltered) % UNFILTERED_TO_UPDATE_PROGRESS == 0)
> updateBytesRead();
> {code}
> and {{getBytesRead()}} returned that field. Two things follow. The field lags
> behind the scanners by up to a hundred unfiltereds, and it is still short of
> the total once the iteration is over, because the last refresh lands on the
> last multiple of a hundred.
> {{ValidationManager}} assigns the value to {{ValidationState.bytesRead}} once
> per partition:
> {code:java}
> state.bytesRead = vi.getBytesRead();
> {code}
> and {{LocalRepairTables}} reports it as the {{bytes_read}} column of
> {{system_views.repair_validations}}. An operator watching a repair therefore
> sees a figure that trails the work done, and a validation of fewer than a
> hundred unfiltereds reports zero throughout.
> {{getBytesRead()}} now sums the scanners on every call. The field remains for
> {{getCompactionInfo}}, which reports from a background thread, so the refresh
> interval still keeps that path off the scanners. The validation path asks
> once per partition rather than once per unfiltered, so the summation is not
> on a hot loop.
> Patch:
> [mck/upstream/stale-bytes-read/5.0|https://github.com/thelastpickle/cassandra/tree/mck/upstream/stale-bytes-read/5.0]
> Provenance:
> [fa9e504c11|https://github.com/datastax/cassandra/commit/fa9e504c1178927d12d40b340048818714f969c7]
> by [~blambov]. That commit reaches the same result, on a class that has
> since diverged: it folds a second accessor, {{getTotalBytesScanned}}, into
> {{bytesRead()}}, and that second accessor does not exist here. That commit
> carries tests for classes that do not exist here; this patch adds one that
> fits the upstream class.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]