[
https://issues.apache.org/jira/browse/HBASE-30459?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
ASF GitHub Bot updated HBASE-30459:
-----------------------------------
Labels: pull-request-available (was: )
> Add an option to account compaction throughput control by actual output
> (post-compression) bytes
> ------------------------------------------------------------------------------------------------
>
> Key: HBASE-30459
> URL: https://issues.apache.org/jira/browse/HBASE-30459
> Project: HBase
> Issue Type: Improvement
> Components: Compaction
> Reporter: Jeongmin Kim
> Priority: Major
> Labels: pull-request-available
>
> Compaction throughput limits are consumed in {{Compactor#performCompaction}}
> by {{cell.getSerializedSize()}} — the pre-encoding, pre-compression size of
> each cell — while the bytes actually written to disk are
> post-DATA_BLOCK_ENCODING / post-compression. Under the same limit, a store
> whose data encodes/compresses R times runs its compactions at roughly 1/R of
> the configured disk write rate.
> This works against both goals of throughput control: keeping compactions
> inside a predictable resource envelope, and letting them actually use the
> resources that envelope grants. The better the encoding and compression work,
> the more of the budget is charged for bytes that never reach the disk, and
> the compaction sleeps instead of writing. In practice — we observed this in
> production — a compaction of a highly compressible store spends most of its
> wall time in throttle sleep while the disk stays nearly idle: even a
> compaction that is small on disk takes a long time, occupies a
> longCompactions thread throughout, and the compaction queue backs up. And
> since limits are per-RegionServer (there is no per-table or per-family
> limit), stores holding incompressible payloads consume the same budget 1:1,
> so mixed workloads additionally become unfair between families.
> Proposal: an opt-in configuration
> {{hbase.hstore.compaction.throughput.control.by.output}} (default
> {{{}false{}}}, current behavior unchanged). When enabled and the sink exposes
> its output position, {{performCompaction}} calls
> {{ThroughputController#control}} with the delta of the writer's output
> position after each cell-batch append, instead of the cells' serialized
> sizes. Notes on the implementation we have been running:
> * {{StoreFileWriter#getPos}} also includes the historical file writer when
> {{{}hbase.enable.historical.compaction.files=true{}}}, so the whole disk
> write load is accounted; {{AbstractMultiFileWriter}} gains a {{getPos}}
> summing its lower writers (stripe / date-tiered compactors).
> * The position only advances when a block is encoded/compressed and flushed
> to the output stream, so the controller sees actual disk bytes. The control
> check interval ({{{}controlPerSize{}}}, by default the throughput lower
> bound) is much larger than block sizes (32K ~ 128K), so block-granularity
> jumps are absorbed by the accounting.
> * Progress accounting, the shipped() cadence and {{CloseChecker}} stay on
> cell serialized sizes — only the throughput accounting unit changes.
> * Sinks that do not expose an output position (e.g.
> {{{}DefaultMobStoreCompactor{}}}, which overrides {{performCompaction}}
> anyway) keep the existing accounting; an unexpected sink falls back with a
> warn log.
> * Bytes written at close time (remaining inline chunks, root index, file
> info, trailer) are not accounted — the existing accounting does not see them
> either.
> * Input-side load (reading and decompressing the source storefiles) is not
> accounted; for major compactions input ≈ output on disk, so the practical gap
> is small.
> * A side benefit: the controller's finish log ("... average throughput is
> ...") then reports disk MB/s, which is directly comparable with device
> throughput.
> With the option enabled, the controller charges only what is actually
> written: compactions of compressible stores can use the full configured disk
> budget, so their duration becomes proportional to their on-disk size, and
> stores with incompressible payloads behave exactly as before — output
> accounting only differs where encoding/compression does. The implementation
> has been running in production with no regressions observed.
> A PR against master follows.
>
--
This message was sent by Atlassian Jira
(v8.20.10#820010)