eomiks opened a new pull request, #8740:
URL: https://github.com/apache/hbase/pull/8740

   Adds an opt-in `hbase.hstore.compaction.throughput.control.by.output` 
(default `false`): when enabled, compaction throughput control is charged with 
the increase of the writer's output position — the post-encoding, 
post-compression bytes actually written — after each cell-batch append, instead 
of the cells' pre-compression serialized sizes.
   
   Without it, a store whose data encodes/compresses R times runs its 
compactions at roughly 1/R of the configured disk write rate: the budget is 
charged for bytes that never reach the disk, so compactions of compressible 
stores spend most of their wall time in throttle sleep while the disk stays 
nearly idle, and the compaction queue backs up.
   
   `StoreFileWriter#getPos` includes the historical file writer when historical 
compaction files are enabled, and `AbstractMultiFileWriter` sums its lower 
writers (stripe / date-tiered compactors). Sinks that do not expose an output 
position (e.g. the MOB compactor) keep the existing accounting, and the default 
path is unchanged.
   
   The new `TestCompactionThroughputAccounting` verifies the accounting unit 
deterministically (no timing assertions) by recording the sizes passed to 
`ThroughputController#control` and comparing them against the on-disk size of 
the compacted output, for the off / on / historical-files cases.
   
   Details: https://issues.apache.org/jira/browse/HBASE-30459
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to