[
https://issues.apache.org/jira/browse/CASSANDRA-21655?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Jon Haddad updated CASSANDRA-21655:
-----------------------------------
Attachment: flame-cpu-forward.html
flame-cpu-reverse.html
> async compaction pipeline
> -------------------------
>
> Key: CASSANDRA-21655
> URL: https://issues.apache.org/jira/browse/CASSANDRA-21655
> Project: Apache Cassandra
> Issue Type: Improvement
> Reporter: Jon Haddad
> Priority: Normal
> Attachments: flame-cpu-forward.html, flame-cpu-reverse.html
>
>
> Compaction compresses each chunk, writes it, and updates its CRC serially on
> the compaction thread, inside \{{SequentialWriter.doFlush}}. A CPU profile
> of a 4.5 GiB compaction puts that write side at 37.2% of the total, so the
> thread that merges rows spends over a third of its time not merging.
> \{{trickle_fsync}} compounds it: the 10 MiB byte interval fires about 450
> times over that compaction, and each \{{fdatasync}} forces the whole growing
> file rather than a range.
> {noformat}
> compaction, 100%
> |<---------------- merge side, 62.8% ---------------->|<---- doFlush, 37.2%
> ---->|
> |<-LZ4 21.1%->|<CRC
> 7.6%>|io|
>
> 8.5%
> {noformat}
> Compression plus CRC is 77% of the write side. Moving the whole 37.2% off
> the compaction thread makes the merge side the slowest stage, a ceiling of
> 1.6x.
> The patch gives each SSTable writer its own writer thread and hands chunks to
> it through a pool of reusable chunk-sized off-heap slots, 4 MiB by default;
> the producer takes a fresh slot instead of waiting for the one it just
> filled. Compression, the write and the checksum all run on that thread. The
> byte-interval fsync is replaced by a periodic force, about six per compaction
> instead of hundreds, which removes work rather than merely overlapping it.
> The pipeline is a collaborator both \{{CompressedSequentialWriter}} and
> \{{DirectCompressedSequentialWriter}} own, so direct IO gets it too, and
> \{{DataComponent.buildWriter}} routes flush, streaming and index builds
> through the same path. It sits behind \{{async_compaction_writer_enabled}}.
> Measured on a 4543 MiB dataset across 16 SSTables, 16 cores, JDK 21,
> throttling off:
> ||path||before||after||change||
> |cursor|15809 ms, 287 MiB/s|5952 ms, 763 MiB/s|+166%|
> |iterator|18476 ms, 246 MiB/s|9129 ms, 498 MiB/s|+102%|
> The result beats the 1.6x ceiling because that ceiling assumed a fixed total;
> the redundant fsyncs were part of the problem.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]