Jon Haddad created CASSANDRA-21555:
--------------------------------------
Summary: Reduce repair overhead on tables with large partitions
(20GB+)
Key: CASSANDRA-21555
URL: https://issues.apache.org/jira/browse/CASSANDRA-21555
Project: Apache Cassandra
Issue Type: Improvement
Components: Feature/Repair
Reporter: Jon Haddad
When repair streams a large partition (20GB+) to rebuild it on the receiving
side, the SSTable writer that constructs the partition's row index grows its
internal buffer incrementally rather than sizing it up front. For a partition
that size, this results in a very large amount of unnecessary memory copying
and discarded garbage, adding significant CPU time and GC pressure on top of
the write itself.
We can avoid that overhead by sizing the buffer correctly ahead of time and
backing it with reused off-heap memory instead of repeatedly growing and
discarding on-heap arrays. It does not change what data is streamed or how
repair diffs partitions — it only reduces the cost of writing out a large
partition once that data is already being transferred.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]