[ 
https://issues.apache.org/jira/browse/CASSANDRA-13241?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16656821#comment-16656821
 ] 

Ariel Weisberg edited comment on CASSANDRA-13241 at 10/22/18 3:44 PM:
----------------------------------------------------------------------

Summary as charts

Load:
||Chunk size|Time|
|64k|39:27|
|64k|36:37|
|32k|37:29|
|16k|39:25|
|16k|38:15|
|8k|37:47|
|4k|39:33|

Read:
||Chunk size|Time|Speedup|
|64k|25:20|1.005x|
|64k|25:33|1.000x|
|32k|20:01|1.265x|
|16k|19:19|1.319x|
|16k|19:14|1.323x|
|8k|16:51|1.534x|
|4k|15:39|1.645x|

||Chunk size|Compression Ratio|Improvement|Compared to 64k improvement|
|64k|0.607208|39.3%|1.000|
|32k|0.634735|36.6%|0.931|
|16k|0.667236|33.3%|0.847|
|8k|0.709473|29.1%|0.740|


was (Author: aweisberg):
Summary as charts

Load:
||Chunk size|Time||
|64k|39:27|
|64k|36:37|
|32k|37:29|
|16k|39:25|
|16k|38:15|
|8k|37:47|
|4k|39:33|

Read:
||Chunk size|Time||
|64k|25:20|
|64k|25:33|
|32k|20:01|
|16k|19:19|
|16k|19:14|
|8k|16:51|
|4k|15:39|

> Lower default chunk_length_in_kb from 64kb to 4kb
> -------------------------------------------------
>
>                 Key: CASSANDRA-13241
>                 URL: https://issues.apache.org/jira/browse/CASSANDRA-13241
>             Project: Cassandra
>          Issue Type: Wish
>          Components: Core
>            Reporter: Benjamin Roth
>            Assignee: Ariel Weisberg
>            Priority: Major
>         Attachments: CompactIntegerSequence.java, 
> CompactIntegerSequenceBench.java, CompactSummingIntegerSequence.java
>
>
> Having a too low chunk size may result in some wasted disk space. A too high 
> chunk size may lead to massive overreads and may have a critical impact on 
> overall system performance.
> In my case, the default chunk size lead to peak read IOs of up to 1GB/s and 
> avg reads of 200MB/s. After lowering chunksize (of course aligned with read 
> ahead), the avg read IO went below 20 MB/s, rather 10-15MB/s.
> The risk of (physical) overreads is increasing with lower (page cache size) / 
> (total data size) ratio.
> High chunk sizes are mostly appropriate for bigger payloads pre request but 
> if the model consists rather of small rows or small resultsets, the read 
> overhead with 64kb chunk size is insanely high. This applies for example for 
> (small) skinny rows.
> Please also see here:
> https://groups.google.com/forum/#!topic/scylladb-dev/j_qXSP-6-gY
> To give you some insights what a difference it can make (460GB data, 128GB 
> RAM):
> - Latency of a quite large CF: https://cl.ly/1r3e0W0S393L
> - Disk throughput: https://cl.ly/2a0Z250S1M3c
> - This shows, that the request distribution remained the same, so no "dynamic 
> snitch magic": https://cl.ly/3E0t1T1z2c0J



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to