Mikhail Erofeev created SPARK-22784:
---------------------------------------
Summary: Increase reading buffer size in Spark History Server
Key: SPARK-22784
URL: https://issues.apache.org/jira/browse/SPARK-22784
Project: Spark
Issue Type: Improvement
Components: Spark Core
Affects Versions: 2.0.0
Reporter: Mikhail Erofeev
Priority: Minor
Motivation:
Our Spark History Server spends most of its warm-up time inside BufferedReader
and StringBuffer. It happens because average line size of our events is
~1.500.000 chars (due to a lot of partitions and iterations), whereas the
default buffer size is 2048 bytes. See the attached flame graph.
Implementation:
I've added logging of spent time and line size for each job.
Parametrised ReplayListenerBus with new buffer size parameter.
Measured best buffer size. x20 of average line size (30mb) gives 32% speedup in
a local test.
Result:
Warm-up of Spark History and reading to cache will be up to 30% faster after
tuning.
--
This message was sent by Atlassian JIRA
(v6.4.14#64029)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]