Mikhail Erofeev created SPARK-22784:
---------------------------------------

             Summary: Increase reading buffer size in Spark History Server
                 Key: SPARK-22784
                 URL: https://issues.apache.org/jira/browse/SPARK-22784
             Project: Spark
          Issue Type: Improvement
          Components: Spark Core
    Affects Versions: 2.0.0
            Reporter: Mikhail Erofeev
            Priority: Minor


Motivation:
Our Spark History Server spends most of its warm-up time inside BufferedReader 
and StringBuffer. It happens because average line size of our events is 
~1.500.000 chars (due to a lot of partitions and iterations), whereas the 
default buffer size is 2048 bytes. See the attached flame graph.

Implementation:
I've added logging of spent time and line size for each job.
Parametrised ReplayListenerBus with new buffer size parameter. 
Measured best buffer size. x20 of average line size (30mb) gives 32% speedup in 
a local test.

Result:
Warm-up of Spark History and reading to cache will be up to 30% faster after 
tuning.



--
This message was sent by Atlassian JIRA
(v6.4.14#64029)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to