[jira] [Updated] (MAPREDUCE-2349) speed up list[located]status calls from input formats

Vinod Kumar Vavilapalli (JIRA) Wed, 19 Mar 2014 13:58:57 -0700

     [ 
https://issues.apache.org/jira/browse/MAPREDUCE-2349?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]


Vinod Kumar Vavilapalli updated MAPREDUCE-2349:
-----------------------------------------------

    Status: Open  (was: Patch Available)

Started looking at it, nicely done patch! Few minor comments
 - Add the configs to mapred-default.xml as documentation?
 - LIST_STATUS_NUM_THREADS_DEFAULT -> DEFAULT_LIST_STATUS_NUM_THREADS
 - oldListStatus() -> singleThreadedListStatus()
 - Can you add a bit of javadoc to all the new classes and methods in 
LocatedFileStatusFetcher? Also to the main LocatedFileStatusFetcher class 
itself.
 - Synchronization needed for ProcessInitialInputPathResult.addError()?
 - Can you group the callable, result and call-back for each type of operation 
together in two classes?
 - The 'result' variable doesn't need to be a class field of 
ProcessInputDirCallable. Similarly the one in ProcessInitialInputPathCallable.

The tests look good!

> speed up list[located]status calls from input formats
> -----------------------------------------------------
>
>                 Key: MAPREDUCE-2349
>                 URL: https://issues.apache.org/jira/browse/MAPREDUCE-2349
>             Project: Hadoop Map/Reduce
>          Issue Type: Improvement
>          Components: task
>            Reporter: Joydeep Sen Sarma
>            Assignee: Siddharth Seth
>         Attachments: MAPREDUCE-2349.1.wip.txt, MAPREDUCE-2349.2.txt, 
> MAPREDUCE-2349.3.txt, MAPREDUCE-2349.4.txt
>
>
> when a job has many input paths - listStatus - or the improved 
> listLocatedStatus - calls (invoked from the getSplits() method) can take a 
> long time. Most of the time is spent waiting for the previous call to 
> complete and then dispatching the next call. 
> This can be greatly speeded up by dispatching multiple calls at once (via 
> executors). If the same filesystem client is used - then the calls are much 
> better pipelined (since calls are serialized) and don't impose extra burden 
> on the namenode while at the same time greatly reducing the latency to the 
> client. In a simple test on non-peak hours, this resulted in the getSplits() 
> time reducing from about 3s to about 0.5s.



--
This message was sent by Atlassian JIRA
(v6.2#6252)

[jira] [Updated] (MAPREDUCE-2349) speed up list[located]status calls from input formats

Reply via email to