[ 
https://issues.apache.org/jira/browse/TEZ-2105?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14350112#comment-14350112
 ] 

shanaka edited comment on TEZ-2105 at 3/6/15 8:37 AM:
------------------------------------------------------

Hi

I'm new to Apache Software Foundation and gsoc, so I may not be very familiar 
the way things happening here, but I always wanted to be a part of Apache 
Software Foundation, I hope to start fresh with this project.

This seems like a interesting project to me, as there are many things involved 
in this like Hadoop, YARN and etc. , so there are many thing I can learn doing 
this project. 

I went through your presentation (Pig-on-Tez), but I could not fully understand 
everything in there. So it would  be really helpful if you could help me 
understanding this project.

Thanks.


was (Author: shanaka.kuruwita):
Hi

I'm new to Apache Software Foundation and gsoc, so I may not me very familiar 
the way things happen here, but I always wanted to be a part of Apache Software 
Foundation, I hope to start fresh with this project.

This seems like a interesting project to me, as there are many things involved 
in this like Hadoop, YARN and etc. , so there are many thing I can learn doing 
this project. 

I went through your presentation (Pig-on-Tez), but I could not fully understand 
everything in there. So it would  be really helpful if you could help me 
understanding this project.

Thanks.

> Totally Sorted Edge with auto-parallelism
> -----------------------------------------
>
>                 Key: TEZ-2105
>                 URL: https://issues.apache.org/jira/browse/TEZ-2105
>             Project: Apache Tez
>          Issue Type: New Feature
>            Reporter: Gopal V
>              Labels: gsoc, gsoc2015, hadoop, java, pig, tez
>
> Pig-on-Tez supports an edge configuration using a sampled Output along with a 
> vertex manager  for automatic parallelism estimation.
> This is referred to in the Pig-on-Tez Hadoop Summit presentation.
> http://www.slideshare.net/Hadoop_Summit/pig-on-tez-low-latency-etl-with-big-data/19
> Migrating that plan-model into Tez as a native edge type would allow for much 
> more efficient scheduling of the downstream edges and effectively turn the 
> auto-parallelism implementation into a runtime skew-correcting mechanism 
> within this edge.
> The Tez Edge has enough information to sample, determine partitioning order 
> and correct parallelism.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Reply via email to