[
https://issues.apache.org/jira/browse/SQOOP-2906?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15280207#comment-15280207
]
Joeri Hermans commented on SQOOP-2906:
--------------------------------------
Hi Attila
I'm definitely not an expert on Sqoop internals, but architectural changes are
definitely better! I'm just publishing the results I have obtained by using the
profiling, and providing a fix for this. Of course, if we implement your
suggestion, this would definitely benefit the CPU consumption in general.
However, for now, we could patch this first (so have the increase in
performance already), and then look at what needs to be done in order to
correctly implement your idea(s).
Kind regards,
Joeri
> Optimization of AvroUtil.toAvroIdentifier
> -----------------------------------------
>
> Key: SQOOP-2906
> URL: https://issues.apache.org/jira/browse/SQOOP-2906
> Project: Sqoop
> Issue Type: Improvement
> Reporter: Joeri Hermans
> Assignee: Joeri Hermans
> Labels: avro, hadoop, optimization
> Attachments: diff.txt
>
>
> Hi all
> Our distributed profiler indicated some inefficiencies in the
> AvroUtil.toAvroIdentifier method, more specifically, the use of Regex
> patterns. This can be directly observed from the FlameGraph generated by this
> profiler (https://jhermans.web.cern.ch/jhermans/sqoop_avro_flamegraph.svg).
> We implemented an optimization, and compared this with the original method.
> On our testing machine, the optimization by itself is about 500% (on average)
> more efficient compared to the original implementation. We have yet to test
> how this optimization will influence the performance of user jobs.
> Any suggestions or remarks are welcome.
> Kind regards,
> Joeri
> https://github.com/apache/sqoop/pull/18
> Writeup:
> https://db-blog.web.cern.ch/blog/joeri-hermans/2016-04-hadoop-performance-troubleshooting-stack-tracing-introduction
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)