[ 
https://issues.apache.org/jira/browse/HIVE-30070?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Stamatis Zampetakis updated HIVE-30070:
---------------------------------------
    Attachment: profile.html

> Conversion to Map join is very slow for Iceberg tables on S3
> ------------------------------------------------------------
>
>                 Key: HIVE-30070
>                 URL: https://issues.apache.org/jira/browse/HIVE-30070
>             Project: Hive
>          Issue Type: Bug
>          Components: HiveServer2
>            Reporter: Stamatis Zampetakis
>            Assignee: Stamatis Zampetakis
>            Priority: Major
>         Attachments: convert-join-map-join.png, profile.html
>
>
> The Conversion of a common shuffle join to map join (ConvertJoinMapJoin) is 
> very slow when we are dealing with Iceberg tables on S3. 
>  !convert-join-map-join.png! 
> The flamegraph was obtained via async profiler while running 
> TestTPCDSIcebergCliDriver (HIVE-30034) with hive.auto.convert.sortmerge.join 
> enabled.
> {noformat}
> mvn test -Dtest=TestTPCDSIcebergCliDriver -Dtest.output.overwrite 
> -Dqfile_regex=cbo_query4.*
> {noformat}
>  
> A large amount of time is spent on 
> {{AvgPartitionSizeBasedBigTableSelectorForAutoSMJ.getBigTablePosition}} and 
> in particular inside the {{S3AFileSystem.getContentSummary}} operation.
> Since this slowness came up in a test setup it has to be verified to what 
> extend it affects production deployments.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to