zabetak commented on code in PR #6746:
URL: https://github.com/apache/hive/pull/6746#discussion_r4105115394


##########
ql/src/test/results/clientpositive/cbo_rp_auto_join1.q.out:
##########
@@ -853,40 +856,149 @@ STAGE PLANS:
                 expressions: key (type: int)
                 outputColumnNames: key
                 Statistics: Num rows: 7 Data size: 28 Basic stats: COMPLETE 
Column stats: COMPLETE
-                Sorted Merge Bucket Map Join Operator

Review Comment:
   These tests are all about Map Join and SMB join conversions. For some reason 
after the changes in the PR the conversions are lost. Apart from losing test 
coverage it seems that the changes incur regressions. Why is this happening? Is 
this something that can be fixed?



##########
ql/src/test/results/clientpositive/perf/tpcds30tb/tez/cbo_query16.q.out:
##########
@@ -3,22 +3,21 @@ HiveProject(order count=[$0], total shipping cost=[$1], total 
net profit=[$2])
   HiveAggregate(group=[{}], agg#0=[count(DISTINCT $4)], agg#1=[sum($5)], 
agg#2=[sum($6)])
     HiveAntiJoin(condition=[=($4, $14)], joinType=[anti])
       HiveSemiJoin(condition=[AND(=($4, $14), <>($3, $13))], joinType=[semi])
-        HiveProject(cs_ship_date_sk=[$2], cs_ship_addr_sk=[$3], 
cs_call_center_sk=[$4], cs_warehouse_sk=[$5], cs_order_number=[$6], 
cs_ext_ship_cost=[$7], cs_net_profit=[$8], d_date_sk=[$9], d_date=[$10], 
ca_address_sk=[$0], ca_state=[$1], cc_call_center_sk=[$11], cc_county=[$12])
-          HiveJoin(condition=[=($4, $11)], joinType=[inner], algorithm=[none], 
cost=[not available])
-            HiveJoin(condition=[=($3, $0)], joinType=[inner], 
algorithm=[none], cost=[not available])
-              HiveProject(ca_address_sk=[$0], ca_state=[CAST('NY'):CHAR(2) 
CHARACTER SET "UTF-16LE"])
-                HiveFilter(condition=[=($8, 'NY')])
-                  HiveTableScan(table=[[default, customer_address]], 
table:alias=[customer_address])
-              HiveJoin(condition=[=($0, $7)], joinType=[inner], 
algorithm=[none], cost=[not available])
-                HiveProject(cs_ship_date_sk=[$1], cs_ship_addr_sk=[$9], 
cs_call_center_sk=[$10], cs_warehouse_sk=[$13], cs_order_number=[$16], 
cs_ext_ship_cost=[$27], cs_net_profit=[$32])
-                  HiveFilter(condition=[AND(IS NOT NULL($9), IS NOT NULL($1), 
IS NOT NULL($10))])
-                    HiveTableScan(table=[[default, catalog_sales]], 
table:alias=[cs1])
-                HiveProject(d_date_sk=[$0], d_date=[$2])
-                  HiveFilter(condition=[BETWEEN(false, CAST($2):TIMESTAMP(9), 
2001-04-01 00:00:00:TIMESTAMP(9), 2001-05-31 00:00:00:TIMESTAMP(9))])
-                    HiveTableScan(table=[[default, date_dim]], 
table:alias=[date_dim])
-            HiveProject(cc_call_center_sk=[$0], cc_county=[$25])
-              HiveFilter(condition=[IN($25, 'Daviess County':VARCHAR(30) 
CHARACTER SET "UTF-16LE", 'Franklin Parish':VARCHAR(30) CHARACTER SET 
"UTF-16LE", 'Huron County':VARCHAR(30) CHARACTER SET "UTF-16LE", 'Levy 
County':VARCHAR(30) CHARACTER SET "UTF-16LE", 'Ziebach County':VARCHAR(30) 
CHARACTER SET "UTF-16LE")])
-                HiveTableScan(table=[[default, call_center]], 
table:alias=[call_center])

Review Comment:
   After changing the selectivity estimation there are various join order 
changes in TPC-DS queries like this one (`.q.out` files under 
`perf/tpcds30tb/tez`).
   
   Join(Join(customer_address, Join(cs1, date_dim)), call_center): **Before**
   Join(Join(Join(cs1, date_dim), customer_address), call_center): **After**
   
   Are the new join orders better purely from a plan perspective? In other 
words, are the new selectivity estimations obtained using the uniform 
distribution more precise (closer to the reality) for the TPC-DS queries?
   
   It's probably hard to tell what will be the impact in the actual performance 
of the query without running it but we should at least verify that 
theoretically the plans are better. I went through the plans myself but I would 
like to confirm if others did the same exercise.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to