[ 
https://issues.apache.org/jira/browse/PHOENIX-6752?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17632612#comment-17632612
 ] 

ASF GitHub Bot commented on PHOENIX-6752:
-----------------------------------------

stoty commented on code in PR #1529:
URL: https://github.com/apache/phoenix/pull/1529#discussion_r1020680442


##########
phoenix-core/src/main/java/org/apache/phoenix/compile/WhereOptimizer.java:
##########
@@ -926,7 +932,11 @@ private KeySlots andKeySlots(AndExpression andExpression, 
List<KeySlots> childSl
                             if (slot.getOrderPreserving() != null) {
                                 orderPreserving = 
slot.getOrderPreserving().combine(orderPreserving);
                             }
-                            if (slot.getKeyPart().getExtractNodes() != null) {
+                            // Extract once per iteration, when there are 
large number
+                            // of OR clauses (for e.g N > 100k).
+                            // The extractNodes.addAll method can get called N 
times.
+                            if (!visitedKeyParts.contains(slot.getKeyPart()) 
&& slot.getKeyPart().getExtractNodes() != null) {

Review Comment:
   Does this contains() check improve perf ?
   I'd expect checking then adding an existing element to a hash to take 
similar (or more) time to simply adding it, and letting the hash take care of  
duplicates.





> Duplicate expression nodes in extract nodes during WHERE compilation phase 
> leads to poor performance.
> -----------------------------------------------------------------------------------------------------
>
>                 Key: PHOENIX-6752
>                 URL: https://issues.apache.org/jira/browse/PHOENIX-6752
>             Project: Phoenix
>          Issue Type: Bug
>    Affects Versions: 4.15.0, 5.1.0, 4.16.1, 5.2.0
>            Reporter: Jacob Isaac
>            Assignee: Jacob Isaac
>            Priority: Critical
>             Fix For: 5.2.0
>
>         Attachments: test-case.txt
>
>
> SQL queries using the OR operator were taking a long time during the WHERE 
> clause compilation phase when a large number of OR clauses (~50k) are used.
> The key observation was that during the AND/OR processing, when there are a 
> large number of OR expression nodes the same set of extracted nodes was 
> getting added.
> Thus bloating the set size and slowing down the processing.
> [code|https://github.com/apache/phoenix/blob/0c2008ddf32566c525df26cb94d60be32acc10da/phoenix-core/src/main/java/org/apache/phoenix/compile/WhereOptimizer.java#L930]



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to