[
https://issues.apache.org/jira/browse/PHOENIX-6751?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=17597431#comment-17597431
]
ASF GitHub Bot commented on PHOENIX-6751:
-----------------------------------------
chrajeshbabu commented on code in PR #1495:
URL: https://github.com/apache/phoenix/pull/1495#discussion_r957870057
##########
phoenix-core/src/main/java/org/apache/phoenix/compile/WhereOptimizer.java:
##########
@@ -308,6 +322,30 @@ public static Expression
pushKeyExpressionsToScan(StatementContext context, Set<
slotSpanArray[cnf.size()] = clipLeftSpan-1;
cnf.add(keyRanges);
}
+
+ // Do not use the skipScanFilter when there is a large IN clause
(for e.g > 50k elements)
+ // Since the generation of point keys for skip scan filter will
blow up the memory usage.
+ // See ScanRanges.getPointKeys(...) where using the various slot
key ranges
+ // to generate point keys will lead to combinatorial explosion.
+ // The following check will ensure the cardinality of generated
point keys
+ // is below the configured max (maxInListSkipScanSize).
+ // We shall force a range scan if the configured max is exceeded.
+ // cnfStartPos => is the start slot of this IN list
+ if (checkMaxSkipScanCardinality) {
Review Comment:
Can this conversion of skip scan to range scan configurable because range
scan on bigger data sets is slow?
> Force using range scan vs skip scan when using the IN operator and large
> number of RVC elements
> ------------------------------------------------------------------------------------------------
>
> Key: PHOENIX-6751
> URL: https://issues.apache.org/jira/browse/PHOENIX-6751
> Project: Phoenix
> Issue Type: Bug
> Affects Versions: 4.15.0, 5.1.1, 4.16.0, 5.2.0
> Reporter: Jacob Isaac
> Assignee: Jacob Isaac
> Priority: Critical
> Fix For: 5.2.0, 5.1.3
>
>
> SQL queries using the IN operator using PKs of different SortOrder were
> failing during the WHERE clause compilation phase and causing OOM issues on
> the servers when a large number (~50k) of RVC elements were used in the IN
> operator.
> SQL queries were failing specifically during the skip scan filter generation.
> The skip scan filter is generated using a list of point key ranges.
> [ScanRanges.create|https://git.soma.salesforce.com/bigdata-packaging/phoenix/blob/e0737e0ea7ba7501e78fe23c16e7abca27bfd944/phoenix-core/src/main/java/org/apache/phoenix/compile/ScanRanges.java#L80]
> The following getPointKeys
> [code|https://git.soma.salesforce.com/bigdata-packaging/phoenix/blob/e0737e0ea7ba7501e78fe23c16e7abca27bfd944/phoenix-core/src/main/java/org/apache/phoenix/compile/ScanRanges.java#L541]
> uses the KeyRange sets to create a new list of point-keys. When there are a
> large number of RVC elements the above
--
This message was sent by Atlassian Jira
(v8.20.10#820010)