Zoltán Borók-Nagy created IMPALA-15421:
------------------------------------------
Summary: Iceberg predicate subsetting drops a needed conjunct when
two residuals sanitize to the same string
Key: IMPALA-15421
URL: https://issues.apache.org/jira/browse/IMPALA-15421
Project: IMPALA
Issue Type: Bug
Components: Frontend
Reporter: Zoltán Borók-Nagy
Assignee: Zoltán Borók-Nagy
IcebergScanPlanner collects the residual expression of every FileScanTask into
a set
that uses Iceberg's sanitized string as the comparator:
{code:java}
// IcebergScanPlanner.java:125
private final Set<Expression> residualExpressions_ =
new TreeSet<>(Comparator.comparing(ExpressionUtil::toSanitizedString));
{code}
{{toSanitizedString()}} replaces numeric literals with placeholders, so
different
residuals compare as equal:
{noformat}
id != 12345 -> id != (5-digit-int)
id != 54321 -> id != (5-digit-int)
{noformat}
The set keeps only the first one. With
ICEBERG_PREDICATE_PUSHDOWN_SUBSETTING=true (the
default), trySubsettingPredicatesBeingPushedDown() retains only the Impala
conjuncts that
map to residuals in the set. The other conjunct goes to {{skippedExpressions_}}
and the
scanner never evaluates it, so rows that should be filtered are returned.
This happens when different files have different residuals that sanitize to the
same
string. Non-identity partition transforms (truncate, bucket) are a typical
cause. For
{{TRUNCATE(10000, id)}} and {{id != 12345 AND id != 54321}}, Iceberg's
ResidualEvaluator
returns {{id != 12345}} for partition 10000 and {{id != 54321}} for partition
50000.
String literals are hashed by the sanitizer, so they don't collide.
h3. Repro
{code:sql}
CREATE TABLE ice_residual (id INT)
PARTITIONED BY SPEC (TRUNCATE(10000, id)) STORED AS ICEBERG;
INSERT INTO ice_residual VALUES (12345), (12346), (54321), (54322);
SELECT id FROM ice_residual WHERE id != 12345 AND id != 54321 ORDER BY id;
{code}
Expected: 12346, 54322.
Actual (on master a6af29eee0):
{noformat}
+-------+
| id |
+-------+
| 12346 |
| 54321 |
| 54322 |
+-------+
{noformat}
The EXPLAIN output shows that the second conjunct is no longer evaluated by the
scan:
{noformat}
00:SCAN HDFS [ice_residual_repro.ice_residual, RANDOM]
HDFS partitions=2/2 files=2 size=716B
predicates: id != CAST(12345 AS INT)
skipped Iceberg predicates: id != CAST(54321 AS INT)
{noformat}
Which of the two values leaks depends on which file planFiles() returns first.
With
ICEBERG_PREDICATE_PUSHDOWN_SUBSETTING=false, the same query returns the correct
12346,
54322.
Workaround: SET ICEBERG_PREDICATE_PUSHDOWN_SUBSETTING=false.
h3. Proposed fix
Key the collection by the exact {{Expression.toString()}}, which prints
literals in full.
Use a TreeMap<String, Expression> so the order of the retained conjuncts stays
deterministic for planner tests. Add an identity short-circuit for the common
case where
every file returns the same residual object (e.g. unpartitioned tables). This
also
removes the per-file sanitization cost from planning, about 2 us per file on
unpartitioned tables in a 100k-file microbenchmark. Add a regression test for
the case
above.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]