[
https://issues.apache.org/jira/browse/SPARK-59171?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18111032#comment-18111032
]
Uroš Bojanić commented on SPARK-59171:
--------------------------------------
[~felipepessoto] if you happen to know Goutam Adwant's
([https://github.com/goutamadwant)] Jira user ID, please assign it here so that
we can attribute the ticket accordingly!
(https://github.com/apache/spark/pull/57809)
> [SQL] SchemaPruning is not idempotent when a query reads _metadata and a
> VARIANT column is only referenced below a nondeterministic filter
> ------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: SPARK-59171
> URL: https://issues.apache.org/jira/browse/SPARK-59171
> Project: Spark
> Issue Type: Bug
> Components: SQL
> Affects Versions: 4.2.0
> Reporter: Felipe
> Priority: Major
> Labels: pull-request-available
> Fix For: 4.4.0
>
> Original Estimate: 2h
> Remaining Estimate: 2h
>
> Originally created at https://github.com/apache/spark/issues/57659
> h3. What
> The {{Early Filter and Projection Push-Down}} batch is a {{Once}} batch, and
> it is not idempotent for a particular plan shape. {{RuleExecutor}} detects
> this and throws:
> {{org.apache.spark.SparkRuntimeException: Once strategy's idempotence is
> broken for batch Early Filter and Projection Push-Down
> Aggregate [file_path#4], [file_path#4] Aggregate
> [file_path#4], [file_path#4]
> +- Project [_metadata#3.file_path AS file_path#4] +- Project
> [_metadata#3.file_path AS file_path#4]
> +- Filter UDF() +- Filter UDF()
> ! +- Project [v#9.0 AS v#2, _metadata#3] +- Filter
> (isnotnull(v#9) AND (v#9.1 = 3))
> ! +- Filter (isnotnull(v#9) AND (v#9.1 = 3)) +- Relation
> [v#9,_metadata#3] parquet
> ! +- Relation [v#9,_metadata#3] parquet}}
>
> Applying the batch a second time removes the {{Project [v#9.0 AS v#2,
> _metadata#3]}} node, so the result of the first application is not a fixed
> point.
> h3. Reproduction
> Plain Parquet, no third-party code. This is written as a Spark unit test
> because {{RuleExecutor}} only checks idempotence when {{{}Utils.isTesting{}}}:
> test("SchemaPruning idempotence with variant and _metadata") {
> withTempDir { dir =>val path = new java.io.File(dir, "t").getAbsolutePath
> spark.range(0, 10)
> .selectExpr("parse_json(cast(id as string)) as v")
> .write.parquet(path) val alwaysTrue = udf(() =>
> true).asNondeterministic()
> spark.read.parquet(path)
> .where("v::int = 3") // filter references the
> variant column .select(col("_metadata.file_path")) // the only
> output column .filter(alwaysTrue()) //
> nondeterministic, so it is not pushed down .distinct()
> .collect()
> }
> }
>
> h3. Why the shape matters
> Three ingredients appear to be needed:
> # the query outputs *only* {{{}_metadata.file_path{}}}, so no data column is
> required above the scan;
> # the filter references a *VARIANT* column, so it is pushed below the
> variant reconstruction projection ({{{}v#9.0 AS v#2{}}}), leaving that
> projection unused;
> # a *nondeterministic* filter sits above, which prevents the unused
> projection from being collapsed in the same pass.
> Removing any one of them makes the failure go away. In particular, a plain
> nested {{struct}} column in place of the variant works fine, as does the same
> query without the nondeterministic filter.
> h3. Impact
> {{RuleExecutor}} only runs {{checkBatchIdempotence}} under
> {{{}Utils.isTesting{}}}:
> // Check idempotence for Once batches.if (batch.strategy == Once
> &&Utils.isTesting && !excludedOnceBatches.contains(batch.name)) \{
> checkBatchIdempotence(batch, curPlan)
> }
>
> So this is not a wrong-results bug for end users: outside tests the batch
> runs once and the plan it produces is correct, just with a redundant
> projection left in place. The practical impact is that the optimizer leaves a
> plan it would keep rewriting, and that any project whose test suites run in a
> Spark test JVM and that builds this plan shape fails.
> It surfaced in Delta Lake, whose {{{}UPDATE{}}}/{{{}DELETE{}}} identify the
> files to rewrite with a query of exactly this shape: they read
> {{{}_metadata.file_path{}}}, and they wrap the scan in a nondeterministic
> filter that increments a SQL metric.
> h3. Versions
> Reproduced with the snippet above on Spark {*}4.0{*}, *4.1* and {*}4.2{*}, so
> this does not look like a recent regression.
> (For context, Delta Lake only sees it on 4.1 and 4.2, because the variant
> reconstruction projection in the plan depends on shredded variant read
> support.)
> h3. Workaround
> Disabling nested schema pruning for the affected query avoids it:
> spark.sql.optimizer.nestedSchemaPruning.enabled = false
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]