sam-1112 opened a new issue, #6687:
URL: https://github.com/apache/datafusion-comet/issues/6687
### Describe the bug
### Describe the bug
`RevertNativeForTransitionHeavyStages` assumes result stages always produce
rows, passing `outputColumnar = false` at both result-stage call sites.
On Spark 4.0+, cached queries can require columnar output. When the cache
serializer accepts columnar input, as Comet’s `ArrowCachedBatchSerializer`
does, Spark marks the cached `AdaptiveSparkPlanExec` columnar and AQE plans the
final stage with `outputsColumnar = true`.
Forced transition reversion does not preserve this output format and the
cached query fails with:
```
FilterExec has column support mismatch
```
### Steps to reproduce
Andy reported the following reproduction during his review of \#5957:
1. Enable Comet and configure `ArrowCachedBatchSerializer` as the cache
serializer.
2. Set:
```
spark.sql.adaptive.enabled=true
spark.comet.exec.transitionRevert.enabled=true
spark.comet.exec.transitionRevert.maxTransitions=0
spark.comet.exec.filter.enabled=false
```
3. Run the following query over `tbl`, cache the resulting DataFrame, and
collect it:
```
SELECT _2, s
FROM (
SELECT _2, sum(_1) AS s
FROM tbl
GROUP BY _2
)
WHERE s > 10
```
4. Observe `FilterExec has column support mismatch`.
### Expected behavior
The cached query should execute successfully and match Spark’s results.
Stage reversion should preserve the columnar output required by the cache
consumer.
### Additional context
This is a follow-up to [Andy’s review comment on
\#5957](<https://github.com/apache/datafusion-comet/pull/5957#discussion_r4184571124>).
Andy confirmed that `main` fails the same way, so this is not a regression
introduced by that PR.
The proposed fix is to replace `outputColumnar = false` at both result-stage
call sites with:
```
outputColumnar = plan.supportsColumnar && !plan.supportsRowBased
```
Andy tested this change locally: the reproduction and all tests in
`RevertNativeForTransitionHeavyStagesSuite` passed.
Regression coverage should exercise Spark 4.0+ with AQE, caching, and forced
transition reversion. The test should install the cache serializer as
`CometInMemoryCacheSuite` does, including clearing Spark’s memoized serializer
before and after the suite.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]