dwsmith1983 commented on PR #25895:
URL: https://github.com/apache/datafusion/pull/25895#issuecomment-5922233473

   > Would you mind adding more information how critical is not having this fix 
in the patchset?
   
   Without it the query returns wrong results and raises no error. With 
`coerce_int96` on, struct, list and map fields read from a file that has an 
INT96 column lose their metadata. A reader that matches columns by Parquet 
field id then no longer finds those containers and fills them with nulls. Comet 
enables the coercion for every scan, so with Spark's field id reads on, `where 
s.inner.x = 1` and `where s is not null` return no rows and `count(s)` is 0 
where Spark reads the data (apache/datafusion-comet#6405). Until DataFusion 
carries this, Comet's stopgap (apache/datafusion-comet#6444) sends every scan 
that asks for an id on a nested field back to Spark.
   
   The fix is three `.with_metadata(...)` calls in `schema_coercion.rs`. Files 
without an INT96 column are unaffected.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to