zhuqi-lucas opened a new issue, #26103:
URL: https://github.com/apache/datafusion/issues/26103

   ### Describe the bug
   
   `ListingTable` schema inference over Parquet files marks a column `NOT NULL` 
when every file that *has* the column declares it required, even if other files 
do not have the column at all. Reading those files fills the column with nulls, 
and the scan then fails:
   
   ```
   Arrow error: Invalid argument error: Column 'c' is declared as non-nullable 
but contains null values
   ```
   
   The table is unusable for any query that touches that column, including 
`SELECT *`. The only way out is to declare the schema by hand.
   
   ### To Reproduce
   
   Pure SQL in `datafusion-cli` on `main` (4d167a167, 2026-10-02); literal 
columns are written as required:
   
   ```sql
   COPY (SELECT 1 AS id, 10 AS c) TO '/tmp/evo/a.parquet';
   COPY (SELECT 2 AS id)          TO '/tmp/evo/b.parquet';
   CREATE EXTERNAL TABLE t STORED AS PARQUET LOCATION '/tmp/evo/';
   DESCRIBE t;
   -- id Int64 NO
   -- c  Int64 NO        <-- but b.parquet has no `c`
   SELECT * FROM t ORDER BY id;
   -- Arrow error: Invalid argument error: Column 'c' is declared as 
non-nullable but contains null values
   ```
   
   File listing order does not matter. `SELECT id FROM t` works, so the data is 
fine; only the inferred nullability is wrong.
   
   ### Expected behavior
   
   `c` should be inferred as nullable: a column missing from some files is read 
as null from them, so "required" cannot hold for the table as a whole. 
`DESCRIBE t` should show `c Int64 YES` and `SELECT *` should return `(1, 10), 
(2, NULL)`.
   
   ### Additional context
   
   `ParquetFormat::infer_schema` merges the per-file schemas with 
`Schema::try_merge`, which widens nullability only for fields that appear in 
more than one schema; a field present in a single file keeps that file's 
`required`. The other formats merge the same way, but Parquet is where required 
columns are common.
   
   Before DataFusion 52 the scan path tolerated this 
(`SchemaAdapter::map_batch` rebuilt batches under the table schema without 
nullability validation; removed in #18998). The strict check is the right 
behavior; the inferred schema is what is wrong. #21290 was my earlier attempt 
to describe this without a minimal reproducer; this is the reproducer.
   
   Fix sketch: in `infer_schema`, after merging, mark any field that is absent 
from at least one file as nullable. I have a patch with an slt test and can 
open the PR.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to