thomasp85 opened a new issue, #4844:
URL: https://github.com/apache/arrow-adbc/issues/4844
### What happened?
> Full disclosure. This bug was found while adding datafusion support to
ggsql through ADBC. The minimal example included here has been generated by an
llm.
## Summary
`adbc_datafusion` 0.23's bulk ingest (`bind` + `execute_update`) does not
validate or coerce the bound `RecordBatch` against the target table's
schema. A batch whose arrays don't match the declared column types is stored
as-is. The corruption surfaces later: `SELECT *` returns arrays that
disagree with the declared schema, and aggregates like `MIN`/`MAX` **panic**
(at an `unwrap` in `DataFusionReader::new`) instead of returning an ADBC
error.
## Actual behavior
```
thread 'main' panicked at adbc_datafusion-0.23.0/src/lib.rs:116:41:
called `Result::unwrap()` on an `Err` value:
Internal("MIN/MAX is not expected to receive scalars of incompatible types
(Int64(NULL), Int32(1))")
```
The panic site is `df.collect().await.unwrap()` in `DataFusionReader::new`
(lib.rs:116). DataFusion plans `MIN` from the declared schema (accumulator
state `Int64(NULL)`) but executes over the ingested Int32 arrays, and
`min_max`'s same-type invariant fails.
## Expected behavior
1. Ingest should either coerce the bound batch to the table schema or reject
the append with an ADBC error (e.g. `Status::InvalidArgument`) at ingest
time, the way the C driver implementations behave.
2. Independently, query execution errors should propagate as ADBC errors
rather than panic — the `unwrap` on `df.collect()` means any DataFusion
execution error is a process abort for embedders, including over FFI.
## Additional notes
- Control: the same query over the same Int32 data registered directly as a
`MemTable` in plain DataFusion 53.1.0 works fine, so this is not an
upstream DataFusion aggregate bug.
- Matching declared/bound types (`CREATE TABLE t (id INT)` + Int32 batch)
also works — the failure requires the mismatch.
- The mismatch class is broader than integers: `CREATE TABLE ... (s VARCHAR)`
followed by an ingested `Utf8` batch hits the same problem (DataFusion
declares `Utf8View`, the stored arrays stay `Utf8`), failing string
`MIN`/`MAX` the same way.
- Two smaller adjacent observations: `bind_stream` is still `todo!()`
(panics if used), and ingest into a nonexistent table fails with
`Plan("No table named ...")` because the implementation routes through
`DataFrame::write_table`, which requires an existing table in DataFusion
53 — worth documenting if auto-create isn't planned.
### Stack Trace
_No response_
### How can we reproduce the bug?
## Minimal reproducer
`Cargo.toml`:
```toml
[dependencies]
adbc_core = "0.23"
adbc_datafusion = "0.23"
arrow = "58"
```
`src/main.rs`:
```rust
use std::sync::Arc;
use adbc_core::options::{OptionStatement, OptionValue};
use adbc_core::{Connection, Database, Driver, Optionable, Statement};
use adbc_datafusion::DataFusionDriver;
use arrow::array::{Int32Array, RecordBatch};
use arrow::datatypes::{DataType, Field, Schema};
fn main() {
let mut driver = DataFusionDriver::new(None);
let db = driver.new_database().unwrap();
let mut conn = db.new_connection().unwrap();
// Declare the column as BIGINT (Int64)...
let mut stmt = conn.new_statement().unwrap();
stmt.set_sql_query("CREATE TABLE t (id BIGINT)").unwrap();
stmt.execute_update().unwrap();
// ...then ingest a batch whose array is Int32.
let schema = Arc::new(Schema::new(vec![Field::new("id", DataType::Int32,
false)]));
let batch =
RecordBatch::try_new(schema, vec![Arc::new(Int32Array::from(vec![1,
2, 3]))]).unwrap();
let mut stmt = conn.new_statement().unwrap();
stmt.set_option(OptionStatement::TargetTable,
OptionValue::String("t".into()))
.unwrap();
stmt.bind(batch).unwrap();
stmt.execute_update().unwrap(); // succeeds — no type checking
// SELECT * already returns Int32 arrays under the declared BIGINT
schema.
// Aggregating panics:
let mut stmt = conn.new_statement().unwrap();
stmt.set_sql_query("SELECT MIN(id) FROM t").unwrap();
let _ = stmt.execute().unwrap();
}
```
### Environment/Setup
- `adbc_datafusion` 0.23.0 (which pins `datafusion` 53.1.0)
- `adbc_core` 0.23.0, `arrow` 58
- Reproduced on macOS (aarch64) and Linux (x86_64, GitHub Actions)
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]