eneskeles opened a new issue, #2084:
URL: https://github.com/apache/iceberg-go/issues/2084
### Apache Iceberg version
main (development)
### Please describe the bug 🐞
Filtered scans can silently drop rows. If a Parquet row group has exactly as
many nulls as non-null values for the filtered column, the whole row group gets
pruned.
Repro on current `main`:
```go
// append x = [1, 2, null, null]
out, _ :=
tbl.Scan(table.WithRowFilter(iceberg.EqualTo(iceberg.Reference("x"),
int64(1)))).ToArrowTable(ctx)
// 0 rows, expected 1
```
`[1, null]` also returns nothing. `[1, 2, null]` works fine. Same for `<`,
`>`, `IN` and `IS NOT NULL`.
It looks like the cause is here:
https://github.com/apache/iceberg-go/blob/96acf3731636165ad0f4ae431bb9589e3a360cf2/table/evaluators.go#L873
arrow-go's `stats.NumValues()` excludes nulls (`column_chunk.go` builds it
as `NumValues - NullCount`), but the spec defines `value_counts` as including
nulls. So `containsNullsOnly` (`valCount == nullCount`) treats "2 values + 2
nulls" as all nulls.
It's easy to hit with paired records. I tried a double-entry ledger (one
debit row and one credit row per transaction), and `credit IS NOT NULL`
returned 0 out of 50k rows. Small files (e.g. streaming micro-batches) can hit
it by chance too.
Switching to `colMeta.NumValues()` fixes it (that's what
`table/internal/parquet_files.go` already uses for data file metrics), and
`./table/...` tests pass.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]