eneskeles opened a new issue, #2084:
URL: https://github.com/apache/iceberg-go/issues/2084

   ### Apache Iceberg version
   
   main (development)
   
   ### Please describe the bug 🐞
   
   Filtered scans can silently drop rows. If a Parquet row group has exactly as 
many nulls as non-null values for the filtered column, the whole row group gets 
pruned.
   
   Repro on current `main`:
   
   ```go
   // append x = [1, 2, null, null]
   out, _ := 
tbl.Scan(table.WithRowFilter(iceberg.EqualTo(iceberg.Reference("x"), 
int64(1)))).ToArrowTable(ctx)
   // 0 rows, expected 1
   ```
   
   `[1, null]` also returns nothing. `[1, 2, null]` works fine. Same for `<`, 
`>`, `IN` and `IS NOT NULL`.
   
   It looks like the cause is here:
   
   
https://github.com/apache/iceberg-go/blob/96acf3731636165ad0f4ae431bb9589e3a360cf2/table/evaluators.go#L873
   
   arrow-go's `stats.NumValues()` excludes nulls (`column_chunk.go` builds it 
as `NumValues - NullCount`), but the spec defines `value_counts` as including 
nulls. So `containsNullsOnly` (`valCount == nullCount`) treats "2 values + 2 
nulls" as all nulls.
   
   It's easy to hit with paired records. I tried a double-entry ledger (one 
debit row and one credit row per transaction), and `credit IS NOT NULL` 
returned 0 out of 50k rows. Small files (e.g. streaming micro-batches) can hit 
it by chance too.
   
   Switching to `colMeta.NumValues()` fixes it (that's what 
`table/internal/parquet_files.go` already uses for data file metrics), and 
`./table/...` tests pass.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to