nhobin219 commented on issue #4084:
URL: 
https://github.com/apache/iceberg-python/issues/4084#issuecomment-6026649498

   **Narrowed to pyiceberg.** pyarrow on its own handles this case: 
`RecordBatch.filter` and a `pyarrow.dataset` scanner with the same filter both 
return the right 2 rows from the same file.
   
   The difference is pyiceberg's projection of a map with **struct values**. In 
`ArrowProjectionVisitor.map` (`pyiceberg/io/pyarrow.py`, 0.12.0), that case 
rebuilds the array instead of casting it:
   
   ```python
   if isinstance(value_result, pa.StructArray):
       # Arrow does not allow reordering of fields, therefore we have to copy 
the array :(
       return pa.MapArray.from_arrays(map_array.offsets, key_result, 
value_result, arrow_field)
   ```
   
   After the filter, `map_array` is a filtered (or sliced) `MapArray`. Its 
`offsets` index into the original child arrays, but `key_result` and 
`value_result` come from `map_array.keys` and `map_array.items`, which are the 
already-rebased children. The offsets and the children no longer line up, and 
`MapArray.from_arrays` builds an invalid array. `map<string, string>` never 
takes this path (it goes through `map_array.cast`), which is why it's fine.
   
   pyarrow's part is only that it aborts (`Check failed`) on the invalid input 
instead of raising.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to