qzyu999 opened a new issue, #3812: URL: https://github.com/apache/iceberg-python/issues/3812
## Summary Before decomposing `pyiceberg/io/pyarrow.py` into focused submodules (#3737, #3738), we should consolidate PyArrow-specific logic that currently lives outside the module. This ensures all PyArrow calls route through a single boundary, making the subsequent split clean and enabling future engine substitution. ## Motivation Per discussion in #3737, @rambleraptor noted that the first useful step is ensuring no PyArrow logic occurs outside `pyarrow.py`. Currently several modules import `pyarrow` directly and implement compute logic inline rather than delegating through `pyiceberg.io.pyarrow`. When we later introduce a `ComputeEngine` protocol, any PyArrow logic outside the module boundary bypasses the protocol and prevents clean substitution. ## Audit Grepped `pyiceberg/` (excluding `io/pyarrow.py` and tests) for runtime `import pyarrow` statements (both top-level and inline). Excluded `TYPE_CHECKING`-only imports since those have no runtime dependency. | Location | What it does | Action | |----------|-------------|--------| | `table/upsert_util.py` | PyArrow table joins, group_by, compute, cast, take | **Absorb** | | `table/inspect.py` | Builds pa.schema + pa.Table.from_pylist for metadata inspection | **TBD** | | `transforms.py` | `pyarrow_transform()` dispatch on pa.Array/ChunkedArray | **TBD** | | `table/__init__.py` | Entry points accept pa.Table, delegate to io.pyarrow | **Leave** | | `table/deletion_vector.py` | Single pa.chunked_array() call | **Leave** | | `catalog/__init__.py` | Delegates to io.pyarrow for schema conversion | **Leave** | ## Plan One PR per absorption. Each is a pure refactor: move code into `io/pyarrow.py`, have the caller import from `pyiceberg.io.pyarrow` instead of `pyarrow` directly. No behavior change, all existing tests pass unchanged. - [ ] PR A: Absorb `table/upsert_util.py` PyArrow logic - [ ] PR B: `table/inspect.py` (pending discussion) - [ ] PR C: `transforms.py` (pending discussion) ## Related - #3737 - Decompose io/pyarrow.py into focused modules - #3738 - Extract PyArrowFileIO (first decomposition step) - #3715 / #3716 - Previous pluggable backend attempt (rejected as too large) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
