qzyu999 opened a new issue, #3812:
URL: https://github.com/apache/iceberg-python/issues/3812

   ## Summary
   
   Before decomposing `pyiceberg/io/pyarrow.py` into focused submodules (#3737, 
#3738), we should consolidate PyArrow-specific logic that currently lives 
outside the module. This ensures all PyArrow calls route through a single 
boundary, making the subsequent split clean and enabling future engine 
substitution.
   
   ## Motivation
   
   Per discussion in #3737, @rambleraptor noted that the first useful step is 
ensuring no PyArrow logic occurs outside `pyarrow.py`. Currently several 
modules import `pyarrow` directly and implement compute logic inline rather 
than delegating through `pyiceberg.io.pyarrow`.
   
   When we later introduce a `ComputeEngine` protocol, any PyArrow logic 
outside the module boundary bypasses the protocol and prevents clean 
substitution.
   
   ## Audit
   
   Grepped `pyiceberg/` (excluding `io/pyarrow.py` and tests) for runtime 
`import pyarrow` statements (both top-level and inline). Excluded 
`TYPE_CHECKING`-only imports since those have no runtime dependency.
   
   | Location | What it does | Action |
   |----------|-------------|--------|
   | `table/upsert_util.py` | PyArrow table joins, group_by, compute, cast, 
take | **Absorb** |
   | `table/inspect.py` | Builds pa.schema + pa.Table.from_pylist for metadata 
inspection | **TBD** |
   | `transforms.py` | `pyarrow_transform()` dispatch on pa.Array/ChunkedArray 
| **TBD** |
   | `table/__init__.py` | Entry points accept pa.Table, delegate to io.pyarrow 
| **Leave** |
   | `table/deletion_vector.py` | Single pa.chunked_array() call | **Leave** |
   | `catalog/__init__.py` | Delegates to io.pyarrow for schema conversion | 
**Leave** |
   
   ## Plan
   
   One PR per absorption. Each is a pure refactor: move code into 
`io/pyarrow.py`, have the caller import from `pyiceberg.io.pyarrow` instead of 
`pyarrow` directly. No behavior change, all existing tests pass unchanged.
   
   - [ ] PR A: Absorb `table/upsert_util.py` PyArrow logic
   - [ ] PR B: `table/inspect.py` (pending discussion)
   - [ ] PR C: `transforms.py` (pending discussion)
   
   ## Related
   
   - #3737 - Decompose io/pyarrow.py into focused modules
   - #3738 - Extract PyArrowFileIO (first decomposition step)
   - #3715 / #3716 - Previous pluggable backend attempt (rejected as too large)
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to