quinncheong commented on issue #4079:
URL:
https://github.com/apache/iceberg-python/issues/4079#issuecomment-6008441752
Two details from reproducing the first two items on 0.12.0 and on `main`
(068aae5), in case they help scope the work. Line numbers are from `main`.
**Item 1 (unquoted form) also needs the emit side.** Reading is only half of
it: PyIceberg writes the quoted form, which is not the spec form
(`"geometry(srid:4326)"`, `"geography(srid:4326, spherical)"` in the JSON
serialization table of format/spec.md):
```python
from pyiceberg.schema import Schema
from pyiceberg.types import GeographyType, GeometryType, NestedField
for t in ["geometry('srid:4326')", "geometry(srid:4326)",
"geometry(EPSG:7415)", "geography(srid:4326, spherical)"]:
try:
print(f"{t!r:36} -> {NestedField(1, 'g', t,
required=False).field_type!r}")
except Exception as e:
print(f"{t!r:36} -> {type(e).__name__}: {str(e).splitlines()[0]}")
print("emitted:", GeometryType("srid:4326").model_dump_json(),
GeographyType("srid:4326", "planar").model_dump_json())
Schema.model_validate_json('{"type":"struct","schema-id":0,"fields":[{"id":1,"name":"geom","required":false,"type":"geometry(srid:4326)"}]}')
```
```
"geometry('srid:4326')" -> GeometryType(crs='srid:4326')
'geometry(srid:4326)' -> ValidationError: Could not parse
geometry(srid:4326) into a GeometryType
'geometry(EPSG:7415)' -> ValidationError: Could not parse
geometry(EPSG:7415) into a GeometryType
'geography(srid:4326, spherical)' -> ValidationError: Could not parse
geography(srid:4326, spherical) into a GeographyType
emitted: "geometry('srid:4326')" "geography('srid:4326', 'planar')"
pyiceberg.exceptions.ValidationError: Could not parse geometry(srid:4326)
into a GeometryType
```
Java's `GEOMETRY_PARAMETERS` pattern (`Types.java` lines 67-71 on
apache/iceberg main) does not strip quotes, so it reads PyIceberg's output as
CRS `'srid:4326'` with the quotes included. Java writes `geometry(%s)` /
`geography(%s, %s)` (lines 632 and 709). This matches the split @moomindani
proposed on #3530: read side there, emit side as a follow-up.
**Item 2 is a hard error, not just plain binary.** Because the file column
resolves to `BinaryType`, `_cast_if_needed` calls `promote(BinaryType,
GeometryType)`, and that raises (pyiceberg/io/pyarrow.py line 2036,
pyiceberg/schema.py line 1691). It happens on write as well as read. Without
geoarrow-pyarrow, writing a binary or large_binary WKB column through
PyIceberg's own writer fails like this:
```python
import tempfile, uuid
import pyarrow as pa
from pyiceberg.io.pyarrow import PyArrowFileIO, _dataframe_to_data_files
from pyiceberg.partitioning import UNPARTITIONED_PARTITION_SPEC
from pyiceberg.schema import Schema
from pyiceberg.table.metadata import new_table_metadata
from pyiceberg.table.sorting import UNSORTED_SORT_ORDER
from pyiceberg.types import GeometryType, LongType, NestedField
schema = Schema(NestedField(1, "id", LongType(), required=False),
NestedField(2, "geom", GeometryType(), required=False))
meta = new_table_metadata(schema, UNPARTITIONED_PARTITION_SPEC,
UNSORTED_SORT_ORDER, "file://" + tempfile.mkdtemp(),
properties={"format-version": "3"})
df = pa.table({"id": pa.array([1], pa.int64()),
"geom":
pa.array([bytes.fromhex("0101000000000000000000f03f0000000000000040")],
pa.large_binary())})
list(_dataframe_to_data_files(table_metadata=meta, df=df,
io=PyArrowFileIO(), write_uuid=uuid.uuid4()))
```
```
File "pyiceberg/io/pyarrow.py", line 1937, in _to_requested_schema
File "pyiceberg/io/pyarrow.py", line 2036, in _cast_if_needed
File "pyiceberg/schema.py", line 1691, in _
pyiceberg.exceptions.ResolveError: Cannot promote an binary to geometry
```
Scanning a Parquet file that stores WKB as a plain BYTE_ARRAY column with
the matching field id fails with the same `ResolveError` in
`ArrowScan.to_table`. With geoarrow-pyarrow installed, a `geoarrow.wkb` input
column is rejected earlier with `UnsupportedPyArrowTypeException: Column 'geom'
has an unsupported type: extension<geoarrow.wkb<WkbType>>`. So as far as I can
tell, no input form round-trips today. That differs from
mkdocs/docs/geospatial.md, which says that without geoarrow-pyarrow "geometry
and geography are written as binary in Parquet while the Iceberg schema still
preserves the spatial type". An end-to-end write/read test for each of the
three column forms would catch this. The existing geo tests in
tests/io/test_pyarrow.py only cover the type conversion.
I have a small patch for item 1, read and emit side. It makes the quotes
optional in `GEOMETRY_REGEX` / `GEOGRAPHY_REGEX`, keeps accepting the quoted
form so older type strings still parse, emits the unquoted form, and adds
parametrized tests to tests/test_types.py (308 pass; the new tests fail on
main). It does not yet do the case-insensitive matching. I'm happy to open it
as the emit-side follow-up once #3530 settles the read side, or to leave it to
@moomindani if that's already in progress.
This report was prepared with an AI coding agent (Claude Code) and
reproduced before posting by me.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]