thisisnic commented on issue #30481:
URL: https://github.com/apache/arrow/issues/30481#issuecomment-5868211338
This issue has been open since 2021, so apologies to everyone who's hit it
in the meantime. I've had Claude take a look at it, and the write-up below is
its findings, which it verified with the reprex at the end.
Thanks for the reports and especially for the `gcs.find()` output above,
which is what made this diagnosable. This is an Arrow bug in the fsspec adapter
rather than anything wrong with your datasets or with gcsfs.
When Spark writes a partitioned dataset to GCS, it creates a zero-byte
object for every directory, named with a trailing slash (e.g.
`bucket/dataset/partition_var=some_value/`). You can see these in the `find()`
output above: there's a `type: directory` entry for `<bucket>/partition/dir`
and a separate zero-byte `type: file` entry for `<bucket>/partition/dir/`.
gcsfs's `ls()` and `info()` know how to collapse these markers into
directories, but `find()` returns them as plain files.
PyArrow's `FSSpecHandler.get_file_info_selector` uses `find()` to discover
the dataset and passes every entry through as-is, so it tells the dataset code
that `.../partition_var=some_value/` is a file. Dataset discovery keeps it, and
when it later tries to open it to inspect the schema, `fs.isfile()` returns
`False` for a directory marker and we raise `FileNotFoundError` with that exact
path. That's the traceback shown above, with the path ending in a `/`.
The fix for #37555 in PyArrow 14 only filtered out the marker for the base
directory itself, which is why reading the top level started working for some
people while reading anything with nested partitions still fails. This also
explains why the fastparquet engine works: it doesn't go through this code path.
The native C++ `GcsFileSystem` (`pyarrow.fs.GcsFileSystem`) already treats
any object whose name ends in `/` as a directory, so switching to that instead
of gcsfs is a workaround in the meantime. The fix is to make the fsspec adapter
do the same, and I'll open a PR for that.
Here's a reprex that reproduces this without needing GCS, by making an
in-memory fsspec filesystem report markers the way gcsfs does:
```python
import pyarrow as pa
import pyarrow.parquet as pq
import pyarrow.dataset as ds
from pyarrow.fs import PyFileSystem, FSSpecHandler
from fsspec.implementations.memory import MemoryFileSystem
class MarkerMemoryFileSystem(MemoryFileSystem):
"""Reports zero-byte 'dir/' marker objects from find(), as gcsfs does."""
def find(self, path, maxdepth=None, withdirs=False, detail=False,
**kwargs):
out = super().find(path, maxdepth=maxdepth, withdirs=True,
detail=True, **kwargs)
out = {p.lstrip("/"): {**info, "name": p.lstrip("/")} for p, info in
out.items()}
for p, info in list(out.items()):
if info["type"] == "directory":
marker = p.rstrip("/") + "/"
out[marker] = {"name": marker, "size": 0, "type": "file"}
return out if detail else sorted(out)
fs = PyFileSystem(FSSpecHandler(MarkerMemoryFileSystem()))
with fs.open_output_stream("base/part=a/data.parquet") as f:
pq.write_table(pa.table({"x": [1, 2, 3]}), f)
ds.dataset("base", filesystem=fs, format="parquet",
partitioning="hive").to_table()
# FileNotFoundError: base/part=a/
```
One thing to separate out: a couple of comments here (and #31339) show a
different error, `GetFileInfo() yielded path '...' which is outside base dir
'gs://...'`. That happens when the path passed to `ParquetDataset` or
`dataset()` includes the `gs://` prefix while gcsfs returns paths without it.
It's a distinct bug from the directory marker one, and dropping the `gs://`
prefix from the path avoids it, so I'll keep it tracked in #31339 rather than
in this issue.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]