thisisnic commented on issue #30481:
URL: https://github.com/apache/arrow/issues/30481#issuecomment-5868211338

   This issue has been open since 2021, so apologies to everyone who's hit it 
in the meantime. I've had Claude take a look at it, and the write-up below is 
its findings, which it verified with the reprex at the end.
   
   Thanks for the reports and especially for the `gcs.find()` output above, 
which is what made this diagnosable. This is an Arrow bug in the fsspec adapter 
rather than anything wrong with your datasets or with gcsfs.
   
   When Spark writes a partitioned dataset to GCS, it creates a zero-byte 
object for every directory, named with a trailing slash (e.g. 
`bucket/dataset/partition_var=some_value/`). You can see these in the `find()` 
output above: there's a `type: directory` entry for `<bucket>/partition/dir` 
and a separate zero-byte `type: file` entry for `<bucket>/partition/dir/`. 
gcsfs's `ls()` and `info()` know how to collapse these markers into 
directories, but `find()` returns them as plain files.
   
   PyArrow's `FSSpecHandler.get_file_info_selector` uses `find()` to discover 
the dataset and passes every entry through as-is, so it tells the dataset code 
that `.../partition_var=some_value/` is a file. Dataset discovery keeps it, and 
when it later tries to open it to inspect the schema, `fs.isfile()` returns 
`False` for a directory marker and we raise `FileNotFoundError` with that exact 
path. That's the traceback shown above, with the path ending in a `/`.
   
   The fix for #37555 in PyArrow 14 only filtered out the marker for the base 
directory itself, which is why reading the top level started working for some 
people while reading anything with nested partitions still fails. This also 
explains why the fastparquet engine works: it doesn't go through this code path.
   
   The native C++ `GcsFileSystem` (`pyarrow.fs.GcsFileSystem`) already treats 
any object whose name ends in `/` as a directory, so switching to that instead 
of gcsfs is a workaround in the meantime. The fix is to make the fsspec adapter 
do the same, and I'll open a PR for that.
   
   Here's a reprex that reproduces this without needing GCS, by making an 
in-memory fsspec filesystem report markers the way gcsfs does:
   
   ```python
   import pyarrow as pa
   import pyarrow.parquet as pq
   import pyarrow.dataset as ds
   from pyarrow.fs import PyFileSystem, FSSpecHandler
   from fsspec.implementations.memory import MemoryFileSystem
   
   
   class MarkerMemoryFileSystem(MemoryFileSystem):
       """Reports zero-byte 'dir/' marker objects from find(), as gcsfs does."""
   
       def find(self, path, maxdepth=None, withdirs=False, detail=False, 
**kwargs):
           out = super().find(path, maxdepth=maxdepth, withdirs=True, 
detail=True, **kwargs)
           out = {p.lstrip("/"): {**info, "name": p.lstrip("/")} for p, info in 
out.items()}
           for p, info in list(out.items()):
               if info["type"] == "directory":
                   marker = p.rstrip("/") + "/"
                   out[marker] = {"name": marker, "size": 0, "type": "file"}
           return out if detail else sorted(out)
   
   
   fs = PyFileSystem(FSSpecHandler(MarkerMemoryFileSystem()))
   with fs.open_output_stream("base/part=a/data.parquet") as f:
       pq.write_table(pa.table({"x": [1, 2, 3]}), f)
   
   ds.dataset("base", filesystem=fs, format="parquet", 
partitioning="hive").to_table()
   # FileNotFoundError: base/part=a/
   ```
   
   One thing to separate out: a couple of comments here (and #31339) show a 
different error, `GetFileInfo() yielded path '...' which is outside base dir 
'gs://...'`. That happens when the path passed to `ParquetDataset` or 
`dataset()` includes the `gs://` prefix while gcsfs returns paths without it. 
It's a distinct bug from the directory marker one, and dropping the `gs://` 
prefix from the path avoids it, so I'll keep it tracked in #31339 rather than 
in this issue.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to