zianliu66-me commented on issue #51674:
URL: https://github.com/apache/arrow/issues/51674#issuecomment-5930301778
Apologies. It's not AI, I just have a tendency to write too verbosely and
haven't submitted many bug reports... please definitely let me know if what I
wrote below is not clear.
See minimal case below. This is analogous to the code I used to create the
problematic parquet file in the first place. This script actually fails
earlier, on `scanner.to_batches()`, with the same error message... But the
created parquet file can be read with `ParquetFile()` just fine:
```
import pyarrow as pa
import pyarrow.parquet as pq
import pyarrow.dataset as ds
table = pa.table({'a': range(500000), 'b': [10] * 500000, 'c': [1, 2] *
250000})
pq.write_table(table, 'example.parquet')
dataset = ds.dataset(
"example.parquet",
format="parquet"
)
scanner = dataset.scanner()
writer = None
try:
for batch in scanner.to_batches():
if writer is None:
writer = pq.ParquetWriter(
"example2.parquet",
batch.schema
)
writer.write_batch(batch)
finally:
if writer is not None:
writer.close()
```
The script above **consistently fails** on `for batch in
scanner.to_batches():` with the same errno 14 error message.
However, it consistently **does not fail** if I pre-emptively call
`print(scanner.to_batches())` prior to the while loop. It also does not fail if
the parquet is much smaller (`table = pa.table({'a': range(10), 'b': [10] * 10,
'c': [1, 2] * 5})` works).
Core issue: errno 14 seems to occur whenever I interact with *any*
sufficiently large parquet file on this system, not necessarily via
`pq.ParquetFile()`.
The issue can be bypassed by invoking some type of call to the file... In
the above snippet, pre-emptively calling `print(scanner.to_batches())` prior to
the for loop solves the issue. In my production code, simply calling
`pq.ParqeutFile()` twice bypasses the issue, as the second pq.ParquetFile()
command no longer returns an error.
This seems to work with any sufficiently large parquet file, but smaller
parquets don't have this issue.
PyArrow was installed via pip onto a `venv` by first running `python -m pip
install --upgrade pip setuptools wheel`, and then running `python -m pip
install -r requirements.txt`. `requirements.txt` contains at least the
following:
```
numpy
pandas
pyarrow
dask[array,dataframe,distributed]
zarr
datasets
```
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]