zianliu66-me commented on issue #51674:
URL: https://github.com/apache/arrow/issues/51674#issuecomment-5930301778

   Apologies. It's not AI, I just have a tendency to write too verbosely and 
haven't submitted many bug reports... please definitely let me know if what I 
wrote below is not clear.
   
   See minimal case below. This is analogous to the code I used to create the 
problematic parquet file in the first place. This script actually fails 
earlier, on `scanner.to_batches()`, with the same error message... But the 
created parquet file can be read with `ParquetFile()` just fine:
   
   ```
   import pyarrow as pa
   import pyarrow.parquet as pq
   import pyarrow.dataset as ds
   
   table = pa.table({'a': range(500000), 'b': [10] * 500000, 'c': [1, 2] * 
250000})
   pq.write_table(table, 'example.parquet')
   
   dataset = ds.dataset(
       "example.parquet",
       format="parquet"
   )
   
   scanner = dataset.scanner()
   writer = None
   try:
       for batch in scanner.to_batches():
           if writer is None:
               writer = pq.ParquetWriter(
                   "example2.parquet",
                   batch.schema
               )
           writer.write_batch(batch)
   finally:
       if writer is not None:
           writer.close()
   ```
   
   The script above **consistently fails** on `for batch in 
scanner.to_batches():` with the same errno 14 error message. 
   
   However, it consistently **does not fail** if I pre-emptively call 
`print(scanner.to_batches())` prior to the while loop. It also does not fail if 
the parquet is much smaller (`table = pa.table({'a': range(10), 'b': [10] * 10, 
'c': [1, 2] * 5})` works).
   
   Core issue: errno 14 seems to occur whenever I interact with *any* 
sufficiently large parquet file on this system, not necessarily via 
`pq.ParquetFile()`. 
   
   The issue can be bypassed by invoking some type of call to the file... In 
the above snippet, pre-emptively calling `print(scanner.to_batches())` prior to 
the for loop solves the issue. In my production code, simply calling 
`pq.ParqeutFile()` twice bypasses the issue, as the second pq.ParquetFile() 
command no longer returns an error.
   
   This seems to work with any sufficiently large parquet file, but smaller 
parquets don't have this issue.
   
   PyArrow was installed via pip onto a `venv` by first running `python -m pip 
install --upgrade pip setuptools wheel`, and then running `python -m pip 
install -r requirements.txt`. `requirements.txt` contains at least the 
following:
   ```
   numpy
   pandas
   pyarrow
   dask[array,dataframe,distributed]
   zarr
   datasets
   ```


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to