jorisvandenbossche opened a new issue, #51155:
URL: https://github.com/apache/arrow/issues/51155

   In e.g. `pd.read_parquet` implementation, we currently wrap `pq.read_table`. 
However, I would like to use the dataset implementation to get access to some 
more features (e.g. reading only the first few rows, defining a filter _after_ 
opening (discovering the schema) the dataset).
   
   That runs into some issues:
   
   - `pq.read_table()` uses ParquetDataset under the hood, but still has a 
fallback to plain `ParquetFile().read()` when `pyarrow.dataset` module is not 
available. While this will definitely be uncommon, I have no clue how important 
it is to keep supporting this (also on the pandas side, since this is currently 
a "feature" of `pandas.read_parquet` by means of using `pq.read_table`).
   - I would prefer using `pq.ParquetDataset` over `pyarrow.dataset`, because 
the parquet version is compatible with `pq.read_table` and does all the 
translation to the `pyarrow.dataset` API for us (constructing the file format 
object, passing the various keywords in the correct place, mapping some naming 
differences, etc)
   - But `ParquetDataset` is also limited, and currently I am essentially using 
it as a constructor to then access the underlying `._dataset` (the 
`pyarrow.dataset` dataset object)
   
   
   Some ideas that I was having:
   
   - Add a new `pq.open_dataset()`-like function that also has the boilerplate 
to construct the `pyarrow.dataset` object, but returns that instead of the 
`ParquetDataset` wrapper
   - Would it be technically possible to let `pq.ParquetDataset` inherit from 
`pyarrow.dataset.Dataset`, so you get those methods that way, while keeping it 
back-compat
   - Add new methods to `pq.ParquetDataset` (`to_table()`, `head()`, etc) to 
make it look more like a `pyarrow.dataset.Dataset`, but without actually 
inheriting from it
   - Simply "officially" expose the underlying dataset, so that I don't have to 
use the private `_dataset` attribute
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to