adonm opened a new issue, #3847:
URL: https://github.com/apache/iceberg-python/issues/3847

   ## Feature Request / Improvement
   
   PyIceberg doesn't expose Parquet page index writing, even though PyArrow 
supports it with `ParquetWriter(write_page_index=True)`.
   
   ### Use case / motivation
   
   Page indexes let readers such as ClickHouse skip non-matching pages during 
predicate evaluation, avoiding decoding work for selective queries. Tables 
written by PyIceberg currently lack these indexes, so page-level pruning is not 
possible regardless of reader support.
   
   ### Proposed change
   
   Add an opt-in `write.parquet.page-index-enabled` table property, default 
false, to preserve current output and file sizes. When enabled, PyIceberg 
passes `write_page_index=True` to PyArrow's ParquetWriter.
   
   ### Implementation
   
   PR #3829 adds the table property, threads it through 
`_get_parquet_writer_kwargs`, documents it in mkdocs/docs/configuration.md, and 
includes unit and integration test coverage.
   
   ### Tooling note
   
   I developed this with assistance from DS v4 Pro and reviewed the changes 
myself.
   
   ### References
   
   - #3829
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to