dannyshafer opened a new pull request, #51644:
URL: https://github.com/apache/arrow/pull/51644

   ### Rationale for this change
   
   The Python Parquet reader docstrings (`ParquetFile`, `ParquetDataset`, 
`read_table` and `read_pandas`) each repeated the descriptions of common 
options such as `thrift_string_size_limit`, and the copies had drifted slightly 
apart. Closes #51265.
   
   ### What changes are included in this PR?
   
   - The shared parameter descriptions now live in module-level snippets next 
to the existing `_read_docstring_common`: `_read_docstring_file_options` 
(read_dictionary through buffer_size), `_read_docstring_partitioning`, 
`_read_docstring_filesystem`, `_read_docstring_reader_options` (pre_buffer 
through schema_depth_limit) and `_read_docstring_checksum_options` 
(page_checksum_verification, arrow_extensions_enabled). 
`_read_docstring_common` is kept, built from the first two.
   - `ParquetFile.__doc__` becomes an f-string built from these snippets, the 
same way `ParquetDataset.__doc__` already is. That is why its body is dedented 
in the diff.
   - `ParquetDataset.__doc__`, `read_table.__doc__` and `read_pandas.__doc__` 
use the same snippets.
   - Where copies differed, the most complete wording was kept. For example, 
`read_table` now gets the full `pre_buffer` text (S3, GCS and the memory-usage 
note), `ParquetFile` gets the full `read_dictionary` description (I checked 
that `ParquetFile(..., read_dictionary=["s.x"])` does accept nested column 
paths), and `decryption_properties` mentions both Parquet Modular Encryption 
and `CryptoFactory.file_decryption_properties()`.
   
   ### Are these changes tested?
   
   This is a docstring-only change. I rendered the four docstrings before and 
after and diffed them: apart from the wording unifications above, the content 
is identical. I also ran the `ParquetFile` docstring examples with doctest and 
ran `flake8 --config python/setup.cfg` and `autopep8 --diff` on the file (both 
clean).
   
   ### Are there any user-facing changes?
   
   Only the rendered docstrings. They are now consistent across the four 
readers.
   
   ### Was AI used for this PR?
   
   In accordance to the [AI generation 
guidelines](https://arrow.apache.org/docs/dev/developers/overview.html#ai-generated-code),
 please disclose below whether and how AI was used in this PR.
   
   This PR was prepared with Claude Code, which wrote the refactor, compared 
the rendered docstrings before and after, and drafted this description.
   
   **PR code and description written by:**
   
   - [ ] Human
   - [x] AI
   
   **Reviewed before submission by:**
   
   - [ ] Human
   - [x] AI
   - [ ] Not reviewed
   
   <!-- oss-routine -->
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to