HyukjinKwon opened a new pull request #26958: 
[SPARK-30128][SPARK-27627][SPARK-27990][SPARK-11412][FOLLOW-UP] 
Document/promote 'recursiveFileLookup' and 'pathGlobFilter' in file sources 
'mergeSchema' in ORC
URL: https://github.com/apache/spark/pull/26958
 
 
   ### What changes were proposed in this pull request?
   
   This PR adds the options, 'recursiveFileLookup' and 'pathGlobFilter' in file 
sources 'mergeSchema' in ORC, into documentation.
   
   - `recursiveFileLookup` at file sources: 
https://github.com/apache/spark/pull/24830 
([SPARK-27627](https://issues.apache.org/jira/browse/SPARK-27627))
   - `pathGlobFilter` at file sources: 
https://github.com/apache/spark/pull/24518 
([SPARK-27990](https://issues.apache.org/jira/browse/SPARK-27990))
   - `mergeSchema` at ORC: https://github.com/apache/spark/pull/24043 
([SPARK-11412](https://issues.apache.org/jira/browse/SPARK-11412))
   
   **Note that** `timeZone` option was not moved from `DataFrameReader.options` 
as I assume it will likely affect other datasources as well once DSv2 is 
complete.
   
   ### Why are the changes needed?
   
   To document available options in sources properly.
   
   ### Does this PR introduce any user-facing change?
   
   In PySpark, `pathGlobFilter` can be set via 
`DataFrameReader.(text|orc|parquet|json|csv)` and 
`DataStreamReader.(text|orc|parquet|json|csv)`.
   
   ### How was this patch tested?
   
   Manually built the doc and checked the output. Option setting in PySpark is 
rather a logical change. I manually tested one only:
   
   ```bash
   $ ls -al tmp
   ...
   -rw-r--r--   1 hyukjin.kwon  staff     3 Dec 20 12:19 aa
   -rw-r--r--   1 hyukjin.kwon  staff     3 Dec 20 12:19 ab
   -rw-r--r--   1 hyukjin.kwon  staff     3 Dec 20 12:19 ac
   -rw-r--r--   1 hyukjin.kwon  staff     3 Dec 20 12:19 cc
   ```
   
   ```python
   >>> spark.read.text("tmp", pathGlobFilter="*c").show()
   ```
   
   ```
   +-----+
   |value|
   +-----+
   |   ac|
   |   cc|
   +-----+
   ```

----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
[email protected]


With regards,
Apache Git Services

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to