sumitsu opened a new pull request #23674: [SPARK-26745][SQL] JsonSuite test 
case: empty line -> 0 record count
URL: https://github.com/apache/spark/pull/23674
 
 
   ## What changes were proposed in this pull request?
   
   This PR consists of the `test` components of #23665 only, minus the 
associated patch from that PR.
   
   It adds a new unit test to `JsonSuite` which verifies that the `count()` 
returned from a `DataFrame` loaded from JSON containing empty lines does not 
include those empty lines in the record count. The test runs `count` prior to 
otherwise reading data from the `DataFrame`, so as to catch future cases where 
a pre-parsing optimization might result in `count` results inconsistent with 
existing behavior.
   
   This PR is intended to be deployed alongside #23667; `master` currently 
causes the test to fail, as described in 
[SPARK-26745](https://issues.apache.org/jira/browse/SPARK-26745).
   
   ## How was this patch tested?
   
   Manual testing, existing `JsonSuite` unit tests.

----------------------------------------------------------------
This is an automated message from the Apache Git Service.
To respond to the message, please log on GitHub and use the
URL above to go to the specific comment.
 
For queries about this service, please contact Infrastructure at:
[email protected]


With regards,
Apache Git Services

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to