[ https://issues.apache.org/jira/browse/SPARK-17916?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15591417#comment-15591417 ]
Hyukjin Kwon commented on SPARK-17916: -------------------------------------- FWIW, the codes I ran in Spark (equalvalent with the last example in R) is as blow. {code} spark.read.format("csv") .option("nullValue", "\"-\"") .option("quote", "") .option("header", "true") .load("path") .show() {code} {code} +----+----+ |col1|col2| +----+----+ | 1|null| | 2| ""| +----+----+ {code} > CSV data source treats empty string as null no matter what nullValue option is > ------------------------------------------------------------------------------ > > Key: SPARK-17916 > URL: https://issues.apache.org/jira/browse/SPARK-17916 > Project: Spark > Issue Type: Bug > Components: SQL > Affects Versions: 2.0.1 > Reporter: Hossein Falaki > > When user configures {{nullValue}} in CSV data source, in addition to those > values, all empty string values are also converted to null. > {code} > data: > col1,col2 > 1,"-" > 2,"" > {code} > {code} > spark.read.format("csv").option("nullValue", "-") > {code} > We will find a null in both rows. -- This message was sent by Atlassian JIRA (v6.3.4#6332) --------------------------------------------------------------------- To unsubscribe, e-mail: issues-unsubscr...@spark.apache.org For additional commands, e-mail: issues-h...@spark.apache.org