ghanse commented on code in PR #58175:
URL: https://github.com/apache/spark/pull/58175#discussion_r3905480723


##########
docs/sql-data-sources-csv.md:
##########
@@ -225,7 +225,13 @@ Data source options of CSV can be set via:
   <tr>
     <td><code>singleVariantColumn</code></td>
     <td>(none)</td>
-    <td>If specified, the entire CSV record is parsed and stored as a single 
column of <code>VariantType</code> with the given column name, instead of being 
split into individual fields.</td>
+    <td>If specified, the entire CSV record is parsed and stored as a single 
column of <code>VariantType</code> with the given column name, instead of being 
split into individual fields. By default, scalar values are type-inferred 
inside the variant (for example, <code>"0001"</code> becomes the integer 
<code>1</code>). To keep every scalar as a string instead, set 
<code>variantRespectInferSchema</code> to <code>true</code> and 
<code>inferSchema</code> to <code>false</code>.</td>
+    <td>read</td>
+  </tr>
+  <tr>
+    <td><code>variantRespectInferSchema</code></td>

Review Comment:
   I chose this design as it's similar to the options for XML, but I see that 
the behavior can be unclear.
   
   My question would be how `variantScalarStrings` should interact with 
`inferSchema` if both are specified. 
   
   I would think `variantScalarStrings` should take precedence and convert the 
values to strings when it is set (regardless of the `inferSchema` setting). 
WDYT?



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to