mrutunjay-kinagi opened a new pull request, #58977:
URL: https://github.com/apache/spark/pull/58977

   ### What changes were proposed in this pull request?
   
   The name of the corrupt record column was compared to schema field names 
with exact
   string equality, so `spark.sql.caseSensitive` was ignored. This PR routes 
those
   comparisons through `SQLConf.resolver`, the same resolution the analyzer 
uses for
   every other identifier.
   
   Three helpers are added to `ExprUtils`, next to the existing
   `verifyColumnNameOfCorruptRecord`:
   
   - `isCorruptRecordColumn(fieldName, columnNameOfCorruptRecord)` wraps
     `SQLConf.resolver`.
   - `schemaWithoutCorruptRecordColumn(schema, columnNameOfCorruptRecord)` 
builds the
     schema handed to the record parser, which is the most common use.
   - `corruptRecordFieldIndex(schema, columnNameOfCorruptRecord)` resolves the 
index the
     raw record is written to.
   
   `verifyColumnNameOfCorruptRecord` and `FailureSafeParser` now use these, 
along with
   the CSV and JSON readers in both the v1 and v2 paths and the `from_csv` / 
`from_json`
   expressions.
   
   XML is deliberately left alone. It carries its own case handling from 
SPARK-45844 in
   `StaxXmlParser`, so auditing it belongs in a separate change rather than 
being folded
   in here.
   
   ### Why are the changes needed?
   
   `spark.sql.caseSensitive` is `false` by default, so a schema that spells the 
field
   `_CORRUPT_RECORD` while the option holds the default `_corrupt_record` is a 
natural
   thing to write. Today that combination does not raise an error. Instead:
   
   - The column is not recognised as the corrupt record column, so it is passed 
to the
     record parser as an ordinary data column. For CSV this changes the 
expected token
     count, which makes every record malformed, including well-formed ones.
   - The raw text of records that genuinely failed to parse is dropped rather 
than being
     captured, which is the data loss the user notices.
   - `verifyColumnNameOfCorruptRecord` does not find the field either, so its 
check that
     the column is a nullable string is skipped. A schema declaring
     `_CORRUPT_RECORD INT` is accepted instead of being rejected with
     `INVALID_CORRUPT_RECORD_TYPE`.
   
   Reading a column name case-sensitively is the anomaly here. Header-to-schema 
matching
   in CSV already honours the conf (SPARK-23786), and XML was given the same 
treatment
   in 4.0.0 (SPARK-45844).
   
   ### Does this PR introduce _any_ user-facing change?
   
   Yes, under the default `spark.sql.caseSensitive=false`.
   
   A schema field whose name matches `columnNameOfCorruptRecord` apart from 
case is now
   treated as the corrupt record column. It receives the raw text of malformed 
records
   and is excluded from the schema given to the parser, where previously it was 
parsed
   as data and left null. If such a field is not a nullable string, the query 
now fails
   with `INVALID_CORRUPT_RECORD_TYPE` rather than silently doing nothing.
   
   Setting `spark.sql.caseSensitive=true` keeps the previous exact-match 
behaviour.
   
   ### How was this patch tested?
   
   New tests, all run against both the v1 and v2 paths:
   
   - `CSVSuite`: a schema declaring `_UNPARSED` against the option `_unparsed` 
captures
     the malformed record and leaves the well-formed record untouched under
     `caseSensitive=false`, and does neither under `caseSensitive=true`.
   - `JsonSuite`: the same for JSON, plus a case asserting that a wrongly-cased 
corrupt
     record column declared as `INT` is rejected with 
`INVALID_CORRUPT_RECORD_TYPE`,
     which was previously skipped.
   
   Existing `CSVSuite`, `JsonSuite`, `CsvFunctionsSuite` and 
`JsonFunctionsSuite` were
   run to check the exact-match paths are unaffected.
   
   ### Was this patch authored or co-authored using generative AI tooling?
   
   Assisted by Claude Code (Opus 5), used to investigate the reported 
behaviour, trace it
   to the comparison sites, and write the tests.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to