mrutunjay-kinagi opened a new pull request, #58977:
URL: https://github.com/apache/spark/pull/58977
### What changes were proposed in this pull request?
The name of the corrupt record column was compared to schema field names
with exact
string equality, so `spark.sql.caseSensitive` was ignored. This PR routes
those
comparisons through `SQLConf.resolver`, the same resolution the analyzer
uses for
every other identifier.
Three helpers are added to `ExprUtils`, next to the existing
`verifyColumnNameOfCorruptRecord`:
- `isCorruptRecordColumn(fieldName, columnNameOfCorruptRecord)` wraps
`SQLConf.resolver`.
- `schemaWithoutCorruptRecordColumn(schema, columnNameOfCorruptRecord)`
builds the
schema handed to the record parser, which is the most common use.
- `corruptRecordFieldIndex(schema, columnNameOfCorruptRecord)` resolves the
index the
raw record is written to.
`verifyColumnNameOfCorruptRecord` and `FailureSafeParser` now use these,
along with
the CSV and JSON readers in both the v1 and v2 paths and the `from_csv` /
`from_json`
expressions.
XML is deliberately left alone. It carries its own case handling from
SPARK-45844 in
`StaxXmlParser`, so auditing it belongs in a separate change rather than
being folded
in here.
### Why are the changes needed?
`spark.sql.caseSensitive` is `false` by default, so a schema that spells the
field
`_CORRUPT_RECORD` while the option holds the default `_corrupt_record` is a
natural
thing to write. Today that combination does not raise an error. Instead:
- The column is not recognised as the corrupt record column, so it is passed
to the
record parser as an ordinary data column. For CSV this changes the
expected token
count, which makes every record malformed, including well-formed ones.
- The raw text of records that genuinely failed to parse is dropped rather
than being
captured, which is the data loss the user notices.
- `verifyColumnNameOfCorruptRecord` does not find the field either, so its
check that
the column is a nullable string is skipped. A schema declaring
`_CORRUPT_RECORD INT` is accepted instead of being rejected with
`INVALID_CORRUPT_RECORD_TYPE`.
Reading a column name case-sensitively is the anomaly here. Header-to-schema
matching
in CSV already honours the conf (SPARK-23786), and XML was given the same
treatment
in 4.0.0 (SPARK-45844).
### Does this PR introduce _any_ user-facing change?
Yes, under the default `spark.sql.caseSensitive=false`.
A schema field whose name matches `columnNameOfCorruptRecord` apart from
case is now
treated as the corrupt record column. It receives the raw text of malformed
records
and is excluded from the schema given to the parser, where previously it was
parsed
as data and left null. If such a field is not a nullable string, the query
now fails
with `INVALID_CORRUPT_RECORD_TYPE` rather than silently doing nothing.
Setting `spark.sql.caseSensitive=true` keeps the previous exact-match
behaviour.
### How was this patch tested?
New tests, all run against both the v1 and v2 paths:
- `CSVSuite`: a schema declaring `_UNPARSED` against the option `_unparsed`
captures
the malformed record and leaves the well-formed record untouched under
`caseSensitive=false`, and does neither under `caseSensitive=true`.
- `JsonSuite`: the same for JSON, plus a case asserting that a wrongly-cased
corrupt
record column declared as `INT` is rejected with
`INVALID_CORRUPT_RECORD_TYPE`,
which was previously skipped.
Existing `CSVSuite`, `JsonSuite`, `CsvFunctionsSuite` and
`JsonFunctionsSuite` were
run to check the exact-match paths are unaffected.
### Was this patch authored or co-authored using generative AI tooling?
Assisted by Claude Code (Opus 5), used to investigate the reported
behaviour, trace it
to the comparison sites, and write the tests.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]