L. C. Hsieh created SPARK-59902:
-----------------------------------
Summary: Detect BYTE_STREAM_SPLIT pages with more encoded values
than non-null values in the vectorized reader
Key: SPARK-59902
URL: https://issues.apache.org/jira/browse/SPARK-59902
Project: Spark
Issue Type: Improvement
Components: SQL
Affects Versions: 5.0.0
Reporter: L. C. Hsieh
Follow-up of SPARK-59831.
For a column with definition levels, the vectorized BYTE_STREAM_SPLIT reader
({{VectorizedByteStreamSplitValuesReader}}) cannot check the number of encoded
values against the page header when the page is initialized: the page value
count includes nulls, so it is only an upper bound. Let D be the number of
non-null definition levels of a page and E the number of encoded values.
* E < D fails with a {{ParquetDecodingException}} once a read or skip runs past
the last encoded value (since SPARK-59831).
* E > D is never detected, even when the page is read to its end. The values
are decoded with the wrong stream stride, so a corrupt data section can return
wrong values without an error. For example, an optional FLOAT page with
{{num_values = 10}}, 5 non-null levels and 24 data bytes (E = 6) is read as 5
values with a stride of 6.
This could be detected without {{DataPageV2}}'s {{num_nulls}} (which no reader
validates today): when {{VectorizedColumnReader.readBatch}} has consumed all
levels of a page ({{readState.valuesToReadInPage == 0}} right after the
{{defColumn.readBatch}} / {{repColumn.readBatchRepeated}} calls), all encoded
values must have been consumed too. A no-op default hook on
{{VectorizedValuesReader}}, implemented by this reader as {{offset ==
valueCount}}, would check it once per page. Checking only before the next
{{readPage()}} would miss the last page of a column chunk.
A prototype of this check found no false positive on 1454 pages of the
BYTE_STREAM_SPLIT tests in the Parquet suites (arrays, nullable columns and
column index row ranges, with v1 and v2 pages). Since a false positive would
reject valid files, it should be verified on more layouts (nested structs and
maps, row ranges on nested columns) before it is added. It is still a late
check, so a query that stops before the end of a page does not reach it.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]