L. C. Hsieh created SPARK-59902:
-----------------------------------

             Summary: Detect BYTE_STREAM_SPLIT pages with more encoded values 
than non-null values in the vectorized reader
                 Key: SPARK-59902
                 URL: https://issues.apache.org/jira/browse/SPARK-59902
             Project: Spark
          Issue Type: Improvement
          Components: SQL
    Affects Versions: 5.0.0
            Reporter: L. C. Hsieh


Follow-up of SPARK-59831.

For a column with definition levels, the vectorized BYTE_STREAM_SPLIT reader 
({{VectorizedByteStreamSplitValuesReader}}) cannot check the number of encoded 
values against the page header when the page is initialized: the page value 
count includes nulls, so it is only an upper bound. Let D be the number of 
non-null definition levels of a page and E the number of encoded values.

* E < D fails with a {{ParquetDecodingException}} once a read or skip runs past 
the last encoded value (since SPARK-59831).
* E > D is never detected, even when the page is read to its end. The values 
are decoded with the wrong stream stride, so a corrupt data section can return 
wrong values without an error. For example, an optional FLOAT page with 
{{num_values = 10}}, 5 non-null levels and 24 data bytes (E = 6) is read as 5 
values with a stride of 6.

This could be detected without {{DataPageV2}}'s {{num_nulls}} (which no reader 
validates today): when {{VectorizedColumnReader.readBatch}} has consumed all 
levels of a page ({{readState.valuesToReadInPage == 0}} right after the 
{{defColumn.readBatch}} / {{repColumn.readBatchRepeated}} calls), all encoded 
values must have been consumed too. A no-op default hook on 
{{VectorizedValuesReader}}, implemented by this reader as {{offset == 
valueCount}}, would check it once per page. Checking only before the next 
{{readPage()}} would miss the last page of a column chunk.

A prototype of this check found no false positive on 1454 pages of the 
BYTE_STREAM_SPLIT tests in the Parquet suites (arrays, nullable columns and 
column index row ranges, with v1 and v2 pages). Since a false positive would 
reject valid files, it should be verified on more layouts (nested structs and 
maps, row ranges on nested columns) before it is added. It is still a late 
check, so a query that stops before the end of a page does not reach it.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to