L. C. Hsieh created SPARK-59812:
-----------------------------------

             Summary: Validate value lengths and skip amounts in the vectorized 
PLAIN and DELTA_LENGTH_BYTE_ARRAY readers
                 Key: SPARK-59812
                 URL: https://issues.apache.org/jira/browse/SPARK-59812
             Project: Spark
          Issue Type: Bug
          Components: SQL
    Affects Versions: 5.0.0
            Reporter: L. C. Hsieh


Similar to SPARK-59352, the vectorized PLAIN and DELTA_LENGTH_BYTE_ARRAY 
readers use value lengths and skip amounts that come from the file without 
validating them. The read paths mostly fail on a corrupt page, but the skip 
paths (used when column-index filtering produces row ranges) do not:

* VectorizedDeltaLengthByteArrayReader.skipBinary loops on `remaining -= 
in.skip(remaining)`. When a length runs past the end of the page, 
ByteBufferInputStream.skip returns -1, so `remaining` grows until the int 
overflows (~2^31 iterations per value), then the read silently continues. A 
negative length is silently skipped.
* VectorizedPlainValuesReader.skipBinary calls `in.skip(len)` on a length read 
from the file. With SingleBufferInputStream a negative length moves the stream 
position backwards, so later values in the page are decoded from the wrong 
offset and the query returns wrong results without an error. A too-short skip 
is silently ignored.
* The fixed-width skips in VectorizedPlainValuesReader (skipIntegers, 
skipLongs, skipFloats, skipDoubles, skipBytes, skipShorts, 
skipFixedLenByteArray) ignore the return value of `in.skip`.
* On the read paths, a negative length is not rejected either: slicing a 
negative length also moves the position backwards, or fails with an 
ArrayIndexOutOfBoundsException instead of a ParquetDecodingException.

Note that ByteBufferInputStream.skipFully does not reject a negative length on 
SingleBufferInputStream, so negative lengths need an explicit check.

The fix rejects negative lengths on both the read and skip paths, and makes the 
skip paths fail with a ParquetDecodingException when the page does not have 
enough bytes, matching the read paths.




--
This message was sent by Atlassian Jira
(v8.20.10#820010)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to