stackedsax commented on issue #47666:
URL: https://github.com/apache/arrow/issues/47666#issuecomment-5981279393

   This was fixed by #47741 (GH-47740), released in Arrow 23.0.0, which answers 
the open question on #47667 about where the `nullptr` came from.
   
   `DeltaByteArrayDecoderImpl::DecodeArrowDense` zero-initialised a 
`std::vector<ByteArray>` (so every entry was `{nullptr, 0}`), `GetInternal()` 
legitimately decoded zero values from the corrupted page, and only a 
`DCHECK_EQ` compared the expected and actual counts. In release builds that 
check is compiled out, so the untouched null entries reached 
`ArrowBinaryHelper<FLBAType>::AppendValue`, which memcpy'd `byte_width` bytes 
from address 0. #47741 replaced the `DCHECK` with a thrown 
`ParquetException("Expected to decode N values, but decoded M values")`, which 
rejects the page before any value reaches the helper.
   
   The exact OSS-Fuzz testcase 
(`clusterfuzz-testcase-minimized-parquet-arrow-fuzz-4656328221196288`) was 
added to arrow-testing in apache/arrow-testing#115 and is present in the 
`testing` submodule pinned on main, so it runs as part of the fuzz regression 
test in `arrow_reader_writer_test.cc`. On PyArrow 24.0.0 the file now fails 
cleanly with `OSError: Expected to decode 2084 values, but decoded 0 values.` 
instead of crashing.
   
   Since the root cause is fixed upstream of `AppendValue`, the null guard in 
#47667 would no longer be reachable.
   
   @thisisnic I think this can be closed as fixed by #47740, and #47667 closed 
along with it.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to