ianmcook opened a new issue, #958:
URL: https://github.com/apache/arrow-nanoarrow/issues/958

   > [!NOTE]
   > I discovered this issue and wrote it up with help from Claude Opus 5.5.
   
   #957 fixed a crash when `DictionaryEncoding.indexType` is missing. The same 
kind of crash can still happen elsewhere. `ArrowIpcDecoderVerifyHeader` accepts 
messages where an optional table or union value is missing, and the decoder 
then dereferences NULL. Debug builds stop on flatcc's "null pointer table 
access" assertion; release builds segfault.
   
   Known cases:
   - `Field.type` is missing while `type_type` is set (Int, FloatingPoint, 
Timestamp, Union, Decimal, ...). This crashes in the `ArrowIpcDecoderSetType*` 
setters.
   - `Message.header` is missing while `header_type` is set. 
`ArrowIpcDecoderDecodeHeader` passes NULL to the header decoders, so a 
malformed stream can crash `ArrowIpcArrayStreamReader`.
   - `DictionaryBatch.data` is missing. `RecordBatch_nodes(NULL)` is called 
when the dictionary is decoded. This one is found by reading the code and not 
yet reproduced.
   
   Suggested fix: check `*_is_present` or `NULL` at each point and return 
`EINVAL` with an error message.
   
   Also worth adding: tests for a dictionary-encoded field with no `indexType` 
that go through verification and `ArrowIpcDecoderDecodeSchemaWithDictionaries`, 
and an end-to-end DictionaryBatch plus RecordBatch read with the defaulted 
int32 indices.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to