ianmcook opened a new issue, #958: URL: https://github.com/apache/arrow-nanoarrow/issues/958
> [!NOTE] > I discovered this issue and wrote it up with help from Claude Opus 5.5. #957 fixed a crash when `DictionaryEncoding.indexType` is missing. The same kind of crash can still happen elsewhere. `ArrowIpcDecoderVerifyHeader` accepts messages where an optional table or union value is missing, and the decoder then dereferences NULL. Debug builds stop on flatcc's "null pointer table access" assertion; release builds segfault. Known cases: - `Field.type` is missing while `type_type` is set (Int, FloatingPoint, Timestamp, Union, Decimal, ...). This crashes in the `ArrowIpcDecoderSetType*` setters. - `Message.header` is missing while `header_type` is set. `ArrowIpcDecoderDecodeHeader` passes NULL to the header decoders, so a malformed stream can crash `ArrowIpcArrayStreamReader`. - `DictionaryBatch.data` is missing. `RecordBatch_nodes(NULL)` is called when the dictionary is decoded. This one is found by reading the code and not yet reproduced. Suggested fix: check `*_is_present` or `NULL` at each point and return `EINVAL` with an error message. Also worth adding: tests for a dictionary-encoded field with no `indexType` that go through verification and `ArrowIpcDecoderDecodeSchemaWithDictionaries`, and an end-to-end DictionaryBatch plus RecordBatch read with the defaulted int32 indices. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
