ianmcook opened a new issue, #51779: URL: https://github.com/apache/arrow/issues/51779
> [!NOTE] > I discovered this issue and wrote it up with help from Claude Opus 5.5. ### Describe the bug, including details regarding any error messages, version, and platform. The format allows `DictionaryEncoding.indexType` to be omitted. [`format/Schema.fbs`](https://github.com/apache/arrow/blob/e85e181188ac3c5e8407b2160331faf5e78fd429/format/Schema.fbs#L500-L505) says: > If this field is null, the indices must be signed int32. The C++ IPC reader instead [requires the field](https://github.com/apache/arrow/blob/e85e181188ac3c5e8407b2160331faf5e78fd429/cpp/src/arrow/ipc/metadata_internal.cc#L899-L902): ```cpp auto int_data = encoding->indexType(); CHECK_FLATBUFFERS_NOT_NULL(int_data, "DictionaryEncoding.indexType"); ``` So PyArrow can't read an Arrow IPC stream whose dictionary-encoded field omits `indexType`: ``` OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata ``` #### Reproduction PyArrow and apache-arrow (JavaScript) both set `indexType` when they write, so this script embeds a stream without it. The reader fails on the schema, so the stream holds only a schema: one field, `text`, of type `dictionary<values=string, indices=int32>`, with `indexType` omitted. It was made from a stream that PyArrow wrote; the script that made it is below. ```python """Read an Arrow IPC stream whose DictionaryEncoding omits indexType.""" import base64 import pyarrow as pa # A schema-only stream with one field, text: dictionary-encoded strings, with # DictionaryEncoding.indexType omitted, so the indices are int32 per Schema.fbs. data = base64.b64decode( "/////5gAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAuP///wQAAAABAAAAFAAAABAAGAAI" "AAYABwAMABAAFAAQAAAAAAABBRQAAABEAAAAIAAAAAQAAAAAAAAABAAAAHRleHQAAAAACAAIAAAA" "BADc////DAAAAAgADAAIAAcACAAAAAAAAAEgAAAABAAEAAQAAAAIAAgAAAAAAP////8AAAAA" ) print(pa.ipc.open_stream(data).schema) ``` **Expected:** `text: dictionary<values=string, indices=int32, ordered=0>` **Actual:** `OSError: Unexpected null field DictionaryEncoding.indexType in flatbuffer-encoded metadata` apache-arrow (JavaScript) 21.2.0 reads the same stream as `Dictionary<Int32, Utf8>`, as the spec describes. The error is the same in PyArrow 12.0.1, 16.1.0, 20.0.0, 24.0.0, and 25.0.1 (Python 3.11 and 3.14, macOS), so this isn't a regression. <details> <summary>How the stream was made</summary> This script writes a schema-only stream with PyArrow, then omits `indexType` from its schema message. With PyArrow 25.0.1 it prints exactly the base64 above. ```python """Write the stream above: a schema-only stream from PyArrow, with indexType omitted.""" import base64 import struct import pyarrow as pa def table_at(buf, pos): """Return the (table, vtable) positions for the table referenced at pos.""" table = pos + struct.unpack_from("<I", buf, pos)[0] return table, table - struct.unpack_from("<i", buf, table)[0] def field_offset(buf, table, vtable, slot): """Return a field's offset within its table, or 0 if the field is absent.""" entry = 4 + 2 * slot present = entry < struct.unpack_from("<H", buf, vtable)[0] return struct.unpack_from("<H", buf, vtable + entry)[0] if present else 0 def without_index_type(stream): """Omit DictionaryEncoding.indexType from the first field of the schema message. Flatbuffers share identical vtables between tables (here the Schema's and the DictionaryEncoding's), so the encoding gets its own copy of its vtable, appended to the message, with indexType cleared. """ buf = bytearray(stream) length = struct.unpack_from("<i", buf, 4)[0] # After the 0xFFFFFFFF continuation marker meta = 8 message, message_vt = table_at(buf, meta) schema, schema_vt = table_at(buf, message + field_offset(buf, message, message_vt, 2)) # Message.header fields = schema + field_offset(buf, schema, schema_vt, 1) # Schema.fields field, field_vt = table_at(buf, fields + struct.unpack_from("<I", buf, fields)[0] + 4) # fields[0] encoding, encoding_vt = table_at(buf, field + field_offset(buf, field, field_vt, 4)) # Field.dictionary vtable = bytearray(buf[encoding_vt:encoding_vt + struct.unpack_from("<H", buf, encoding_vt)[0]]) struct.pack_into("<H", vtable, 4 + 2 * 1, 0) # DictionaryEncoding.indexType: absent struct.pack_into("<i", buf, encoding, encoding - (meta + length)) metadata = bytes(buf[meta:meta + length]) + bytes(vtable) metadata += bytes(-len(metadata) % 8) return struct.pack("<Ii", 0xFFFFFFFF, len(metadata)) + metadata + bytes(buf[meta + length:]) schema = pa.schema({"text": pa.dictionary(pa.int32(), pa.string())}) sink = pa.BufferOutputStream() with pa.ipc.new_stream(sink, schema): pass print(base64.b64encode(without_index_type(sink.getvalue().to_pybytes())).decode()) ``` </details> #### Suggested fix Default to int32 when the field is absent: ```cpp std::shared_ptr<DataType> index_type; auto int_data = encoding->indexType(); if (int_data == nullptr) { // Format/Schema.fbs: "If this field is null, the indices must be signed int32." index_type = int32(); } else { RETURN_NOT_OK(IntFromFlatbuffer(int_data, &index_type)); } ``` #### Related - I found this through an apache-arrow (JavaScript) writer bug that drops `indexType`, among other problems; that's reported separately in as https://github.com/apache/arrow-js/issues/496. This issue is only about reading a stream that is valid under the spec. - nanoarrow crashes on a stream like this one; I will report that separately in apache/arrow-nanoarrow. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
