rustyconover opened a new issue, #1308:
URL: https://github.com/apache/arrow-java/issues/1308
### Summary
Reading a spec-canonical **sparse union** produced by another Arrow
implementation and then re-serializing it through arrow-java rewrites the
schema's union **type ids** from the wire values `[0, 1]` to the
`Types.MinorType` ordinals `[24, 3]`, while the `types` (type-id) buffer is
left unchanged at `[0, 1]`. The result is a self-inconsistent union: the buffer
selects type id `0`, which no longer appears in the field's declared type ids
`{24, 3}`.
Strict Arrow consumers reject the output. Arrow C++/pyarrow
`RecordBatch.validate(full=True)` fails with:
```
Invalid: In column 0: Invalid: Union value at position 0 has invalid type id 0
```
(Some readers, e.g. DuckDB, tolerate it by treating the type-id buffer
positionally, which masks the problem until a strict consumer sees it.)
### Version and platform
- arrow-java **18.1.0** and **19.0.0** — both reproduce (identical behavior).
- OpenJDK 25.0.2, macOS arm64.
### Steps to reproduce (pure arrow-java, no third-party runtime)
The input is a canonical sparse union `sparse_union<name: string=0, age:
int16=1>` with a 2-row type-id buffer `[0, 1]`, produced by pyarrow and
serialized to an Arrow IPC stream (base64 below). arrow-java reads it,
transfers the vectors into a fresh root (a generic pass-through), and writes it
back:
```java
byte[] input = Base64.getDecoder().decode(CANONICAL_UNION_IPC_B64);
try (BufferAllocator alloc = new RootAllocator()) {
VectorSchemaRoot outRoot;
try (ArrowStreamReader reader = new ArrowStreamReader(new
ByteArrayInputStream(input), alloc)) {
reader.loadNextBatch();
VectorSchemaRoot in = reader.getVectorSchemaRoot();
ArrowType.Union inType = (ArrowType.Union)
in.getSchema().getFields().get(0).getType();
System.out.println("INPUT Union.typeIds = " +
Arrays.toString(inType.getTypeIds())); // [0, 1]
List<FieldVector> moved = new ArrayList<>();
for (FieldVector v : in.getFieldVectors()) {
TransferPair tp = v.getTransferPair(alloc);
tp.transfer();
moved.add((FieldVector) tp.getTo());
}
outRoot = new VectorSchemaRoot(moved);
outRoot.setRowCount(in.getRowCount());
}
ByteArrayOutputStream sink = new ByteArrayOutputStream();
try (ArrowStreamWriter w = new ArrowStreamWriter(outRoot, null,
Channels.newChannel(sink))) {
w.start(); w.writeBatch(); w.end();
}
outRoot.close();
try (ArrowStreamReader reader = new ArrowStreamReader(new
ByteArrayInputStream(sink.toByteArray()), alloc)) {
reader.loadNextBatch();
Field f = reader.getVectorSchemaRoot().getSchema().getFields().get(0);
ArrowType.Union outType = (ArrowType.Union) f.getType();
System.out.println("OUTPUT Union.typeIds = " +
Arrays.toString(outType.getTypeIds())); // [24, 3]
}
}
```
Output:
```
INPUT Union.typeIds = [0, 1]
OUTPUT Union.typeIds = [24, 3] // Utf8=24, SmallInt=3 — Types.MinorType
ordinals
```
The type-id **buffer** is still `[0, 1]`, so type id `0` is no longer
declared.
`CANONICAL_UNION_IPC_B64` (a valid pyarrow-produced union IPC stream):
```
/////+AAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAyP///wQAAAABAAAABAAAAIj///8AAAEOGAAAACQAAAAEAAAAAgAAAHAAAAAoAAAAAQAAAHUAAAAIAAgAAAAEAAgAAAAEAAAAAgAAAAAAAAABAAAAzP///wAAAQIQAAAAHAAAAAQAAAAAAAAAAwAAAGFnZQAIAAwACAAHAAgAAAAAAAABEAAAABAAFAAIAAYABwAMAAAAEAAQAAAAAAABBRAAAAAcAAAABAAAAAAAAAAEAAAAbmFtZQAAAAAEAAQABAAAAP/////oAAAAFAAAAAAAAAAMABYABgAFAAgADAAMAAAAAAMEABgAAAAoAAAAAAAAAAAACgAYAAwABAAIAAoAAAB8AAAAEAAAAAIAAAAAAAAAAAAAAAYAAAAAAAAAAAAAAAIAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAgAAAAAAAAADAAAAAAAAAAYAAAAAAAAAAYAAAAAAAAAIAAAAAAAAAAAAAAAAAAAACAAAAAAAAAABAAAAAAAAAAAAAAAAwAAAAIAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAgAAAAAAAAAAAAAAAAAAAAABAAAAAAAAAAAAAAUAAAAGAAAAAAAAAEZyYW5reAAACQAFAAAAAAD/////AAAAAA==
```
Regenerate the input (Arrow C++/pyarrow), which round-trips it correctly:
```python
import pyarrow as pa, base64
children = [pa.array(["Frank", "x"], pa.utf8()), pa.array([9, 5],
pa.int16())]
type_ids = pa.array([0, 1], pa.int8())
u = pa.UnionArray.from_sparse(type_ids, children, field_names=["name",
"age"])
batch = pa.RecordBatch.from_arrays([u], names=["u"])
sink = pa.BufferOutputStream()
with pa.ipc.new_stream(sink, batch.schema) as w:
w.write_batch(batch)
print(base64.b64encode(sink.getvalue().to_pybytes()).decode())
# reading arrow-java's OUTPUT back in pyarrow raises:
# In column 0: Invalid: Union value at position 0 has invalid type id 0
```
### Expected
The re-serialized union should preserve the declared type ids `[0, 1]`
(matching the type-id buffer), i.e. `UnionVector.getField()` should report the
union's actual wire type ids rather than deriving them from
`Types.MinorType.ordinal()`.
### Notes
- The corruption is in the schema/`getField()` type-id derivation, not the
data buffers — `getField()` returns type ids based on `MinorType` ordinals
rather than the union's wire type ids.
- Constructing a `UnionVector` fresh via `setType(...)` and calling
`getField()` returns `Union(Sparse, [])` (empty type ids), a related
manifestation of the same root cause.
- Searched existing reports: ARROW-1692 (dense/sparse detection on read,
fixed 1.0.0) and ARROW-6145 (field metadata preservation, fixed 0.15.0) are
adjacent but distinct from this type-id derivation on serialize.
Reported by Rusty Conover (Query Farm — https://query.farm), found while
running a cross-implementation Arrow conformance suite.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]