rustyconover opened a new issue, #1308:
URL: https://github.com/apache/arrow-java/issues/1308

   ### Summary
   
   Reading a spec-canonical **sparse union** produced by another Arrow 
implementation and then re-serializing it through arrow-java rewrites the 
schema's union **type ids** from the wire values `[0, 1]` to the 
`Types.MinorType` ordinals `[24, 3]`, while the `types` (type-id) buffer is 
left unchanged at `[0, 1]`. The result is a self-inconsistent union: the buffer 
selects type id `0`, which no longer appears in the field's declared type ids 
`{24, 3}`.
   
   Strict Arrow consumers reject the output. Arrow C++/pyarrow 
`RecordBatch.validate(full=True)` fails with:
   
   ```
   Invalid: In column 0: Invalid: Union value at position 0 has invalid type id 0
   ```
   
   (Some readers, e.g. DuckDB, tolerate it by treating the type-id buffer 
positionally, which masks the problem until a strict consumer sees it.)
   
   ### Version and platform
   
   - arrow-java **18.1.0** and **19.0.0** — both reproduce (identical behavior).
   - OpenJDK 25.0.2, macOS arm64.
   
   ### Steps to reproduce (pure arrow-java, no third-party runtime)
   
   The input is a canonical sparse union `sparse_union<name: string=0, age: 
int16=1>` with a 2-row type-id buffer `[0, 1]`, produced by pyarrow and 
serialized to an Arrow IPC stream (base64 below). arrow-java reads it, 
transfers the vectors into a fresh root (a generic pass-through), and writes it 
back:
   
   ```java
   byte[] input = Base64.getDecoder().decode(CANONICAL_UNION_IPC_B64);
   try (BufferAllocator alloc = new RootAllocator()) {
     VectorSchemaRoot outRoot;
     try (ArrowStreamReader reader = new ArrowStreamReader(new 
ByteArrayInputStream(input), alloc)) {
       reader.loadNextBatch();
       VectorSchemaRoot in = reader.getVectorSchemaRoot();
       ArrowType.Union inType = (ArrowType.Union) 
in.getSchema().getFields().get(0).getType();
       System.out.println("INPUT  Union.typeIds = " + 
Arrays.toString(inType.getTypeIds())); // [0, 1]
   
       List<FieldVector> moved = new ArrayList<>();
       for (FieldVector v : in.getFieldVectors()) {
         TransferPair tp = v.getTransferPair(alloc);
         tp.transfer();
         moved.add((FieldVector) tp.getTo());
       }
       outRoot = new VectorSchemaRoot(moved);
       outRoot.setRowCount(in.getRowCount());
     }
     ByteArrayOutputStream sink = new ByteArrayOutputStream();
     try (ArrowStreamWriter w = new ArrowStreamWriter(outRoot, null, 
Channels.newChannel(sink))) {
       w.start(); w.writeBatch(); w.end();
     }
     outRoot.close();
     try (ArrowStreamReader reader = new ArrowStreamReader(new 
ByteArrayInputStream(sink.toByteArray()), alloc)) {
       reader.loadNextBatch();
       Field f = reader.getVectorSchemaRoot().getSchema().getFields().get(0);
       ArrowType.Union outType = (ArrowType.Union) f.getType();
       System.out.println("OUTPUT Union.typeIds = " + 
Arrays.toString(outType.getTypeIds())); // [24, 3]
     }
   }
   ```
   
   Output:
   
   ```
   INPUT  Union.typeIds = [0, 1]
   OUTPUT Union.typeIds = [24, 3]     // Utf8=24, SmallInt=3 — Types.MinorType 
ordinals
   ```
   
   The type-id **buffer** is still `[0, 1]`, so type id `0` is no longer 
declared.
   
   `CANONICAL_UNION_IPC_B64` (a valid pyarrow-produced union IPC stream):
   
   ```
   
/////+AAAAAQAAAAAAAKAAwABgAFAAgACgAAAAABBAAEAAAAyP///wQAAAABAAAABAAAAIj///8AAAEOGAAAACQAAAAEAAAAAgAAAHAAAAAoAAAAAQAAAHUAAAAIAAgAAAAEAAgAAAAEAAAAAgAAAAAAAAABAAAAzP///wAAAQIQAAAAHAAAAAQAAAAAAAAAAwAAAGFnZQAIAAwACAAHAAgAAAAAAAABEAAAABAAFAAIAAYABwAMAAAAEAAQAAAAAAABBRAAAAAcAAAABAAAAAAAAAAEAAAAbmFtZQAAAAAEAAQABAAAAP/////oAAAAFAAAAAAAAAAMABYABgAFAAgADAAMAAAAAAMEABgAAAAoAAAAAAAAAAAACgAYAAwABAAIAAoAAAB8AAAAEAAAAAIAAAAAAAAAAAAAAAYAAAAAAAAAAAAAAAIAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAgAAAAAAAAADAAAAAAAAAAYAAAAAAAAAAYAAAAAAAAAIAAAAAAAAAAAAAAAAAAAACAAAAAAAAAABAAAAAAAAAAAAAAAAwAAAAIAAAAAAAAAAAAAAAAAAAACAAAAAAAAAAAAAAAAAAAAAgAAAAAAAAAAAAAAAAAAAAABAAAAAAAAAAAAAAUAAAAGAAAAAAAAAEZyYW5reAAACQAFAAAAAAD/////AAAAAA==
   ```
   
   Regenerate the input (Arrow C++/pyarrow), which round-trips it correctly:
   
   ```python
   import pyarrow as pa, base64
   children = [pa.array(["Frank", "x"], pa.utf8()), pa.array([9, 5], 
pa.int16())]
   type_ids = pa.array([0, 1], pa.int8())
   u = pa.UnionArray.from_sparse(type_ids, children, field_names=["name", 
"age"])
   batch = pa.RecordBatch.from_arrays([u], names=["u"])
   sink = pa.BufferOutputStream()
   with pa.ipc.new_stream(sink, batch.schema) as w:
       w.write_batch(batch)
   print(base64.b64encode(sink.getvalue().to_pybytes()).decode())
   # reading arrow-java's OUTPUT back in pyarrow raises:
   #   In column 0: Invalid: Union value at position 0 has invalid type id 0
   ```
   
   ### Expected
   
   The re-serialized union should preserve the declared type ids `[0, 1]` 
(matching the type-id buffer), i.e. `UnionVector.getField()` should report the 
union's actual wire type ids rather than deriving them from 
`Types.MinorType.ordinal()`.
   
   ### Notes
   
   - The corruption is in the schema/`getField()` type-id derivation, not the 
data buffers — `getField()` returns type ids based on `MinorType` ordinals 
rather than the union's wire type ids.
   - Constructing a `UnionVector` fresh via `setType(...)` and calling 
`getField()` returns `Union(Sparse, [])` (empty type ids), a related 
manifestation of the same root cause.
   - Searched existing reports: ARROW-1692 (dense/sparse detection on read, 
fixed 1.0.0) and ARROW-6145 (field metadata preservation, fixed 0.15.0) are 
adjacent but distinct from this type-id derivation on serialize.
   
   Reported by Rusty Conover (Query Farm — https://query.farm), found while 
running a cross-implementation Arrow conformance suite.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to