CurtHagenlocher opened a new pull request, #445: URL: https://github.com/apache/arrow-dotnet/pull/445
## What's Changed The shredding pipeline had no way to represent a SQL-NULL row: every entry point took `VariantValue`, and `ShreddedVariantArrayBuilder.Build` always produced a storage struct with no validity. Callers had to use `VariantValue.Null` as a placeholder, which is stored as a present variant null and changes what `IS NULL` means. This follows the `VariantValue?` convention already used by `VariantArray.Builder`: - `ShredSchemaInferer.Infer(IEnumerable<VariantValue?>, ShredOptions)` ignores null rows, so they don't count toward frequency thresholds. An all-null column infers as unshredded. - `VariantShredder.Shred(IEnumerable<VariantValue?>, ShredSchema)` leaves null rows out of the shared metadata and returns a `null` entry for each one. - `ShreddedVariantArrayBuilder.Build` treats a `null` entry in `rows` as a null element of the storage struct. Its children are emitted as missing, and `metadata` is still populated because that field is required. No validity buffer is allocated when there are no nulls. Existing signatures are unchanged. The one source-level catch is that a call passing a literal `null` for `values` (e.g. `Shred(null, schema)`) is now ambiguous between the two overloads. This lays the groundwork for the array-level entry points proposed in #399, which can read validity directly from the input `VariantArray`. Closes #398. 🤖 Generated with [Claude Code](https://claude.com/claude-code) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
