I think it's valuable to differentiate optional from "nullable" or "effectively nullable", the latter includes nested required fields whose parent is optional.
+1 on the Spec change PR: https://github.com/apache/iceberg/pull/17609 Also Eduard's PR above looks very interesting, we have the opportunity to finalize the terms here to make it specific Best, Shawn On Tue, Aug 11, 2026 at 9:27 AM Xiening Dai <[email protected]> wrote: > As a side effect of this bug fix PR - > https://github.com/apache/iceberg/pull/17560, the null_value_count will > be recorded for nested required fields. > > On 2026/08/11 07:48:27 Eduard Tudenhöfner wrote: > > Coincidentally I ran into this as well while working on the evaluator for > > content stats (see > > > https://github.com/apache/iceberg/pull/17413/changes/53456855fbf126502a7e6a15ce39e1833f40d43f > > from > > https://github.com/apache/iceberg/pull/17413) but didn't have time yet > to > > discuss this issue with the community. I think that including the > > *null_value_count* for required fields inside optional parents makes > sense. > > > > On Mon, Aug 10, 2026 at 11:41 PM Xiening Dai <[email protected]> wrote: > > > > > Hi Anoop, > > > > > > I like your "effective nullability" definition. The current spec > wording > > > definitely needs update: > > > > > > ``` > > > A struct is a tuple of typed values. Each field in the tuple is named > and > > > has an integer id that is unique in the table schema. Each field can be > > > either optional or required, meaning that values can (or cannot) be > null. > > > Fields may be any type. > > > ``` > > > > > > And I agree we should remove the "optional null_value_count" comment, > and > > > writer should record the null count whenever possible for nested fields > > > (regardless required or optional). Checking the null count against 0 is > > > much easier than going through and checking the ancestors nullability. > > > > > > On 2026/08/10 20:09:00 Anoop Johnson wrote: > > > > Hi, Xiening - > > > > > > > > What you brought up is an important subtlety. I view it a bit > differently > > > > though: what matters is *effective* nullability along the whole path > > > from > > > > the root, not the leaf field's own *required* flag. In your > > > > specific example, `s.a` is declared NOT NULL, but its parent `s` is > > > > optional. So effectively `s.a` is nullable. A field is only truly > > > > non-nullable if it is required and every ancestor struct up to the > root > > > is > > > > also required. A genuinely non-nullable field like this does not > require > > > > `null_value_count` (because it will always be 0). > > > > > > > > So `s.a` here is not a required field in the spec perspective, and > > > > null_value_count *should* be recorded for it. I believe the current > > > > Iceberg stats collector and evaluator handles this correctly. e.g. > the > > > > InclusiveMetricsEvaluator relies > > > > < > > > > https://github.com/apache/iceberg/blob/89e2f887491c1b5fa9f8b9de81b3aa8b31fa6974/api/src/main/java/org/apache/iceberg/expressions/InclusiveMetricsEvaluator.java#L100-L102 > > > > > > > > on the explicit count and not the field-level required flag. > > > > > > > > @Override > > > > protected boolean mayContainNull(int id) { > > > > return nullCounts == null || !nullCounts.containsKey(id) || > > > > nullCounts.get(id) != 0; > > > > } > > > > > > > > Overall, I agree with the spirit of your suggestion. We could > clarify > > > the > > > > spec wording so "optional fields" reads as *effectively required*, > and > > > make > > > > explicit that a field is non-nullable only when it and all of its > > > ancestors > > > > are required. > > > > > > > > Best, > > > > Anoop > > > > > > > > On Mon, Aug 10, 2026 at 10:45 AM Xiening Dai <[email protected]> > wrote: > > > > > > > > > Hi all, > > > > > > > > > > I see this in V4 spec: "null_value_count is only used for optional > > > fields" > > > > > ( > > > > > > > > > https://github.com/apache/iceberg/blob/89e2f887491c1b5fa9f8b9de81b3aa8b31fa6974/format/spec.md?plain=1#L838 > > > > > ) > > > > > > > > > > This only makes sense if required field never contains nulls. But > that > > > > > actually is not true. In table schema, we can have a struct field > > > which is > > > > > optional and has a child field that is required. In that case, > when the > > > > > parent struct is null, the required field values is treated as NULL > > > too. > > > > > > > > > > For example, below works in Spark: > > > > > > > > > > CREATE TABLE t (i INT, s STRUCT<a: INT NOT NULL, b: STRING>) USING > > > iceberg; > > > > > INSERT INTO t VALUES (1, NULL); > > > > > SELECT COUNT (*) WHERE s.a IS NULL; -- returns 1 > > > > > > > > > > Because of this, I don't think the query engine can optimize the > > > execution > > > > > plan and assume a field doesn't contain null when it's marked as > > > required. > > > > > It would need to rely on the null_value_count explicitly being 0. > If > > > you > > > > > agree on this, then we should always record the null_value_count no > > > matter > > > > > if it's required or optionals field. > > > > > > > > > > Thoughts? > > > > > > > > > > > > > > >
