Hi folks,

I'm starting this thread to discuss how we add and enable new sort orders
in the format. There was a recent change [1] to add a new
IEEE_754_TOTAL_ORDER sort order and a nan_count field to row-group/page
stats. TLDR for this change: NaN values invalidate stats, so writers leave
them out of min/max stats and instead set nan_count to signal whether NaN's
are present.

The merged parquet-java implementation [2] emits nan_count from the writer,
but not the new sort order (this is valid per the spec). The arrow-rs
implementation [3] emits both nan_count and the sort order, but is not
merged and is marked with "api-release" and "next-major-release" labels.
There seems to be a disconnect between the format spec and the
implementations: the change went into to parquet-format v2.13.0 and was
considered forward-compatible, but both reference implementations are
considering use of the new sort order to be forward-incompatible.
Additionally, adopting nan_count but not IEEE_754_TOTAL_ORDER can cause
correctness as described parquet-java PR [4].

Given that the writer changes will go out in parquet-java 1.18 and there is
another new sort order (INT96_TIMESTAMP_ORDER) merged but not yet released
in parquet-format, we need to decide how to handle these cases. I see two
options:

1. Consider new sort orders to be forward-compatible and update the
implementations to adopt IEEE_754_TOTAL_ORDER. Per the spec, readers should
gracefully handle unrecognized sort orders by ignoring min/max stats [5].
This is the simplest approach, but risks older readers breaking. Per
discussion on the initial nan_count thread [6] the Java, C++, and Python
implementations have been verified to gracefully ignore unrecognized union
values. Older versions of arrow-rs would fail, but this has been fixed in
newer versions.
2. Consider new sort orders to be forward-incompatible at the format level.
Revert the IEEE_754_TOTAL_ORDER/nan_count change as well as the
INT96_TIMESTAMP_ORDER in advance of parquet-format v2.14.0. Revert the
parquet-java reference implementation before releasing v1.18 as well. This
is the safest approach, but it would mean both new sort orders would be
blocked by the ongoing versioning discussion and released only in the next
format major version.

Given past discussion and verification of different implementations, my
opinion is that this is largely a forward compatible change. Regardless of
what we decide, this discussion should block the ongoing parquet-java
release as the current state (emitting stats with TYPE_DEFINED_ORDER and
nan_count) is not forward compatible and poses correctness risks.

-- Div

[1]
https://github.com/apache/parquet-format/commit/662cdac7ffb7e71e10e0e0c519b215791ff2d1aa
[2] https://github.com/apache/parquet-java/pull/3393
[3] https://github.com/apache/arrow-rs/pull/9619
[4] https://github.com/apache/parquet-java/pull/3393#issuecomment-5125834732
[5]
https://github.com/apache/parquet-format/blob/2076361bb64e2de9ca6a8d06eda025a6fa4e9df6/src/main/thrift/parquet.thrift#L1100
[6] https://lists.apache.org/thread/t93kw60gkn3lq7vt4n6c4nkmsr7dtyn1

Reply via email to