> Just wanted to confirm if you have both low and high precision values and
outliers.

Yes I think the file covers these cases, though it would be great if you
could double check. The cases are described in more detail in [1]

> Quick question : If we commit the parquet file how does one test that a
different language reader reads the values correctly?

The first two columns in the file have the same float and double values,
encoded using PLAIN+zstd. So the verification looks like reading the first
two columns and comparing them to the in the ALP encoded columns

I think this strategy is more accurate than a CSV as there are no potential
ambiguity (e.g. exactly what bit pattern does NaN or Inf mean, or a decimal
value that can't be stored exactly as floating point)

Andrew

[1]:
https://github.com/alamb/parquet-testing/blob/alamb/alp_test_data/data/README.md#alp-encoding


On Tue, Aug 4, 2026 at 9:20 PM PRATEEK GAUR <[email protected]> wrote:

> Thanks Andrew,
>
> Yes in our PR [1] we ended up adding all datasets which we needed to prove
> the value of ALP. (also we didn't trim the number of rows)
>
> We tried to cover scenarios like
> 1) Low precision values
> 2) High precision values
> 3) Values with outliers
> 4) Datasets with different vector sizes
>
> And I think you've captured all of that in your dataset.
> ```
>
> message schema {
>   OPTIONAL FLOAT float_plain;
>   OPTIONAL DOUBLE double_plain;
>   OPTIONAL FLOAT float_alp_1024;
>   OPTIONAL DOUBLE double_alp_1024;
>   OPTIONAL FLOAT float_alp_4096;
>   OPTIONAL DOUBLE double_alp_4096;
>   OPTIONAL FLOAT float_alp_32;
>   OPTIONAL DOUBLE double_alp_32;
> }
> ```
>
> Just wanted to confirm if you have both low and high precision values and
> outliers.
>
> Quick question : If we commit the parquet file how does one test that a
> different language reader reads the values correctly?
> I thought we would want to write a CSV file too against which we can
> compare. (or maybe a hash?)
>
> [1] : https://github.com/apache/parquet-testing/pull/100
>
> On Tue, Aug 4, 2026 at 10:38 AM Andrew Lamb <[email protected]>
> wrote:
>
> > Interop testing is something I am definitely interested in too
> >
> > We had a prior discussion[1] and there is an issue about this [2] that
> > maybe it is time to revive
> >
> > [1]: https://lists.apache.org/thread/kd3k4q691lp5c4q3r767zb8jltrm9z33
> > [2]: https://github.com/apache/parquet-format/issues/441
> >
> >
> > On Tue, Aug 4, 2026 at 1:04 PM Curt Hagenlocher <[email protected]>
> > wrote:
> >
> > > Thanks! I validated that my own implementation
> > > (https://github.com/clast-project/engineered-wood/pull/64) is working
> > > with this data.
> > >
> > > Better interop testing in general is a recurring topic, and perhaps
> > > should be taken up again once the versioning work winds down.
> > >
> > >
> > >
> > > On Tue, Aug 4, 2026 at 8:42 AM Andrew Lamb <[email protected]>
> > wrote:
> > > >
> > > > As part of rolling out ALP to the ecosystem, I think it important to
> > have
> > > > an example data set encoded with ALP in parquet-testing for readers
> to
> > > test
> > > > against. Prateek and Vinoo created files like this as part of their
> > C/C++
> > > > and Java implementations, but the files were quite large (multiple
> MB)
> > > >
> > > > While working to get the ALP Rust implementation ready to merge[1], I
> > > spent
> > > > some time creating a smaller example ALP dataset (211KB) for testing
> > for
> > > > consideration[2].
> > > >
> > > > I created this file using the ALP C++ implementation and was able to
> > read
> > > > it with the ALP Rust implementation.
> > > >
> > > > Any feedback would be appreciated,
> > > > Andrew
> > > >
> > > >
> > > > [1]: https://github.com/apache/arrow-rs/pull/9372
> > > > [2]: https://github.com/apache/parquet-testing/pull/119
> > >
> >
>

Reply via email to