I do think we need to have a new release of parquet-format that includes ALP; I can help with this if needed once we have completed work on the blog and the example file.
However, I am not sure if this is what you mean by "track a release of parquet-format" Andrew On Mon, Aug 17, 2026 at 7:07 PM Vinoo Ganesh <[email protected]> wrote: > One quick question here - should we also track a release of parquet-format > given that the ALP spec has merged? Apologies if this was covered at the > last sync and I'm out of date. > > > <[email protected]> > > > On Sat, Aug 15, 2026 at 5:56 AM Andrew Lamb <[email protected]> > wrote: > > > Thank you Russell for the review -- it was very helpful and I have > > incorporated all your feedback. > > > > I now have reports that this file can be read by multiple ALP > > implementations (C++, Java, Rust, as well as a project in C# from Curt, > so > > I am feeling pretty good about it. > > > > Please anyone else that is interested (or has an initial ALP > > implementation), take a look and try it out. > > > > Anrew > > > > On Thu, Aug 13, 2026 at 10:35 AM Russell Spitzer < > > [email protected]> > > wrote: > > > > > Great! I'll take a look now > > > > > > On Thu, Aug 13, 2026 at 5:42 AM Andrew Lamb <[email protected]> > > > wrote: > > > > > > > Hello there, > > > > > > > > Following up on our sync status update yesterday, I would like to ask > > for > > > > some more reviews / verification of my proposed test dataset for the > > ALP > > > > encoding[1] . It covers a variety of vector sizes and data > > distributions > > > / > > > > exceptions and I think will help ensure compatibility across > > > > implementations. > > > > > > > > Thank you to Curt and Vinoo who have verified their implementations > (C# > > > and > > > > Java) can read this file. I have verified it can be read with our > Rust > > > > implementation. > > > > > > > > I would like to merge this relatively soon as I believe it will > unlock > > > full > > > > on implementation in the ecosystem. > > > > > > > > Thank you for your time, > > > > Andrew > > > > > > > > p.s. Kosta, Prateek and I are also working on a blog post [2] to > > > introduce > > > > this feature and how it works, which hopefully will also accelerate > its > > > > rollout > > > > > > > > > > > > [1]: https://github.com/apache/parquet-testing/pull/119 > > > > [2]: https://github.com/apache/parquet-site/pull/195 > > > > > > > > On Wed, Aug 5, 2026 at 5:49 AM Andrew Lamb <[email protected]> > > > wrote: > > > > > > > > > > Just wanted to confirm if you have both low and high precision > > values > > > > > and > > > > > outliers. > > > > > > > > > > Yes I think the file covers these cases, though it would be great > if > > > you > > > > > could double check. The cases are described in more detail in [1] > > > > > > > > > > > Quick question : If we commit the parquet file how does one test > > > that a > > > > > different language reader reads the values correctly? > > > > > > > > > > The first two columns in the file have the same float and double > > > values, > > > > > encoded using PLAIN+zstd. So the verification looks like reading > the > > > > first > > > > > two columns and comparing them to the in the ALP encoded columns > > > > > > > > > > I think this strategy is more accurate than a CSV as there are no > > > > > potential ambiguity (e.g. exactly what bit pattern does NaN or Inf > > > mean, > > > > or > > > > > a decimal value that can't be stored exactly as floating point) > > > > > > > > > > Andrew > > > > > > > > > > [1]: > > > > > > > > > > > > > > > https://github.com/alamb/parquet-testing/blob/alamb/alp_test_data/data/README.md#alp-encoding > > > > > > > > > > > > > > > On Tue, Aug 4, 2026 at 9:20 PM PRATEEK GAUR <[email protected]> > > > wrote: > > > > > > > > > >> Thanks Andrew, > > > > >> > > > > >> Yes in our PR [1] we ended up adding all datasets which we needed > to > > > > prove > > > > >> the value of ALP. (also we didn't trim the number of rows) > > > > >> > > > > >> We tried to cover scenarios like > > > > >> 1) Low precision values > > > > >> 2) High precision values > > > > >> 3) Values with outliers > > > > >> 4) Datasets with different vector sizes > > > > >> > > > > >> And I think you've captured all of that in your dataset. > > > > >> ``` > > > > >> > > > > >> message schema { > > > > >> OPTIONAL FLOAT float_plain; > > > > >> OPTIONAL DOUBLE double_plain; > > > > >> OPTIONAL FLOAT float_alp_1024; > > > > >> OPTIONAL DOUBLE double_alp_1024; > > > > >> OPTIONAL FLOAT float_alp_4096; > > > > >> OPTIONAL DOUBLE double_alp_4096; > > > > >> OPTIONAL FLOAT float_alp_32; > > > > >> OPTIONAL DOUBLE double_alp_32; > > > > >> } > > > > >> ``` > > > > >> > > > > >> Just wanted to confirm if you have both low and high precision > > values > > > > and > > > > >> outliers. > > > > >> > > > > >> Quick question : If we commit the parquet file how does one test > > that > > > a > > > > >> different language reader reads the values correctly? > > > > >> I thought we would want to write a CSV file too against which we > can > > > > >> compare. (or maybe a hash?) > > > > >> > > > > >> [1] : https://github.com/apache/parquet-testing/pull/100 > > > > >> > > > > >> On Tue, Aug 4, 2026 at 10:38 AM Andrew Lamb < > [email protected] > > > > > > > >> wrote: > > > > >> > > > > >> > Interop testing is something I am definitely interested in too > > > > >> > > > > > >> > We had a prior discussion[1] and there is an issue about this > [2] > > > that > > > > >> > maybe it is time to revive > > > > >> > > > > > >> > [1]: > > > https://lists.apache.org/thread/kd3k4q691lp5c4q3r767zb8jltrm9z33 > > > > >> > [2]: https://github.com/apache/parquet-format/issues/441 > > > > >> > > > > > >> > > > > > >> > On Tue, Aug 4, 2026 at 1:04 PM Curt Hagenlocher < > > > [email protected] > > > > > > > > > >> > wrote: > > > > >> > > > > > >> > > Thanks! I validated that my own implementation > > > > >> > > (https://github.com/clast-project/engineered-wood/pull/64) is > > > > working > > > > >> > > with this data. > > > > >> > > > > > > >> > > Better interop testing in general is a recurring topic, and > > > perhaps > > > > >> > > should be taken up again once the versioning work winds down. > > > > >> > > > > > > >> > > > > > > >> > > > > > > >> > > On Tue, Aug 4, 2026 at 8:42 AM Andrew Lamb < > > > [email protected]> > > > > >> > wrote: > > > > >> > > > > > > > >> > > > As part of rolling out ALP to the ecosystem, I think it > > > important > > > > to > > > > >> > have > > > > >> > > > an example data set encoded with ALP in parquet-testing for > > > > readers > > > > >> to > > > > >> > > test > > > > >> > > > against. Prateek and Vinoo created files like this as part > of > > > > their > > > > >> > C/C++ > > > > >> > > > and Java implementations, but the files were quite large > > > (multiple > > > > >> MB) > > > > >> > > > > > > > >> > > > While working to get the ALP Rust implementation ready to > > > > merge[1], > > > > >> I > > > > >> > > spent > > > > >> > > > some time creating a smaller example ALP dataset (211KB) for > > > > testing > > > > >> > for > > > > >> > > > consideration[2]. > > > > >> > > > > > > > >> > > > I created this file using the ALP C++ implementation and was > > > able > > > > to > > > > >> > read > > > > >> > > > it with the ALP Rust implementation. > > > > >> > > > > > > > >> > > > Any feedback would be appreciated, > > > > >> > > > Andrew > > > > >> > > > > > > > >> > > > > > > > >> > > > [1]: https://github.com/apache/arrow-rs/pull/9372 > > > > >> > > > [2]: https://github.com/apache/parquet-testing/pull/119 > > > > >> > > > > > > >> > > > > > >> > > > > > > > > > > > > > > >
