Thank you Arnav -- I am sorry for the delay. Now that we finally got ALP out the door, I plan to spend more time on FSST/OnPair.
I just went through the proposal again, and left some comments. Overall it is looking quite good I think -- my personal opinion is there we should also support FSST encoding dictionary pages (I have a minor suggestion here[1], basically to include an optional for an "inline symbol table" in a FSST encoded page, and remove the SYMBOL_TABLE_PAGE page type) Andrew [1]: https://docs.google.com/document/d/1Xg2b8HR19QnI3nhtQUDWZJhCLwJzW6y9tU1ziiLFZrM/edit?disco=AAACElKssY4 On Sat, Aug 15, 2026 at 10:45 AM Arnav Balyan <[email protected]> wrote: > Hi Gunnar and Curt, > > Thank you for going through the spec and the comments. > > > One question for my understanding: have you considered, instead of adding > a > > new dedicated symbol page type, re-using the existing dict page > mechanism? > > Yes this was initially considered, The FSST symbol table is analogous to a > dictionary page in dictionary encoding, with one page per column chunk. We > have removed unnecessary parts like the dictionary page fallback. Also FSST > symbol table requires storing (symbol data + length histograms + count), > given that the symbol table is always 255 symbols and 8 bytes, it provides > assumptions which can be better used to encode the symbol table. Curt's > point is additionally valid, length histogram is also required. > > > I've published some test data in > https://github.com/apache/parquet-testing/pull/121 which I believe > corresponds > to the current version of the spec. > > Thank you Curt, really appreciate it. This will be valuable for the cross > compat tests. > > > FWIW I don't think we normally finalize / vote on a spec until we have 2 > example > implementations > > Thanks for the feedback Andrew, I have added the updated implementation > <https://github.com/apache/arrow/pull/50872> for Arrow here which conforms > to the latest spec (cc Curt). Additionally will add Parquet java shortly > which gives us 2 implementations in addition to Devan's great work on arrow > rust. > > > My main concern is that the spec is overly complicated to implement on > decode. > > Thanks, based on the comments on the spec, I see 2 points which we can > discuss/update in the spec. > > 1. Length Histogram: The length histogram was initially proposed in > the FSST implementation/paper to allow the decoder to reconstruct symbol > table efficiently. Infact we had a brief discussion on the spec where we > aligned to use the length histogram since it can encode information in less > space with no effect on the decode speed. > > 2. Offset array: The offset array encode is an optional feature which uses > the existing parquet encoding to encode the offsets. FSST with offset array > encoding enabled beats other encodings on the decode speed, the numbers in > the spec include this decoding cost and use compressed offsets. However > this remains an optional choice upto the writer, as of today the > implementation supports both. > > Additionally, I've addressed the comments on the spec, thank you everyone > for the comments. I will monitor the document for additional comments in > the coming days. > > Warm Regards, > Arnav > > > On Fri, Aug 14, 2026 at 2:14 AM Andrew Lamb <[email protected]> > wrote: > > > Thanks Arnav, > > > > I read the spec in detail again and I think the FSST symbol table > placement > > is looking good to me. My main concern is that the spec is overly > > complicated to implement on decode. > > > > I understand the primary motivation for adding FSST-like encodings is > > cases where existing codecs don't do well -- namely fast decoding speed > and > > random access. > > > > There appear to be several features in the current spec that trade off > > decoding speed and random access to improve the compression ratio (e.g. > > length_histogram in the symbol table encoding, and delta encoding > offsets, > > vs a simpler but less space efficient versions) > > > > I think we need to carefully evaluate if that is the right tradeoff (as > > parquet can already get very high compression using zstd and other > general > > purpose compression algorithms) > > > > Andrew > > > > On Wed, Aug 12, 2026 at 2:45 AM Arnav Balyan <[email protected]> > > wrote: > > > > > Hello everyone, > > > > > > I wanted to provide a quick update on the FSST encoding proposal for > > > Parquet. > > > > > > Over the last several months, we have incorporated the feedback > received > > > from the community and finalized the design. Thank you to everyone who > > > reviewed the proposal and helped resolve the open questions. > > > > > > At this point, we are planning to move the proposal from the > Design/Draft > > > phase to the Implementation phase. Before doing so, I would like to ask > > for > > > one final round of feedback and ensure that everyone who wanted to > review > > > the specification has had an opportunity to do so. > > > > > > Implementation work is also underway: > > > - Devan Benz has started an Arrow Rust > > > <https://github.com/apache/arrow-rs/pull/10153> implementation. > > > - I have an existing Arrow C++ < > > https://github.com/apache/arrow/pull/48232 > > > > > > > PoC for FSST. > > > > > > For the initial interoperability work, it would be particularly helpful > > to > > > have implementations in Parquet Java and Arrow Go as well. > > > > > > If anyone is interested in owning or contributing to the Parquet Java > or > > > Arrow Go implementation, please let us know. Contributions to the > > existing > > > Rust and C++ implementations and to cross language interoperability > > testing > > > would also be greatly appreciated. > > > > > > Artifacts: > > > > > > - FSST specification: > > > > > > > > > > > > https://docs.google.com/document/d/1Xg2b8HR19QnI3nhtQUDWZJhCLwJzW6y9tU1ziiLFZrM/edit > > > > > > - Tracking issue: > > > https://github.com/apache/parquet-format/issues/531 > > > > > > - Original FSST mailing-list discussion: > > > https://lists.apache.org/thread/15w66hvm516641btrljmtqhb4m38wn3s > > > > > > - FSST and OnPair discussion: > > > https://lists.apache.org/thread/hfoltdl6o6txc3zp4680nns1mh29h0r8 > > > > > > - Arrow C++ PoC: > > > https://github.com/apache/arrow/pull/48232 > > > > > > - Arrow Rust implementation: > > > https://github.com/apache/arrow-rs/pull/10153 > > > > > > Based on any new comments/questions, the next steps will be to: > > > 1. Address any feedback from this final review. > > > 2. Open a PR updating the proposal status from Draft/PoC to > > Implementation. > > > 3. Continue the java/go implementations and establish cross language > > tests. > > > 4. Prepare for the formal vote once the implementation requirements > have > > > been met. > > > > > > Please share any remaining design feedback or implementation interest > on > > > this thread or the spec doc. > > > > > > Thanks again to everyone involved for the thoughtful reviews, > benchmarks, > > > and implementation work. I really appreciate all the help in bringing > the > > > proposal to this stage. > > > > > > Warm regards, > > > Arnav > > > > > >
