The interleaved column on top of zstd skippable frames is a truly *epic*
hack. I love it. Well done.

It is probably worth noting in the blogpost that LZ4 also support skippable
frames if one wants to play the same trickery on top of a faster general
compressor.

On Sat, Aug 1, 2026 at 10:09 AM Andrew Lamb <[email protected]> wrote:

> Thank you for sharing this Will -- I really like the proposed term "Random
> Access Parquet" (RAP) - and several of the techniques are quite clever.
>
> The interleaving column technique I found especially innovative, and solves
> one of the major remaining issues of random access using parquet (at the
> cost of increased storage, of course). This usecase I also think motivated
> Dan's prior discussion on potential data page headers[1], and is frequently
> cited by Weston Pace as a benefit of Lance (writers can cluster the data
> for one row together), so your proposed method to avoid spec changes is
> really neat.
>
> I also wrote a potentially related blog[2] with some more details on how to
> implement external indexes for Parquet, in case that is interesting.
>
>
> Andrew
>
> [1]: https://lists.apache.org/thread/jgq7wk3641ss27y851zdok1v2nskyvhd
> [2]:
> https://datafusion.apache.org/blog/2025/08/15/external-parquet-indexes/
>
> On Fri, Jul 31, 2026 at 4:41 AM Will Edwards via dev <
> [email protected]>
> wrote:
>
> > Howdy folks :)
> >
> > At Spotify we have been exploring how to efficiently use the data lake -
> > something normally organised and optimised for big batch workloads - for
> > fast point queries too.  Eg the AI agent that needs to answer your
> question
> > about what you did last summer.
> >
> > We found that if we extract metadata and put it in a fast key value store
> > we can know exactly which parts of which files to read without doing the
> > normal loading footers and searching, which dramatically changes the
> > performance and cost profile. And that there are some tricks you can do
> > when writing parquet files that helps this access patterns.  It’s
> basically
> > not so different from the metadata store that speeds up analytic
> workloads
> > but by key instead.
> >
> > Here’s blog we just posted that is hopefully interesting:
> >
> >
> >
> https://engineering.atspotify.com/2026/7/indexing-the-data-lake-for-online-point-queries
> >
> > Happy to go into more details and share perf tips and tricks if you have
> > any thoughts or just wanna chat about Parquet perf!
> >
> > Best,
> > Will
> >
>

Reply via email to