Hi Ryan, thx for clarifying :thumb-up: So, I support "option 0," which I'll place at the top of the list as: 'Changes in Parquet files are versioned in the current way.'
Using path_in_schema as the example big braking change I could imagine getting traction: When we make path-in-schema optional we might bump the version number of the parquet-format package as a hint for implementers to read the release notes, but we won't put a version gate on the files themselves. We expect it to be opted-in only by those most affected by footer bloat and only on those files most affected by that bloat. With the passage of time more users might be reaching for it and the reader base will have matured and we might contemplate some writers making it the default or a threshold-driven default. But that'll take time. As you mentioned, Ryan, readers may not correctly follow Thrift versioning and could be buggy or implementers might not read release notes. This is true but I reason that those upgrading parquet-format without reading release notes would also be those who don't know to gate on a new version field, etc., in the other options on the list. Also, I know of one mainstream widely-used Parquet reader that doesn't check the trailing file magic for PAR1 either. So, it is what it is and the same problems will be faced by any version-gate-in-file approach. When this was voted on previously I had thought the versioning discussion would focus on making it easier to reason about what readers and writers support in feature matrices and those kinds of decisions that users opting into features want to know about, perhaps by bumping the version of parquet-format. I wasn't thinking it'd be a version gate on the produced Parquet files at rest themselves :D Best, Will On Thu, 10 Sept 2026 at 21:25, Ryan Blue <[email protected]> wrote: > I'm also replying to Andrew here, but I think it helps to keep the replies > separate and focused on one topic. > > > From what I can tell, the current state of parquet is implicitly Option 2 > > (as readers can and do try to read any file and skip or error when > > encountering unsupported features), and don't check version numbers. > > > [With option 2] evolving the format would have > > additional requirements (but the same requirements already implicitly > > exist) > > I don't agree that option 2 is the current state. I think the current state > is that we don't make these breaking changes, and that's the problem I want > to fix. > > We've released new encodings and compression that guarantee failure because > they use a new enum symbol, but will only break readers when those columns > are read. However, we have structured all of the other changes to be > forward-compatible. For instance, when we found the sort bug for string > columns, we introduced a second set of fields for lower and upper bounds. > We also use structs to mimic enums when we want them to be forward > compatible (like logical type annotations). > > Without a way to make and coordinate breaking changes, I think we must > adopt a guarantee like the one in Option 2. If we do that, I think we've > made it even harder to evolve the format and I don't see why we would > decide to make a breaking change to fix something like path_in_schema. > > There's also another way to look at this: if we are confident that > path_in_schema will break all older readers, why not use that compatibility > break to get other cleanup features in? Doing that is one of the advantages > of bundling. > > On Thu, Sep 10, 2026 at 2:12 AM Andrew Lamb <[email protected]> > wrote: > > > Here is my attempt to summarize the tradeoffs (Ryan's explanation during > > the call was very helpful for me). > > > > The core tradeoff is in requirements for future changes to the Parquet > spec > > itself vs how many files particular readers can read. > > > > If the spec requires that readers fail fast on unknown versions ("Option > > 1"), > > * Pro: New changes to the spec don't have to consider existing readers, > and > > are thus in theory are easier/faster to make (e.g. Ryan's example of > > relocatable page headers) > > * Con: requires readers to fail on all files with newer versions, even > > those files that the reader could have read correctly > > > > If the spec allows readers to attempt to read unknown versions ("Option > 2") > > * Pro: Readers will be able to read more files, though it will be harder > to > > reason up front if a reader can read any file (may have to test it) > > * Con: There is a requirement on any future changes to the spec to ensure > > old readers don't interpret new features ("additional guarantee that > > reading future formats will either fail or produce correct results"). > This > > requirement is hard to define precisely given the wide and unknown > variety > > of readers > > > > From what I can tell, the current state of parquet is implicitly Option 2 > > (as readers can and do try to read any file and skip or error when > > encountering unsupported features), and don't check version numbers. > > > > My personal opinion is that allowing readers to read unknown versions > > (option 2 / option 3) is the most practical: Changing the current > implicit > > behavior would be quite confusing, and making the spec harder to change > for > > wider read interoperability is the right tradeoff in my mind. > > > > Andrew > > > > > > p.s. > > > > > I think path_in_schema is a good example of why option 2 will > inevitably > > produce > > correctness bugs > > > > It seems to me that making path_in_schema optional will simply cause old > > readers to fail if they need it, so is not a good example of why option 2 > > would necessarily correctness issues compared to option 1. The example > > about java hash sets could be avoided with adequate testing, for example, > > and I don't see how gating that code behind reading a new version number > > would make that bug any more/less likely. > > > > > My second argument for why we should not attempt to read all future > > versions > > of Parquet is that it ends up limiting how we can evolve the format. > > > > This makes a lot of sense to me -- evolving the format would have > > additional requirements (but the same requirements already implicitly > > exist); Adding relocatable pages, for example could be achieved by > adding a > > new DataPageHeaderV3 rather than modifying the existing structure. That > is > > more complicated to be sure, but not impossible. > > > > > > > > > > On Thu, Sep 10, 2026 at 4:39 AM Antoine Pitrou <[email protected]> > wrote: > > > > > Le 10/09/2026 à 00:38, Ryan Blue a écrit : > > > > > > > > There are 3 main options: > > > > 1. A reader should fail because it does not support the version > > > > 2. A reader should attempt to read the file > > > > 3. This choice is left up to implementations > > > > > > > > I'll cover each option in more detail below, but first I want to > > clarify > > > > that we are not talking about "preview" features like encodings or > > > > forward-compatible changes like new logical types. Preview features > > will > > > > break readers that do not support them and only affect specific > columns > > > > using the feature. For preview features, the expectation is that > > readers > > > > will attempt to read the file and will fail if they need to project a > > > > column that cannot be read. > > > > > > "Preview features" is a very weird terminology. It sounds like > > > "unfinished" or "experimental". > > > > > > > The choice of how to handle an unsupported format version primarily > > > affects > > > > changes that add, remove, or modify the semantics of metadata fields. > > For > > > > example: > > > > - Changing `path_in_schema` from required to optional > > > > > > This depends whether the Thrift parser checks that required fields are > > > actually present in the serialized payload? Do we know what their > > > current behavior is? > > > > > > i.e., does a Thrift parser generated with a required `path_in_schema` > > > specification accept a serialized payload without that field? > > > > > > It also depends what the Parquet reader actually *does* with the > > > `path_in_schema`? AFAICT, the Parquet C++ reader isn't doing anything > > > specific with it. > > > > > > > Option 3 would mean that readers may choose to attempt to read, but > do > > > not > > > > have additional guarantees. > > > > > > Option 3 can also mean "the reader is exposing an option to let the > user > > > choose the behavior (reject up front or attempt to read anyway)". > > > > > > (this would be more costly to implement, so I'm not sure any > > > implementation would actually do that; but it's at least conceptually > > > possible) > > > > > > Regards > > > > > > Antoine. > > > > > > > > > > > >
