julienledem commented on code in PR #258: URL: https://github.com/apache/parquet-format/pull/258#discussion_r1676359531
########## CONTRIBUTING.md: ########## @@ -29,3 +29,142 @@ Recommendations and requirements for how to best contribute to Parquet. We striv ### License By contributing your code, you agree to license your contribution under the terms of the APLv2: https://github.com/apache/parquet-format/blob/master/LICENSE + +### Additions/Changes to the Format + +Note: This section applies to actual functional changes to the specification. +Fixing typos, grammar, and clarifying concepts that would not change the +semantics of the specification can be done as long a comitter feels comfortable +to merge them. When in doubt starting a discussion on the dev mailing list is +encouraged. + +The general steps for adding features to the format are as follows: + +1. Discuss changes on the developer mailing list ([email protected]). + Often times it is helpful to link to a draft pull request to make the + discussion concrete. This step is complete when there is lazy consensus. Part + of the consensus is whether it sufficient to provide 2 working + implementations as outlined in step 2 or if demonstration of the feature with + a down-stream query engine is necessary to justify the feature (e.g. + demonstrate performance improvements in Arrow's DataSet library or Apache + Data Fusion or another open source engine). + +2. Once a change has lazy consensus, two implementations of the feature + demonstrating interopability must also be provided. One implementation MUST + be [`parquet-java`](http://github.com/apache/parquet-java). It is preferred + that the second implementation be + [`parquet-cpp`](https://github.com/apache/arrow) or + [`parquet-rs`](https://github.com/apache/arrow-rs), however at the discretion + of the PMC any open source Parquet implementation may be acceptable. + Implementations whose contributors actively participate in the community + (e.g. keep their feature matrix up-to-date on the Parquet website) are more + likely to be considered. If discussed as a requirement in step one, + demonstration of integration with a query engine is also required for this + step. The implementations must be made available publicly, and they should be + fit for inclusion (for example, they were submitted as a pull request against + the target repository and committers gave positive reviews). + +Unless otherwise discussed, it is expected the implementations will be developed +from their respective main branch (i.e. backporting is not expected). + +3. After the first two steps are complete a formal vote is held on + [email protected] to officially ratify the feature. After the vote + passes the format change is merged into the `parquet-format` repository and + it is expected the changes from step 2 will also be merged soon after + (implementations should not be merged until the addition has been merged to + `parquet-format`). + +#### General guidelines/preferences on additions. + +1. To the greatest extent possible changes should have an option for forward + compatibility (old readers can still read files). + +2. New encodings should be fully specified in this repository and ideally not + rely on an external dependencies for implementation (i.e. `parquet-format` is + the source of truth for the encoding). + +3. New compression mechanisms must have a pure Java implementation that can be + used as a dependency in `parquet-java`. + +### Releases + +The Parquet PMC aims to do releases of the format package only as needed when +new features are introduced. If multiple new features are being proposed +simultaneously some features might be consolidated into the same release. +Guidance is provided below on when implementations should enable features added +to the specification. Due to confusion in the past over Parquet versioning it +is not expected that there will be a 3.x release of the specification in the +foreseeable future. + +### Compatibility and Feature Enablement + +For the purposes of this discussion we classify features into the following buckets: + +1. Backward compatible. A file written under an older version of the format + should be readable under a newer version of the format. + +2. Forward compatible. A file written under a newer version of the format with + the feature enabled can be read under an older version of the format, but + some information might be missing or performance might be suboptimal. + +3. Forward incompatible. A file written under a newer version of the format with + the feature enabled cannot be read under an older version of the format (e.g. + adding and using a new compression algorithm). It is expected any feature in + this category will provide a signal to older readers, so they can + unambiguously determine that they cannot properly read the file (e.g. via + adding a new value to an existing enum). + +New features are intended to be widely beneficial to users of Parquet, and +therefore it is hoped third-party implementations will adopt them quickly after +they are introduced. It is assumed that writing new parts of the format, and +especially forward incompatible features, will be configured with a feature flag +defaulted to "off", and at some future point the feature is turned on by default +(reading of the new feature will typically be enabled without configuration or +defaulted to on). Some amount of lead time is desirable to ensure a critical +mass of Parquet implementations support a feature to avoid compatibility issues +across the ecosystem. Therefore, the Parquet PMC gives the following +recommendations for managing features: + +1. Backward compatibility is the concern of implementations but given the + ubiquity of Parquet and the length of time it has been used, libraries should + support reading older versions of the format to the greatest extent possible. + +2. Forward compatible features/changes may be enabled and used by default in + implementations once the parquet-format containing those changes has been + formally released. For features that may pose a significant performance + regression to older format readers, libaries should consider delaying default + enablement until 1 year after the release of the parquet-java implementation + that contains the feature implementation. + +3. Forward incompatible features/changes should not be turned on by default + until 2 years after the parquet-java implementation containing the feature is + released. It is recommended that changing the default value for a forward + incompatible feature flag should be clearly advertised to consumers (e.g. via + a major version release if using Semantic Versioning, or highlighed in + release notes). + +For forward compatible changes which have a high chance of performance +regression for older readers and forward incompatible changes, implementations +should clearly document the compatibility issues. Additionally, while it is up +to maintainers of individual implementations to make the best decision to serve +their ecosystem, they are encouraged to start enabling features by default along +the same timelines as `parquet-java`. Parquet-java will wait to enable features +by default until the most conservative timelines outlined above have been +exceeded. + +For features released prior to October 2024, target dates for each of these Review Comment: Sorry, I'm not following which part of the conversation "per above" refers to here. Could you point to it and explain what you mean? I was assuming for example that once we define a new footer in a new format we would increment the format spec. Would that be wrong? ########## CONTRIBUTING.md: ########## @@ -29,3 +29,142 @@ Recommendations and requirements for how to best contribute to Parquet. We striv ### License By contributing your code, you agree to license your contribution under the terms of the APLv2: https://github.com/apache/parquet-format/blob/master/LICENSE + +### Additions/Changes to the Format + +Note: This section applies to actual functional changes to the specification. +Fixing typos, grammar, and clarifying concepts that would not change the +semantics of the specification can be done as long as a committer feels comfortable +to merge them. When in doubt starting a discussion on the dev mailing list is +encouraged. + +The general steps for adding features to the format are as follows: + +1. Discuss changes on the developer mailing list ([email protected]). + Often times it is helpful to link to a draft pull request to make the + discussion concrete. This step is complete when there is lazy consensus. Part + of the consensus is whether it is sufficient to provide two working + implementations as outlined in step 2, or if demonstration of the feature with + a downstream query engine is necessary to justify the feature (e.g. + demonstrate performance improvements in the Apache Arrow C++ Dataset library, + the Apache DataFusion query engine, or any other open source engine). + +2. Once a change has lazy consensus, two implementations of the feature + demonstrating interopability must also be provided. One implementation MUST + be [`parquet-java`](http://github.com/apache/parquet-java). It is preferred + that the second implementation be + [`parquet-cpp`](https://github.com/apache/arrow) or + [`parquet-rs`](https://github.com/apache/arrow-rs), however at the discretion + of the PMC any open source Parquet implementation may be acceptable. + Implementations whose contributors actively participate in the community + (e.g. keep their feature matrix up-to-date on the Parquet website) are more + likely to be considered. If discussed as a requirement in step 1 above, + demonstration of integration with a query engine is also required for this + step. The implementations must be made available publicly, and they should be + fit for inclusion (for example, they were submitted as a pull request against + the target repository and committers gave positive reviews). + +Unless otherwise discussed, it is expected the implementations will be developed +from their respective main branch (i.e. backporting is not expected). + +3. After the first two steps are complete a formal vote is held on + [email protected] to officially ratify the feature. After the vote + passes the format change is merged into the `parquet-format` repository and + it is expected the changes from step 2 will also be merged soon after + (implementations should not be merged until the addition has been merged to + `parquet-format`). + +#### General guidelines/preferences on additions. + +1. To the greatest extent possible changes should have an option for forward + compatibility (old readers can still read files). Review Comment: We should avoid language like "To the greatest extent possible" which leaves to interpretation what the exceptions are. I would just write "Changes should have an option for forward compatibility (old readers can still read files)." and list specific exceptions: Non forward compatible changes happen in two ways: 1) after a deprecation cycle we remove a feature 2) after being default to off for some time we enable a non forward compatible change to on by default (for example a new encoding). If you prefer to add the detailes later later in the doc, then I would add a link here to point to the specifics. ########## CONTRIBUTING.md: ########## @@ -29,3 +29,138 @@ Recommendations and requirements for how to best contribute to Parquet. We striv ### License By contributing your code, you agree to license your contribution under the terms of the APLv2: https://github.com/apache/parquet-format/blob/master/LICENSE + +### Additions/Changes to the Format + +Note: This section applies to actual functional changes to the specification. +Fixing typos, grammar, and clarifying concepts that would not change the +semantics of the specification can be done as long a comitter feels comfortable +to merge them. When in doubt starting a discussion on the dev mailing list is +encouraged. + +The general steps for adding features to the format are as follows: + +1. Discuss changes on the developer mailing list ([email protected]). + Often times it is helpful to link to a draft pull request to make the + discussion concrete. This step is complete when there is lazy consensus. Part + of the consensus is whether it sufficient to provide 2 working + implementations as outlined in step 2 or if demonstration of the feature with + a down-stream query engine is necessary to justify the feature (e.g. + demonstrate performance improvements in Arrow's DataSet library or Apache + Data Fusion or another open source engine). + +2. Once a change has lazy consensus, two implementations of the feature + demonstrating interopability must also be provided. One implementation MUST + be [parquet-java](http://github.com/apache/parquet-java). It is preferred + that the second implementation be + [parquet-cpp](https://github.com/apache/arrow) or + [parquet-rs](https://github.com/apache/arrow-rs), however at the discretion + of the PMC any open source Parquet implementation may be acceptable. + Implementations whose contributors actively participate in the community + (e.g. keep their feature matrix up-to-date on parquet-site) are more likely + to be considered. If discussed as a requirement in step one, demonstration + of integration with a query engine is also required for this step. The + implementations must be made available publicly (e.g. as a pull request + against the target repository). + +Unless otherwise discussed, it is expected the implementations will develop from +the main branch (i.e. backporting is not expected). + +3. After the first two steps are complete a formal vote is held on the Parquet + mailing list to officially ratify the feature. After the vote passes the + format change is merged into the parquet-format repository and it is expected + the changes from step 2 will also be merged soon after. Before merging into + Parquet-java a parquet-format release must be performed. + +#### General guidelines/preferences on additions. + +1. To the greatest extent possible changes should have an option for forwards + compatibility (old readers can still read files). + +2. New encodings should be fully specified in this repository and ideally not + rely on an external dependencies for implementation (i.e. `parquet-format` is + the source of truth for the encoding). + +3. New compression mechanisms must have a pure Java implementation that can be + used as dependency in parquet-java. + +### Releases + +The Parquet PMC aims to do releases of the format package only as needed when +new features are introduced. If multiple new features are being proposed +simultaneously some features might be consolidated into the same release. +Guidance is provided below on when implementations should enable features added +to the specification. Due to confusion in the past over parquet versioning it +is not expected that there will be a 3.x release of the specification in the +foreseeable future. + +### Compatibility and Feature Enablement + +For the purposes of this discussion we classify features into the following buckets: + +1. Backwards compatible. A file written under an older version of the format + should be readable under a newer version of the format. + +2. Forwards compatible. A file written under a newer version of the format with + the enabled feature can be read under an older version of the format, but + some information might be missing or performance might be suboptimal. + +3. Forward incompatible. A file written under a new version of the format with + the feature enabled cannot be read under an older version of the format (e.g. + Adding a new compression algorithm). + +New features are intended to be widely beneficial to users of Parquet, and +therefore it is hoped third-party implementations will adopt them quickly after +they are introduced. It is assumed that writing new parts of the format, and +especially forward incompatible features, will be configured with feature flag +defaulted to "off" and at some future point the features are turned on by default +(reading of the new feature will typically be enabled without configuration or +defaulted to on). Some amount of lead time is desirable to ensure a critical +mass of Parquet implementations support a feature to avoid compatability issues +across the ecosystem. Therefore, the Parquet PMC gives the following +recommendations for managing features: + +1. Backwards compatibility is the concern of implementations but given the + ubiquity of Parquet and the length of time it has been used, libraries should + support reading older version of the format to the greatest extent possible. + +2. Forward compatible features/changes may be used by default in implementations + once the parquet-format containing those changes has been formally released. + For features that may pose a significant performance regression to older + format readers, libaries should consider delaying default enablement until 1 + year after the release of the parquet-java implementation that contains the + feature implementation. + +3. Forwards incompatible features/changes should not be turned on by default + until 2 years after the parquet-java implementation containing the feature is Review Comment: I would suggest the following guidance: In the Parquet reference implementations mentioned above (java, cpp, rust, ...) - forward incompatible features (ex: new encoding) are always off by default. - there is a setting to turn them on: enableEncodingFoo(true) - there is a setting to change the default: setForwardIncompatibleFeaturesOnByDefault(true) - We will change the default to on on a given feature, when a minimum amount of time is elapsed (say 6 months, 2 releases) and there is enough adoption (we should make a list here? ex: latest Flink, Spark, Trino releases updated to a version of parquet that supports it) Third party implementations are advised to: - have forward incompatible features off by default. - have a setting to turn them on: config.encoding_foo=enabled - They can decide to enable them immediately if no other system is consuming their files or if they know that all the readers are compatible. That's what "setForwardIncompatibleFeaturesOnByDefault(true)" is for if they use reference libraries. My goal here is to not artificially slow down adoption of new features when it is not needed. We need a transition period, hence the always off by default. We also want the shortest possible transition period. ########## CONTRIBUTING.md: ########## @@ -29,3 +29,142 @@ Recommendations and requirements for how to best contribute to Parquet. We striv ### License By contributing your code, you agree to license your contribution under the terms of the APLv2: https://github.com/apache/parquet-format/blob/master/LICENSE + +### Additions/Changes to the Format + +Note: This section applies to actual functional changes to the specification. +Fixing typos, grammar, and clarifying concepts that would not change the +semantics of the specification can be done as long as a committer feels comfortable +to merge them. When in doubt starting a discussion on the dev mailing list is +encouraged. + +The general steps for adding features to the format are as follows: + +1. Discuss changes on the developer mailing list ([email protected]). + Often times it is helpful to link to a draft pull request to make the + discussion concrete. This step is complete when there is lazy consensus. Part + of the consensus is whether it is sufficient to provide two working + implementations as outlined in step 2, or if demonstration of the feature with + a downstream query engine is necessary to justify the feature (e.g. + demonstrate performance improvements in the Apache Arrow C++ Dataset library, + the Apache DataFusion query engine, or any other open source engine). + +2. Once a change has lazy consensus, two implementations of the feature + demonstrating interopability must also be provided. One implementation MUST + be [`parquet-java`](http://github.com/apache/parquet-java). It is preferred + that the second implementation be + [`parquet-cpp`](https://github.com/apache/arrow) or + [`parquet-rs`](https://github.com/apache/arrow-rs), however at the discretion + of the PMC any open source Parquet implementation may be acceptable. + Implementations whose contributors actively participate in the community + (e.g. keep their feature matrix up-to-date on the Parquet website) are more + likely to be considered. If discussed as a requirement in step 1 above, + demonstration of integration with a query engine is also required for this + step. The implementations must be made available publicly, and they should be + fit for inclusion (for example, they were submitted as a pull request against + the target repository and committers gave positive reviews). + +Unless otherwise discussed, it is expected the implementations will be developed +from their respective main branch (i.e. backporting is not expected). + +3. After the first two steps are complete a formal vote is held on + [email protected] to officially ratify the feature. After the vote + passes the format change is merged into the `parquet-format` repository and + it is expected the changes from step 2 will also be merged soon after + (implementations should not be merged until the addition has been merged to + `parquet-format`). + +#### General guidelines/preferences on additions. + +1. To the greatest extent possible changes should have an option for forward + compatibility (old readers can still read files). + +2. New encodings should be fully specified in this repository and ideally not + rely on an external dependencies for implementation (i.e. `parquet-format` is + the source of truth for the encoding). + +3. New compression mechanisms must have a pure Java implementation that can be + used as a dependency in `parquet-java`. + +### Releases + +The Parquet PMC aims to do releases of the format package only as needed when +new features are introduced. If multiple new features are being proposed +simultaneously some features might be consolidated into the same release. +Guidance is provided below on when implementations should enable features added +to the specification. Due to confusion in the past over Parquet versioning it +is not expected that there will be a 3.x release of the specification in the +foreseeable future. + +### Compatibility and Feature Enablement + +For the purposes of this discussion we classify features into the following buckets: + +1. Backward compatible. A file written under an older version of the format + should be readable under a newer version of the format. + +2. Forward compatible. A file written under a newer version of the format with + the feature enabled can be read under an older version of the format, but + some information might be missing or performance might be suboptimal. + +3. Forward incompatible. A file written under a newer version of the format with + the feature enabled cannot be read under an older version of the format (e.g. + adding and using a new compression algorithm). It is expected any feature in + this category will provide a signal to older readers, so they can + unambiguously determine that they cannot properly read the file (e.g. via + adding a new value to an existing enum). + +New features are intended to be widely beneficial to users of Parquet, and +therefore it is hoped third-party implementations will adopt them quickly after +they are introduced. It is assumed that writing new parts of the format, and +especially forward incompatible features, will be configured with a feature flag +defaulted to "off", and at some future point the feature is turned on by default +(reading of the new feature will typically be enabled without configuration or +defaulted to on). Some amount of lead time is desirable to ensure a critical +mass of Parquet implementations support a feature to avoid compatibility issues +across the ecosystem. Therefore, the Parquet PMC gives the following +recommendations for managing features: + +1. Backward compatibility is the concern of implementations but given the + ubiquity of Parquet and the length of time it has been used, libraries should + support reading older versions of the format to the greatest extent possible. + +2. Forward compatible features/changes may be enabled and used by default in + implementations once the parquet-format containing those changes has been + formally released. For features that may pose a significant performance + regression to older format readers, libaries should consider delaying default + enablement until 1 year after the release of the parquet-java implementation + that contains the feature implementation. + +3. Forward incompatible features/changes should not be turned on by default + until 2 years after the parquet-java implementation containing the feature is + released. It is recommended that changing the default value for a forward + incompatible feature flag should be clearly advertised to consumers (e.g. via + a major version release if using Semantic Versioning, or highlighed in + release notes). + +For forward compatible changes which have a high chance of performance +regression for older readers and forward incompatible changes, implementations +should clearly document the compatibility issues. Additionally, while it is up +to maintainers of individual implementations to make the best decision to serve +their ecosystem, they are encouraged to start enabling features by default along +the same timelines as `parquet-java`. Parquet-java will wait to enable features +by default until the most conservative timelines outlined above have been +exceeded. Review Comment: > > I think they should be encourage implementations to enable it as soon as they are comfortable with it and no later than the library unless they have legacy constraints. > Is the library here parquet-java. Here, I meant the library they're using. So either parquet-java, parquet-rust or parquet-cpp depending on the context. > I agree for early adopters it is good, but these are most likely people who understand there dependency stack very well. I think the issue is even with release dates, these libraries take time to percolate through the ecosystem so being more conservative with language here makes sense to me. > If you can think of language that might strike a better balance here, I am open to suggestions. (Here, I'm considering that most third party implementation are actually integrated in a broader system or query engine) How about: _Additionally, while it is up to maintainers of individual implementations to make the best decision to serve their ecosystem, they are encouraged to start enabling features by default as soon as potential consumers of the files they produce will be able to read them. They can follow the same timelines as `parquet-java` but should consider that it follows a conservative timeline._ Is this pushing too much the other way? ########## CONTRIBUTING.md: ########## @@ -29,3 +29,142 @@ Recommendations and requirements for how to best contribute to Parquet. We striv ### License By contributing your code, you agree to license your contribution under the terms of the APLv2: https://github.com/apache/parquet-format/blob/master/LICENSE + +### Additions/Changes to the Format + +Note: This section applies to actual functional changes to the specification. +Fixing typos, grammar, and clarifying concepts that would not change the +semantics of the specification can be done as long as a committer feels comfortable +to merge them. When in doubt starting a discussion on the dev mailing list is +encouraged. + +The general steps for adding features to the format are as follows: + +1. Discuss changes on the developer mailing list ([email protected]). + Often times it is helpful to link to a draft pull request to make the + discussion concrete. This step is complete when there is lazy consensus. Part + of the consensus is whether it is sufficient to provide two working + implementations as outlined in step 2, or if demonstration of the feature with + a downstream query engine is necessary to justify the feature (e.g. + demonstrate performance improvements in the Apache Arrow C++ Dataset library, + the Apache DataFusion query engine, or any other open source engine). + +2. Once a change has lazy consensus, two implementations of the feature + demonstrating interopability must also be provided. One implementation MUST + be [`parquet-java`](http://github.com/apache/parquet-java). It is preferred + that the second implementation be + [`parquet-cpp`](https://github.com/apache/arrow) or + [`parquet-rs`](https://github.com/apache/arrow-rs), however at the discretion + of the PMC any open source Parquet implementation may be acceptable. + Implementations whose contributors actively participate in the community + (e.g. keep their feature matrix up-to-date on the Parquet website) are more + likely to be considered. If discussed as a requirement in step 1 above, + demonstration of integration with a query engine is also required for this + step. The implementations must be made available publicly, and they should be + fit for inclusion (for example, they were submitted as a pull request against + the target repository and committers gave positive reviews). + +Unless otherwise discussed, it is expected the implementations will be developed +from their respective main branch (i.e. backporting is not expected). + +3. After the first two steps are complete a formal vote is held on + [email protected] to officially ratify the feature. After the vote + passes the format change is merged into the `parquet-format` repository and + it is expected the changes from step 2 will also be merged soon after + (implementations should not be merged until the addition has been merged to + `parquet-format`). + +#### General guidelines/preferences on additions. + +1. To the greatest extent possible changes should have an option for forward + compatibility (old readers can still read files). + +2. New encodings should be fully specified in this repository and ideally not + rely on an external dependencies for implementation (i.e. `parquet-format` is + the source of truth for the encoding). + +3. New compression mechanisms must have a pure Java implementation that can be + used as a dependency in `parquet-java`. Review Comment: The PCodec discussion seems to be more of an Encoding in this instance? This statement makes more sense to me if we say "New encodings must have a pure Java implementation that can be used as a dependency in `parquet-java`." ########## CONTRIBUTING.md: ########## @@ -29,3 +29,142 @@ Recommendations and requirements for how to best contribute to Parquet. We striv ### License By contributing your code, you agree to license your contribution under the terms of the APLv2: https://github.com/apache/parquet-format/blob/master/LICENSE + +### Additions/Changes to the Format + +Note: This section applies to actual functional changes to the specification. +Fixing typos, grammar, and clarifying concepts that would not change the +semantics of the specification can be done as long as a committer feels comfortable +to merge them. When in doubt starting a discussion on the dev mailing list is +encouraged. + +The general steps for adding features to the format are as follows: + +1. Discuss changes on the developer mailing list ([email protected]). + Often times it is helpful to link to a draft pull request to make the + discussion concrete. This step is complete when there is lazy consensus. Part + of the consensus is whether it is sufficient to provide two working + implementations as outlined in step 2, or if demonstration of the feature with + a downstream query engine is necessary to justify the feature (e.g. + demonstrate performance improvements in the Apache Arrow C++ Dataset library, + the Apache DataFusion query engine, or any other open source engine). + +2. Once a change has lazy consensus, two implementations of the feature + demonstrating interopability must also be provided. One implementation MUST + be [`parquet-java`](http://github.com/apache/parquet-java). It is preferred + that the second implementation be + [`parquet-cpp`](https://github.com/apache/arrow) or + [`parquet-rs`](https://github.com/apache/arrow-rs), however at the discretion + of the PMC any open source Parquet implementation may be acceptable. + Implementations whose contributors actively participate in the community + (e.g. keep their feature matrix up-to-date on the Parquet website) are more + likely to be considered. If discussed as a requirement in step 1 above, + demonstration of integration with a query engine is also required for this + step. The implementations must be made available publicly, and they should be + fit for inclusion (for example, they were submitted as a pull request against + the target repository and committers gave positive reviews). + +Unless otherwise discussed, it is expected the implementations will be developed +from their respective main branch (i.e. backporting is not expected). + +3. After the first two steps are complete a formal vote is held on + [email protected] to officially ratify the feature. After the vote + passes the format change is merged into the `parquet-format` repository and + it is expected the changes from step 2 will also be merged soon after + (implementations should not be merged until the addition has been merged to + `parquet-format`). + +#### General guidelines/preferences on additions. + +1. To the greatest extent possible changes should have an option for forward + compatibility (old readers can still read files). + +2. New encodings should be fully specified in this repository and ideally not + rely on an external dependencies for implementation (i.e. `parquet-format` is + the source of truth for the encoding). + +3. New compression mechanisms must have a pure Java implementation that can be + used as a dependency in `parquet-java`. + +### Releases + +The Parquet PMC aims to do releases of the format package only as needed when +new features are introduced. If multiple new features are being proposed +simultaneously some features might be consolidated into the same release. +Guidance is provided below on when implementations should enable features added +to the specification. Due to confusion in the past over Parquet versioning it +is not expected that there will be a 3.x release of the specification in the +foreseeable future. + +### Compatibility and Feature Enablement + +For the purposes of this discussion we classify features into the following buckets: + +1. Backward compatible. A file written under an older version of the format + should be readable under a newer version of the format. + +2. Forward compatible. A file written under a newer version of the format with + the feature enabled can be read under an older version of the format, but + some information might be missing or performance might be suboptimal. + +3. Forward incompatible. A file written under a newer version of the format with + the feature enabled cannot be read under an older version of the format (e.g. + adding and using a new compression algorithm). It is expected any feature in + this category will provide a signal to older readers, so they can + unambiguously determine that they cannot properly read the file (e.g. via + adding a new value to an existing enum). + +New features are intended to be widely beneficial to users of Parquet, and +therefore it is hoped third-party implementations will adopt them quickly after +they are introduced. It is assumed that writing new parts of the format, and +especially forward incompatible features, will be configured with a feature flag +defaulted to "off", and at some future point the feature is turned on by default +(reading of the new feature will typically be enabled without configuration or +defaulted to on). Some amount of lead time is desirable to ensure a critical +mass of Parquet implementations support a feature to avoid compatibility issues +across the ecosystem. Therefore, the Parquet PMC gives the following +recommendations for managing features: + +1. Backward compatibility is the concern of implementations but given the + ubiquity of Parquet and the length of time it has been used, libraries should + support reading older versions of the format to the greatest extent possible. + +2. Forward compatible features/changes may be enabled and used by default in + implementations once the parquet-format containing those changes has been + formally released. For features that may pose a significant performance + regression to older format readers, libaries should consider delaying default + enablement until 1 year after the release of the parquet-java implementation + that contains the feature implementation. + +3. Forward incompatible features/changes should not be turned on by default + until 2 years after the parquet-java implementation containing the feature is + released. It is recommended that changing the default value for a forward + incompatible feature flag should be clearly advertised to consumers (e.g. via + a major version release if using Semantic Versioning, or highlighed in + release notes). Review Comment: It sounds reasonable to me that we use major version bumps in this way. I assumed we were still talking about "Parquet V3". I think we should allow ourselves to make major releases when needed and make it " at least once a year" rather than exactly once a year. Let's say this year, we release: - a new footer - a new encoding for time series - a new encoding for strings I would want to avoid having to wait more than a year to turn the default on if something is added too close to the yearly release. On the discussion of V2, do we have a clear list of what would be turned on by default and what encodings might not? From your analysis earlier, it seemed that some encodings maybe should not be on by default. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
