pitrou commented on code in PR #258: URL: https://github.com/apache/parquet-format/pull/258#discussion_r1676148578
########## CONTRIBUTING.md: ########## @@ -29,3 +29,142 @@ Recommendations and requirements for how to best contribute to Parquet. We striv ### License By contributing your code, you agree to license your contribution under the terms of the APLv2: https://github.com/apache/parquet-format/blob/master/LICENSE + +### Additions/Changes to the Format + +Note: This section applies to actual functional changes to the specification. +Fixing typos, grammar, and clarifying concepts that would not change the +semantics of the specification can be done as long as a committer feels comfortable +to merge them. When in doubt starting a discussion on the dev mailing list is +encouraged. + +The general steps for adding features to the format are as follows: + +1. Discuss changes on the developer mailing list ([email protected]). + Often times it is helpful to link to a draft pull request to make the + discussion concrete. This step is complete when there is lazy consensus. Part + of the consensus is whether it is sufficient to provide two working + implementations as outlined in step 2, or if demonstration of the feature with + a downstream query engine is necessary to justify the feature (e.g. + demonstrate performance improvements in the Apache Arrow C++ Dataset library, + the Apache DataFusion query engine, or any other open source engine). + +2. Once a change has lazy consensus, two implementations of the feature + demonstrating interopability must also be provided. One implementation MUST + be [`parquet-java`](http://github.com/apache/parquet-java). It is preferred + that the second implementation be + [`parquet-cpp`](https://github.com/apache/arrow) or + [`parquet-rs`](https://github.com/apache/arrow-rs), however at the discretion + of the PMC any open source Parquet implementation may be acceptable. + Implementations whose contributors actively participate in the community + (e.g. keep their feature matrix up-to-date on the Parquet website) are more + likely to be considered. If discussed as a requirement in step 1 above, + demonstration of integration with a query engine is also required for this + step. The implementations must be made available publicly, and they should be + fit for inclusion (for example, they were submitted as a pull request against + the target repository and committers gave positive reviews). + +Unless otherwise discussed, it is expected the implementations will be developed +from their respective main branch (i.e. backporting is not expected). + +3. After the first two steps are complete a formal vote is held on + [email protected] to officially ratify the feature. After the vote + passes the format change is merged into the `parquet-format` repository and + it is expected the changes from step 2 will also be merged soon after + (implementations should not be merged until the addition has been merged to + `parquet-format`). + +#### General guidelines/preferences on additions. + +1. To the greatest extent possible changes should have an option for forward + compatibility (old readers can still read files). + +2. New encodings should be fully specified in this repository and ideally not + rely on an external dependencies for implementation (i.e. `parquet-format` is + the source of truth for the encoding). + +3. New compression mechanisms must have a pure Java implementation that can be + used as a dependency in `parquet-java`. + +### Releases + +The Parquet PMC aims to do releases of the format package only as needed when +new features are introduced. If multiple new features are being proposed +simultaneously some features might be consolidated into the same release. +Guidance is provided below on when implementations should enable features added +to the specification. Due to confusion in the past over Parquet versioning it +is not expected that there will be a 3.x release of the specification in the +foreseeable future. + +### Compatibility and Feature Enablement + +For the purposes of this discussion we classify features into the following buckets: + +1. Backward compatible. A file written under an older version of the format + should be readable under a newer version of the format. + +2. Forward compatible. A file written under a newer version of the format with + the feature enabled can be read under an older version of the format, but + some information might be missing or performance might be suboptimal. Review Comment: If the data should be readable back "without loss", does it imply that new logical types are not considered forward compatible? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
