RussellSpitzer commented on code in PR #17198: URL: https://github.com/apache/iceberg/pull/17198#discussion_r3582389206
########## format/spec.md: ########## @@ -700,6 +700,10 @@ The `sequence_number` field represents the data sequence number and must never c The `file_sequence_number` field represents the sequence number of the snapshot that added the file and must also remain unchanged upon assigning at commit. The file sequence number can't be used for pruning delete files as the data within the file may have an older data sequence number. The data and file sequence numbers are inherited only if the entry status is 1 (added). If the entry status is 0 (existing) or 2 (deleted), the entry must include both sequence numbers explicitly. +#### Content file uniqueness + +Within a snapshot, each live manifest entry must be uniquely defined by `file_path` across all manifest files. A snapshot with multiple live entries for the same `file_path` has undefined behavior. Writers are not required to validate uniqueness because doing so can be expensive at commit time. Review Comment: I'm trying to write this in a way in that no one actually has to validate uniqueness. But you should be doing what you can to guarantee you are writing unique paths. A writer that names all it's files "foo.parquet" would be obviously noncompliant, or a writer that names all it's files "fileN.parquet" where N is a number starting at 0 per process, would also be broken. The statement here is more of a way to clearly be able to have something to point to and say "this implementation is broken" -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
