RussellSpitzer commented on code in PR #17822: URL: https://github.com/apache/iceberg/pull/17822#discussion_r4136674230
########## format/spec.md: ########## @@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the following behavior: `-NaN` < ` A data or delete file is associated with a sort order by the sort order's id within [a manifest](#manifests). Therefore, the table must declare all the sort orders for lookup. A table could also be configured with a default sort order id, indicating how the new data should be sorted by default. Writers should use this default sort order to sort the data on write, but are not required to if the default order is prohibitively expensive, as it would be for streaming writes. +### Constraints + +Constraints are added in v4 and are not supported in v3 or earlier. + +A **constraint** declares a property that a table's rows are expected to satisfy. A constraint's definition is stored in table metadata. Whether a constraint holds is recorded for each snapshot, see [Constraint Validation](#constraint-validation). + +Iceberg does not evaluate constraints. Enforcement and validation are performed by engines that write to a table. Iceberg stores constraint definitions and records the status that a writer reports for a commit without verifying it. + +Three constraint types are defined: + +* `check` -- every row must satisfy a predicate +* `unique` -- the non-null values of a set of fields must be distinct across all rows; more than one row may have a null value +* `primary-key` -- the values of a set of fields must be distinct across all rows and must not be null + +Constraints are stored separately from schemas because they span multiple fields and evolve independently. Every constraint references the fields that it applies to by field ID, so a constraint continues to apply to the same columns after a column is renamed or reordered. + +#### Constraint Fields + +A constraint consists of the following fields: + +| Requirement | Field name | Type | Description | +|-------------|---------------------------|-----------|-------------| +| _required_ | **`constraint-id`** | `int` | ID of the constraint; unique within the table | +| _required_ | **`type`** | `string` | The constraint type: `check`, `unique`, or `primary-key` | +| _required_ | **`name`** | `string` | A name for the constraint that is unique within the table. Names are for human consumption and must not be used to identify a constraint in metadata | +| _required_ | **`enforced`** | `boolean` | Whether writers must verify that the rows they add satisfy the constraint | +| _required_ | **`timestamp-ms`** | `long` | Timestamp in milliseconds from the unix epoch when the constraint was created or last modified. The timestamp is informational and must not be used to determine whether a constraint applies to a snapshot or whether it holds | +| _optional_ | **`expression`** | `expression` | The predicate that every row must satisfy, see [Check Constraint Expressions](#check-constraint-expressions). Required for a `check` constraint and must not be set for other types | +| _optional_ | **`field-ids`** | `list<int>` | The list of field IDs that the constraint applies to. Required for a `unique` or `primary-key` constraint and must not be set for a `check` constraint | + +The fields that define what a constraint requires are embedded directly in the constraint based on its `type`. Each type carries only the metadata that it requires: a `check` constraint has an `expression` and must not declare `field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and must not declare an `expression`. This keeps a single source of truth for the fields that a constraint references. + +The `field-ids` of a `unique` or `primary-key` constraint must reference primitive fields that are either top-level fields or nested in required structs, and must not reference fields within a `list` or a `map`. These are the same restrictions that apply to [identifier fields](#identifier-field-ids). + +When a constraint is `enforced`, writers must verify that the rows they add satisfy the constraint and must fail the write if they do not. A writer that cannot verify an enforced constraint must reject writes to the table. When a constraint is not enforced, writers are not required to verify the rows they add. + +Whether to trust a constraint that is not enforced is left to engines and is not tracked in table metadata. + +A table may have at most one `primary-key` constraint. A key that spans several fields is expressed as a single `primary-key` constraint over multiple `field-ids`. + +A `primary-key` constraint replaces [identifier field IDs](#identifier-field-ids), which express the same concept: a set of fields that identifies a row, without a uniqueness guarantee. When a table is upgraded to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint that is not enforced. Identifier field IDs are not used in v4. + +Only a constraint's `name` and `enforced` fields may be changed in place. Changing the `expression` of a `check` constraint or the `field-ids` of a `unique` or `primary-key` constraint changes what the constraint requires, so it must be done by removing the constraint and adding a new one with a new `constraint-id`, so that statuses recorded for the old definition are not read as applying to the new one. Review Comment: This probably can go in a little section on evolution? -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
