huaxingao commented on code in PR #17822: URL: https://github.com/apache/iceberg/pull/17822#discussion_r4163297184
########## format/spec.md: ########## @@ -655,6 +658,123 @@ Sorting floating-point numbers should produce the following behavior: `-NaN` < ` A data or delete file is associated with a sort order by the sort order's id within [a manifest](#manifests). Therefore, the table must declare all the sort orders for lookup. A table could also be configured with a default sort order id, indicating how the new data should be sorted by default. Writers should use this default sort order to sort the data on write, but are not required to if the default order is prohibitively expensive, as it would be for streaming writes. +### Constraints + +Constraints are added in v4 and are not supported in v3 or earlier. + +A **constraint** declares a condition that a table's rows are expected to satisfy. Its definition is stored in table metadata. Each snapshot contains the writer-reported status of each constraint that existed when the snapshot was created; see [Constraint Validation](#constraint-validation). + +Three constraint types are defined: + +* `check` -- every row must satisfy a predicate +* `unique` -- the values of a set of fields must be distinct across the rows in which every field in the set is non-null; a row in which any field in the set is null is not compared, so more than one such row may exist +* `primary-key` -- the values of a set of fields must be distinct across all rows and must not be null + +A constraint is stored separately from the table schema and must refer to fields by field ID. Constraints may evolve independently from the schema, being added, modified, or removed. + +#### Constraint Fields + +A constraint consists of the following fields: + +| Requirement | Field name | Type | Description | +|-------------|---------------------------|-----------|-------------| +| _required_ | **`constraint-id`** | `int` | ID of the constraint; unique within the table | +| _required_ | **`type`** | `string` | The constraint type: `check`, `unique`, or `primary-key` | +| _required_ | **`name`** | `string` | A name for the constraint that is unique within the table. Names are for human consumption and must not be used to identify a constraint in metadata | +| _required_ | **`enforced`** | `boolean` | Whether writers must only add rows that satisfy the constraint | +| _required_ | **`timestamp-ms`** | `long` | Timestamp in milliseconds from the unix epoch when the constraint was created or last modified. The timestamp is informational and must not be used to determine whether a constraint applies to a snapshot or whether it holds | +| _optional_ | **`expression`** | `expression` | The predicate that every row must satisfy, see [Check Constraint Expressions](#check-constraint-expressions). Required for a `check` constraint and must not be set for other types | +| _optional_ | **`field-ids`** | `list<int>` | The list of field IDs that the constraint applies to. Required for a `unique` or `primary-key` constraint and must not be set for a `check` constraint | + +The fields that define what a constraint requires are embedded directly in the constraint based on its `type`. Each type carries only the metadata that it requires: a `check` constraint has an `expression` and must not declare `field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and must not declare an `expression`. This keeps a single source of truth for the fields that a constraint references. + +The `field-ids` of a `unique` or `primary-key` constraint must reference primitive fields and must not reference fields within a `list` or a `map`. Fields of type `float`, `double`, `geometry`, or `geography` must not be used because Iceberg does not define equality for them. Fields of type `unknown` must not be used because their values are always null. + +The fields of a `primary-key` constraint must be `required` and must not be nested in an optional struct, so that the schema rules out null keys even when the constraint is not enforced. These are the same restrictions that apply to [identifier fields](#identifier-field-ids). The fields of a `unique` constraint may be optional or nested in optional structs, because a row in which any field is null is not compared. + +When a constraint is `enforced`, writers must only add rows they can prove follow the constraint. A constraint that is not enforced can be written to by any writer, regardless of their ability to prove that the constraint is followed by new rows. + +When a commit is retried against a new parent snapshot, the rows that a writer adds must still satisfy the table's enforced constraints. A `check` constraint is evaluated for each row on its own, so a retry cannot change its result. A `unique` or `primary-key` constraint compares rows to each other, so a writer must verify the rows that it adds against the new parent; verifying them against the original parent is not sufficient. A writer must also verify the rows that it adds against any constraint that a concurrent commit added or changed to enforced, and must recompute `constraint-statuses` from the new parent. + +Whether to trust a constraint that is not enforced is left to engines and is not tracked in table metadata. + +A table may have at most one `primary-key` constraint. A `primary-key` constraint replaces [identifier field IDs](#identifier-field-ids), which express the same concept: a set of fields that identifies a row. Iceberg never guaranteed uniqueness for identifier fields and did not track whether it held. A `primary-key` constraint makes both expressible, through `enforced` and `constraint-statuses`. Identifier field IDs are not used in v4. + +When a table is upgraded to v4, the `identifier-field-ids` of the table's current schema are rewritten as a single `primary-key` constraint that is not enforced. The constraint is assigned a `constraint-id` from `last-constraint-id` in the same way as any other constraint, its `name` is `pk`, and its `timestamp-ms` is the time of the upgrade. A table whose current schema has no identifier fields has no constraints until they are added. Writers must not set `identifier-field-ids` in a schema that is added to a v4 table, and readers must ignore `identifier-field-ids` in a v4 table. + +Constraint IDs are assigned from the table's `last-constraint-id`, which is treated as 0 when it is not present. Writers must assign a new constraint an ID that is higher than the table's current `last-constraint-id` and must update `last-constraint-id` to the highest assigned ID. Constraint IDs must not be reused after the constraint that used an ID is removed, because retained snapshots may still reference the removed ID. Readers must not assume that every `constraint-id` referenced by a snapshot is present in `constraints`. + +#### Check Constraint Expressions + +The `expression` of a `check` constraint is serialized as described in the [Iceberg expressions spec](expressions-spec.md) and must use ID references so that it remains bound to the same fields when columns are renamed or reordered. + +A check expression is evaluated for each row over the values of that row. An expression may reference more than one field of the row, such as `start_date <= end_date`. Expressions that depend on more than one row, such as aggregates and window functions, and expressions that depend on another table, such as subqueries, must not be used. + +A check expression must produce the same result every time it is evaluated for the same row. A function that depends on anything other than its arguments, such as the current time or a random value, must not be called, because the status recorded for a snapshot describes the table's data and an expression whose result can change on its own would make a recorded status wrong without any write. A [user-defined function](udf-spec.md) must not be called unless it declares `deterministic` as true. + +Changing the definition of a function that a check expression calls changes what the constraint requires even though the constraint itself is unchanged. Statuses recorded before the change do not describe the new definition, so a writer that changes such a function should validate the constraint again. + +Iceberg predicates use two-valued logic and null-safe comparisons, as defined in the [expressions spec](expressions-spec.md#boolean-logic). This differs from SQL `CHECK`, where a row satisfies a constraint unless the predicate produces false. + +To express SQL `CHECK` semantics for an optional field, the stored expression must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an optional `price` field is stored as the expression for `price >= 0 OR price IS NULL`. This is unnecessary for required fields, which can never be null. + +#### Constraints and Schema Evolution + +A constraint references fields by ID, so schema changes interact with constraints as follows. The referenced fields of a `check` constraint are the field IDs in its `expression`; the referenced fields of a `unique` or `primary-key` constraint are its `field-ids`. + +* Renaming or reordering a referenced field is allowed; the constraint continues to apply to the same fields. +* Every field referenced by a constraint must exist in the table's current schema. A schema change that removes a referenced field must be rejected unless the constraint is removed in the same change. +* The type of a field referenced by a `check` constraint must not be changed, even for type promotions that are otherwise allowed, because the constraint's `expression` contains literals that are bound to the field's type. The type of a field referenced by a `unique` or `primary-key` constraint may be changed by a type promotion, which preserves values and therefore preserves distinctness. +* A field referenced by a `primary-key` constraint must not be made optional, and a struct that contains such a field must not be made optional, because a `primary-key` requires its fields to be non-null. + +Only a constraint's `name` and `enforced` fields may be changed in place, and a writer that changes either must update `timestamp-ms`. Changing the `expression` of a `check` constraint or the `field-ids` of a `unique` or `primary-key` constraint changes what the constraint requires, so it must be done by removing the constraint and adding a new one with a new `constraint-id`, so that statuses recorded for the old definition are not read as applying to the new one. + +#### Constraint Validation + +Enforcement and validation are separate properties. Whether writers must verify the rows that they add is a property of a constraint, tracked by `enforced`. Whether a table is known to satisfy a constraint is a property of a table's data, tracked per snapshot by `constraint-statuses`. Review Comment: Thanks for your suggestion! Applied. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
