RussellSpitzer commented on code in PR #17822:
URL: https://github.com/apache/iceberg/pull/17822#discussion_r4136806632


##########
format/spec.md:
##########
@@ -654,6 +655,127 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a property that a table's rows are expected to 
satisfy. A constraint's definition is stored in table metadata. Whether a 
constraint holds is recorded for each snapshot, see [Constraint 
Validation](#constraint-validation).
+
+Iceberg does not evaluate constraints. Enforcement and validation are 
performed by engines that write to a table. Iceberg stores constraint 
definitions and records the status that a writer reports for a commit without 
verifying it.
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the non-null values of a set of fields must be distinct across 
all rows; more than one row may have a null value
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+Constraints are stored separately from schemas because they span multiple 
fields and evolve independently. Every constraint references the fields that it 
applies to by field ID, so a constraint continues to apply to the same columns 
after a column is renamed or reordered.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+
+When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A key that spans 
several fields is expressed as a single `primary-key` constraint over multiple 
`field-ids`.
+
+A `primary-key` constraint replaces [identifier field 
IDs](#identifier-field-ids), which express the same concept: a set of fields 
that identifies a row, without a uniqueness guarantee. When a table is upgraded 
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint 
that is not enforced. Identifier field IDs are not used in v4.
+
+Only a constraint's `name` and `enforced` fields may be changed in place. 
Changing the `expression` of a `check` constraint or the `field-ids` of a 
`unique` or `primary-key` constraint changes what the constraint requires, so 
it must be done by removing the constraint and adding a new one with a new 
`constraint-id`, so that statuses recorded for the old definition are not read 
as applying to the new one.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is 
treated as 0 when it is not present. Writers must assign a new constraint an ID 
that is higher than the table's current `last-constraint-id` and must update 
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be 
reused after the constraint that used an ID is removed, because retained 
snapshots may still reference the removed ID. Readers must not assume that 
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the 
[Iceberg expressions spec](expressions-spec.md) and must use ID references so 
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An 
expression may reference more than one field of the row, such as `start_date <= 
end_date`. Expressions that depend on more than one row, such as aggregates and 
window functions, and expressions that depend on another table, such as 
subqueries, must not be used.
+
+Iceberg predicates use two-valued logic: a predicate always produces true or 
false and never produces null, so a comparison with a null operand produces 
false. This differs from SQL `CHECK`, where a row satisfies a constraint unless 
the predicate produces false and a null value therefore satisfies the 
constraint.
+
+To express SQL `CHECK` semantics for an optional field, the stored expression 
must make the null case explicit. For example, SQL `CHECK (price >= 0)` for an 
optional `price` field is stored as the expression for `price >= 0 OR price IS 
NULL`. This is unnecessary for required fields, which can never be null.
+
+#### Constraints and Schema Evolution
+
+A constraint references fields by ID, so schema changes interact with 
constraints as follows. The referenced fields of a `check` constraint are the 
field IDs in its `expression`; the referenced fields of a `unique` or 
`primary-key` constraint are its `field-ids`.
+
+* Renaming or reordering a referenced field is allowed; the constraint 
continues to apply to the same fields.
+* If a dropped field is referenced only by single-column constraints, the drop 
is allowed and those constraints are removed automatically. If a dropped field 
is referenced by a multi-column constraint, the writer must reject the drop 
unless that constraint is removed in the same change.

Review Comment:
   I would also rephrase this to be proscriptive rather than reactive and 
describe valid states rather than the path to getting to them. For example it's 
an engine decision on whether or not to automatically remove a constraint when 
a field is removed or fail. The Spec invariant here I believe is 
   
   "All fields referred to by constraints MUST exist in the table's current 
schema. Field removals that would make it impossible to evaluate a constraint 
must be rejected unless the constraint is modified or removed."
   
   Then if we want to elaborate on engine behaviors
   
   "Writers may decided wether to automatically remove constraints when their 
corresponding fields are removed or may throw an error and require manual 
removal of the referencing constraints first."
   
   Although I probably wouldn't recommend an engines automatically removing 
constraints without manual intervention.



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to