huaxingao commented on code in PR #17822:
URL: https://github.com/apache/iceberg/pull/17822#discussion_r4161877902


##########
format/spec.md:
##########
@@ -655,6 +658,123 @@ Sorting floating-point numbers should produce the 
following behavior: `-NaN` < `
 
 A data or delete file is associated with a sort order by the sort order's id 
within [a manifest](#manifests). Therefore, the table must declare all the sort 
orders for lookup. A table could also be configured with a default sort order 
id, indicating how the new data should be sorted by default. Writers should use 
this default sort order to sort the data on write, but are not required to if 
the default order is prohibitively expensive, as it would be for streaming 
writes.
 
+### Constraints
+
+Constraints are added in v4 and are not supported in v3 or earlier.
+
+A **constraint** declares a condition that a table's rows are expected to 
satisfy. Its definition is stored in table metadata. Each snapshot contains the 
writer-reported status of each constraint that existed when the snapshot was 
created; see [Constraint Validation](#constraint-validation).
+
+Three constraint types are defined:
+
+* `check` -- every row must satisfy a predicate
+* `unique` -- the values of a set of fields must be distinct across the rows 
in which every field in the set is non-null; a row in which any field in the 
set is null is not compared, so more than one such row may exist
+* `primary-key` -- the values of a set of fields must be distinct across all 
rows and must not be null
+
+A constraint is stored separately from the table schema and must refer to 
fields by field ID. Constraints may evolve independently from the schema, being 
added, modified, or removed.
+
+#### Constraint Fields
+
+A constraint consists of the following fields:
+
+| Requirement | Field name                | Type      | Description |
+|-------------|---------------------------|-----------|-------------|
+| _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
+| _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
+| _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
only add rows that satisfy the constraint |
+| _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
+| _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
+| _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
+
+The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
+
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields and must not reference fields within a `list` or a `map`. 
Fields of type `float`, `double`, `geometry`, or `geography` must not be used 
because Iceberg does not define equality for them. Fields of type `unknown` 
must not be used because their values are always null.
+
+The fields of a `primary-key` constraint must be `required` and must not be 
nested in an optional struct, so that the schema rules out null keys even when 
the constraint is not enforced. These are the same restrictions that apply to 
[identifier fields](#identifier-field-ids). The fields of a `unique` constraint 
may be optional or nested in optional structs, because a row in which any field 
is null is not compared.
+
+When a constraint is `enforced`, writers must only add rows they can prove 
follow the constraint. A constraint that is not enforced can be written to by 
any writer, regardless of their ability to prove that the constraint is 
followed by new rows.
+
+When a commit is retried against a new parent snapshot, the rows that a writer 
adds must still satisfy the table's enforced constraints. A `check` constraint 
is evaluated for each row on its own, so a retry cannot change its result. A 
`unique` or `primary-key` constraint compares rows to each other, so a writer 
must verify the rows that it adds against the new parent; verifying them 
against the original parent is not sufficient. A writer must also verify the 
rows that it adds against any constraint that a concurrent commit added or 
changed to enforced, and must recompute `constraint-statuses` from the new 
parent.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+
+A table may have at most one `primary-key` constraint. A `primary-key` 
constraint replaces [identifier field IDs](#identifier-field-ids), which 
express the same concept: a set of fields that identifies a row, without a 
uniqueness guarantee. Identifier field IDs are not used in v4.
+
+When a table is upgraded to v4, the `identifier-field-ids` of the table's 
current schema are rewritten as a single `primary-key` constraint that is not 
enforced. The constraint is assigned a `constraint-id` from 
`last-constraint-id` in the same way as any other constraint, its `name` is 
`pk`, and its `timestamp-ms` is the time of the upgrade. A table whose current 
schema has no identifier fields has no constraints until they are added. 
Writers must not set `identifier-field-ids` in a schema that is added to a v4 
table, and readers must ignore `identifier-field-ids` in a v4 table.
+
+Constraint IDs are assigned from the table's `last-constraint-id`, which is 
treated as 0 when it is not present. Writers must assign a new constraint an ID 
that is higher than the table's current `last-constraint-id` and must update 
`last-constraint-id` to the highest assigned ID. Constraint IDs must not be 
reused after the constraint that used an ID is removed, because retained 
snapshots may still reference the removed ID. Readers must not assume that 
every `constraint-id` referenced by a snapshot is present in `constraints`.
+
+#### Check Constraint Expressions
+
+The `expression` of a `check` constraint is serialized as described in the 
[Iceberg expressions spec](expressions-spec.md) and must use ID references so 
that it remains bound to the same fields when columns are renamed or reordered.
+
+A check expression is evaluated for each row over the values of that row. An 
expression may reference more than one field of the row, such as `start_date <= 
end_date`. Expressions that depend on more than one row, such as aggregates and 
window functions, and expressions that depend on another table, such as 
subqueries, must not be used.
+
+A check expression must produce the same result every time it is evaluated for 
the same row. A function that depends on anything other than its arguments, 
such as the current time or a random value, must not be called, because the 
status recorded for a snapshot describes the table's data and an expression 
whose result can change on its own would make a recorded status wrong without 
any write. A [user-defined function](udf-spec.md) records whether it is 
deterministic.

Review Comment:
   Applied the change. Thanks!



##########
format/spec.md:
##########
@@ -680,24 +681,26 @@ A constraint consists of the following fields:
 | _required_ | **`constraint-id`**       | `int`     | ID of the constraint; 
unique within the table |
 | _required_ | **`type`**                | `string`  | The constraint type: 
`check`, `unique`, or `primary-key` |
 | _required_ | **`name`**                | `string`  | A name for the 
constraint that is unique within the table. Names are for human consumption and 
must not be used to identify a constraint in metadata |
-| _required_ | **`enforced`**            | `boolean` | Whether writers must 
verify that the rows they add satisfy the constraint |
+| _required_ | **`enforced`**            | `boolean` | Whether writers must 
only add rows that satisfy the constraint |
 | _required_ | **`timestamp-ms`**        | `long`    | Timestamp in 
milliseconds from the unix epoch when the constraint was created or last 
modified. The timestamp is informational and must not be used to determine 
whether a constraint applies to a snapshot or whether it holds |
 | _optional_ | **`expression`**          | `expression` | The predicate that 
every row must satisfy, see [Check Constraint 
Expressions](#check-constraint-expressions). Required for a `check` constraint 
and must not be set for other types |
 | _optional_ | **`field-ids`**           | `list<int>`  | The list of field 
IDs that the constraint applies to. Required for a `unique` or `primary-key` 
constraint and must not be set for a `check` constraint |
 
 The fields that define what a constraint requires are embedded directly in the 
constraint based on its `type`. Each type carries only the metadata that it 
requires: a `check` constraint has an `expression` and must not declare 
`field-ids`, and a `unique` or `primary-key` constraint has `field-ids` and 
must not declare an `expression`. This keeps a single source of truth for the 
fields that a constraint references.
 
-The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields that are either top-level fields or nested in required 
structs, and must not reference fields within a `list` or a `map`. These are 
the same restrictions that apply to [identifier fields](#identifier-field-ids).
+The `field-ids` of a `unique` or `primary-key` constraint must reference 
primitive fields and must not reference fields within a `list` or a `map`. 
Fields of type `float`, `double`, `geometry`, or `geography` must not be used 
because Iceberg does not define equality for them. Fields of type `unknown` 
must not be used because their values are always null.
 
-When a constraint is `enforced`, writers must verify that the rows they add 
satisfy the constraint and must fail the write if they do not. A writer that 
cannot verify an enforced constraint must reject writes to the table. When a 
constraint is not enforced, writers are not required to verify the rows they 
add.
+The fields of a `primary-key` constraint must be `required` and must not be 
nested in an optional struct, so that the schema rules out null keys even when 
the constraint is not enforced. These are the same restrictions that apply to 
[identifier fields](#identifier-field-ids). The fields of a `unique` constraint 
may be optional or nested in optional structs, because a row in which any field 
is null is not compared.
 
-Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
+When a constraint is `enforced`, writers must only add rows they can prove 
follow the constraint. A constraint that is not enforced can be written to by 
any writer, regardless of their ability to prove that the constraint is 
followed by new rows.
 
-A table may have at most one `primary-key` constraint. A key that spans 
several fields is expressed as a single `primary-key` constraint over multiple 
`field-ids`.
+When a commit is retried against a new parent snapshot, the rows that a writer 
adds must still satisfy the table's enforced constraints. A `check` constraint 
is evaluated for each row on its own, so a retry cannot change its result. A 
`unique` or `primary-key` constraint compares rows to each other, so a writer 
must verify the rows that it adds against the new parent; verifying them 
against the original parent is not sufficient. A writer must also verify the 
rows that it adds against any constraint that a concurrent commit added or 
changed to enforced, and must recompute `constraint-statuses` from the new 
parent.
+
+Whether to trust a constraint that is not enforced is left to engines and is 
not tracked in table metadata.
 
-A `primary-key` constraint replaces [identifier field 
IDs](#identifier-field-ids), which express the same concept: a set of fields 
that identifies a row, without a uniqueness guarantee. When a table is upgraded 
to v4, its `identifier-field-ids` are rewritten as a `primary-key` constraint 
that is not enforced. Identifier field IDs are not used in v4.
+A table may have at most one `primary-key` constraint. A `primary-key` 
constraint replaces [identifier field IDs](#identifier-field-ids), which 
express the same concept: a set of fields that identifies a row, without a 
uniqueness guarantee. Identifier field IDs are not used in v4.

Review Comment:
   Fixed. Thanks!



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to