stevenzwu commented on code in PR #16025:
URL: https://github.com/apache/iceberg/pull/16025#discussion_r4020598232
##########
format/spec.md:
##########
@@ -656,15 +662,33 @@ A data or delete file is associated with a sort order by
the sort order's id wit
### Manifests
-A manifest is an immutable Avro file that lists data files or delete files,
along with each file’s partition data tuple, metrics, and tracking information.
One or more manifest files are used to store a [snapshot](#snapshots), which
tracks all of the files in a table at some point in time. Manifests are tracked
by a [manifest list](#manifest-lists) for each table snapshot.
+A manifest is an immutable file that lists data files or delete files, along
with each file’s partition data, metrics, and tracking information. One or more
manifest files are used to store a [snapshot](#snapshots), which tracks all of
the files in a table at some point in time. Manifests are tracked by a snapshot
root for each table snapshot. In v4, the snapshot root is a root manifest that
may track data files in addition to leaf manifest files.
Review Comment:
should we complete this paragraph with "In v1-v3, the snapshot root is a
manifest list file that tracks manifest files."?
##########
format/spec.md:
##########
@@ -676,13 +700,19 @@ A manifest file must store the partition spec and other
metadata as properties i
| _optional_ | _required_ | `format-version` | Table format version
number of the manifest as a string
|
| | _required_ | `content` | Type of content files
tracked by the manifest: "data" or "deletes"
|
+=== "v4"
+ | Requirement | Key | Value
|
+
|-------------|---------------------|---------------------------------------------------------------------------------------------------------------------------------------------|
+ | _optional_ | `schema-id` | ID of the schema used to write the
manifest as a string
|
+ | _optional_ | `format-version` | Table format version number of the
manifest as a string
|
+
#### Content file uniqueness
Within a snapshot, each content file must be referenced by at most one live
manifest entry across all manifests; otherwise, the snapshot has undefined
behavior. Writers should not produce multiple manifest entries for the same
content file in a snapshot (for example, both ADDED and DELETED entries for the
same file). Writers are not required to validate uniqueness at commit time.
-#### Manifest Entry Fields
+#### Entries in Manifests
-The schema of a manifest file is defined by the `manifest_entry` struct, which
consists of the following fields:
+In v1-v3, manifest entries are described by the `manifest_entry` struct. In
v4, entries are called tracked files and are described by the `tracked_file`
struct. In v4, `data_file` struct fields are flattened directly into the
tracked file, and tracking fields are grouped into a nested `tracking` struct.
Review Comment:
> In v4, `data_file` struct fields are flattened directly into the tracked
file, and tracking fields are grouped into a nested `tracking` struct.
not sure if we need to say "data_file struct files are flattened directly
into ..." and the "tracking fields ...". It probably won't hurt to remove this
sentence.
##########
format/spec.md:
##########
@@ -656,15 +662,33 @@ A data or delete file is associated with a sort order by
the sort order's id wit
### Manifests
-A manifest is an immutable Avro file that lists data files or delete files,
along with each file’s partition data tuple, metrics, and tracking information.
One or more manifest files are used to store a [snapshot](#snapshots), which
tracks all of the files in a table at some point in time. Manifests are tracked
by a [manifest list](#manifest-lists) for each table snapshot.
+A manifest is an immutable file that lists data files or delete files, along
with each file’s partition data, metrics, and tracking information. One or more
manifest files are used to store a [snapshot](#snapshots), which tracks all of
the files in a table at some point in time. Manifests are tracked by a snapshot
root for each table snapshot. In v4, the snapshot root is a root manifest that
may track data files in addition to leaf manifest files.
A manifest is a valid Iceberg data file: files must use valid Iceberg formats,
schemas, and column projection.
-A manifest may store either data files or delete files, but not both because
manifests that contain delete files are scanned first during job planning.
Whether a manifest is a data manifest or a delete manifest is stored in
manifest metadata.
+Each manifest type contains the following content:
-A manifest stores files for a single partition spec. When a table’s partition
spec changes, old files remain in the older manifest and newer files are
written to a new manifest. This is required because a manifest file’s schema is
based on its partition spec (see below). The partition spec of each manifest is
also used to transform predicates on the table's data rows into predicates on
partition values that are used during job planning to select files from a
manifest.
+| Manifest type | Contents |
+|----------------|----------|
+| v1-v3 data manifest | Data files |
+| v2-v3 delete manifest | Delete files |
+| v4 root manifest (snapshot root) | Data files, data manifests, delete
manifests |
+| v4 data manifest | Data files and their colocated deletion vectors |
Review Comment:
do we also need to include column files, like `their colocated deletion
vectors and column files`?
##########
format/spec.md:
##########
@@ -656,15 +662,33 @@ A data or delete file is associated with a sort order by
the sort order's id wit
### Manifests
-A manifest is an immutable Avro file that lists data files or delete files,
along with each file’s partition data tuple, metrics, and tracking information.
One or more manifest files are used to store a [snapshot](#snapshots), which
tracks all of the files in a table at some point in time. Manifests are tracked
by a [manifest list](#manifest-lists) for each table snapshot.
+A manifest is an immutable file that lists data files or delete files, along
with each file’s partition data, metrics, and tracking information. One or more
manifest files are used to store a [snapshot](#snapshots), which tracks all of
the files in a table at some point in time. Manifests are tracked by a snapshot
root for each table snapshot. In v4, the snapshot root is a root manifest that
may track data files in addition to leaf manifest files.
A manifest is a valid Iceberg data file: files must use valid Iceberg formats,
schemas, and column projection.
-A manifest may store either data files or delete files, but not both because
manifests that contain delete files are scanned first during job planning.
Whether a manifest is a data manifest or a delete manifest is stored in
manifest metadata.
+Each manifest type contains the following content:
-A manifest stores files for a single partition spec. When a table’s partition
spec changes, old files remain in the older manifest and newer files are
written to a new manifest. This is required because a manifest file’s schema is
based on its partition spec (see below). The partition spec of each manifest is
also used to transform predicates on the table's data rows into predicates on
partition values that are used during job planning to select files from a
manifest.
+| Manifest type | Contents |
+|----------------|----------|
+| v1-v3 data manifest | Data files |
+| v2-v3 delete manifest | Delete files |
+| v4 root manifest (snapshot root) | Data files, data manifests, delete
manifests |
+| v4 data manifest | Data files and their colocated deletion vectors |
-A manifest file must store the partition spec and other metadata as properties
in the Avro file's key-value metadata:
+In v2-v3, data and delete files are kept in separate manifests because
manifests that contain delete files are scanned first during job planning.
Whether a manifest is a data manifest or a delete manifest is stored in
manifest metadata.
Review Comment:
> because manifests that contain delete files are scanned first during job
planning.
This is a result, not a cause for separate data and delete manifests.
##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
| | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
| | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
+ 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
+ 3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 4. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ **Tracked Files**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4:
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+ | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* |
Writer format version. |
+ | 100 | **`location`** | `string` | *required* | Location of the file or
manifest. |
+ | 101 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, `parquet`, or `puffin` |
+ | 147 | **`tracking`** | `tracking` struct | *required* | Groups status,
snapshot, and sequence number. See tracking struct below. |
+ | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used
to write this manifest or data file. |
+ | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort
order for this file. If missing or unknown, the order is assumed to be
unsorted. |
+ | 103 | **`record_count`** | `long` | *required* | Number of records in
this file. |
+ | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size
in bytes. |
+ | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column
stats. See [Content Stats](#content-stats). |
+ | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See
manifest_info struct below. |
+ | 131 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split
offsets for the data file. Must be sorted ascending. |
+ | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* |
Row-level deletion vector for a data file. |
+ | 158 | **`column_files`** | `list<159: column_file>` | *optional* |
Column update files associated with this entry. |
+
+ **`tracking` struct (field 147)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3:
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions,
replacements, and modifications. Deletes are not used in scans. |
+ | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file
was added or deleted. Inherited when null. |
+ | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the
deletion vector was added. |
+ | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* |
Snapshot ID where the latest column file was added. |
+ | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number
of the file. Inherited when null and status is 1 (ADDED). |
+ | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence
number indicating when the file was added. Inherited when null and status is
ADDED. |
+ | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the
`_row_id` for its first row. For a data manifest, the starting `_row_id` to
assign to rows added by ADDED data files. See [First Row ID
Inheritance](#first-row-id-inheritance). |
+ | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted
in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 7 | **`replaced_positions`** | `binary` | *optional* | Positions
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+
+ **`deletion_vector` struct (field 148)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 155 | **`location`** | `string` | *required* | Location of the Puffin
file. |
+ | 144 | **`offset`** | `long` | *required* | Offset in the file where the
content starts. |
+ | 145 | **`size_in_bytes`** | `long` | *required* | Length of the
referenced content stored in the file. |
+ | 156 | **`cardinality`** | `long` | *required* | Cardinality of the
deletion vector. |
+ | 149 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+
+ **`manifest_info` struct (field 150)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 504 | **`added_files_count`** | `int` | *required* | Count of entries
with status ADDED in the manifest. |
+ | 505 | **`existing_files_count`** | `int` | *required* | Count of entries
with status EXISTING in the manifest. |
+ | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries
with status DELETED in the manifest. |
+ | 520 | **`replaced_files_count`** | `int` | *required* | Count of entries
with status REPLACED in the manifest. |
+ | 524 | **`modified_files_count`** | `int` | *required* | Count of entries
with status MODIFIED in the manifest. |
+ | 512 | **`added_rows_count`** | `long` | *required* | Total number of
rows in ADDED entries. |
+ | 513 | **`existing_rows_count`** | `long` | *required* | Total number of
rows in EXISTING entries. |
+ | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of
rows in DELETED entries. |
+ | 521 | **`replaced_rows_count`** | `long` | *required* | Total number of
rows in REPLACED entries. |
+ | 525 | **`modified_rows_count`** | `long` | *required* | Total number of
rows in MODIFIED entries. |
+ | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data
sequence number of all live entries in the manifest. |
+ | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced
leaf manifest that are not live. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 523 | **`dv_cardinality`** | `long` | *optional* | Cardinality of the
manifest deletion vector. |
+
+ **`column_file` struct (element 159 of `column_files`, field 158)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 161 | **`format_version`** | `int` | *required* | Format version of this
column file. |
+ | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs
stored in this column file. |
+ | 164 | **`location`** | `string` | *required* | Location of the column
file. |
+ | 165 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, or `parquet`. |
+ | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file
size in bytes. |
+ | 167 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split
offsets for the column file. Must be sorted ascending. |
+
+ **Tracked File Requirements**
+
+ - `content_type` must not be 1 (POSITION_DELETES) or 2 (EQUALITY DELETES).
+ - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must
exactly match the `offset` and `length` stored in the Puffin footer for the
deletion vector blob.
+ - A leaf manifest may only contain data files.
+ - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf
manifest must have `format_version` PRE-V4.
+ - Other v4 tracked files must have `format_version` V4.
+ - `manifest_info` must be set if and only if the tracked file is a
manifest.
+ - `deletion_vector` may only be set if the tracked file is a data file.
+ - `column_files` may only be set if the tracked file is a data file or a
data manifest.
+ - `tracking.deleted_positions` and `tracking.replaced_positions` may only
be set if the tracked file is a manifest.
+ - `tracking.snapshot_id` and `tracking.sequence_number` are required for
the tracked file in the root manifest.
Review Comment:
can these be inherited from snapshot metadata? But I guess this is probably
following the manifest list practice of persisted values.
##########
format/spec.md:
##########
@@ -820,14 +944,18 @@ Each stats struct holds statistics for one table field.
It may contain the follo
| Requirement | Offset | Name | Type
| Included for | Description |
|-------------|--------|---------------------------|---------------------------|-----------------------------------------------|-------------|
-| _optional_ | 1 | `lower_bound` | Field type or `geo_lower`
| all primitives or `variant` | Lower bound stored as the
field's type, or `geo_lower` for geo types |
-| _optional_ | 2 | `upper_bound` | Field type or `geo_upper`
| all primitives or `variant` | Upper bound stored as the
field's type, or `geo_upper` for geo types |
+| _optional_ | 1 | `lower_bound` | Field type or `geo_lower`
| all primitives or `variant` | Lower bound stored as the
field's type, or `geo_lower` for geo types [1] |
Review Comment:
is the `[1]` referencing the first bullet point (float and double) in the
Note below? seems a bit weird.
##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
| | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
| | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
+ 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
+ 3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 4. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ **Tracked Files**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4:
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+ | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* |
Writer format version. |
+ | 100 | **`location`** | `string` | *required* | Location of the file or
manifest. |
+ | 101 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, `parquet`, or `puffin` |
+ | 147 | **`tracking`** | `tracking` struct | *required* | Groups status,
snapshot, and sequence number. See tracking struct below. |
+ | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used
to write this manifest or data file. |
+ | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort
order for this file. If missing or unknown, the order is assumed to be
unsorted. |
+ | 103 | **`record_count`** | `long` | *required* | Number of records in
this file. |
+ | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size
in bytes. |
+ | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column
stats. See [Content Stats](#content-stats). |
+ | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See
manifest_info struct below. |
+ | 131 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split
offsets for the data file. Must be sorted ascending. |
+ | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* |
Row-level deletion vector for a data file. |
+ | 158 | **`column_files`** | `list<159: column_file>` | *optional* |
Column update files associated with this entry. |
+
+ **`tracking` struct (field 147)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3:
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions,
replacements, and modifications. Deletes are not used in scans. |
+ | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file
was added or deleted. Inherited when null. |
+ | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the
deletion vector was added. |
+ | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* |
Snapshot ID where the latest column file was added. |
+ | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number
of the file. Inherited when null and status is 1 (ADDED). |
+ | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence
number indicating when the file was added. Inherited when null and status is
ADDED. |
+ | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the
`_row_id` for its first row. For a data manifest, the starting `_row_id` to
assign to rows added by ADDED data files. See [First Row ID
Inheritance](#first-row-id-inheritance). |
+ | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted
in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 7 | **`replaced_positions`** | `binary` | *optional* | Positions
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+
+ **`deletion_vector` struct (field 148)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 155 | **`location`** | `string` | *required* | Location of the Puffin
file. |
+ | 144 | **`offset`** | `long` | *required* | Offset in the file where the
content starts. |
+ | 145 | **`size_in_bytes`** | `long` | *required* | Length of the
referenced content stored in the file. |
+ | 156 | **`cardinality`** | `long` | *required* | Cardinality of the
deletion vector. |
+ | 149 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+
+ **`manifest_info` struct (field 150)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 504 | **`added_files_count`** | `int` | *required* | Count of entries
with status ADDED in the manifest. |
+ | 505 | **`existing_files_count`** | `int` | *required* | Count of entries
with status EXISTING in the manifest. |
+ | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries
with status DELETED in the manifest. |
+ | 520 | **`replaced_files_count`** | `int` | *required* | Count of entries
with status REPLACED in the manifest. |
+ | 524 | **`modified_files_count`** | `int` | *required* | Count of entries
with status MODIFIED in the manifest. |
+ | 512 | **`added_rows_count`** | `long` | *required* | Total number of
rows in ADDED entries. |
+ | 513 | **`existing_rows_count`** | `long` | *required* | Total number of
rows in EXISTING entries. |
+ | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of
rows in DELETED entries. |
+ | 521 | **`replaced_rows_count`** | `long` | *required* | Total number of
rows in REPLACED entries. |
+ | 525 | **`modified_rows_count`** | `long` | *required* | Total number of
rows in MODIFIED entries. |
+ | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data
sequence number of all live entries in the manifest. |
+ | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced
leaf manifest that are not live. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 523 | **`dv_cardinality`** | `long` | *optional* | Cardinality of the
manifest deletion vector. |
+
+ **`column_file` struct (element 159 of `column_files`, field 158)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 161 | **`format_version`** | `int` | *required* | Format version of this
column file. |
+ | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs
stored in this column file. |
+ | 164 | **`location`** | `string` | *required* | Location of the column
file. |
+ | 165 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, or `parquet`. |
+ | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file
size in bytes. |
+ | 167 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split
offsets for the column file. Must be sorted ascending. |
+
+ **Tracked File Requirements**
+
+ - `content_type` must not be 1 (POSITION_DELETES) or 2 (EQUALITY DELETES).
Review Comment:
this seems redundant. the description cell for this row already clarified it.
##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
| | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
| | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
+ 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
+ 3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 4. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ **Tracked Files**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4:
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+ | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* |
Writer format version. |
+ | 100 | **`location`** | `string` | *required* | Location of the file or
manifest. |
+ | 101 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, `parquet`, or `puffin` |
+ | 147 | **`tracking`** | `tracking` struct | *required* | Groups status,
snapshot, and sequence number. See tracking struct below. |
+ | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used
to write this manifest or data file. |
+ | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort
order for this file. If missing or unknown, the order is assumed to be
unsorted. |
+ | 103 | **`record_count`** | `long` | *required* | Number of records in
this file. |
+ | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size
in bytes. |
+ | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column
stats. See [Content Stats](#content-stats). |
+ | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See
manifest_info struct below. |
+ | 131 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split
offsets for the data file. Must be sorted ascending. |
+ | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* |
Row-level deletion vector for a data file. |
+ | 158 | **`column_files`** | `list<159: column_file>` | *optional* |
Column update files associated with this entry. |
+
+ **`tracking` struct (field 147)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3:
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions,
replacements, and modifications. Deletes are not used in scans. |
+ | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file
was added or deleted. Inherited when null. |
Review Comment:
just to confirm `snapshot_id` won't change for replaced and modified entries?
##########
format/spec.md:
##########
@@ -1051,7 +1193,9 @@ A simple and valid approach is to estimate the number of
rows in data files that
### Scan Planning
-Scans are planned by reading the manifest files for the current snapshot.
Deleted entries in data and delete manifests (those marked with status
"DELETED") are not used in a scan.
+Scans are planned by reading the manifests referenced by the snapshot root for
the current snapshot; starting in v4, the snapshot root may also contain data
files.
+
+Deleted entries in data and delete manifests (those marked with status
"DELETED") are not used in a scan; starting in v4, an entry is also not live if
its status is REPLACED or, for a leaf-manifest entry, if its position is set in
the referencing root manifest entry's `manifest_info.dv` (see [Manifest
Deletion Vectors](#manifest-deletion-vectors)).
Review Comment:
should we replace `Deleted entries` with `Retired entries` or `Non-live
entries` which would also include `replaced` entries
##########
format/spec.md:
##########
@@ -656,15 +662,33 @@ A data or delete file is associated with a sort order by
the sort order's id wit
### Manifests
-A manifest is an immutable Avro file that lists data files or delete files,
along with each file’s partition data tuple, metrics, and tracking information.
One or more manifest files are used to store a [snapshot](#snapshots), which
tracks all of the files in a table at some point in time. Manifests are tracked
by a [manifest list](#manifest-lists) for each table snapshot.
+A manifest is an immutable file that lists data files or delete files, along
with each file’s partition data, metrics, and tracking information. One or more
manifest files are used to store a [snapshot](#snapshots), which tracks all of
the files in a table at some point in time. Manifests are tracked by a snapshot
root for each table snapshot. In v4, the snapshot root is a root manifest that
may track data files in addition to leaf manifest files.
A manifest is a valid Iceberg data file: files must use valid Iceberg formats,
schemas, and column projection.
-A manifest may store either data files or delete files, but not both because
manifests that contain delete files are scanned first during job planning.
Whether a manifest is a data manifest or a delete manifest is stored in
manifest metadata.
+Each manifest type contains the following content:
-A manifest stores files for a single partition spec. When a table’s partition
spec changes, old files remain in the older manifest and newer files are
written to a new manifest. This is required because a manifest file’s schema is
based on its partition spec (see below). The partition spec of each manifest is
also used to transform predicates on the table's data rows into predicates on
partition values that are used during job planning to select files from a
manifest.
+| Manifest type | Contents |
+|----------------|----------|
+| v1-v3 data manifest | Data files |
+| v2-v3 delete manifest | Delete files |
+| v4 root manifest (snapshot root) | Data files, data manifests, delete
manifests |
+| v4 data manifest | Data files and their colocated deletion vectors |
-A manifest file must store the partition spec and other metadata as properties
in the Avro file's key-value metadata:
+In v2-v3, data and delete files are kept in separate manifests because
manifests that contain delete files are scanned first during job planning.
Whether a manifest is a data manifest or a delete manifest is stored in
manifest metadata.
+
+**Partition Spec Binding:**
+
+- v1-v3: A manifest stores files for a single partition spec. When a table’s
partition spec changes, old files remain in the older manifest and newer files
are written to a new manifest. This is required because a manifest file’s
schema is based on its partition spec.
+- v4: A manifest may store files written with different partition specs.
Review Comment:
should we mention `using unionized schema from all partition specs`?
##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
| | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
| | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
+ 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
+ 3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 4. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ **Tracked Files**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4:
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+ | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* |
Writer format version. |
+ | 100 | **`location`** | `string` | *required* | Location of the file or
manifest. |
+ | 101 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, `parquet`, or `puffin` |
+ | 147 | **`tracking`** | `tracking` struct | *required* | Groups status,
snapshot, and sequence number. See tracking struct below. |
+ | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used
to write this manifest or data file. |
Review Comment:
v4 uses unionized spec for tracked file. do we need `spec_id` at all?
##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
| | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
| | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
+ 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
+ 3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 4. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ **Tracked Files**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4:
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+ | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* |
Writer format version. |
+ | 100 | **`location`** | `string` | *required* | Location of the file or
manifest. |
+ | 101 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, `parquet`, or `puffin` |
+ | 147 | **`tracking`** | `tracking` struct | *required* | Groups status,
snapshot, and sequence number. See tracking struct below. |
+ | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used
to write this manifest or data file. |
+ | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort
order for this file. If missing or unknown, the order is assumed to be
unsorted. |
+ | 103 | **`record_count`** | `long` | *required* | Number of records in
this file. |
+ | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size
in bytes. |
+ | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column
stats. See [Content Stats](#content-stats). |
+ | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See
manifest_info struct below. |
+ | 131 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split
offsets for the data file. Must be sorted ascending. |
+ | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* |
Row-level deletion vector for a data file. |
+ | 158 | **`column_files`** | `list<159: column_file>` | *optional* |
Column update files associated with this entry. |
+
+ **`tracking` struct (field 147)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3:
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions,
replacements, and modifications. Deletes are not used in scans. |
+ | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file
was added or deleted. Inherited when null. |
+ | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the
deletion vector was added. |
+ | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* |
Snapshot ID where the latest column file was added. |
+ | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number
of the file. Inherited when null and status is 1 (ADDED). |
+ | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence
number indicating when the file was added. Inherited when null and status is
ADDED. |
+ | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the
`_row_id` for its first row. For a data manifest, the starting `_row_id` to
assign to rows added by ADDED data files. See [First Row ID
Inheritance](#first-row-id-inheritance). |
+ | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted
in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 7 | **`replaced_positions`** | `binary` | *optional* | Positions
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+
+ **`deletion_vector` struct (field 148)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 155 | **`location`** | `string` | *required* | Location of the Puffin
file. |
+ | 144 | **`offset`** | `long` | *required* | Offset in the file where the
content starts. |
+ | 145 | **`size_in_bytes`** | `long` | *required* | Length of the
referenced content stored in the file. |
+ | 156 | **`cardinality`** | `long` | *required* | Cardinality of the
deletion vector. |
+ | 149 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+
+ **`manifest_info` struct (field 150)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 504 | **`added_files_count`** | `int` | *required* | Count of entries
with status ADDED in the manifest. |
+ | 505 | **`existing_files_count`** | `int` | *required* | Count of entries
with status EXISTING in the manifest. |
+ | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries
with status DELETED in the manifest. |
+ | 520 | **`replaced_files_count`** | `int` | *required* | Count of entries
with status REPLACED in the manifest. |
+ | 524 | **`modified_files_count`** | `int` | *required* | Count of entries
with status MODIFIED in the manifest. |
+ | 512 | **`added_rows_count`** | `long` | *required* | Total number of
rows in ADDED entries. |
+ | 513 | **`existing_rows_count`** | `long` | *required* | Total number of
rows in EXISTING entries. |
+ | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of
rows in DELETED entries. |
+ | 521 | **`replaced_rows_count`** | `long` | *required* | Total number of
rows in REPLACED entries. |
+ | 525 | **`modified_rows_count`** | `long` | *required* | Total number of
rows in MODIFIED entries. |
+ | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data
sequence number of all live entries in the manifest. |
+ | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced
leaf manifest that are not live. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 523 | **`dv_cardinality`** | `long` | *optional* | Cardinality of the
manifest deletion vector. |
+
+ **`column_file` struct (element 159 of `column_files`, field 158)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 161 | **`format_version`** | `int` | *required* | Format version of this
column file. |
+ | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs
stored in this column file. |
+ | 164 | **`location`** | `string` | *required* | Location of the column
file. |
+ | 165 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, or `parquet`. |
+ | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file
size in bytes. |
+ | 167 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split
offsets for the column file. Must be sorted ascending. |
+
+ **Tracked File Requirements**
+
+ - `content_type` must not be 1 (POSITION_DELETES) or 2 (EQUALITY DELETES).
+ - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must
exactly match the `offset` and `length` stored in the Puffin footer for the
deletion vector blob.
+ - A leaf manifest may only contain data files.
+ - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf
manifest must have `format_version` PRE-V4.
+ - Other v4 tracked files must have `format_version` V4.
+ - `manifest_info` must be set if and only if the tracked file is a
manifest.
+ - `deletion_vector` may only be set if the tracked file is a data file.
+ - `column_files` may only be set if the tracked file is a data file or a
data manifest.
+ - `tracking.deleted_positions` and `tracking.replaced_positions` may only
be set if the tracked file is a manifest.
+ - `tracking.snapshot_id` and `tracking.sequence_number` are required for
the tracked file in the root manifest.
+ - For manifests, `tracking.sequence_number` must equal
`tracking.file_sequence_number`.
+ - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or
`manifest_info.dv` is set.
+ - `tracking.latest_column_file_snapshot_id` may only be set if
`column_files` is set.
+ - `manifest_info.dv_cardinality` must be set if and only if
`manifest_info.dv` is non-null.
+
+ When a file is added to the dataset, its tracked file must set status to
ADDED and store the snapshot ID in which the file was added.
+
+ When a data file's deletion vector or column files are updated, the writer
records a MODIFIED entry for the live version and marks the prior version as
replaced, either with a REPLACED entry or in a [manifest deletion
vector](#manifest-deletion-vectors). The resulting entries' `dv_snapshot_id` or
`latest_column_file_snapshot_id` must record the snapshot in which the deletion
vector or column files, respectively, last changed. For leaf manifest entries,
MODIFIED marks a live manifest whose `dv` changed.
Review Comment:
> in a [manifest deletion vector](#manifest-deletion-vectors)
should we just say `in the replaced_positions bitmap`?
> live manifest whose `dv` changed.
`manifest_info.dv`?
##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
| | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
| | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
+ 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
+ 3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 4. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ **Tracked Files**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4:
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+ | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* |
Writer format version. |
+ | 100 | **`location`** | `string` | *required* | Location of the file or
manifest. |
+ | 101 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, `parquet`, or `puffin` |
+ | 147 | **`tracking`** | `tracking` struct | *required* | Groups status,
snapshot, and sequence number. See tracking struct below. |
+ | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used
to write this manifest or data file. |
+ | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort
order for this file. If missing or unknown, the order is assumed to be
unsorted. |
+ | 103 | **`record_count`** | `long` | *required* | Number of records in
this file. |
+ | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size
in bytes. |
+ | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column
stats. See [Content Stats](#content-stats). |
+ | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See
manifest_info struct below. |
+ | 131 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split
offsets for the data file. Must be sorted ascending. |
+ | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* |
Row-level deletion vector for a data file. |
+ | 158 | **`column_files`** | `list<159: column_file>` | *optional* |
Column update files associated with this entry. |
+
+ **`tracking` struct (field 147)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3:
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions,
replacements, and modifications. Deletes are not used in scans. |
+ | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file
was added or deleted. Inherited when null. |
+ | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the
deletion vector was added. |
+ | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* |
Snapshot ID where the latest column file was added. |
+ | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number
of the file. Inherited when null and status is 1 (ADDED). |
+ | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence
number indicating when the file was added. Inherited when null and status is
ADDED. |
+ | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the
`_row_id` for its first row. For a data manifest, the starting `_row_id` to
assign to rows added by ADDED data files. See [First Row ID
Inheritance](#first-row-id-inheritance). |
+ | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted
in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 7 | **`replaced_positions`** | `binary` | *optional* | Positions
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+
+ **`deletion_vector` struct (field 148)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 155 | **`location`** | `string` | *required* | Location of the Puffin
file. |
+ | 144 | **`offset`** | `long` | *required* | Offset in the file where the
content starts. |
+ | 145 | **`size_in_bytes`** | `long` | *required* | Length of the
referenced content stored in the file. |
+ | 156 | **`cardinality`** | `long` | *required* | Cardinality of the
deletion vector. |
+ | 149 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+
+ **`manifest_info` struct (field 150)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 504 | **`added_files_count`** | `int` | *required* | Count of entries
with status ADDED in the manifest. |
+ | 505 | **`existing_files_count`** | `int` | *required* | Count of entries
with status EXISTING in the manifest. |
+ | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries
with status DELETED in the manifest. |
+ | 520 | **`replaced_files_count`** | `int` | *required* | Count of entries
with status REPLACED in the manifest. |
+ | 524 | **`modified_files_count`** | `int` | *required* | Count of entries
with status MODIFIED in the manifest. |
+ | 512 | **`added_rows_count`** | `long` | *required* | Total number of
rows in ADDED entries. |
+ | 513 | **`existing_rows_count`** | `long` | *required* | Total number of
rows in EXISTING entries. |
+ | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of
rows in DELETED entries. |
+ | 521 | **`replaced_rows_count`** | `long` | *required* | Total number of
rows in REPLACED entries. |
+ | 525 | **`modified_rows_count`** | `long` | *required* | Total number of
rows in MODIFIED entries. |
+ | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data
sequence number of all live entries in the manifest. |
+ | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced
leaf manifest that are not live. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 523 | **`dv_cardinality`** | `long` | *optional* | Cardinality of the
manifest deletion vector. |
+
+ **`column_file` struct (element 159 of `column_files`, field 158)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 161 | **`format_version`** | `int` | *required* | Format version of this
column file. |
+ | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs
stored in this column file. |
+ | 164 | **`location`** | `string` | *required* | Location of the column
file. |
+ | 165 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, or `parquet`. |
+ | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file
size in bytes. |
+ | 167 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split
offsets for the column file. Must be sorted ascending. |
+
+ **Tracked File Requirements**
+
+ - `content_type` must not be 1 (POSITION_DELETES) or 2 (EQUALITY DELETES).
+ - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must
exactly match the `offset` and `length` stored in the Puffin footer for the
deletion vector blob.
+ - A leaf manifest may only contain data files.
+ - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf
manifest must have `format_version` PRE-V4.
+ - Other v4 tracked files must have `format_version` V4.
+ - `manifest_info` must be set if and only if the tracked file is a
manifest.
+ - `deletion_vector` may only be set if the tracked file is a data file.
+ - `column_files` may only be set if the tracked file is a data file or a
data manifest.
+ - `tracking.deleted_positions` and `tracking.replaced_positions` may only
be set if the tracked file is a manifest.
+ - `tracking.snapshot_id` and `tracking.sequence_number` are required for
the tracked file in the root manifest.
+ - For manifests, `tracking.sequence_number` must equal
`tracking.file_sequence_number`.
+ - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or
`manifest_info.dv` is set.
+ - `tracking.latest_column_file_snapshot_id` may only be set if
`column_files` is set.
+ - `manifest_info.dv_cardinality` must be set if and only if
`manifest_info.dv` is non-null.
+
+ When a file is added to the dataset, its tracked file must set status to
ADDED and store the snapshot ID in which the file was added.
+
+ When a data file's deletion vector or column files are updated, the writer
records a MODIFIED entry for the live version and marks the prior version as
replaced, either with a REPLACED entry or in a [manifest deletion
vector](#manifest-deletion-vectors). The resulting entries' `dv_snapshot_id` or
`latest_column_file_snapshot_id` must record the snapshot in which the deletion
vector or column files, respectively, last changed. For leaf manifest entries,
MODIFIED marks a live manifest whose `dv` changed.
+
+ When a file is deleted from the dataset, its tracked file must set status
to DELETED and store the snapshot ID in which the file was deleted. Writers
must include DELETED entries in the manifest for the snapshot that deletes the
file. The next manifest written for those entries must omit the DELETED entries.
+
+The file may be deleted from the file system when the snapshot in which it was
deleted is garbage collected, assuming that older snapshots have also been
garbage collected [1].
+
+Iceberg v2 adds data and file sequence numbers to the entry and makes the
snapshot ID optional. Values for these fields are inherited from manifest
metadata when `null`. That is, if the field is `null` for an entry, then the
entry must inherit its value from the manifest file's metadata, stored in the
snapshot root.
+The `sequence_number` field represents the data sequence number and must never
change after a file is added to the dataset, except during the addition of a
column file. The data sequence number represents a relative age of the file
content and should be used for planning which delete files apply to a data file.
+The `file_sequence_number` field represents the sequence number of the
snapshot that added the file and must also remain unchanged upon assigning at
commit. The file sequence number can't be used for pruning delete files as the
data within the file may have an older data sequence number.
+The data and file sequence numbers are inherited only if the entry status is 1
(added). If the entry status is 0 (existing) or 2 (deleted), the entry must
include both sequence numbers explicitly.
Notes:
-1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
-2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in the
IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
-3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
-4. Position delete metadata can use `referenced_data_file` when all deletes
tracked by the entry are in a single data file. Setting the referenced file is
required for deletion vectors.
-5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
-6. The following field ids are reserved on `data_file`: 141.
+1. Technically, data files can be deleted when the last snapshot that contains
the file as "live" data is garbage collected. But this is harder to detect and
requires finding the diff of multiple snapshots. It is easier to track what
files are deleted in a snapshot and delete them when that snapshot expires. It
is not recommended to add a deleted file back to a table. Adding a deleted file
can lead to edge cases where incremental deletes can break table snapshots.
Review Comment:
> delete them when that snapshot expires
This seems to recommend only delete retired entries when the snapshot
expires.
I think the v4 requirement is that retired entries must be kept for one
snapshot. should we spell it out here>
##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
| | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
| | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ Notes:
+
+ 1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
+ 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
+ 3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 4. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ **Tracked Files**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4:
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+ | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* |
Writer format version. |
+ | 100 | **`location`** | `string` | *required* | Location of the file or
manifest. |
+ | 101 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, `parquet`, or `puffin` |
+ | 147 | **`tracking`** | `tracking` struct | *required* | Groups status,
snapshot, and sequence number. See tracking struct below. |
+ | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used
to write this manifest or data file. |
+ | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort
order for this file. If missing or unknown, the order is assumed to be
unsorted. |
+ | 103 | **`record_count`** | `long` | *required* | Number of records in
this file. |
+ | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size
in bytes. |
+ | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column
stats. See [Content Stats](#content-stats). |
+ | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See
manifest_info struct below. |
+ | 131 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split
offsets for the data file. Must be sorted ascending. |
+ | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* |
Row-level deletion vector for a data file. |
+ | 158 | **`column_files`** | `list<159: column_file>` | *optional* |
Column update files associated with this entry. |
+
+ **`tracking` struct (field 147)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3:
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions,
replacements, and modifications. Deletes are not used in scans. |
Review Comment:
I would suggest another version.
```
Retired entries (DELETED and REPLACED) are not used in the scans.
```
We can also drop this sentence completely and cover this in the scan
planning section.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]