amogh-jahagirdar commented on code in PR #16025: URL: https://github.com/apache/iceberg/pull/16025#discussion_r4118310752
########## format/spec.md: ########## @@ -1387,6 +1534,20 @@ At most one deletion vector is allowed per data file in a snapshot. If a DV is w [puffin-spec]: https://iceberg.apache.org/puffin-spec/ +#### Manifest Deletion Vectors + +A manifest deletion vector marks entries in a leaf manifest as not live by encoding their positions in a bitmap. A set bit at position P indicates that the entry at position P in the referenced leaf manifest is not live. + +Manifest deletion vectors are encoded using the [Mumbling bitmap spec][mumbling-spec] and stored inline on the root manifest entry that references the leaf manifest. The snapshot in which the vector last changed is recorded in `tracking.dv_snapshot_id`; the three bitmaps are: + +* `manifest_info.dv`: every position not live as of that snapshot. `manifest_info.dv_cardinality` is its cardinality. +* `tracking.deleted_positions`: the positions deleted in that snapshot. +* `tracking.replaced_positions`: the positions replaced in that snapshot. + +`deleted_positions` and `replaced_positions` are disjoint. Review Comment: Most of this was already there: deletes through a manifest DV must set `tracking.deleted_positions`, `manifest_info.dv`, and `dv_snapshot_id`, liveness is defined by `manifest_info.dv` alone, and `dv_snapshot_id` must record where the DV last changed, so a writer that only stores the cumulative bitmap isn't conformant. I was missing the `replaced_positions` requirements though, so I updated it so that replacing through a manifest DV must set `tracking.replaced_positions` and `manifest_info.dv`. I also added that the deltas should only be set in the snapshot that changes `manifest_info.dv`. That's a should rather than a must because carrying them forward is still correct, since `dv_snapshot_id` tells readers which snapshot they belong to; dropping them just saves space, and is what we'd typically expect writers to do anyways. ########## format/spec.md: ########## @@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields: | | | _optional_ | **`144 content_offset`** | `long` | The offset in the file where the content starts [5] | | | | _optional_ | **`145 content_size_in_bytes`** | `long` | The length of a referenced content stored in the file; required if `content_offset` is present [5] | -The `partition` struct stores the tuple of partition values for each file. Its type is derived from the partition fields of the partition spec used to write the manifest file. In v2, the partition struct's field ids must match the ids from the partition spec. + The `partition` struct stores the tuple of partition values for each file. Its type is derived from the partition fields of the partition spec used to write the manifest file. In v2, the partition struct's field ids must match the ids from the partition spec. -The v4 `content_stats` container struct stores field-level metrics. Unlike the metrics maps, the type of `content_stats` is based on table metadata, like schema. Similar to the `partition` struct, the same type is used for all files tracked in a manifest. + Notes: + + 1. Single-value serialization for lower and upper bounds is detailed in Appendix D. + 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper bounds. + 3. If sort order ID is missing or unknown, then the order is assumed to be unsorted. Only data files and equality delete files should be written with a non-null order id. [Position deletes](#position-delete-files) are required to be sorted by file and position, not a table order, and should set sort order id to null. Readers must ignore sort order id for position delete files. + 4. Position delete metadata can use `referenced_data_file` when all deletes tracked by the entry are in a single data file. Setting the referenced file is required for deletion vectors. + 5. The `content_offset` and `content_size_in_bytes` fields are used to reference a specific blob for direct access to a deletion vector. For deletion vectors, these values are required and must exactly match the `offset` and `length` stored in the Puffin footer for the deletion vector blob. + 6. The following field ids are reserved on `data_file`: 141. + +=== "v4" + **Tracked Files** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: DELETE_MANIFEST) | *required* | Type of content stored in the entry. | + | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | Writer format version. | + | 100 | **`location`** | `string` | *required* | Location of the file or manifest. | + | 101 | **`file_format`** | `string` | *required* | String file format name: `avro`, `orc`, `parquet`, or `puffin` | + | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, snapshot, and sequence number. See tracking struct below. | + | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used to write this manifest or data file. | + | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort order for this file. If missing or unknown, the order is assumed to be unsorted. | + | 103 | **`record_count`** | `long` | *required* | Number of records in this file. | + | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size in bytes. | + | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column stats. See [Content Stats](#content-stats). | + | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See manifest_info struct below. | + | 131 | **`key_metadata`** | `binary` | *optional* | Implementation-specific key metadata for encryption. | + | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split offsets for the data file. Must be sorted ascending. | + | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | Row-level deletion vector for a data file. | + | 158 | **`column_files`** | `list<159: column_file>` | *optional* | Column update files associated with this entry. | + + **`tracking` struct (field 147)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, replacements, and modifications. Deletes are not used in scans. | + | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file was added or deleted. Inherited when null. | + | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the deletion vector was added. | + | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | Snapshot ID where the latest column file was added. | + | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number of the file. Inherited when null and status is 1 (ADDED). | + | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence number indicating when the file was added. Inherited when null and status is ADDED. | + | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the `_row_id` for its first row. For a data manifest, the starting `_row_id` to assign to rows added by ADDED data files. See [First Row ID Inheritance](#first-row-id-inheritance). | + | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted in the referenced leaf manifest this snapshot. See [Manifest Deletion Vectors](#manifest-deletion-vectors). | + | 7 | **`replaced_positions`** | `binary` | *optional* | Positions replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion Vectors](#manifest-deletion-vectors). | + + **`deletion_vector` struct (field 148)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 155 | **`location`** | `string` | *required* | Location of the Puffin file. | + | 144 | **`offset`** | `long` | *required* | Offset in the file where the content starts. | + | 145 | **`size_in_bytes`** | `long` | *required* | Length of the referenced content stored in the file. | + | 156 | **`cardinality`** | `long` | *required* | Cardinality of the deletion vector. | + | 149 | **`key_metadata`** | `binary` | *optional* | Implementation-specific key metadata for encryption. | + + **`manifest_info` struct (field 150)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 504 | **`added_files_count`** | `int` | *required* | Count of entries with status ADDED in the manifest. | + | 505 | **`existing_files_count`** | `int` | *required* | Count of entries with status EXISTING in the manifest. | + | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries with status DELETED in the manifest. | + | 520 | **`replaced_files_count`** | `int` | *required* | Count of entries with status REPLACED in the manifest. | + | 524 | **`modified_files_count`** | `int` | *required* | Count of entries with status MODIFIED in the manifest. | + | 512 | **`added_rows_count`** | `long` | *required* | Total number of rows in ADDED entries. | + | 513 | **`existing_rows_count`** | `long` | *required* | Total number of rows in EXISTING entries. | + | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of rows in DELETED entries. | + | 521 | **`replaced_rows_count`** | `long` | *required* | Total number of rows in REPLACED entries. | + | 525 | **`modified_rows_count`** | `long` | *required* | Total number of rows in MODIFIED entries. | + | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data sequence number of all live entries in the manifest. | + | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced leaf manifest that are not live. See [Manifest Deletion Vectors](#manifest-deletion-vectors). | + | 523 | **`dv_cardinality`** | `long` | *optional* | Cardinality of the manifest deletion vector. | + + **`column_file` struct (element 159 of `column_files`, field 158)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 161 | **`format_version`** | `int` | *required* | Format version of this column file. | + | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs stored in this column file. | + | 164 | **`location`** | `string` | *required* | Location of the column file. | + | 165 | **`file_format`** | `string` | *required* | String file format name: `avro`, `orc`, or `parquet`. | + | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file size in bytes. | + | 167 | **`key_metadata`** | `binary` | *optional* | Implementation-specific key metadata for encryption. | + | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split offsets for the column file. Must be sorted ascending. | + + **Tracked File Requirements** + + - `content_type` must not be 1 (POSITION_DELETES) or 2 (EQUALITY DELETES). + - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must exactly match the `offset` and `length` stored in the Puffin footer for the deletion vector blob. + - A leaf manifest may only contain data files. + - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf manifest must have `format_version` PRE-V4. + - Other v4 tracked files must have `format_version` V4. + - `manifest_info` must be set if and only if the tracked file is a manifest. + - `deletion_vector` may only be set if the tracked file is a data file. + - `column_files` may only be set if the tracked file is a data file or a data manifest. + - `tracking.deleted_positions` and `tracking.replaced_positions` may only be set if the tracked file is a manifest. + - `tracking.snapshot_id` and `tracking.sequence_number` are required for the tracked file in the root manifest. + - For manifests, `tracking.sequence_number` must equal `tracking.file_sequence_number`. + - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or `manifest_info.dv` is set. + - `tracking.latest_column_file_snapshot_id` may only be set if `column_files` is set. + - `manifest_info.dv_cardinality` must be set if and only if `manifest_info.dv` is non-null. + + When a file is added to the dataset, its tracked file must set status to ADDED and store the snapshot ID in which the file was added. + + When a data file's deletion vector or column files are updated, the writer records a MODIFIED entry for the live version and marks the prior version as replaced, either with a REPLACED entry or in a [manifest deletion vector](#manifest-deletion-vectors). The resulting entries' `dv_snapshot_id` or `latest_column_file_snapshot_id` must record the snapshot in which the deletion vector or column files, respectively, last changed. For leaf manifest entries, MODIFIED marks a live manifest whose `dv` changed. Review Comment: Updated, the sentence now names `tracking.replaced_positions` and `manifest_info.dv` explicitly. ########## format/spec.md: ########## Review Comment: Fixed, scoped it to v1-v3. I also updated "tracked by a manifest list" to "snapshot root" and scoped the partition predicate conversion paragraph in scan planning to v1-v3. ########## format/spec.md: ########## @@ -1051,7 +1193,9 @@ A simple and valid approach is to estimate the number of rows in data files that ### Scan Planning -Scans are planned by reading the manifest files for the current snapshot. Deleted entries in data and delete manifests (those marked with status "DELETED") are not used in a scan. +Scans are planned by reading the manifests referenced by the snapshot root for the current snapshot; starting in v4, the snapshot root may also contain data files. + +Deleted entries in data and delete manifests (those marked with status "DELETED") are not used in a scan; starting in v4, an entry is also not live if its status is REPLACED or, for a leaf-manifest entry, if its position is set in the referencing root manifest entry's `manifest_info.dv` (see [Manifest Deletion Vectors](#manifest-deletion-vectors)). Manifests that contain no matching files, determined using either file counts or partition summaries, may be skipped. Review Comment: Partition summaries are now scoped to v1-v3, with column stats for v4. For the counts, skipping a fully deleted manifest is just an optimization, not something needed for correctness. Entries covered by a manifest DV are still counted, so the counts can only overcount, and skipping is a may, so the worst case is reading a manifest that could have been skipped. `dv_cardinality` isn't in `manifest_info` anymore, but a reader can still get the cardinality from the `dv` bitmap itself. And in practice I'd expect manifest DVs to get compacted well before they cover the whole manifest. For row lineage, `first_row_id` only has to be greater than or equal to the last one plus the rows that need IDs, so overcounting still satisfies it. MODIFIED/REPLACED entries carry `first_row_id` explicitly, so they don't need to be counted. ########## format/spec.md: ########## @@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields: | | | _optional_ | **`144 content_offset`** | `long` | The offset in the file where the content starts [5] | | | | _optional_ | **`145 content_size_in_bytes`** | `long` | The length of a referenced content stored in the file; required if `content_offset` is present [5] | -The `partition` struct stores the tuple of partition values for each file. Its type is derived from the partition fields of the partition spec used to write the manifest file. In v2, the partition struct's field ids must match the ids from the partition spec. + The `partition` struct stores the tuple of partition values for each file. Its type is derived from the partition fields of the partition spec used to write the manifest file. In v2, the partition struct's field ids must match the ids from the partition spec. -The v4 `content_stats` container struct stores field-level metrics. Unlike the metrics maps, the type of `content_stats` is based on table metadata, like schema. Similar to the `partition` struct, the same type is used for all files tracked in a manifest. + Notes: + + 1. Single-value serialization for lower and upper bounds is detailed in Appendix D. + 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper bounds. + 3. If sort order ID is missing or unknown, then the order is assumed to be unsorted. Only data files and equality delete files should be written with a non-null order id. [Position deletes](#position-delete-files) are required to be sorted by file and position, not a table order, and should set sort order id to null. Readers must ignore sort order id for position delete files. + 4. Position delete metadata can use `referenced_data_file` when all deletes tracked by the entry are in a single data file. Setting the referenced file is required for deletion vectors. + 5. The `content_offset` and `content_size_in_bytes` fields are used to reference a specific blob for direct access to a deletion vector. For deletion vectors, these values are required and must exactly match the `offset` and `length` stored in the Puffin footer for the deletion vector blob. + 6. The following field ids are reserved on `data_file`: 141. + +=== "v4" + **Tracked Files** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: DELETE_MANIFEST) | *required* | Type of content stored in the entry. | + | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | Writer format version. | + | 100 | **`location`** | `string` | *required* | Location of the file or manifest. | + | 101 | **`file_format`** | `string` | *required* | String file format name: `avro`, `orc`, `parquet`, or `puffin` | + | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, snapshot, and sequence number. See tracking struct below. | + | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used to write this manifest or data file. | + | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort order for this file. If missing or unknown, the order is assumed to be unsorted. | + | 103 | **`record_count`** | `long` | *required* | Number of records in this file. | + | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size in bytes. | + | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column stats. See [Content Stats](#content-stats). | + | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See manifest_info struct below. | + | 131 | **`key_metadata`** | `binary` | *optional* | Implementation-specific key metadata for encryption. | + | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split offsets for the data file. Must be sorted ascending. | + | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | Row-level deletion vector for a data file. | + | 158 | **`column_files`** | `list<159: column_file>` | *optional* | Column update files associated with this entry. | + + **`tracking` struct (field 147)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, replacements, and modifications. Deletes are not used in scans. | + | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file was added or deleted. Inherited when null. | + | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the deletion vector was added. | + | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | Snapshot ID where the latest column file was added. | + | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number of the file. Inherited when null and status is 1 (ADDED). | + | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence number indicating when the file was added. Inherited when null and status is ADDED. | + | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the `_row_id` for its first row. For a data manifest, the starting `_row_id` to assign to rows added by ADDED data files. See [First Row ID Inheritance](#first-row-id-inheritance). | + | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted in the referenced leaf manifest this snapshot. See [Manifest Deletion Vectors](#manifest-deletion-vectors). | + | 7 | **`replaced_positions`** | `binary` | *optional* | Positions replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion Vectors](#manifest-deletion-vectors). | + + **`deletion_vector` struct (field 148)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 155 | **`location`** | `string` | *required* | Location of the Puffin file. | + | 144 | **`offset`** | `long` | *required* | Offset in the file where the content starts. | + | 145 | **`size_in_bytes`** | `long` | *required* | Length of the referenced content stored in the file. | + | 156 | **`cardinality`** | `long` | *required* | Cardinality of the deletion vector. | + | 149 | **`key_metadata`** | `binary` | *optional* | Implementation-specific key metadata for encryption. | + + **`manifest_info` struct (field 150)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 504 | **`added_files_count`** | `int` | *required* | Count of entries with status ADDED in the manifest. | + | 505 | **`existing_files_count`** | `int` | *required* | Count of entries with status EXISTING in the manifest. | + | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries with status DELETED in the manifest. | + | 520 | **`replaced_files_count`** | `int` | *required* | Count of entries with status REPLACED in the manifest. | + | 524 | **`modified_files_count`** | `int` | *required* | Count of entries with status MODIFIED in the manifest. | + | 512 | **`added_rows_count`** | `long` | *required* | Total number of rows in ADDED entries. | + | 513 | **`existing_rows_count`** | `long` | *required* | Total number of rows in EXISTING entries. | + | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of rows in DELETED entries. | + | 521 | **`replaced_rows_count`** | `long` | *required* | Total number of rows in REPLACED entries. | + | 525 | **`modified_rows_count`** | `long` | *required* | Total number of rows in MODIFIED entries. | + | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data sequence number of all live entries in the manifest. | + | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced leaf manifest that are not live. See [Manifest Deletion Vectors](#manifest-deletion-vectors). | + | 523 | **`dv_cardinality`** | `long` | *optional* | Cardinality of the manifest deletion vector. | + + **`column_file` struct (element 159 of `column_files`, field 158)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 161 | **`format_version`** | `int` | *required* | Format version of this column file. | + | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs stored in this column file. | + | 164 | **`location`** | `string` | *required* | Location of the column file. | + | 165 | **`file_format`** | `string` | *required* | String file format name: `avro`, `orc`, or `parquet`. | + | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file size in bytes. | + | 167 | **`key_metadata`** | `binary` | *optional* | Implementation-specific key metadata for encryption. | + | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split offsets for the column file. Must be sorted ascending. | + + **Tracked File Requirements** + + - `content_type` must not be 1 (POSITION_DELETES) or 2 (EQUALITY DELETES). + - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must exactly match the `offset` and `length` stored in the Puffin footer for the deletion vector blob. + - A leaf manifest may only contain data files. + - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf manifest must have `format_version` PRE-V4. + - Other v4 tracked files must have `format_version` V4. + - `manifest_info` must be set if and only if the tracked file is a manifest. + - `deletion_vector` may only be set if the tracked file is a data file. + - `column_files` may only be set if the tracked file is a data file or a data manifest. + - `tracking.deleted_positions` and `tracking.replaced_positions` may only be set if the tracked file is a manifest. + - `tracking.snapshot_id` and `tracking.sequence_number` are required for the tracked file in the root manifest. + - For manifests, `tracking.sequence_number` must equal `tracking.file_sequence_number`. + - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or `manifest_info.dv` is set. + - `tracking.latest_column_file_snapshot_id` may only be set if `column_files` is set. + - `manifest_info.dv_cardinality` must be set if and only if `manifest_info.dv` is non-null. + + When a file is added to the dataset, its tracked file must set status to ADDED and store the snapshot ID in which the file was added. + + When a data file's deletion vector or column files are updated, the writer records a MODIFIED entry for the live version and marks the prior version as replaced, either with a REPLACED entry or in a [manifest deletion vector](#manifest-deletion-vectors). The resulting entries' `dv_snapshot_id` or `latest_column_file_snapshot_id` must record the snapshot in which the deletion vector or column files, respectively, last changed. For leaf manifest entries, MODIFIED marks a live manifest whose `dv` changed. + + When a file is deleted from the dataset, its tracked file must set status to DELETED and store the snapshot ID in which the file was deleted. Writers must include DELETED entries in the manifest for the snapshot that deletes the file. The next manifest written for those entries must omit the DELETED entries. + +The file may be deleted from the file system when the snapshot in which it was deleted is garbage collected, assuming that older snapshots have also been garbage collected [1]. + +Iceberg v2 adds data and file sequence numbers to the entry and makes the snapshot ID optional. Values for these fields are inherited from manifest metadata when `null`. That is, if the field is `null` for an entry, then the entry must inherit its value from the manifest file's metadata, stored in the snapshot root. +The `sequence_number` field represents the data sequence number and must never change after a file is added to the dataset, except during the addition of a column file. The data sequence number represents a relative age of the file content and should be used for planning which delete files apply to a data file. Review Comment: Agreed this has to be written down somewhere, but I'd like to keep it out of this baseline PR. Column update write rules, including rewriting equality deletes as DVs first, should get their own section in a follow-up imo. There's enough other specific rules for column updates that it warrants a separate focused PR. ########## format/spec.md: ########## @@ -676,13 +700,19 @@ A manifest file must store the partition spec and other metadata as properties i | _optional_ | _required_ | `format-version` | Table format version number of the manifest as a string | | | _required_ | `content` | Type of content files tracked by the manifest: "data" or "deletes" | +=== "v4" + | Requirement | Key | Value | + |-------------|---------------------|---------------------------------------------------------------------------------------------------------------------------------------------| + | _optional_ | `schema-id` | ID of the schema used to write the manifest as a string | + | _optional_ | `format-version` | Table format version number of the manifest as a string | + #### Content file uniqueness Within a snapshot, each content file must be referenced by at most one live manifest entry across all manifests; otherwise, the snapshot has undefined behavior. Writers should not produce multiple manifest entries for the same content file in a snapshot (for example, both ADDED and DELETED entries for the same file). Writers are not required to validate uniqueness at commit time. -#### Manifest Entry Fields +#### Entries in Manifests -The schema of a manifest file is defined by the `manifest_entry` struct, which consists of the following fields: +In v1-v3, manifest entries are described by the `manifest_entry` struct. In v4, entries are called tracked files and are described by the `tracked_file` struct. In v4, `data_file` struct fields are flattened directly into the tracked file, and tracking fields are grouped into a nested `tracking` struct. Review Comment: If we feel strongly about it I can remove it, but I was mainly trying to capture the evolution from the previous version to this one in a single statement. ########## format/spec.md: ########## @@ -676,13 +700,19 @@ A manifest file must store the partition spec and other metadata as properties i | _optional_ | _required_ | `format-version` | Table format version number of the manifest as a string | | | _required_ | `content` | Type of content files tracked by the manifest: "data" or "deletes" | +=== "v4" + | Requirement | Key | Value | + |-------------|---------------------|---------------------------------------------------------------------------------------------------------------------------------------------| + | _optional_ | `schema-id` | ID of the schema used to write the manifest as a string | + | _optional_ | `format-version` | Table format version number of the manifest as a string | + #### Content file uniqueness Within a snapshot, each content file must be referenced by at most one live manifest entry across all manifests; otherwise, the snapshot has undefined behavior. Writers should not produce multiple manifest entries for the same content file in a snapshot (for example, both ADDED and DELETED entries for the same file). Writers are not required to validate uniqueness at commit time. -#### Manifest Entry Fields +#### Entries in Manifests -The schema of a manifest file is defined by the `manifest_entry` struct, which consists of the following fields: +In v1-v3, manifest entries are described by the `manifest_entry` struct. In v4, entries are called tracked files and are described by the `tracked_file` struct. In v4, `data_file` struct fields are flattened directly into the tracked file, and tracking fields are grouped into a nested `tracking` struct. Review Comment: That's true, but I think it's worth establishing a term we can use for the overall entry without confusing it with the previous format version's entry. `tracked_file` is mainly there as a distinguishing term for v4 entries. ########## format/spec.md: ########## @@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields: | | | _optional_ | **`144 content_offset`** | `long` | The offset in the file where the content starts [5] | | | | _optional_ | **`145 content_size_in_bytes`** | `long` | The length of a referenced content stored in the file; required if `content_offset` is present [5] | -The `partition` struct stores the tuple of partition values for each file. Its type is derived from the partition fields of the partition spec used to write the manifest file. In v2, the partition struct's field ids must match the ids from the partition spec. + The `partition` struct stores the tuple of partition values for each file. Its type is derived from the partition fields of the partition spec used to write the manifest file. In v2, the partition struct's field ids must match the ids from the partition spec. -The v4 `content_stats` container struct stores field-level metrics. Unlike the metrics maps, the type of `content_stats` is based on table metadata, like schema. Similar to the `partition` struct, the same type is used for all files tracked in a manifest. + Notes: + + 1. Single-value serialization for lower and upper bounds is detailed in Appendix D. + 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper bounds. + 3. If sort order ID is missing or unknown, then the order is assumed to be unsorted. Only data files and equality delete files should be written with a non-null order id. [Position deletes](#position-delete-files) are required to be sorted by file and position, not a table order, and should set sort order id to null. Readers must ignore sort order id for position delete files. + 4. Position delete metadata can use `referenced_data_file` when all deletes tracked by the entry are in a single data file. Setting the referenced file is required for deletion vectors. + 5. The `content_offset` and `content_size_in_bytes` fields are used to reference a specific blob for direct access to a deletion vector. For deletion vectors, these values are required and must exactly match the `offset` and `length` stored in the Puffin footer for the deletion vector blob. + 6. The following field ids are reserved on `data_file`: 141. + +=== "v4" + **Tracked Files** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: DELETE_MANIFEST) | *required* | Type of content stored in the entry. | + | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | Writer format version. | + | 100 | **`location`** | `string` | *required* | Location of the file or manifest. | + | 101 | **`file_format`** | `string` | *required* | String file format name: `avro`, `orc`, `parquet`, or `puffin` | + | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, snapshot, and sequence number. See tracking struct below. | + | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used to write this manifest or data file. | + | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort order for this file. If missing or unknown, the order is assumed to be unsorted. | + | 103 | **`record_count`** | `long` | *required* | Number of records in this file. | + | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size in bytes. | + | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column stats. See [Content Stats](#content-stats). | + | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See manifest_info struct below. | + | 131 | **`key_metadata`** | `binary` | *optional* | Implementation-specific key metadata for encryption. | + | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split offsets for the data file. Must be sorted ascending. | + | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | Row-level deletion vector for a data file. | + | 158 | **`column_files`** | `list<159: column_file>` | *optional* | Column update files associated with this entry. | + + **`tracking` struct (field 147)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, replacements, and modifications. Deletes are not used in scans. | + | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file was added or deleted. Inherited when null. | + | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the deletion vector was added. | + | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | Snapshot ID where the latest column file was added. | + | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number of the file. Inherited when null and status is 1 (ADDED). | + | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence number indicating when the file was added. Inherited when null and status is ADDED. | + | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the `_row_id` for its first row. For a data manifest, the starting `_row_id` to assign to rows added by ADDED data files. See [First Row ID Inheritance](#first-row-id-inheritance). | + | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted in the referenced leaf manifest this snapshot. See [Manifest Deletion Vectors](#manifest-deletion-vectors). | + | 7 | **`replaced_positions`** | `binary` | *optional* | Positions replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion Vectors](#manifest-deletion-vectors). | + + **`deletion_vector` struct (field 148)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 155 | **`location`** | `string` | *required* | Location of the Puffin file. | + | 144 | **`offset`** | `long` | *required* | Offset in the file where the content starts. | + | 145 | **`size_in_bytes`** | `long` | *required* | Length of the referenced content stored in the file. | + | 156 | **`cardinality`** | `long` | *required* | Cardinality of the deletion vector. | + | 149 | **`key_metadata`** | `binary` | *optional* | Implementation-specific key metadata for encryption. | + + **`manifest_info` struct (field 150)** + + | Field id | Name | Type | Required | Description | + |----------|------|------|----------|-------------| + | 504 | **`added_files_count`** | `int` | *required* | Count of entries with status ADDED in the manifest. | + | 505 | **`existing_files_count`** | `int` | *required* | Count of entries with status EXISTING in the manifest. | + | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries with status DELETED in the manifest. | + | 520 | **`replaced_files_count`** | `int` | *required* | Count of entries with status REPLACED in the manifest. | + | 524 | **`modified_files_count`** | `int` | *required* | Count of entries with status MODIFIED in the manifest. | + | 512 | **`added_rows_count`** | `long` | *required* | Total number of rows in ADDED entries. | + | 513 | **`existing_rows_count`** | `long` | *required* | Total number of rows in EXISTING entries. | + | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of rows in DELETED entries. | + | 521 | **`replaced_rows_count`** | `long` | *required* | Total number of rows in REPLACED entries. | + | 525 | **`modified_rows_count`** | `long` | *required* | Total number of rows in MODIFIED entries. | Review Comment: `replaced_files_count` is already in `ManifestInfo`. `modified_files_count` is being added by @stevenzwu in #18212. ########## format/spec.md: ########## @@ -676,13 +700,19 @@ A manifest file must store the partition spec and other metadata as properties i | _optional_ | _required_ | `format-version` | Table format version number of the manifest as a string | | | _required_ | `content` | Type of content files tracked by the manifest: "data" or "deletes" | +=== "v4" + | Requirement | Key | Value | + |-------------|---------------------|---------------------------------------------------------------------------------------------------------------------------------------------| + | _optional_ | `schema-id` | ID of the schema used to write the manifest as a string | Review Comment: I'm not sure the k/v metadata buys us much in v4. In v1-v3 the argument was that it could be a source for recovery if metadata got corrupted, since each manifest was bound to a single spec. In v4 manifests aren't bound to a particular spec, so that notion mostly goes away, and the format version is already on the tracked file that references the manifest. So I'd keep them optional. I guess there's a stronger argument for format version but we're already tracking it per entry and so it'd only be useful for leaf manifests. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
