rdblue commented on code in PR #16025:
URL: https://github.com/apache/iceberg/pull/16025#discussion_r4161367768
##########
format/spec.md:
##########
@@ -742,18 +758,127 @@ The `data_file` struct consists of the following fields:
| | | _optional_ | **`144 content_offset`**
| `long` |
The offset in the file where the content starts [5] |
| | | _optional_ | **`145 content_size_in_bytes`**
| `long` |
The length of a referenced content stored in the file; required if
`content_offset` is present [5] |
-The `partition` struct stores the tuple of partition values for each file. Its
type is derived from the partition fields of the partition spec used to write
the manifest file. In v2, the partition struct's field ids must match the ids
from the partition spec.
+ The `partition` struct stores the tuple of partition values for each file.
Its type is derived from the partition fields of the partition spec used to
write the manifest file. In v2, the partition struct's field ids must match the
ids from the partition spec.
+
+ Notes:
-The v4 `content_stats` container struct stores field-level metrics. Unlike the
metrics maps, the type of `content_stats` is based on table metadata, like
schema. Similar to the `partition` struct, the same type is used for all files
tracked in a manifest.
+ 1. Single-value serialization for lower and upper bounds is detailed in
Appendix D.
+ 2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper
bounds.
+ 3. If sort order ID is missing or unknown, then the order is assumed to be
unsorted. Only data files and equality delete files should be written with a
non-null order id. [Position deletes](#position-delete-files) are required to
be sorted by file and position, not a table order, and should set sort order id
to null. Readers must ignore sort order id for position delete files.
+ 4. Position delete metadata can use `referenced_data_file` when all
deletes tracked by the entry are in a single data file. Setting the referenced
file is required for deletion vectors.
+ 5. The `content_offset` and `content_size_in_bytes` fields are used to
reference a specific blob for direct access to a deletion vector. For deletion
vectors, these values are required and must exactly match the `offset` and
`length` stored in the Puffin footer for the deletion vector blob.
+ 6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+ **Tracked Files**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4:
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+ | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* |
Writer format version. |
+ | 100 | **`location`** | `string` | *required* | Location of the file. |
+ | 101 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, or `parquet` |
+ | 147 | **`tracking`** | `tracking` struct | *required* | Tracking
metadata like status, snapshot ID, and sequence number. See tracking struct
below. |
+ | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used
to partition the file; null if unpartitioned |
+ | 102 | **`partition`** | `struct<...>` | *optional* | Partition data
tuple for the file; null if unpartitioned. |
+ | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort
order for this file. If missing or unknown, the order is assumed to be
unsorted. |
+ | 103 | **`record_count`** | `long` | *required* | Number of records in
this file. |
+ | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size
in bytes. |
+ | 146 | **`content_stats`** | `content_stats` struct | *optional* |
Field-level stats. See [Content Stats](#content-stats). |
+ | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See
manifest_info struct below. |
+ | 131 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split
offsets for the data file. Must be sorted ascending. |
+ | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* |
Row-level deletion vector for a data file. |
+ | 158 | **`column_files`** | `list<159: column_file>` | *optional* |
Column files associated with this file. |
+
+ **`tracking` struct (field 147)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3:
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions,
replacements, and modifications. |
+ | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file
was added or deleted. Inherited when null. |
+ | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the
deletion vector was added. |
+ | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* |
Snapshot ID where the latest column file was added. |
+ | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number
of the file. Inherited when null. See [Sequence Number
Inheritance](#sequence-number-inheritance). |
+ | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence
number indicating when the file was added. Inherited when null. See [Sequence
Number Inheritance](#sequence-number-inheritance). |
+ | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the
`_row_id` for its first row. For a data manifest, the starting `_row_id` to
assign to rows added by ADDED data files. See [First Row ID
Inheritance](#first-row-id-inheritance). |
+ | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted
in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+ | 7 | **`replaced_positions`** | `binary` | *optional* | Positions
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+
+ **`deletion_vector` struct (field 148)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 155 | **`location`** | `string` | *required* | Location of the Puffin
file. |
+ | 144 | **`offset`** | `long` | *required* | Offset in the file where the
content starts. |
+ | 145 | **`size_in_bytes`** | `long` | *required* | Length of the
referenced content stored in the file. |
+ | 156 | **`cardinality`** | `long` | *required* | Cardinality of the
deletion vector. |
+ | 149 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+
+ **`manifest_info` struct (field 150)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 504 | **`added_files_count`** | `int` | *required* | Count of entries
with status ADDED in the manifest. |
+ | 505 | **`existing_files_count`** | `int` | *required* | Count of entries
with status EXISTING in the manifest. |
+ | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries
with status DELETED in the manifest. |
+ | 523 | **`replaced_files_count`** | `int` | *required* | Count of entries
with status REPLACED in the manifest. |
+ | 525 | **`modified_files_count`** | `int` | *required* | Count of entries
with status MODIFIED in the manifest. |
+ | 512 | **`added_rows_count`** | `long` | *required* | Total number of
rows in ADDED entries. |
+ | 513 | **`existing_rows_count`** | `long` | *required* | Total number of
rows in EXISTING entries. |
+ | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of
rows in DELETED entries. |
+ | 524 | **`replaced_rows_count`** | `long` | *required* | Total number of
rows in REPLACED entries. |
+ | 526 | **`modified_rows_count`** | `long` | *required* | Total number of
rows in MODIFIED entries. |
+ | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data
sequence number of all live entries in the manifest. |
+ | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced
leaf manifest that are not live. See [Manifest Deletion
Vectors](#manifest-deletion-vectors). |
+
+ **`column_file` struct (element 159 of `column_files`, field 158)**
+
+ | Field id | Name | Type | Required | Description |
+ |----------|------|------|----------|-------------|
+ | 161 | **`format_version`** | `int` (4: V4) | *required* | Format version
of this column file. |
+ | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs
stored in this column file. |
+ | 164 | **`location`** | `string` | *required* | Location of the column
file. |
+ | 165 | **`file_format`** | `string` | *required* | String file format
name: `avro`, `orc`, or `parquet`. |
+ | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file
size in bytes. |
+ | 167 | **`key_metadata`** | `binary` | *optional* |
Implementation-specific key metadata for encryption. |
+ | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split
offsets for the column file. Must be sorted ascending. |
+
+ **Tracked File Requirements**
+
+ - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must
exactly match the `offset` and `length` stored in the Puffin footer for the
deletion vector blob.
+ - A leaf manifest written in v4 may only contain data files.
+ - A v1-v3 delete manifest referenced by a root manifest may contain v2-v3
delete files.
+ - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf
manifest must have `format_version` PRE-V4.
+ - Other v4 tracked files must have `format_version` V4.
+ - `manifest_info` must be set if and only if the tracked file is a
manifest.
+ - `deletion_vector` may only be set if the tracked file is a data file.
+ - `column_files` may only be set if the tracked file is a data file or a
data manifest.
+ - `tracking.deleted_positions` and `tracking.replaced_positions` may only
be set if the tracked file is a manifest.
+ - `tracking.snapshot_id` and `tracking.sequence_number` are required for
the tracked file in the root manifest.
+ - For manifests, `tracking.sequence_number` must equal
`tracking.file_sequence_number`.
+ - `tracking.dv_snapshot_id` may only be set if `deletion_vector` or
`manifest_info.dv` is set.
+ - `tracking.latest_column_file_snapshot_id` may only be set if
`column_files` is set.
+
+ When a file is added to the dataset, its tracked file must set status to
ADDED and store the snapshot ID in which the file was added.
+
+ When a data file's deletion vector or column files are updated, the writer
must record a MODIFIED entry for the live version and must mark the prior
version as replaced with a REPLACED entry or in a [manifest deletion
vector](#manifest-deletion-vectors). When using a manifest deletion vector, the
writer must set the position in the leaf manifest's
`tracking.replaced_positions` and `manifest_info.dv`. The resulting entries'
`dv_snapshot_id` or `latest_column_file_snapshot_id` must record the snapshot
in which their deletion vector, manifest deletion vector, or column files last
changed.
Review Comment:
```suggestion
When a data file's deletion vector or column files are updated, the
writer produces two entries: a MODIFIED entry for the updated live version and
a REPLACED entry with the previous metadata (for change detection). The
MODIFIED entry's `dv_snapshot_id` or `column_file_snapshot_id` is used to
record the snapshot ID in which the change occurred. The REPLACED entry can be
produced in place by setting its position in the leaf manifest's
`manifest_info.dv` (to remove it from planning) and
`tracking.replaced_positions` (for change detection). The REPLACED entry's
`snapshot_id` is used to record the snapshot where the entry was replaced; when
reading changes using `tracking.replaced_positions`, this snapshot ID comes
from the `dv_snapshot_id`.
```
I'm trying to edit this to be a bit more clear. I think that worked by
phrasing this to state that entries with these changes produce a `MODIFIED` and
`REPLACED` pair.
While editing and adding the part about how to produce changes, I hit an
issue: where does the `REPLACED` snapshot ID come from? For DV updates it
should be `dv_snapshot_id` and for column file changes it should be
`column_file_snapshot_id`. But we don't know which it should be without knowing
what was changed.
For the `REPLACED` entry, I think it would make sense to combine the
`dv_snapshot_id` with the `column_file_snapshot_id` so we only have one. But
for a normal entry, does this still make sense? We would only be able to detect
that the entry had changed, not whether it was a column change or a DV change
unless we find the `REPLACED` entry. I think this is okay, but we should get
more people thinking about this.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]