nastra commented on code in PR #16025:
URL: https://github.com/apache/iceberg/pull/16025#discussion_r4120400180


##########
format/spec.md:
##########
@@ -742,18 +758,127 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
+
+    Notes:
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, 
snapshot, and sequence number. See tracking struct below. |
+    | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used 
to write this manifest or data file. |
+    | 102 | **`partition`** | `struct<...>` | *optional* | Partition data 
tuple for the file. |
+    | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort 
order for this file. If missing or unknown, the order is assumed to be 
unsorted. |
+    | 103 | **`record_count`** | `long` | *required* | Number of records in 
this file. |
+    | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size 
in bytes. |
+    | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column 
stats. See [Content Stats](#content-stats). |
+    | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See 
manifest_info struct below. |
+    | 131 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split 
offsets for the data file. Must be sorted ascending. |
+    | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | 
Row-level deletion vector for a data file. |
+    | 158 | **`column_files`** | `list<159: column_file>` | *optional* | 
Column files associated with this file. |
+
+    **`tracking` struct (field 147)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: 
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, 
replacements, and modifications. |
+    | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the 
deletion vector was added. |
+    | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | 
Snapshot ID where the latest column file was added. |
+    | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number 
of the file. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence 
number indicating when the file was added. Inherited when null. See [Sequence 
Number Inheritance](#sequence-number-inheritance). |
+    | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the 
`_row_id` for its first row. For a data manifest, the starting `_row_id` to 
assign to rows added by ADDED data files. See [First Row ID 
Inheritance](#first-row-id-inheritance). |
+    | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted 
in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | 7 | **`replaced_positions`** | `binary` | *optional* | Positions 
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`deletion_vector` struct (field 148)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 155 | **`location`** | `string` | *required* | Location of the Puffin 
file. |
+    | 144 | **`offset`** | `long` | *required* | Offset in the file where the 
content starts. |
+    | 145 | **`size_in_bytes`** | `long` | *required* | Length of the 
referenced content stored in the file. |
+    | 156 | **`cardinality`** | `long` | *required* | Cardinality of the 
deletion vector. |

Review Comment:
   nit: maybe `Number of set bits (deleted rows) in the deletion vector`?



##########
format/spec.md:
##########
@@ -546,7 +552,7 @@ Note that:
 
 ### Partitioning
 
-Data files are stored in manifests with a tuple of partition values that are 
used in scans to filter out files that cannot contain records that match the 
scan’s filter predicate. Partition values for a data file must be the same for 
all records stored in the data file. (Manifests store data files from any 
partition, as long as the partition spec is the same for the data files.)
+Data files are stored in manifests with a tuple of partition values that are 
used in scans to filter out files that cannot contain records that match the 
scan’s filter predicate. Partition values for a data file must be the same for 
all records stored in the data file. Manifests store data files from any 
partition. v4 manifests may store partitions from any spec, but manifests in v3 
and earlier store files for a single spec. A manifest is considered partitioned 
by a spec if every entry in the manifest is partitioned by that spec.

Review Comment:
   ```suggestion
   Data files are stored in manifests with a tuple of partition values that are 
used in scans to filter out files that cannot contain records that match the 
scan’s filter predicate. Partition values for a data file must be the same for 
all records stored in the data file. Manifests store data files from any 
partition. v4 manifests may store partitions from any partition spec, but 
manifests in v3 and earlier store files for a single partition spec. A manifest 
is considered partitioned by a spec if every entry in the manifest is 
partitioned by that spec.
   ```



##########
format/spec.md:
##########
@@ -61,7 +61,10 @@ The full set of changes are listed in [Appendix 
E](#version-3).
 
 Version 4 of the Iceberg spec restructures metadata for improved performance 
and new capabilities:
 
+* Support for an [adaptive metadata tree](#manifests), enabling efficient 
small commits, column updates, and columnar statistics representation
+* Data files and deletion vectors are stored in a combined entry 
representation, removing the need for two phase planning

Review Comment:
   ```suggestion
   * Data files and deletion vectors are stored in a combined entry 
representation, removing the need for two-phase planning
   ```



##########
format/spec.md:
##########
@@ -656,15 +662,33 @@ A data or delete file is associated with a sort order by 
the sort order's id wit
 
 ### Manifests
 
-A manifest is an immutable Avro file that lists data files or delete files, 
along with each file’s partition data tuple, metrics, and tracking information. 
One or more manifest files are used to store a [snapshot](#snapshots), which 
tracks all of the files in a table at some point in time. Manifests are tracked 
by a [manifest list](#manifest-lists) for each table snapshot.
+A manifest is an immutable file that lists data files or delete files, along 
with each file’s partition data, metrics, and tracking information. One or more 
manifest files are used to store a [snapshot](#snapshots), which tracks all of 
the files in a table at some point in time. Manifests are tracked by a snapshot 
root for each table snapshot. In v4, the snapshot root is a root manifest that 
may track data files in addition to leaf manifest files.
 
 A manifest is a valid Iceberg data file: files must use valid Iceberg formats, 
schemas, and column projection.
 
-A manifest may store either data files or delete files, but not both because 
manifests that contain delete files are scanned first during job planning. 
Whether a manifest is a data manifest or a delete manifest is stored in 
manifest metadata.
+Each manifest type contains the following content:
 
-A manifest stores files for a single partition spec. When a table’s partition 
spec changes, old files remain in the older manifest and newer files are 
written to a new manifest. This is required because a manifest file’s schema is 
based on its partition spec (see below). The partition spec of each manifest is 
also used to transform predicates on the table's data rows into predicates on 
partition values that are used during job planning to select files from a 
manifest.
+| Manifest type | Contents |
+|----------------|----------|
+| v1-v3 data manifest | Data files |
+| v2-v3 delete manifest | Delete files |
+| v4 root manifest (snapshot root) | Data files, data manifests, delete 
manifests |
+| v4 data manifest | Data files and their colocated deletion vectors |
 
-A manifest file must store the partition spec and other metadata as properties 
in the Avro file's key-value metadata:
+In v2-v3, data and delete files are kept in separate manifests because 
manifests that contain delete files are scanned first during job planning. 
Whether a manifest is a data manifest or a delete manifest is stored in 
manifest metadata.
+
+**Partition Spec Binding:**
+
+- v1-v3: A manifest stores files for a single partition spec. When a table’s 
partition spec changes, old files remain in the older manifest and newer files 
are written to a new manifest. This is required because a manifest file’s 
schema is based on its partition spec.
+- v4: A manifest may store files written with different partition specs.

Review Comment:
   nit: or maybe just saying v4+ might be already enough



##########
format/spec.md:
##########
@@ -742,18 +758,127 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
+
+    Notes:
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, 
snapshot, and sequence number. See tracking struct below. |
+    | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used 
to write this manifest or data file. |
+    | 102 | **`partition`** | `struct<...>` | *optional* | Partition data 
tuple for the file. |
+    | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort 
order for this file. If missing or unknown, the order is assumed to be 
unsorted. |
+    | 103 | **`record_count`** | `long` | *required* | Number of records in 
this file. |
+    | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size 
in bytes. |
+    | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column 
stats. See [Content Stats](#content-stats). |
+    | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See 
manifest_info struct below. |
+    | 131 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split 
offsets for the data file. Must be sorted ascending. |
+    | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | 
Row-level deletion vector for a data file. |
+    | 158 | **`column_files`** | `list<159: column_file>` | *optional* | 
Column files associated with this file. |
+
+    **`tracking` struct (field 147)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: 
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, 
replacements, and modifications. |
+    | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the 
deletion vector was added. |
+    | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | 
Snapshot ID where the latest column file was added. |
+    | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number 
of the file. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence 
number indicating when the file was added. Inherited when null. See [Sequence 
Number Inheritance](#sequence-number-inheritance). |
+    | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the 
`_row_id` for its first row. For a data manifest, the starting `_row_id` to 
assign to rows added by ADDED data files. See [First Row ID 
Inheritance](#first-row-id-inheritance). |
+    | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted 
in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | 7 | **`replaced_positions`** | `binary` | *optional* | Positions 
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`deletion_vector` struct (field 148)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 155 | **`location`** | `string` | *required* | Location of the Puffin 
file. |
+    | 144 | **`offset`** | `long` | *required* | Offset in the file where the 
content starts. |
+    | 145 | **`size_in_bytes`** | `long` | *required* | Length of the 
referenced content stored in the file. |
+    | 156 | **`cardinality`** | `long` | *required* | Cardinality of the 
deletion vector. |
+    | 149 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+
+    **`manifest_info` struct (field 150)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 504 | **`added_files_count`** | `int` | *required* | Count of entries 
with status ADDED in the manifest. |
+    | 505 | **`existing_files_count`** | `int` | *required* | Count of entries 
with status EXISTING in the manifest. |
+    | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries 
with status DELETED in the manifest. |
+    | 523 | **`replaced_files_count`** | `int` | *required* | Count of entries 
with status REPLACED in the manifest. |
+    | 525 | **`modified_files_count`** | `int` | *required* | Count of entries 
with status MODIFIED in the manifest. |
+    | 512 | **`added_rows_count`** | `long` | *required* | Total number of 
rows in ADDED entries. |
+    | 513 | **`existing_rows_count`** | `long` | *required* | Total number of 
rows in EXISTING entries. |
+    | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of 
rows in DELETED entries. |
+    | 524 | **`replaced_rows_count`** | `long` | *required* | Total number of 
rows in REPLACED entries. |
+    | 526 | **`modified_rows_count`** | `long` | *required* | Total number of 
rows in MODIFIED entries. |
+    | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data 
sequence number of all live entries in the manifest. |
+    | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced 
leaf manifest that are not live. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`column_file` struct (element 159 of `column_files`, field 158)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 161 | **`format_version`** | `int` (4: V4) | *required* | Format version 
of this column file. |
+    | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs 
stored in this column file. |
+    | 164 | **`location`** | `string` | *required* | Location of the column 
file. |
+    | 165 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet`. |
+    | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file 
size in bytes. |
+    | 167 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split 
offsets for the column file. Must be sorted ascending. |
+
+    **Tracked File Requirements**
+
+    - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must 
exactly match the `offset` and `length` stored in the Puffin footer for the 
deletion vector blob.
+    - A leaf manifest written in v4 may only contain data files.
+    - A v1-v3 delete manifest referenced by a root manifest may contain v2-v3 
delete files.
+    - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf 
manifest must have `format_version` PRE-V4.

Review Comment:
   nit: instead of PRE-v4 maybe just < v4?



##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    Notes:
+
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file or 
manifest. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, `parquet`, or `puffin` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, 
snapshot, and sequence number. See tracking struct below. |
+    | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used 
to write this manifest or data file. |
+    | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort 
order for this file. If missing or unknown, the order is assumed to be 
unsorted. |
+    | 103 | **`record_count`** | `long` | *required* | Number of records in 
this file. |
+    | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size 
in bytes. |
+    | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column 
stats. See [Content Stats](#content-stats). |
+    | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See 
manifest_info struct below. |
+    | 131 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split 
offsets for the data file. Must be sorted ascending. |
+    | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | 
Row-level deletion vector for a data file. |
+    | 158 | **`column_files`** | `list<159: column_file>` | *optional* | 
Column update files associated with this entry. |
+
+    **`tracking` struct (field 147)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: 
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, 
replacements, and modifications. Deletes are not used in scans. |
+    | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the 
deletion vector was added. |
+    | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | 
Snapshot ID where the latest column file was added. |
+    | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number 
of the file. Inherited when null and status is 1 (ADDED). |
+    | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence 
number indicating when the file was added. Inherited when null and status is 
ADDED. |
+    | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the 
`_row_id` for its first row. For a data manifest, the starting `_row_id` to 
assign to rows added by ADDED data files. See [First Row ID 
Inheritance](#first-row-id-inheritance). |
+    | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted 
in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |

Review Comment:
   same for `replaced_positions`



##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    Notes:
+
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file or 
manifest. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, `parquet`, or `puffin` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, 
snapshot, and sequence number. See tracking struct below. |
+    | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used 
to write this manifest or data file. |
+    | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort 
order for this file. If missing or unknown, the order is assumed to be 
unsorted. |
+    | 103 | **`record_count`** | `long` | *required* | Number of records in 
this file. |
+    | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size 
in bytes. |
+    | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column 
stats. See [Content Stats](#content-stats). |
+    | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See 
manifest_info struct below. |
+    | 131 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split 
offsets for the data file. Must be sorted ascending. |
+    | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | 
Row-level deletion vector for a data file. |
+    | 158 | **`column_files`** | `list<159: column_file>` | *optional* | 
Column update files associated with this entry. |
+
+    **`tracking` struct (field 147)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: 
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, 
replacements, and modifications. Deletes are not used in scans. |
+    | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the 
deletion vector was added. |
+    | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | 
Snapshot ID where the latest column file was added. |
+    | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number 
of the file. Inherited when null and status is 1 (ADDED). |
+    | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence 
number indicating when the file was added. Inherited when null and status is 
ADDED. |
+    | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the 
`_row_id` for its first row. For a data manifest, the starting `_row_id` to 
assign to rows added by ADDED data files. See [First Row ID 
Inheritance](#first-row-id-inheritance). |
+    | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted 
in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |

Review Comment:
   I agree that the sentence is slightly off and it's not entirely clear what 
it tries to say. Maybe `Positions deleted in this snapshot in the referenced 
leaf manifest`



##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    Notes:
+
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |

Review Comment:
   we're calling this `content_type` in the implementation and I also feel like 
`content_type` is slightly more descriptive



##########
format/spec.md:
##########
@@ -656,15 +662,33 @@ A data or delete file is associated with a sort order by 
the sort order's id wit
 
 ### Manifests
 
-A manifest is an immutable Avro file that lists data files or delete files, 
along with each file’s partition data tuple, metrics, and tracking information. 
One or more manifest files are used to store a [snapshot](#snapshots), which 
tracks all of the files in a table at some point in time. Manifests are tracked 
by a [manifest list](#manifest-lists) for each table snapshot.
+A manifest is an immutable file that lists data files or delete files, along 
with each file’s partition data, metrics, and tracking information. One or more 
manifest files are used to store a [snapshot](#snapshots), which tracks all of 
the files in a table at some point in time. Manifests are tracked by a snapshot 
root for each table snapshot. In v4, the snapshot root is a root manifest that 
may track data files in addition to leaf manifest files.
 
 A manifest is a valid Iceberg data file: files must use valid Iceberg formats, 
schemas, and column projection.
 
-A manifest may store either data files or delete files, but not both because 
manifests that contain delete files are scanned first during job planning. 
Whether a manifest is a data manifest or a delete manifest is stored in 
manifest metadata.
+Each manifest type contains the following content:
 
-A manifest stores files for a single partition spec. When a table’s partition 
spec changes, old files remain in the older manifest and newer files are 
written to a new manifest. This is required because a manifest file’s schema is 
based on its partition spec (see below). The partition spec of each manifest is 
also used to transform predicates on the table's data rows into predicates on 
partition values that are used during job planning to select files from a 
manifest.
+| Manifest type | Contents |
+|----------------|----------|
+| v1-v3 data manifest | Data files |
+| v2-v3 delete manifest | Delete files |
+| v4 root manifest (snapshot root) | Data files, data manifests, delete 
manifests |
+| v4 data manifest | Data files and their colocated deletion vectors |
 

Review Comment:
   I had the same question as @gaborkaszab here. Maybe it would help to mention 
that delete manifests are only v2+v3 in L154?



##########
format/spec.md:
##########
@@ -742,18 +758,127 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
+
+    Notes:
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, 
snapshot, and sequence number. See tracking struct below. |
+    | 141 | **`spec_id`** | `int` | *optional* | ID of the partition spec used 
to write this manifest or data file. |
+    | 102 | **`partition`** | `struct<...>` | *optional* | Partition data 
tuple for the file. |
+    | 140 | **`sort_order_id`** | `int` | *optional* | ID representing sort 
order for this file. If missing or unknown, the order is assumed to be 
unsorted. |
+    | 103 | **`record_count`** | `long` | *required* | Number of records in 
this file. |
+    | 104 | **`file_size_in_bytes`** | `long` | *required* | Total file size 
in bytes. |
+    | 146 | **`content_stats`** | `content_stats` struct | *optional* | Column 
stats. See [Content Stats](#content-stats). |
+    | 150 | **`manifest_info`** | `manifest_info` struct | *optional* | See 
manifest_info struct below. |
+    | 131 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 132 | **`split_offsets`** | `list<133: long>` | *optional* | Split 
offsets for the data file. Must be sorted ascending. |
+    | 148 | **`deletion_vector`** | `deletion_vector` struct | *optional* | 
Row-level deletion vector for a data file. |
+    | 158 | **`column_files`** | `list<159: column_file>` | *optional* | 
Column files associated with this file. |
+
+    **`tracking` struct (field 147)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 0 | **`status`** | `int` (0: EXISTING, 1: ADDED, 2: DELETED, 3: 
REPLACED, 4: MODIFIED) | *required* | Used to track additions, deletions, 
replacements, and modifications. |
+    | 1 | **`snapshot_id`** | `long` | *optional* | Snapshot ID where the file 
was added or deleted. Inherited when null. |
+    | 5 | **`dv_snapshot_id`** | `long` | *optional* | Snapshot ID where the 
deletion vector was added. |
+    | 160 | **`latest_column_file_snapshot_id`** | `long` | *optional* | 
Snapshot ID where the latest column file was added. |
+    | 3 | **`sequence_number`** | `long` | *optional* | Data sequence number 
of the file. Inherited when null. See [Sequence Number 
Inheritance](#sequence-number-inheritance). |
+    | 4 | **`file_sequence_number`** | `long` | *optional* | File sequence 
number indicating when the file was added. Inherited when null. See [Sequence 
Number Inheritance](#sequence-number-inheritance). |
+    | 142 | **`first_row_id`** | `long` | *optional* | For a data file, the 
`_row_id` for its first row. For a data manifest, the starting `_row_id` to 
assign to rows added by ADDED data files. See [First Row ID 
Inheritance](#first-row-id-inheritance). |
+    | 6 | **`deleted_positions`** | `binary` | *optional* | Positions deleted 
in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+    | 7 | **`replaced_positions`** | `binary` | *optional* | Positions 
replaced in the referenced leaf manifest this snapshot. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`deletion_vector` struct (field 148)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 155 | **`location`** | `string` | *required* | Location of the Puffin 
file. |
+    | 144 | **`offset`** | `long` | *required* | Offset in the file where the 
content starts. |
+    | 145 | **`size_in_bytes`** | `long` | *required* | Length of the 
referenced content stored in the file. |
+    | 156 | **`cardinality`** | `long` | *required* | Cardinality of the 
deletion vector. |
+    | 149 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+
+    **`manifest_info` struct (field 150)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 504 | **`added_files_count`** | `int` | *required* | Count of entries 
with status ADDED in the manifest. |
+    | 505 | **`existing_files_count`** | `int` | *required* | Count of entries 
with status EXISTING in the manifest. |
+    | 506 | **`deleted_files_count`** | `int` | *required* | Count of entries 
with status DELETED in the manifest. |
+    | 523 | **`replaced_files_count`** | `int` | *required* | Count of entries 
with status REPLACED in the manifest. |
+    | 525 | **`modified_files_count`** | `int` | *required* | Count of entries 
with status MODIFIED in the manifest. |
+    | 512 | **`added_rows_count`** | `long` | *required* | Total number of 
rows in ADDED entries. |
+    | 513 | **`existing_rows_count`** | `long` | *required* | Total number of 
rows in EXISTING entries. |
+    | 514 | **`deleted_rows_count`** | `long` | *required* | Total number of 
rows in DELETED entries. |
+    | 524 | **`replaced_rows_count`** | `long` | *required* | Total number of 
rows in REPLACED entries. |
+    | 526 | **`modified_rows_count`** | `long` | *required* | Total number of 
rows in MODIFIED entries. |
+    | 516 | **`min_sequence_number`** | `long` | *required* | Minimum data 
sequence number of all live entries in the manifest. |
+    | 522 | **`dv`** | `binary` | *optional* | Positions in the referenced 
leaf manifest that are not live. See [Manifest Deletion 
Vectors](#manifest-deletion-vectors). |
+
+    **`column_file` struct (element 159 of `column_files`, field 158)**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 161 | **`format_version`** | `int` (4: V4) | *required* | Format version 
of this column file. |
+    | 162 | **`field_ids`** | `list<163: int>` | *required* | Live field IDs 
stored in this column file. |
+    | 164 | **`location`** | `string` | *required* | Location of the column 
file. |
+    | 165 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, or `parquet`. |
+    | 166 | **`file_size_in_bytes`** | `long` | *required* | Total column file 
size in bytes. |
+    | 167 | **`key_metadata`** | `binary` | *optional* | 
Implementation-specific key metadata for encryption. |
+    | 168 | **`split_offsets`** | `list<169: long>` | *optional* | Split 
offsets for the column file. Must be sorted ascending. |
+
+    **Tracked File Requirements**
+
+    - `deletion_vector.offset` and `deletion_vector.size_in_bytes` must 
exactly match the `offset` and `length` stored in the Puffin footer for the 
deletion vector blob.
+    - A leaf manifest written in v4 may only contain data files.
+    - A v1-v3 delete manifest referenced by a root manifest may contain v2-v3 
delete files.
+    - A root manifest may reference v1-v3 manifests; a referenced v1-v3 leaf 
manifest must have `format_version` PRE-V4.
+    - Other v4 tracked files must have `format_version` V4.

Review Comment:
   nit: most places use lowercase `v4`, but I see a few places across the 
changes that use `V4`. Maybe we should align those?



##########
format/spec.md:
##########
@@ -742,18 +758,126 @@ The `data_file` struct consists of the following fields:
     |            |            | _optional_ | **`144  content_offset`**         
| `long`                                                                      | 
The offset in the file where the content starts [5] |
     |            |            | _optional_ | **`145  content_size_in_bytes`**  
| `long`                                                                      | 
The length of a referenced content stored in the file; required if 
`content_offset` is present [5] |
 
-The `partition` struct stores the tuple of partition values for each file. Its 
type is derived from the partition fields of the partition spec used to write 
the manifest file. In v2, the partition struct's field ids must match the ids 
from the partition spec.
+    The `partition` struct stores the tuple of partition values for each file. 
Its type is derived from the partition fields of the partition spec used to 
write the manifest file. In v2, the partition struct's field ids must match the 
ids from the partition spec.
 
-The v4 `content_stats` container struct stores field-level metrics. Unlike the 
metrics maps, the type of `content_stats` is based on table metadata, like 
schema. Similar to the `partition` struct, the same type is used for all files 
tracked in a manifest.
+    Notes:
+
+    1. Single-value serialization for lower and upper bounds is detailed in 
Appendix D.
+    2. For `float` and `double`, the value `-0.0` must precede `+0.0`, as in 
the IEEE 754 `totalOrder` predicate. NaNs are not permitted as lower or upper 
bounds.
+    3. If sort order ID is missing or unknown, then the order is assumed to be 
unsorted. Only data files and equality delete files should be written with a 
non-null order id. [Position deletes](#position-delete-files) are required to 
be sorted by file and position, not a table order, and should set sort order id 
to null. Readers must ignore sort order id for position delete files.
+    4. Position delete metadata can use `referenced_data_file` when all 
deletes tracked by the entry are in a single data file. Setting the referenced 
file is required for deletion vectors.
+    5. The `content_offset` and `content_size_in_bytes` fields are used to 
reference a specific blob for direct access to a deletion vector. For deletion 
vectors, these values are required and must exactly match the `offset` and 
`length` stored in the Puffin footer for the deletion vector blob.
+    6. The following field ids are reserved on `data_file`: 141.
+
+=== "v4"
+    **Tracked Files**
+
+    | Field id | Name | Type | Required | Description |
+    |----------|------|------|----------|-------------|
+    | 134 | **`content_type`** | `int` (0: DATA, 3: DATA_MANIFEST, 4: 
DELETE_MANIFEST) | *required* | Type of content stored in the entry. |
+    | 157 | **`format_version`** | `int` (0: PRE-V4, 4: V4) | *required* | 
Writer format version. |
+    | 100 | **`location`** | `string` | *required* | Location of the file or 
manifest. |
+    | 101 | **`file_format`** | `string` | *required* | String file format 
name: `avro`, `orc`, `parquet`, or `puffin` |
+    | 147 | **`tracking`** | `tracking` struct | *required* | Groups status, 
snapshot, and sequence number. See tracking struct below. |

Review Comment:
   +1



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to