sevbanbayrak opened a new issue, #2070:
URL: https://github.com/apache/iceberg-go/issues/2070
### Apache Iceberg version
main (development) — reproduced on v0.7.0-rc0 (and v0.6.0, where DV read is
not implemented at all)
### Please describe the bug 🐞
**Summary**
`table/dv.ReadDVs` (v0.7.0-rc0) opens the DV file with `puffin.NewReader`
and locates the blob through the Puffin footer. Databricks Unity Catalog tables
with `delta.enableIcebergCompatV3` + deletion vectors write the DV as a Delta
`deletion_vector_<uuid>.bin` file: one version byte followed by the
`deletion-vector-v1` blob (4-byte length, magic `D1 D3 39 64`, roaring bitmap,
CRC32). There is no Puffin header/footer, so the read fails with `puffin:
invalid header magic`.
The manifest entry is otherwise well-formed: `file_format=PUFFIN`,
`referenced_data_file` set, `content_offset=1`, `content_size_in_bytes=<blob
length>`, and the bytes at that range decode as a valid `deletion-vector-v1`
blob (CRC verified).
The reference Java implementation (`BaseDeleteLoader.readDV`) never reads
the footer: it reads `content_size_in_bytes` bytes at `content_offset` and
deserializes the blob, so Spark/Trino read these tables. iceberg-go is stricter
than the reference reader and cannot read them.
**Expected**
Read the DV blob directly from `content_offset` / `content_size_in_bytes`
(the spec says these fields exist "for direct access to a deletion vector", and
that is what Java does), and fall back to the Puffin footer only when the
offset fields are absent. Alternatively, document that a Puffin container is
required and fail with a clearer error.
**Reproduce**
1. Databricks (Unity Catalog, table on an external location):
```sql
CREATE TABLE t (day DATE, golden_id STRING, row_hash STRING) USING DELTA
PARTITIONED BY (day)
TBLPROPERTIES ('delta.universalFormat.enabledFormats'='iceberg',
'delta.enableIcebergCompatV3'='true',
'delta.enableRowTracking'='true',
'delta.enableDeletionVectors'='true');
-- insert ~1M rows, then a small delete (a few thousand rows) so
Databricks writes a DV instead of rewriting the file
DELETE FROM t WHERE ...;
```
2. `rest.NewCatalog` against
`https://<workspace>/api/2.1/unity-catalog/iceberg-rest` with header
`X-Iceberg-Access-Delegation: vended-credentials`, `LoadTable`,
`Scan().ToArrowRecords`.
3. Errors:
- v0.6.0: `not implemented: deletion vector read is not yet implemented,
data file … has 1 deletion vector(s)`
- v0.7.0-rc0: `read deletion vectors from …/deletion_vector_<uuid>.bin:
create puffin reader for …: puffin: invalid header magic`
Manifest entry as seen through `FileScanTask.DeletionVectorFiles`:
`size=4233 offset=1 content_size=4232 format=PUFFIN`. Bytes at offset 1:
`00001080 d1d33964 …` (length 4224, magic, bitmap, CRC ok). Snapshot summary:
`added-delete-files=1, added-position-deletes=1020`. An earlier snapshot of the
same table without DVs (copy-on-write MERGE) reads correctly with both
versions, including the v3 metadata and column-mapped Parquet.
I'd like to work on this if the offset-based read is the preferred fix.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]