sevbanbayrak opened a new issue, #2070:
URL: https://github.com/apache/iceberg-go/issues/2070

   ### Apache Iceberg version
   
   main (development) — reproduced on v0.7.0-rc0 (and v0.6.0, where DV read is 
not implemented at all)
   
   ### Please describe the bug 🐞
   
   **Summary**
   
   `table/dv.ReadDVs` (v0.7.0-rc0) opens the DV file with `puffin.NewReader` 
and locates the blob through the Puffin footer. Databricks Unity Catalog tables 
with `delta.enableIcebergCompatV3` + deletion vectors write the DV as a Delta 
`deletion_vector_<uuid>.bin` file: one version byte followed by the 
`deletion-vector-v1` blob (4-byte length, magic `D1 D3 39 64`, roaring bitmap, 
CRC32). There is no Puffin header/footer, so the read fails with `puffin: 
invalid header magic`.
   
   The manifest entry is otherwise well-formed: `file_format=PUFFIN`, 
`referenced_data_file` set, `content_offset=1`, `content_size_in_bytes=<blob 
length>`, and the bytes at that range decode as a valid `deletion-vector-v1` 
blob (CRC verified).
   
   The reference Java implementation (`BaseDeleteLoader.readDV`) never reads 
the footer: it reads `content_size_in_bytes` bytes at `content_offset` and 
deserializes the blob, so Spark/Trino read these tables. iceberg-go is stricter 
than the reference reader and cannot read them.
   
   **Expected**
   
   Read the DV blob directly from `content_offset` / `content_size_in_bytes` 
(the spec says these fields exist "for direct access to a deletion vector", and 
that is what Java does), and fall back to the Puffin footer only when the 
offset fields are absent. Alternatively, document that a Puffin container is 
required and fail with a clearer error.
   
   **Reproduce**
   
   1. Databricks (Unity Catalog, table on an external location):
      ```sql
      CREATE TABLE t (day DATE, golden_id STRING, row_hash STRING) USING DELTA 
PARTITIONED BY (day)
      TBLPROPERTIES ('delta.universalFormat.enabledFormats'='iceberg', 
'delta.enableIcebergCompatV3'='true',
                     'delta.enableRowTracking'='true', 
'delta.enableDeletionVectors'='true');
      -- insert ~1M rows, then a small delete (a few thousand rows) so 
Databricks writes a DV instead of rewriting the file
      DELETE FROM t WHERE ...;
      ```
   2. `rest.NewCatalog` against 
`https://<workspace>/api/2.1/unity-catalog/iceberg-rest` with header 
`X-Iceberg-Access-Delegation: vended-credentials`, `LoadTable`, 
`Scan().ToArrowRecords`.
   3. Errors:
      - v0.6.0: `not implemented: deletion vector read is not yet implemented, 
data file … has 1 deletion vector(s)`
      - v0.7.0-rc0: `read deletion vectors from …/deletion_vector_<uuid>.bin: 
create puffin reader for …: puffin: invalid header magic`
   
   Manifest entry as seen through `FileScanTask.DeletionVectorFiles`: 
`size=4233 offset=1 content_size=4232 format=PUFFIN`. Bytes at offset 1: 
`00001080 d1d33964 …` (length 4224, magic, bitmap, CRC ok). Snapshot summary: 
`added-delete-files=1, added-position-deletes=1020`. An earlier snapshot of the 
same table without DVs (copy-on-write MERGE) reads correctly with both 
versions, including the v3 metadata and column-mapped Parquet.
   
   I'd like to work on this if the offset-based read is the preferred fix.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to