nssalian commented on code in PR #18344: URL: https://github.com/apache/iceberg/pull/18344#discussion_r4162502240
########## site/docs/blog/posts/2026-10-02-iceberg-1.12.0-release.md: ########## @@ -0,0 +1,208 @@ +--- +date: 2026-10-02 +title: Apache Iceberg 1.12.0 Release +slug: apache-iceberg-1.12.0-release +authors: + - iceberg-pmc +categories: + - release +--- + +<!-- + - Licensed to the Apache Software Foundation (ASF) under one or more + - contributor license agreements. See the NOTICE file distributed with + - this work for additional information regarding copyright ownership. + - The ASF licenses this file to You under the Apache License, Version 2.0 + - (the "License"); you may not use this file except in compliance with + - the License. You may obtain a copy of the License at + - + - http://www.apache.org/licenses/LICENSE-2.0 + - + - Unless required by applicable law or agreed to in writing, software + - distributed under the License is distributed on an "AS IS" BASIS, + - WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. + - See the License for the specific language governing permissions and + - limitations under the License. + --> + +The Apache Iceberg community is pleased to announce the release of Apache Iceberg 1.12.0. This release is the result of **777 commits** from **142 contributors** since 1.11.0. See the [release notes](https://iceberg.apache.org/releases/#1120-release) for the complete list of changes. + +<!-- more --> + +## Release Highlights + +### Data Types: Variant and Geospatial + +This release improves Variant read performance, extends shredded Variant writes beyond Spark, and adds read and write support for the geospatial types. + +**Variant.** Reading unshredded Variant data is now faster in Spark. On both Spark 4.0 and 4.1, unshredded Variant columns are [read through the vectorized Parquet path](https://github.com/apache/iceberg/pull/16292) instead of row-at-a-time decoding. + +Variant shredding also uses a stricter layout rule. Shredding now [requires type uniformity](https://github.com/apache/iceberg/pull/17424): a field is shredded into a typed column only when all of its values fall into a single type family (after numeric widening); fields that mix types stay in the untyped residual. Unlike the previous majority-based rule, which could shred a field while covering only a fraction of its rows, every typed column now fully covers its field. Single-type fields are unaffected. + +Shredded Variant writes are no longer Spark-only. Flink can now [write shredded Variant](https://github.com/apache/iceberg/pull/15596), and the [Kafka Connect sink and the generic record writer](https://github.com/apache/iceberg/pull/17520) can produce shredded Variant as well, so semi-structured data ingested through those paths benefits from the same read-time pushdown. Flink also gains [Variant support in Avro readers and writers](https://github.com/apache/iceberg/pull/17737). + +Several correctness and hardening fixes also landed for Variant: + +- Shredded-column string bounds are computed in [UTF-8 byte order](https://github.com/apache/iceberg/pull/17397) and binary upper bounds [truncate up](https://github.com/apache/iceberg/pull/16880) so pruning stays correct; bounds also honor the column's [configured truncation length](https://github.com/apache/iceberg/pull/17342) +- A [crash computing metrics for a value column with no statistics](https://github.com/apache/iceberg/pull/16585) is fixed, and [large-decimal shredding (precision > 18)](https://github.com/apache/iceberg/pull/17002) is corrected +- The Variant classes are made [serializable](https://github.com/apache/iceberg/pull/17260), and binary parsing is [hardened against malformed input](https://github.com/apache/iceberg/pull/16568) +- [ORC filter pushdown on tables with a Variant column](https://github.com/apache/iceberg/pull/17998) is fixed + +**Geospatial.** The `geometry` and `geography` types gain read and write support. Both are stored as Well-Known Binary (WKB) and can now be [read and written in Avro](https://github.com/apache/iceberg/pull/17119) and [in Parquet](https://github.com/apache/iceberg/pull/16982), where they map to the [Parquet geometry and geography logical types](https://github.com/apache/iceberg/pull/16765) so files are self-describing; [single-value binary serialization](https://github.com/apache/iceberg/pull/16607) is also in place for defaults and metadata. + +[Spark 4.1](https://github.com/apache/iceberg/pull/17073) is the first engine with an end-to-end geospatial path: it reads and writes both types in Parquet and supports row-level `DELETE`, `UPDATE`, and `MERGE` on tables with geospatial columns, including the merge-on-read, deletion-vector path on format version 3. Current limitations: + +- Support is limited to Spark 4.1 (not Spark 3.5 or 4.0, Flink, or ORC) +- Reads use the row-based reader; there is no Arrow geospatial vector yet +- There are no spatial predicates yet, so filters are expressed against non-geospatial columns + +### Deletion Vectors and Streaming Deletes + +Streaming pipelines that upsert into Iceberg write *equality deletes*: markers that say "remove every row whose key matches these values." They are cheap to write but expensive to read, because every query has to re-open data files and compare values to work out which rows still exist. 1.12.0 adds a Flink-native maintenance task that resolves those deletes once, instead of on every scan. + +**`ConvertEqualityDeletes`.** This new maintenance task resolves the equality deletes produced during streaming ingest into row-position deletion vectors and commits them alongside the data files. After conversion, readers apply deletes by position rather than re-scanning and comparing values, so queries no longer pay this cost. + +It pairs with `IcebergSink`, which stages new data files and equality deletes on a source branch; the converter resolves those into deletion vectors and commits to the target branch, or converts in place. Because deletion vectors are a v3 feature, the task requires table format version 3 or later and runs on Flink 1.20, 2.1, 2.2, and 2.3. It landed across several changes, including the [core data model](https://github.com/apache/iceberg/pull/16831) and [integration with `IcebergSink`](https://github.com/apache/iceberg/pull/17142); a follow-up [ensures deleted rows do not reappear after a failed conversion cycle](https://github.com/apache/iceberg/pull/17630). + +**Correctness.** Deletion vectors also gain [co-located access through `DataFile.deletionVector()`](https://github.com/apache/iceberg/pull/17928), and get fixes for [references when they share a Puffin file](https://github.com/apache/iceberg/pull/17497) and [preserved encryption metadata on merge](https://github.com/apache/iceberg/pull/15911). + +### Data Layout + +**Hilbert clustering.** `rewrite_data_files` gains [Hilbert-curve clustering](https://github.com/apache/iceberg/pull/16827), a new multi-dimensional sort strategy alongside Z-order. Both map several columns onto a single space-filling curve so rows with similar values in those columns are stored together, improving file skipping for multi-column filters. Hilbert typically preserves locality better than Z-order because neighboring points on the curve are always adjacent in the data, without the large "jumps" across the space that Z-order makes. You select it through the sort strategy: + +```sql +CALL system.rewrite_data_files( + table => 'db.tbl', + strategy => 'sort', + sort_order => 'hilbert(c1, c2)' +); +``` + +Hilbert clustering ships for Spark 4.1. + +**Other maintenance.** The [`RepairTable` action interface](https://github.com/apache/iceberg/pull/17399) is defined (a standard way to repair manifest-entry statistics that disagree with the files they describe), and Spark's `rewrite_data_files` now accepts the [`max-file-group-input-files`](https://github.com/apache/iceberg/pull/17544) option to cap the input files in a single group. + +### REST Catalog + +The REST catalog protocol picks up several additions, most at the specification and OpenAPI layer. + +The largest is [finer-grained read restrictions](https://github.com/apache/iceberg/pull/13879) on `loadTable`. A catalog can return a `ReadRestrictions` object in the load response describing required column projections (column-masking actions such as showing only the last four characters, replacing a value with null, truncating a timestamp, hashing, or alphanumeric masking) together with a required row filter modeled as an Iceberg expression. The contract is client-enforced and fail-closed: a reader that supports read restrictions must apply every returned action and filter in full, and if it cannot apply one it must fail the query rather than return raw, partial, or empty rows. This is a spec and OpenAPI contract only; there is no engine-side enforcement in 1.12.0. + +The [`VariantType`](https://github.com/apache/iceberg/pull/17256) is now representable in the OpenAPI spec so Variant columns can travel in schemas over the protocol. + +Two endpoints are added: read-only [list and load function](https://github.com/apache/iceberg/pull/15180) endpoints, and an [unregister-table endpoint](https://github.com/apache/iceberg/pull/16400) that detaches a table from a catalog without deleting its data or metadata. [`CatalogObjectIdentifier`](https://github.com/apache/iceberg/pull/16160) adds a shared way to name catalog objects. Remote signing configuration is also [formalized in the spec](https://github.com/apache/iceberg/pull/16822), with a corresponding [client implementation](https://github.com/apache/iceberg/pull/17709). A [`labels` field for catalog metadata enrichment](https://github.com/apache/iceberg/pull/15750) is added to the spec, [read on load responses](https://github.com/apache/iceberg/pull/18045) and [exposed via `SupportsLabels`](https://github.com/apache/iceberg/pull/18046). + +### Spec Changes + +The [expressions specification is adopted](https://github.com/apache/iceberg/pull/16652), giving Iceberg a formal, engine-independent definition of predicate expressions; the [REST OpenAPI spec is aligned](https://github.com/apache/iceberg/pull/17138) to it in the same release, so read-restriction row filters are portable across clients. + +[Content-file uniqueness](https://github.com/apache/iceberg/pull/17198) is now spelled out: within a snapshot, each content file must be referenced by at most one live manifest entry across all manifests, and a snapshot that violates this has undefined behavior. Smaller clarifications round out the spec work: [Variant type classification](https://github.com/apache/iceberg/pull/16836) and [decimal type serialization](https://github.com/apache/iceberg/pull/16798). + +### Format V4 Foundations + +Work toward Table Format V4 continues. V4 is under active development and has not been formally adopted: Iceberg 1.12.0 cannot read or write a V4 table, and everything here is groundwork the next ratified spec version will build on. + +**Manifests.** The release adds a [V4 manifest reader](https://github.com/apache/iceberg/pull/16958) that reads V4 manifests written in both Parquet and Avro, the [ability to write Parquet and Avro manifests in the V4 layout](https://github.com/apache/iceberg/pull/15634), and an [adapter layer](https://github.com/apache/iceberg/pull/17932) with a [tracked-file metadata model](https://github.com/apache/iceberg/pull/16100) that surfaces V4-tracked files through the existing `ManifestFile` and `DataFile` interfaces. These operate at the manifest level for testing and future use; tables still cannot be created or committed in V4. + +**Foundations.** Several structural pieces land: [relative paths in metadata](https://github.com/apache/iceberg/pull/15630) (backed by [relativization utilities](https://github.com/apache/iceberg/pull/16174) and [reader-side resolution](https://github.com/apache/iceberg/pull/17434)), a step toward relocatable tables; a new [content-statistics spec addition](https://github.com/apache/iceberg/pull/14234); and a [read-only Mumbling bitmap implementation](https://github.com/apache/iceberg/pull/16747). + +### Performance and Reliability + +**Faster planning and scans.** [Manifest list files are now cached](https://github.com/apache/iceberg/pull/16762) in the manifest content cache, and Parquet gains [adaptive bloom filter sizing](https://github.com/apache/iceberg/pull/16363) and [per-column dictionary encoding](https://github.com/apache/iceberg/pull/16713). + +**Vectorized reader fixes.** The vectorized Parquet reader picks up several correctness fixes: [`int`-to-`long` promotion](https://github.com/apache/iceberg/pull/16343) that previously threw a `ClassCastException`, [decimal columns with default values](https://github.com/apache/iceberg/pull/16501), [decimals with precision greater than 18](https://github.com/apache/iceberg/pull/16627), [dictionary-encoded `VARCHAR`/`VARBINARY` through direct byte buffers](https://github.com/apache/iceberg/pull/17055), [all-null `DELTA`-encoded pages](https://github.com/apache/iceberg/pull/17017), an [INT96 dictionary-decode offset](https://github.com/apache/iceberg/pull/16435), and a [direct-memory leak in the row-lineage readers](https://github.com/apache/iceberg/pull/17296). + +**Sharper pruning.** Predicate pushdown on nanosecond timestamps is fixed in [Parquet](https://github.com/apache/iceberg/pull/16619) and [ORC](https://github.com/apache/iceberg/pull/17750), [`notStartsWith` no longer skips row groups containing nulls](https://github.com/apache/iceberg/pull/17656), and [null counting is correct for Parquet files written without `null_count`](https://github.com/apache/iceberg/pull/17557). + +**Core fixes.** [FileIO leaks are fixed and `close()` is standardized across catalog implementations](https://github.com/apache/iceberg/pull/16862), [Z-order byte encoding of floating-point values is corrected](https://github.com/apache/iceberg/pull/17071), [`EncryptingFileIO` is reworked as a `DelegateFileIO`](https://github.com/apache/iceberg/pull/14876) so encrypted tables keep access to underlying I/O capabilities such as bulk operations, and a [manifest-pruning and residual-evaluation bug](https://github.com/apache/iceberg/pull/17443) is resolved. + +### Security + +This release fixes two HIGH-severity `jackson-databind` CVEs, CVE-2026-54512 and CVE-2026-54513, by [aligning Jackson versions across the runtimes and bundles](https://github.com/apache/iceberg/pull/16954); a further [Jackson bump](https://github.com/apache/iceberg/pull/17336) fixes GHSA-r7wm-3cxj-wff9. Three Jackson findings remain in the Kafka Connect runtime because they come from the copy of Jackson shaded inside `parquet-jackson`, which cannot be upgraded independently of Apache Parquet. + +The [shaded-jar LICENSE and NOTICE files were cleaned up](https://github.com/apache/iceberg/pull/16543): duplicate license metadata was stripped from the cloud and engine bundles, and missing third-party notices were added. + +### Engine Updates + +#### Spark + +Spark support in 1.12.0 is Spark 3.5, 4.0, and 4.1; [Spark 3.4 support is removed](https://github.com/apache/iceberg/pull/14122). Notable additions: + +- **Delegated PURGE**: `DROP TABLE ... PURGE` can be [delegated to REST catalogs](https://github.com/apache/iceberg/pull/15614) via the `rest-catalog-purge` property, on Spark 3.5, 4.0, and 4.1 +- **Tolerant migration**: the [`snapshot`](https://github.com/apache/iceberg/pull/16710) and [`migrate`](https://github.com/apache/iceberg/pull/16643) procedures accept `ignore_missing_files` +- **Manifest rewrite by sort key**: the [`rewrite_manifests` procedure gains a `sort_by` parameter](https://github.com/apache/iceberg/pull/18065) on Spark 3.5 and 4.0, matching Spark 4.1 +- **Streaming merge-append**: a [write config](https://github.com/apache/iceberg/pull/17347) enables merge-append for Structured Streaming writes, [ported to Spark 3.5 and 4.0](https://github.com/apache/iceberg/pull/17403) + +Spark 4.1 also gains geospatial support, described in the Data Types section above. + +#### Flink + +Flink support in 1.12.0 is Flink 1.20, 2.1, 2.2, and 2.3; Flink 2.2 and 2.3 are added and Flink 2.0 is removed. + +Beyond the equality-delete-to-deletion-vector conversion above, Flink gains Iceberg view support: it can [read views in SQL](https://github.com/apache/iceberg/pull/17859) and [create, drop, and rename them](https://github.com/apache/iceberg/pull/17873) through `FlinkCatalog` (both backported to 2.2, 2.1, and 1.20). + +The Dynamic Sink gains [fine-grained slot-sharing-group control](https://github.com/apache/iceberg/pull/16065) and several fixes: + +- [Honors schema identifier fields when routing records](https://github.com/apache/iceberg/pull/16243) +- [Avoids duplicate commits when the Flink job id changes on restart](https://github.com/apache/iceberg/pull/16011) + +#### Kafka Connect + +The sink connector now [surfaces commit failures instead of swallowing them](https://github.com/apache/iceberg/pull/16237), adds [bounded retry for transient commit exceptions](https://github.com/apache/iceberg/pull/16434) and a [metric for partial commit failures](https://github.com/apache/iceberg/pull/16433), and improves offset handling by [only committing offsets greater than the existing ones](https://github.com/apache/iceberg/pull/17552) and [tracking control-topic offsets as a high-water mark](https://github.com/apache/iceberg/pull/17933). + +## Breaking Changes + +Users upgrading from 1.11.0 should review these before upgrading. + +**Removals:** + +- **Spark 3.4 removed**: [Spark 3.4 support has been removed](https://github.com/apache/iceberg/pull/14122); use Spark 3.5, 4.0, or 4.1 +- **Flink 2.0 removed**: Flink 2.0 support has been removed (Flink 2.2 and 2.3 are added), so 1.12.0 supports Flink 1.20, 2.1, 2.2, and 2.3 +- **Position delete files with row data removed**: [writing position deletes that carry row data is removed](https://github.com/apache/iceberg/pull/17706), consistent with the move to deletion vectors on format version 3 +- **`DataReader` removed**: [removed in favor of `PlannedDataReader`](https://github.com/apache/iceberg/pull/17699) + +**Deprecated APIs removed:** + +- [Partition-stats read functionality](https://github.com/apache/iceberg/pull/14998) (use the Partition Stats Scan API) +- Spark [`SparkFilters`](https://github.com/apache/iceberg/pull/17702), [`SparkTableUtil` methods](https://github.com/apache/iceberg/pull/17703), and [`SparkReadConf`/`SparkWriteConf`/`SparkSchemaUtil` methods](https://github.com/apache/iceberg/pull/17626) +- Data [`GenericAppenderFactory` and `BaseFileWriterFactory`](https://github.com/apache/iceberg/pull/17696) +- AWS [S3 signer classes and properties](https://github.com/apache/iceberg/pull/17627) +- Flink [`RewriteDataFiles.Builder.filter(Expression)`](https://github.com/apache/iceberg/pull/17624) +- Core [REST namespace encoding helpers](https://github.com/apache/iceberg/pull/17697) and the [`HadoopFileIO(SerializableSupplier)` constructor](https://github.com/apache/iceberg/pull/17704) +- Kafka Connect [`TableReference` and `IcebergWriterResult` members](https://github.com/apache/iceberg/pull/17623) +- BigQuery [catalog property constants](https://github.com/apache/iceberg/pull/17625) +- A [final sweep of methods and fields scheduled for 1.12.0 removal](https://github.com/apache/iceberg/pull/17700) + +**Behavior changes:** + +- **AWS HTTP client**: the default AWS SDK HTTP client [migrated to Apache HttpClient 5](https://github.com/apache/iceberg/pull/18195); users who provide AWS dependencies separately must switch from `software.amazon.awssdk:apache-client` to `software.amazon.awssdk:apache5-client` +- **Geospatial `toString()`**: [`GeometryType` and `GeographyType` `toString()`](https://github.com/apache/iceberg/pull/16765) now include the resolved CRS, and for geography the edge algorithm (for example `geometry(OGC:CRS84)` and `geography(OGC:CRS84, spherical)`), instead of the bare type name +- **REST idempotency retries**: the REST client now [retries POST requests carrying an `Idempotency-Key`](https://github.com/apache/iceberg/pull/17947) on retriable errors (408, 500, 502, 503, 504) + +## Dependency Updates + +Notable dependency updates in 1.12.0: + +- **AWS SDK (bom)**: 2.44.4 -> 2.54.17 +- **Jackson (bom)**: 2.21.3 -> 2.22.2 +- **Netty**: 4.2.13.Final -> 4.2.18.Final +- **Nessie**: 0.107.5 -> 0.108.8 +- **RoaringBitmap**: 1.6.14 -> 1.6.23 +- **Avro**: 1.12.1 -> 1.12.2 +- **ORC**: 1.9.8 -> 1.9.9 +- **Guava**: 33.6.0-jre -> 33.7.1-jre +- **Caffeine**: 2.9.3 -> 3.2.4 + +The [Caffeine upgrade](https://github.com/apache/iceberg/pull/16680) is a major-line jump that drops the deprecated `sun.misc.Unsafe` dependency in favor of `VarHandle`; because Caffeine 3.x migrated its nullness annotations to JSpecify, the runtime bundles now pull `org.jspecify:jspecify` in place of `org.checkerframework:checker-qual`, and the Spark and Flink runtime LICENSE files were updated accordingly. Review Comment: will remove -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
