Thanks Xiening, that's a fair concern, and I agree the fast update/delete
need doesn't go away.

I'd frame it less as shifting the burden and more as changing where it's
paid. Equality deletes push the cost onto every read: each reader re-joins
the delete files against candidate rows, on every query, for the life of
the table. Resolving position once (at write time, or once during
background conversion) pays that cost a single time and makes all later
reads cheap, an O(1) DV check. And the streaming writer stays cheap either
way: background conversion (or write-time DV resolution via the index)
moves the position-resolution cost off the hot write path, so the write
stays cheap while every downstream reader gets fast, position-based deletes
instead of re-joining equality-delete files on each query.

On index cost: I agree keeping a key -> position index current isn't free,
but the evidence so far is that it's manageable. Max's
ConvertEqualityDeletes already maintains a persistent RocksDB PK index
incrementally at streaming scale, not rebuilt each cycle, so this is
running in practice, not just in theory.

On tooling and adoption: tooling is central to the proposal, not an
afterthought. Background conversion already lets writers keep working
unchanged, they write equality deletes as usual, and the conversion job
resolves them to DVs using its own internal index. The persistent
key-lookup index is the shared tooling that makes write-time elimination
practical and engine-agnostic. Because forbidding is a V4-table property,
streaming-upsert workloads that don't yet have a write-time DV path can
keep running on V3 (writing equality deletes, with background conversion
keeping reads fast), and move to V4 once their engine can emit DVs
directly. Non-streaming workloads can adopt V4 right away. So it shouldn't
force anyone into a worse position or block adoption.

Finally, removing equality deletes isn't only about read cost. They also
block CDC, row lineage, and incremental maintenance of indexes and
materialized views, so it's less "shift the burden" and more "unblock
features that equality deletes currently make impossible."

Thanks,
Huaxin

On Mon, Jul 20, 2026 at 3:20 PM Xiening Dai <[email protected]> wrote:

> Hi Huaxin,
>
> Thanks for bringing this up. Equality delete is indeed a pain point we
> have seen in many customer use cases.
>
> That been said the scenario of fast update/delete still exists no matter
> which technology or table spec we choose. The proposal is just going to
> shift the burden from the reader to the writer. To achieve fast
> update/delete, customer can build index structure like you mentioned, but
> building such index and keeping it up to date all time can be very
> expensive too (especially given that we are tackling the fast update/delete
> scenario). So although it feels like a right direction as we are saying
> that we don't want to handle this complexity on the table spec, the
> underlying problem is not solved. Without a good alternative or tooling
> support, i am afraid this could become an adoption issue for v4 going
> forward.
>
> On 2026/07/17 17:13:55 huaxin gao wrote:
> > Thanks all for the discussion. I've thought this over, and I'd like to
> > change my view to forbidding equality-delete writes for V4 tables, rather
> > than the softer "deprecated but permitted."
> >
> > My earlier hesitation was about gating V4 on a write-time DV
> implementation
> > that isn't built yet. I think that concern goes away once we separate the
> > spec decision from engine adoption:
> >
> > - Forbidding equality deletes is a V4 format decision. It defines what a
> V4
> > table allows; it does not require every engine to have the write-time DV
> > path on day one.
> > - Workloads that still rely on streaming upserts can stay on V3 until
> their
> > engine's write-time path is ready, then adopt V4. Readers continue to
> > support equality deletes for existing V2 and V3 tables.
> > - So we don't need the write-time implementation finished to forbid
> > equality deletes in V4. What we need is a clear, credible path, and I
> think
> > we have it: Flink's ConvertEqualityDeletes already maintains a PK index
> in
> > Flink state and does key -> position -> DV today, and the next step is to
> > persist that index into Iceberg so any engine can resolve positions and
> > emit DVs directly at write time.
> >
> > This also addresses Max's concern that "deprecated but permitted" is too
> > soft. Forbidding for V4 gives engines a real incentive to move, while
> > keeping existing tables fully readable.
> >
> > On Xin's point, agreed that some perf data comparing DVs on V4 against
> > equality deletes on V3 would be useful to have as we go.
> >
> > Thanks,
> > Huaxin
> >
> > On Thu, Jul 16, 2026 at 8:43 AM Xin Huang via dev <
> [email protected]>
> > wrote:
> >
> > > Conceptually +1 as equality deletes really complicates the format and
> > > implementation.
> > >
> > > However given the concern is around performance side. Is there a way to
> > > make the decision making more data driven — having some benchmarking on
> > > perf comparison between dv+single file commit in v4 vs equality delete
> in
> > > v3 could help making a call.
> > >
> > > Thanks
> > > Xin
> > >
> > > On Wed, Jul 15, 2026 at 11:01 PM Maximilian Michels <[email protected]>
> > > wrote:
> > >
> > >> I understood "deprecate equality deletes" as not forbidding engines to
> > >> write them, but rather discouraging them. IMHO this is long overdue,
> > >> but it is also a very soft transition. Perhaps too soft, because it
> > >> doesn't give engines who write them a real incentive to stop writing
> > >> equality deletes.
> > >>
> > >> Engines will likely be quicker to move away from writing equality
> > >> deletes if we disallow writing them in V4. Regardless, we will have to
> > >> support reading equality deletes for V2 and V3 tables.
> > >>
> > >> I'm leaning more towards removing equality deletes for V4 tables, but
> > >> I would like to hear what others think.
> > >>
> > >> -Max
> > >>
> > >>
> > >> On Wed, Jul 15, 2026 at 10:03 PM Steven Wu <[email protected]>
> wrote:
> > >> >
> > >> > > So I'd separate two things: deprecating in V4 (signal and
> direction,
> > >> safe to do now) versus forbidding equality-delete writes (gated on the
> > >> engine-agnostic path being ready). I'm only proposing the first for
> V4.
> > >> >
> > >> > I thought we wanted to forbid equality-delete writes for v4 tables,
> > >> which would really simplify the v4 adaptive metadata tree along with
> other
> > >> benefits that Huaxin already outlined,
> > >> >
> > >> > > deprecation path in v4
> > >> >
> > >> > I heard the Kafka connector in the Iceberg repo doesn't produce
> > >> equality deletes. We would need to migrate the Flink sink to leverage
> the
> > >> index to produce DVs only in v4.
> > >> >
> > >> > On Wed, Jul 15, 2026 at 12:13 PM huaxin gao <[email protected]
> >
> > >> wrote:
> > >> >>
> > >> >> Thanks Max and Manu.
> > >> >>
> > >> >> Max, thanks for the added detail. It's a good point that the index
> in
> > >> ConvertEqualityDeletes is persisted in Flink state (RocksDB) and
> updated
> > >> incrementally. That strengthens the case, since it shows the key to
> > >> position index is already durable, just scoped to Flink today. And
> your
> > >> closing point is exactly the plan I have in mind: deprecate equality
> > >> deletes in V4, and once the spec has an index, persist that PK index
> into
> > >> Iceberg so it can be shared across engines.
> > >> >>
> > >> >> Manu, good question on how this works in practice. To be clear, I'm
> > >> proposing "deprecated but permitted," not removal. In V4, writers
> > >> (including Kafka Connect) could keep emitting equality deletes and
> readers
> > >> would keep applying them. Deprecation just declares deletion vectors
> the
> > >> going-forward mechanism and stops new investment in equality deletes.
> So V4
> > >> would not be gated on the index landing.
> > >> >>
> > >> >> On the cleanup path today: ConvertEqualityDeletes runs against the
> > >> table, not a specific writer, so a Kafka Connect pipeline can already
> be
> > >> cleaned up. It keeps writing equality deletes, and a standalone Flink
> > >> ConvertEqualityDeletes job converts them to DVs. The one friction is
> that
> > >> the conversion runtime is Flink today, so a Kafka-only shop would
> have to
> > >> run Flink just for maintenance. A Spark action would be a natural
> follow-up
> > >> here, since Spark is the usual Iceberg maintenance engine and most
> batch
> > >> shops already run it.
> > >> >>
> > >> >> Longer term, the next step is to persist the key to position index
> > >> into Iceberg. Once it's shared, engines can look up positions and
> write DVs
> > >> directly at write time, so a writer can stop producing equality
> deletes
> > >> entirely, and neither the Flink job nor a Spark action needs to
> rebuild the
> > >> index each run. Each engine (Kafka Connect, Spark, Flink) would adopt
> > >> write-time DVs on its own schedule; until then it keeps writing
> equality
> > >> deletes and relies on background conversion. So V4 deprecation doesn't
> > >> require re-implementing every writer up front.
> > >> >>
> > >> >> So I'd separate two things: deprecating in V4 (signal and
> direction,
> > >> safe to do now) versus forbidding equality-delete writes (gated on the
> > >> engine-agnostic path being ready). I'm only proposing the first for
> V4.
> > >> >>
> > >> >> Thanks,
> > >> >> Huaxin
> > >> >>
> > >> >> On Wed, Jul 15, 2026 at 3:10 AM Manu Zhang <
> [email protected]>
> > >> wrote:
> > >> >>>
> > >> >>> Hi Huaxin,
> > >> >>>
> > >> >>> +1 for deprecating equality deletes, but how would this
> deprecation
> > >> work practically in V4?
> > >> >>> As Max pointed out, we still lack an engine-agnostic solution for
> > >> streaming use cases. For example, how would we handle equality deletes
> > >> written by Kafka Connect?
> > >> >>> While the index proposal looks promising, I don't see a clear path
> > >> for deprecating equality deletes in V4 before that index work actually
> > >> lands.
> > >> >>>
> > >> >>> Thanks,
> > >> >>> Manu
> > >> >>>
> > >> >>>
> > >> >>> On Wed, Jul 15, 2026 at 5:30 PM Maximilian Michels <
> [email protected]>
> > >> wrote:
> > >> >>>>
> > >> >>>> Hi Huaxin,
> > >> >>>>
> > >> >>>> Thanks for reviving the discussion on deprecating equality
> deletes.
> > >> >>>> Equality deletes are the number one pain for streaming use cases.
> > >> Many
> > >> >>>> users give up when they see the merge-on-read costs, or they
> build
> > >> >>>> custom solutions which move them further away from core Iceberg.
> That
> > >> >>>> said, we've made great progress since the initial conversation in
> > >> >>>> 2024.
> > >> >>>>
> > >> >>>> Just to add what you said: The index we maintain in
> > >> >>>> ConvertEqualityDeletes is not ephemeral. The index is persisted
> in
> > >> >>>> Flink's managed state (RocksDB). It is continuously updated as
> new
> > >> >>>> data arrives and checkpointed periodically. However, even though
> the
> > >> >>>> conversion works for data written by any engine, we currently
> require
> > >> >>>> Flink for the conversion itself. Storing the index directly in
> > >> Iceberg
> > >> >>>> and enabling all engines access would be the next logical step
> > >> towards
> > >> >>>> a fully engine-agnostic solution.
> > >> >>>>
> > >> >>>> The reality is that we don't yet have a working solution to avoid
> > >> >>>> writing equality deletes across all engines, but given the recent
> > >> >>>> progress, the proposed plan seems realistic. So +1 for
> deprecating
> > >> >>>> equality deletes in V4.
> > >> >>>>
> > >> >>>> Cheers,
> > >> >>>> Max
> > >> >>>>
> > >> >>>>
> > >> >>>>
> > >> >>>> On Tue, Jul 14, 2026 at 3:25 AM huaxin gao <
> [email protected]>
> > >> wrote:
> > >> >>>> >
> > >> >>>> > Hi all,
> > >> >>>> >
> > >> >>>> > I'd like to restart the conversation about deprecating equality
> > >> deletes, now in the context of the V4 spec.
> > >> >>>> >
> > >> >>>> > Background
> > >> >>>> >
> > >> >>>> > This isn't a new idea. Russell proposed deprecating equality
> > >> deletes in V3 and removing them from the spec in V4, back in October
> 2024
> > >> in "[DISCUSS] - Deprecate Equality Deletes". The main blocker at the
> time
> > >> was that equality deletes served real use cases (especially Flink
> streaming
> > >> upserts) with no efficient alternative. Two developments since then
> make
> > >> the V4 removal worth acting on now.
> > >> >>>> >
> > >> >>>> > Why equality deletes are costly
> > >> >>>> >
> > >> >>>> > Equality deletes are cheap to write but expensive to read: a
> > >> reader must load the equality-delete files and join them against every
> > >> candidate row in the delete's sequence-number range. Positional
> deletes
> > >> skip that per-row join by marking exact positions, so they have
> always read
> > >> faster, and V3 deletion vectors make them faster still, one compact
> bitmap
> > >> per data file, applied by an O(1) position check, instead of V2's many
> > >> position-delete files. So equality deletes' only real edge is the
> cheap
> > >> write, and both background conversion and a write-time key-lookup
> index can
> > >> recover that.
> > >> >>>> >
> > >> >>>> > Beyond performance
> > >> >>>> >
> > >> >>>> > Equality deletes also block other features. CDC and row lineage
> > >> are effectively impossible while they are in use, because the true
> state of
> > >> the table can only be determined with a full scan. That same property
> means
> > >> differential structures such as materialized views and secondary
> indexes
> > >> have to be fully rebuilt whenever an equality delete is added, rather
> than
> > >> maintained incrementally. So removing equality deletes is close to a
> > >> prerequisite for the index work to stay incrementally maintainable.
> > >> >>>> >
> > >> >>>> > Evidence the alternatives are practical
> > >> >>>> >
> > >> >>>> > 1. Converting equality deletes to DVs works today. Max Michels'
> > >> ConvertEqualityDeletes maintenance task (16831, 16844, 16858, 16874,
> 16889,
> > >> 16948) rewrites equality deletes into deletion vectors as a background
> > >> Flink job: the writer keeps appending equality deletes to a staging
> branch,
> > >> and the task converts them to DVs on the target branch so reads apply
> > >> deletes by position. Notably, the task resolves each delete to a
> position
> > >> using a primary-key index that it builds and maintains inside the job,
> > >> demonstrating the full "key -> position -> DV" path end to end.
> > >> >>>> >
> > >> >>>> > 2. A persistent key-lookup index removes the need to write
> them at
> > >> all. The secondary index spec we're working on (#16961) includes a
> > >> key-lookup index mapping a key to its data file and row position.
> This is
> > >> essentially the persistent, catalog-managed form of the index Max's
> task
> > >> builds ephemerally. With it, a writer can resolve positions at write
> time
> > >> and emit DVs directly, without ever producing an equality delete.
> > >> >>>> >
> > >> >>>> > How these two efforts fit together
> > >> >>>> >
> > >> >>>> > They're complementary, and they cover the two things we need to
> > >> deprecate equality deletes:
> > >> >>>> >
> > >> >>>> > Migration (existing data): ConvertEqualityDeletes cleans up
> tables
> > >> that already contain equality deletes, and supports writers that
> still emit
> > >> them, converting them to DVs in the background.
> > >> >>>> > Going forward (new writes): the persistent key-lookup index
> lets
> > >> writers skip equality deletes entirely by looking up positions
> directly.
> > >> >>>> > The connection is that Max's task already proves the core
> > >> mechanism (resolve key -> position, write a DV); it just rebuilds a
> > >> throwaway index each cycle. A durable, shared index both enables
> write-time
> > >> elimination and removes that rebuild cost from the conversion path.
> > >> >>>> >
> > >> >>>> >
> > >> >>>> > Proposal
> > >> >>>> >
> > >> >>>> > I propose that we deprecate equality deletes in V4. The blocker
> > >> from 2024 was the lack of a viable alternative, and we now have the
> pieces:
> > >> background conversion to DVs works today, and the key-lookup index
> gives us
> > >> a path to eliminating them at write time. Deletion vectors should be
> the
> > >> going-forward mechanism for row-level deletes and upserts, produced by
> > >> background conversion now and directly by writers once the index is
> > >> available. Readers would continue to support equality deletes for
> backward
> > >> compatibility with existing V2/V3 tables.
> > >> >>>> >
> > >> >>>> > Migration path
> > >> >>>> >
> > >> >>>> > Existing tables keep working; readers continue to apply
> equality
> > >> deletes.
> > >> >>>> > ConvertEqualityDeletes (Flink) rewrites existing equality
> deletes
> > >> into DVs so tables can be cleared of them over time.
> > >> >>>> >
> > >> >>>> >
> > >> >>>> > I'd love people's thoughts, especially from those running large
> > >> streaming-upsert workloads.
> > >> >>>> >
> > >> >>>> > Thanks,
> > >> >>>> > Huaxin
> > >>
> > >
> >
>

Reply via email to