I also agree that we don't need to fully replace equality deletes before moving forward. Demonstrating that they are not required is a reasonable standard, since v3 continues to work. We don't want to tightly couple Flink equality delete work with v4.
On Fri, Jul 24, 2026 at 10:15 AM Steven Wu <[email protected]> wrote: > I explicitly support disallowing the writing of new equality deletes in > v4. Reading existing equality deletes would still be supported in the > reference implementation as others mentioned. > > I echo Dan's comment on the benchmark result Viquar shared. The > *IcebergSourceParquetEqDeleteBenchmark* (from the Java repo) doesn't > fully capture the read performance problem of equality deletes, as it only > applies one equality file to the data file. Note that equality delete files > apply to all previously added data files (in the same partition scope or > globally). In production workloads, each data file can have hundreds, > thousands, or even more equality delete files attached to it, which makes > both scan planning and file reading very expensive. > > I agree with Russell that as long as we see a clear alternative path for > the streaming writer like Flink, we don't have to complete the replacement > implementation before v4 certification. It would certainly be great to have > though. > > On Fri, Jul 24, 2026 at 9:46 AM Ryan Blue <[email protected]> wrote: > >> > Do you mean we should allow a V3 table that contains equality delete to >> upgrade to V4? Sound like you'd suggest the V4 table spec would still allow >> the presence of equality delete. In that case, there's no way to enforce >> writer not to generate EQ deletes for V4. >> >> That's correct: I think that an upgrade can leave existing equality >> deletes in place. >> >> I don't agree that there is no way to enforce writers not generating >> equality deletes. If we don't allow equality deletes to be written, then we >> don't need to support storing them in v4 manifest files. We would keep the >> v3 delete manifests around until the delete files age off or are compacted >> into the data files. Not being able to store new equality deletes in v4 >> manifests would be effective. >> >> Going further, I think what you're suggesting is that as long as we allow >> v3 delete manifests, a client could produce delete files and write them to >> a v3 manifest file. It's true that misbehaving clients could see this as a >> hack and do it. But spec-compliant implementations would not allow this and >> can easily check for it. The same goes for catalogs, so someone doing this >> would need to control their client and their catalog. At that point, it's a >> lot of work to maintain essentially a fork of the format. I don't think >> anyone would do this, but if we find it's a problem we can always disallow >> upgrading the next version (v5) if there are v3 delete manifests present. >> >> People can do things that aren't allowed by the format, but that >> shouldn't stop us from moving forward when it is a healthy choice to drop >> support for something that has been a persistent problem. >> >> Ryan >> >> On Thu, Jul 23, 2026 at 3:38 PM Xiening Dai <[email protected]> wrote: >> >>> > I think we can continue to support equality deletes on upgrade. >>> >>> Do you mean we should allow a V3 table that contains equality delete to >>> upgrade to V4? Sound like you'd suggest the V4 table spec would still allow >>> the presence of equality delete. In that case, there's no way to enforce >>> writer not to generate EQ deletes for V4. >>> >>> Or should we require the engine to perform a synchronized EQ delete >>> re-write during the upgrade? >>> >>> On 2026/07/23 21:34:13 Ryan Blue wrote: >>> > Also, I forgot to include how to handle existing tables. I think we can >>> > continue to support equality deletes on upgrade. The main benefit of >>> > getting rid of them would be in scan planning, where we have to run >>> 2-phase >>> > planning to match delete files to data files. But we can't get rid of >>> > 2-phase planning yet because position deletes (DVs) also require >>> 2-phase >>> > planning. >>> > >>> > For v4, we should check for delete manifests and perform 2-phase >>> planning >>> > if there are any. Once all DVs are co-located with data files and >>> equality >>> > deletes are rewritten, there will be no delete manifests and we can >>> skip it. >>> > >>> > On Thu, Jul 23, 2026 at 2:22 PM Ryan Blue <[email protected]> wrote: >>> > >>> > > I strongly support this proposal. We should deprecate equality >>> deletes and >>> > > not allow writing them in v4 tables. >>> > > >>> > > The Flink work shows that it is reasonable to maintain an index so >>> the >>> > > application can create DVs rather than using equality deletes. This >>> is >>> > > independent of the on-going index work where the index can be used >>> for >>> > > other purposes. I don't think moving to DVs is blocked by the index >>> work, >>> > > although it will be great when Flink can produce and share its index. >>> > > >>> > > Now that the index-based solution is demonstrated, I think that >>> > > disallowing equality deletes in v4 is the right path forward. >>> Equality >>> > > deletes are much, much more expensive to apply than DVs. In the v3 >>> release, >>> > > we moved from multiple position delete files to a single DV and I >>> see this >>> > > decision as very similar: we thought that table maintenance would be >>> viable >>> > > to mitigate the downside of having multiple delete files, but that >>> just >>> > > wasn't the case so we should move the responsibility. Instead of >>> relying on >>> > > async maintenance that didn't materialize, we need to make this a >>> writer >>> > > responsibility. >>> > > >>> > > I agree with the points that Huaxin wrote. Existing workloads can >>> continue >>> > > to use v3 until they are moved over to produce DVs. The DV-based >>> write can >>> > > happen in v3 tables so there is a clean upgrade path for the writer >>> before >>> > > upgrading a table. >>> > > >>> > > The thread also brought up Kafka Connect, where producing DVs is a >>> bit >>> > > harder because KC doesn't shuffle data. I don't think that this is a >>> > > blocker for a few reasons: >>> > > 1. The Apache Iceberg KC bundle deliberately does not support upsert >>> > > because of the problems caused by equality deletes. We chose to limit >>> > > functionality rather than make tables unusable at read time. >>> > > 2. Anyone self-supporting their own KC build that writes equality >>> deletes >>> > > can continue to use v3 without issues >>> > > 3. There are solutions for writing DVs from KC. I think the primary >>> > > challenge is first repartitioning the data. Just because KC doesn't >>> do this >>> > > today doesn't mean we shouldn't move forward when moving this >>> > > responsibility to writers is the right trade-off. >>> > > >>> > > Thanks for all your work on this, Huaxin and Max! I think this is >>> going to >>> > > be a major step forward. >>> > > >>> > > Ryan >>> > > >>> > > On Thu, Jul 23, 2026 at 2:28 AM Junwang Zhao <[email protected]> >>> wrote: >>> > > >>> > >> Hi, >>> > >> >>> > >> On Wed, Jul 22, 2026 at 12:12 AM huaxin gao <[email protected] >>> > >>> > >> wrote: >>> > >> > >>> > >> > Thanks Xiening, that's a fair concern, and I agree the fast >>> > >> update/delete need doesn't go away. >>> > >> > >>> > >> > I'd frame it less as shifting the burden and more as changing >>> where >>> > >> it's paid. Equality deletes push the cost onto every read: each >>> reader >>> > >> re-joins the delete files against candidate rows, on every query, >>> for the >>> > >> life of the table. Resolving position once (at write time, or once >>> during >>> > >> background conversion) pays that cost a single time and makes all >>> later >>> > >> reads cheap, an O(1) DV check. And the streaming writer stays cheap >>> either >>> > >> way: background conversion (or write-time DV resolution via the >>> index) >>> > >> moves the position-resolution cost off the hot write path, so the >>> write >>> > >> stays cheap while every downstream reader gets fast, position-based >>> deletes >>> > >> instead of re-joining equality-delete files on each query. >>> > >> > >>> > >> > On index cost: I agree keeping a key -> position index current >>> isn't >>> > >> free, but the evidence so far is that it's manageable. Max's >>> > >> ConvertEqualityDeletes already maintains a persistent RocksDB PK >>> index >>> > >> incrementally at streaming scale, not rebuilt each cycle, so this is >>> > >> running in practice, not just in theory. >>> > >> >>> > >> I'm a little confused about the index part: >>> > >> 1. If we already have a key -> position index, why can't the reader >>> > >> use it to resolve equality deletes into row positions at read time, >>> > >> assuming the equality_ids match the index key? >>> > >> 2. Regarding the RocksDB-based primary-key index mentioned above, >>> how >>> > >> would we maintain such an index efficiently in cloud object storage? >>> > >> Is the idea that the index remains local to a specific engine, >>> rather >>> > >> than being stored and shared as part of the Iceberg table? >>> > >> >>> > >> > >>> > >> > On tooling and adoption: tooling is central to the proposal, not >>> an >>> > >> afterthought. Background conversion already lets writers keep >>> working >>> > >> unchanged, they write equality deletes as usual, and the conversion >>> job >>> > >> resolves them to DVs using its own internal index. The persistent >>> > >> key-lookup index is the shared tooling that makes write-time >>> elimination >>> > >> practical and engine-agnostic. Because forbidding is a V4-table >>> property, >>> > >> streaming-upsert workloads that don't yet have a write-time DV path >>> can >>> > >> keep running on V3 (writing equality deletes, with background >>> conversion >>> > >> keeping reads fast), and move to V4 once their engine can emit DVs >>> > >> directly. Non-streaming workloads can adopt V4 right away. So it >>> shouldn't >>> > >> force anyone into a worse position or block adoption. >>> > >> > >>> > >> > Finally, removing equality deletes isn't only about read cost. >>> They >>> > >> also block CDC, row lineage, and incremental maintenance of indexes >>> and >>> > >> materialized views, so it's less "shift the burden" and more >>> "unblock >>> > >> features that equality deletes currently make impossible." >>> > >> > >>> > >> > Thanks, >>> > >> > Huaxin >>> > >> > >>> > >> > On Mon, Jul 20, 2026 at 3:20 PM Xiening Dai <[email protected]> >>> wrote: >>> > >> >> >>> > >> >> Hi Huaxin, >>> > >> >> >>> > >> >> Thanks for bringing this up. Equality delete is indeed a pain >>> point we >>> > >> have seen in many customer use cases. >>> > >> >> >>> > >> >> That been said the scenario of fast update/delete still exists no >>> > >> matter which technology or table spec we choose. The proposal is >>> just going >>> > >> to shift the burden from the reader to the writer. To achieve fast >>> > >> update/delete, customer can build index structure like you >>> mentioned, but >>> > >> building such index and keeping it up to date all time can be very >>> > >> expensive too (especially given that we are tackling the fast >>> update/delete >>> > >> scenario). So although it feels like a right direction as we are >>> saying >>> > >> that we don't want to handle this complexity on the table spec, the >>> > >> underlying problem is not solved. Without a good alternative or >>> tooling >>> > >> support, i am afraid this could become an adoption issue for v4 >>> going >>> > >> forward. >>> > >> >> >>> > >> >> On 2026/07/17 17:13:55 huaxin gao wrote: >>> > >> >> > Thanks all for the discussion. I've thought this over, and I'd >>> like >>> > >> to >>> > >> >> > change my view to forbidding equality-delete writes for V4 >>> tables, >>> > >> rather >>> > >> >> > than the softer "deprecated but permitted." >>> > >> >> > >>> > >> >> > My earlier hesitation was about gating V4 on a write-time DV >>> > >> implementation >>> > >> >> > that isn't built yet. I think that concern goes away once we >>> > >> separate the >>> > >> >> > spec decision from engine adoption: >>> > >> >> > >>> > >> >> > - Forbidding equality deletes is a V4 format decision. It >>> defines >>> > >> what a V4 >>> > >> >> > table allows; it does not require every engine to have the >>> > >> write-time DV >>> > >> >> > path on day one. >>> > >> >> > - Workloads that still rely on streaming upserts can stay on V3 >>> > >> until their >>> > >> >> > engine's write-time path is ready, then adopt V4. Readers >>> continue to >>> > >> >> > support equality deletes for existing V2 and V3 tables. >>> > >> >> > - So we don't need the write-time implementation finished to >>> forbid >>> > >> >> > equality deletes in V4. What we need is a clear, credible >>> path, and >>> > >> I think >>> > >> >> > we have it: Flink's ConvertEqualityDeletes already maintains a >>> PK >>> > >> index in >>> > >> >> > Flink state and does key -> position -> DV today, and the next >>> step >>> > >> is to >>> > >> >> > persist that index into Iceberg so any engine can resolve >>> positions >>> > >> and >>> > >> >> > emit DVs directly at write time. >>> > >> >> > >>> > >> >> > This also addresses Max's concern that "deprecated but >>> permitted" is >>> > >> too >>> > >> >> > soft. Forbidding for V4 gives engines a real incentive to >>> move, while >>> > >> >> > keeping existing tables fully readable. >>> > >> >> > >>> > >> >> > On Xin's point, agreed that some perf data comparing DVs on V4 >>> > >> against >>> > >> >> > equality deletes on V3 would be useful to have as we go. >>> > >> >> > >>> > >> >> > Thanks, >>> > >> >> > Huaxin >>> > >> >> > >>> > >> >> > On Thu, Jul 16, 2026 at 8:43 AM Xin Huang via dev < >>> > >> [email protected]> >>> > >> >> > wrote: >>> > >> >> > >>> > >> >> > > Conceptually +1 as equality deletes really complicates the >>> format >>> > >> and >>> > >> >> > > implementation. >>> > >> >> > > >>> > >> >> > > However given the concern is around performance side. Is >>> there a >>> > >> way to >>> > >> >> > > make the decision making more data driven — having some >>> > >> benchmarking on >>> > >> >> > > perf comparison between dv+single file commit in v4 vs >>> equality >>> > >> delete in >>> > >> >> > > v3 could help making a call. >>> > >> >> > > >>> > >> >> > > Thanks >>> > >> >> > > Xin >>> > >> >> > > >>> > >> >> > > On Wed, Jul 15, 2026 at 11:01 PM Maximilian Michels < >>> > >> [email protected]> >>> > >> >> > > wrote: >>> > >> >> > > >>> > >> >> > >> I understood "deprecate equality deletes" as not forbidding >>> > >> engines to >>> > >> >> > >> write them, but rather discouraging them. IMHO this is long >>> > >> overdue, >>> > >> >> > >> but it is also a very soft transition. Perhaps too soft, >>> because >>> > >> it >>> > >> >> > >> doesn't give engines who write them a real incentive to stop >>> > >> writing >>> > >> >> > >> equality deletes. >>> > >> >> > >> >>> > >> >> > >> Engines will likely be quicker to move away from writing >>> equality >>> > >> >> > >> deletes if we disallow writing them in V4. Regardless, we >>> will >>> > >> have to >>> > >> >> > >> support reading equality deletes for V2 and V3 tables. >>> > >> >> > >> >>> > >> >> > >> I'm leaning more towards removing equality deletes for V4 >>> tables, >>> > >> but >>> > >> >> > >> I would like to hear what others think. >>> > >> >> > >> >>> > >> >> > >> -Max >>> > >> >> > >> >>> > >> >> > >> >>> > >> >> > >> On Wed, Jul 15, 2026 at 10:03 PM Steven Wu < >>> [email protected]> >>> > >> wrote: >>> > >> >> > >> > >>> > >> >> > >> > > So I'd separate two things: deprecating in V4 (signal >>> and >>> > >> direction, >>> > >> >> > >> safe to do now) versus forbidding equality-delete writes >>> (gated >>> > >> on the >>> > >> >> > >> engine-agnostic path being ready). I'm only proposing the >>> first >>> > >> for V4. >>> > >> >> > >> > >>> > >> >> > >> > I thought we wanted to forbid equality-delete writes for >>> v4 >>> > >> tables, >>> > >> >> > >> which would really simplify the v4 adaptive metadata tree >>> along >>> > >> with other >>> > >> >> > >> benefits that Huaxin already outlined, >>> > >> >> > >> > >>> > >> >> > >> > > deprecation path in v4 >>> > >> >> > >> > >>> > >> >> > >> > I heard the Kafka connector in the Iceberg repo doesn't >>> produce >>> > >> >> > >> equality deletes. We would need to migrate the Flink sink to >>> > >> leverage the >>> > >> >> > >> index to produce DVs only in v4. >>> > >> >> > >> > >>> > >> >> > >> > On Wed, Jul 15, 2026 at 12:13 PM huaxin gao < >>> > >> [email protected]> >>> > >> >> > >> wrote: >>> > >> >> > >> >> >>> > >> >> > >> >> Thanks Max and Manu. >>> > >> >> > >> >> >>> > >> >> > >> >> Max, thanks for the added detail. It's a good point that >>> the >>> > >> index in >>> > >> >> > >> ConvertEqualityDeletes is persisted in Flink state >>> (RocksDB) and >>> > >> updated >>> > >> >> > >> incrementally. That strengthens the case, since it shows >>> the key >>> > >> to >>> > >> >> > >> position index is already durable, just scoped to Flink >>> today. >>> > >> And your >>> > >> >> > >> closing point is exactly the plan I have in mind: deprecate >>> > >> equality >>> > >> >> > >> deletes in V4, and once the spec has an index, persist that >>> PK >>> > >> index into >>> > >> >> > >> Iceberg so it can be shared across engines. >>> > >> >> > >> >> >>> > >> >> > >> >> Manu, good question on how this works in practice. To be >>> > >> clear, I'm >>> > >> >> > >> proposing "deprecated but permitted," not removal. In V4, >>> writers >>> > >> >> > >> (including Kafka Connect) could keep emitting equality >>> deletes >>> > >> and readers >>> > >> >> > >> would keep applying them. Deprecation just declares deletion >>> > >> vectors the >>> > >> >> > >> going-forward mechanism and stops new investment in equality >>> > >> deletes. So V4 >>> > >> >> > >> would not be gated on the index landing. >>> > >> >> > >> >> >>> > >> >> > >> >> On the cleanup path today: ConvertEqualityDeletes runs >>> against >>> > >> the >>> > >> >> > >> table, not a specific writer, so a Kafka Connect pipeline >>> can >>> > >> already be >>> > >> >> > >> cleaned up. It keeps writing equality deletes, and a >>> standalone >>> > >> Flink >>> > >> >> > >> ConvertEqualityDeletes job converts them to DVs. The one >>> friction >>> > >> is that >>> > >> >> > >> the conversion runtime is Flink today, so a Kafka-only shop >>> would >>> > >> have to >>> > >> >> > >> run Flink just for maintenance. A Spark action would be a >>> natural >>> > >> follow-up >>> > >> >> > >> here, since Spark is the usual Iceberg maintenance engine >>> and >>> > >> most batch >>> > >> >> > >> shops already run it. >>> > >> >> > >> >> >>> > >> >> > >> >> Longer term, the next step is to persist the key to >>> position >>> > >> index >>> > >> >> > >> into Iceberg. Once it's shared, engines can look up >>> positions and >>> > >> write DVs >>> > >> >> > >> directly at write time, so a writer can stop producing >>> equality >>> > >> deletes >>> > >> >> > >> entirely, and neither the Flink job nor a Spark action >>> needs to >>> > >> rebuild the >>> > >> >> > >> index each run. Each engine (Kafka Connect, Spark, Flink) >>> would >>> > >> adopt >>> > >> >> > >> write-time DVs on its own schedule; until then it keeps >>> writing >>> > >> equality >>> > >> >> > >> deletes and relies on background conversion. So V4 >>> deprecation >>> > >> doesn't >>> > >> >> > >> require re-implementing every writer up front. >>> > >> >> > >> >> >>> > >> >> > >> >> So I'd separate two things: deprecating in V4 (signal and >>> > >> direction, >>> > >> >> > >> safe to do now) versus forbidding equality-delete writes >>> (gated >>> > >> on the >>> > >> >> > >> engine-agnostic path being ready). I'm only proposing the >>> first >>> > >> for V4. >>> > >> >> > >> >> >>> > >> >> > >> >> Thanks, >>> > >> >> > >> >> Huaxin >>> > >> >> > >> >> >>> > >> >> > >> >> On Wed, Jul 15, 2026 at 3:10 AM Manu Zhang < >>> > >> [email protected]> >>> > >> >> > >> wrote: >>> > >> >> > >> >>> >>> > >> >> > >> >>> Hi Huaxin, >>> > >> >> > >> >>> >>> > >> >> > >> >>> +1 for deprecating equality deletes, but how would this >>> > >> deprecation >>> > >> >> > >> work practically in V4? >>> > >> >> > >> >>> As Max pointed out, we still lack an engine-agnostic >>> solution >>> > >> for >>> > >> >> > >> streaming use cases. For example, how would we handle >>> equality >>> > >> deletes >>> > >> >> > >> written by Kafka Connect? >>> > >> >> > >> >>> While the index proposal looks promising, I don't see a >>> clear >>> > >> path >>> > >> >> > >> for deprecating equality deletes in V4 before that index >>> work >>> > >> actually >>> > >> >> > >> lands. >>> > >> >> > >> >>> >>> > >> >> > >> >>> Thanks, >>> > >> >> > >> >>> Manu >>> > >> >> > >> >>> >>> > >> >> > >> >>> >>> > >> >> > >> >>> On Wed, Jul 15, 2026 at 5:30 PM Maximilian Michels < >>> > >> [email protected]> >>> > >> >> > >> wrote: >>> > >> >> > >> >>>> >>> > >> >> > >> >>>> Hi Huaxin, >>> > >> >> > >> >>>> >>> > >> >> > >> >>>> Thanks for reviving the discussion on deprecating >>> equality >>> > >> deletes. >>> > >> >> > >> >>>> Equality deletes are the number one pain for streaming >>> use >>> > >> cases. >>> > >> >> > >> Many >>> > >> >> > >> >>>> users give up when they see the merge-on-read costs, >>> or they >>> > >> build >>> > >> >> > >> >>>> custom solutions which move them further away from core >>> > >> Iceberg. That >>> > >> >> > >> >>>> said, we've made great progress since the initial >>> > >> conversation in >>> > >> >> > >> >>>> 2024. >>> > >> >> > >> >>>> >>> > >> >> > >> >>>> Just to add what you said: The index we maintain in >>> > >> >> > >> >>>> ConvertEqualityDeletes is not ephemeral. The index is >>> > >> persisted in >>> > >> >> > >> >>>> Flink's managed state (RocksDB). It is continuously >>> updated >>> > >> as new >>> > >> >> > >> >>>> data arrives and checkpointed periodically. However, >>> even >>> > >> though the >>> > >> >> > >> >>>> conversion works for data written by any engine, we >>> > >> currently require >>> > >> >> > >> >>>> Flink for the conversion itself. Storing the index >>> directly >>> > >> in >>> > >> >> > >> Iceberg >>> > >> >> > >> >>>> and enabling all engines access would be the next >>> logical >>> > >> step >>> > >> >> > >> towards >>> > >> >> > >> >>>> a fully engine-agnostic solution. >>> > >> >> > >> >>>> >>> > >> >> > >> >>>> The reality is that we don't yet have a working >>> solution to >>> > >> avoid >>> > >> >> > >> >>>> writing equality deletes across all engines, but given >>> the >>> > >> recent >>> > >> >> > >> >>>> progress, the proposed plan seems realistic. So +1 for >>> > >> deprecating >>> > >> >> > >> >>>> equality deletes in V4. >>> > >> >> > >> >>>> >>> > >> >> > >> >>>> Cheers, >>> > >> >> > >> >>>> Max >>> > >> >> > >> >>>> >>> > >> >> > >> >>>> >>> > >> >> > >> >>>> >>> > >> >> > >> >>>> On Tue, Jul 14, 2026 at 3:25 AM huaxin gao < >>> > >> [email protected]> >>> > >> >> > >> wrote: >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Hi all, >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > I'd like to restart the conversation about >>> deprecating >>> > >> equality >>> > >> >> > >> deletes, now in the context of the V4 spec. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Background >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > This isn't a new idea. Russell proposed deprecating >>> > >> equality >>> > >> >> > >> deletes in V3 and removing them from the spec in V4, back in >>> > >> October 2024 >>> > >> >> > >> in "[DISCUSS] - Deprecate Equality Deletes". The main >>> blocker at >>> > >> the time >>> > >> >> > >> was that equality deletes served real use cases (especially >>> Flink >>> > >> streaming >>> > >> >> > >> upserts) with no efficient alternative. Two developments >>> since >>> > >> then make >>> > >> >> > >> the V4 removal worth acting on now. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Why equality deletes are costly >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Equality deletes are cheap to write but expensive to >>> read: >>> > >> a >>> > >> >> > >> reader must load the equality-delete files and join them >>> against >>> > >> every >>> > >> >> > >> candidate row in the delete's sequence-number range. >>> Positional >>> > >> deletes >>> > >> >> > >> skip that per-row join by marking exact positions, so they >>> have >>> > >> always read >>> > >> >> > >> faster, and V3 deletion vectors make them faster still, one >>> > >> compact bitmap >>> > >> >> > >> per data file, applied by an O(1) position check, instead >>> of V2's >>> > >> many >>> > >> >> > >> position-delete files. So equality deletes' only real edge >>> is the >>> > >> cheap >>> > >> >> > >> write, and both background conversion and a write-time >>> key-lookup >>> > >> index can >>> > >> >> > >> recover that. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Beyond performance >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Equality deletes also block other features. CDC and >>> row >>> > >> lineage >>> > >> >> > >> are effectively impossible while they are in use, because >>> the >>> > >> true state of >>> > >> >> > >> the table can only be determined with a full scan. That same >>> > >> property means >>> > >> >> > >> differential structures such as materialized views and >>> secondary >>> > >> indexes >>> > >> >> > >> have to be fully rebuilt whenever an equality delete is >>> added, >>> > >> rather than >>> > >> >> > >> maintained incrementally. So removing equality deletes is >>> close >>> > >> to a >>> > >> >> > >> prerequisite for the index work to stay incrementally >>> > >> maintainable. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Evidence the alternatives are practical >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > 1. Converting equality deletes to DVs works today. >>> Max >>> > >> Michels' >>> > >> >> > >> ConvertEqualityDeletes maintenance task (16831, 16844, >>> 16858, >>> > >> 16874, 16889, >>> > >> >> > >> 16948) rewrites equality deletes into deletion vectors as a >>> > >> background >>> > >> >> > >> Flink job: the writer keeps appending equality deletes to a >>> > >> staging branch, >>> > >> >> > >> and the task converts them to DVs on the target branch so >>> reads >>> > >> apply >>> > >> >> > >> deletes by position. Notably, the task resolves each delete >>> to a >>> > >> position >>> > >> >> > >> using a primary-key index that it builds and maintains >>> inside the >>> > >> job, >>> > >> >> > >> demonstrating the full "key -> position -> DV" path end to >>> end. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > 2. A persistent key-lookup index removes the need to >>> write >>> > >> them at >>> > >> >> > >> all. The secondary index spec we're working on (#16961) >>> includes a >>> > >> >> > >> key-lookup index mapping a key to its data file and row >>> position. >>> > >> This is >>> > >> >> > >> essentially the persistent, catalog-managed form of the >>> index >>> > >> Max's task >>> > >> >> > >> builds ephemerally. With it, a writer can resolve positions >>> at >>> > >> write time >>> > >> >> > >> and emit DVs directly, without ever producing an equality >>> delete. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > How these two efforts fit together >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > They're complementary, and they cover the two things >>> we >>> > >> need to >>> > >> >> > >> deprecate equality deletes: >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Migration (existing data): ConvertEqualityDeletes >>> cleans >>> > >> up tables >>> > >> >> > >> that already contain equality deletes, and supports writers >>> that >>> > >> still emit >>> > >> >> > >> them, converting them to DVs in the background. >>> > >> >> > >> >>>> > Going forward (new writes): the persistent key-lookup >>> > >> index lets >>> > >> >> > >> writers skip equality deletes entirely by looking up >>> positions >>> > >> directly. >>> > >> >> > >> >>>> > The connection is that Max's task already proves the >>> core >>> > >> >> > >> mechanism (resolve key -> position, write a DV); it just >>> rebuilds >>> > >> a >>> > >> >> > >> throwaway index each cycle. A durable, shared index both >>> enables >>> > >> write-time >>> > >> >> > >> elimination and removes that rebuild cost from the >>> conversion >>> > >> path. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Proposal >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > I propose that we deprecate equality deletes in V4. >>> The >>> > >> blocker >>> > >> >> > >> from 2024 was the lack of a viable alternative, and we now >>> have >>> > >> the pieces: >>> > >> >> > >> background conversion to DVs works today, and the key-lookup >>> > >> index gives us >>> > >> >> > >> a path to eliminating them at write time. Deletion vectors >>> should >>> > >> be the >>> > >> >> > >> going-forward mechanism for row-level deletes and upserts, >>> > >> produced by >>> > >> >> > >> background conversion now and directly by writers once the >>> index >>> > >> is >>> > >> >> > >> available. Readers would continue to support equality >>> deletes for >>> > >> backward >>> > >> >> > >> compatibility with existing V2/V3 tables. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Migration path >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Existing tables keep working; readers continue to >>> apply >>> > >> equality >>> > >> >> > >> deletes. >>> > >> >> > >> >>>> > ConvertEqualityDeletes (Flink) rewrites existing >>> equality >>> > >> deletes >>> > >> >> > >> into DVs so tables can be cleared of them over time. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > I'd love people's thoughts, especially from those >>> running >>> > >> large >>> > >> >> > >> streaming-upsert workloads. >>> > >> >> > >> >>>> > >>> > >> >> > >> >>>> > Thanks, >>> > >> >> > >> >>>> > Huaxin >>> > >> >> > >> >>> > >> >> > > >>> > >> >> > >>> > >> >>> > >> >>> > >> >>> > >> -- >>> > >> Regards >>> > >> Junwang Zhao >>> > >> >>> > > >>> > >>> >>
