"newAppend" is the alternate to "fastAppend" for this purpose :)

On Fri, Jul 31, 2026 at 2:31 PM Hongyue Zhang <[email protected]>
wrote:

> Want to add that, not just other streaming implementations, the
> table.newAppend in iceberg-core today also default to mergeAppend. In
> my opinion, we probably need more reasons to justify why we go out of our
> way to use fast append.
>
> On Fri, Jul 31, 2026 at 11:27 AM Russell Spitzer <
> [email protected]> wrote:
>
>> I think that's reasonable. If we already default this way in other
>> engines we can make it consistent. I do think we should probably look
>> toward a longer-term default that behaves more like V4. Fast, fast, fast,
>> MERGE or something like that but that could definitely wait.
>>
>> On Fri, Jul 31, 2026 at 12:59 PM Daniel Weeks <[email protected]> wrote:
>>
>>> I agree with making the default change and would argue against needing
>>> to validate the change through some sort of benchmarking.
>>>
>>> We know this behavior results in poor metadata management and iirc, this
>>> is the same path we took with Flink many years ago.  The intent was to make
>>> commits lighter weight, but the result was poorly organized metadata that
>>> impacted readers.  Without active metadata maintenance, this can
>>> severely degrade performance and requires aggressive metadata management
>>> that is better performed inline.
>>>
>>> The other streaming implementations (Kafka Connect, Flink) use this by
>>> default, so I think we don't need to undergo exhaustive efforts to justify
>>> the change; it is viable since we have those other implementations as
>>> reference.
>>>
>>> -Dan
>>>
>>>
>>> On Thu, Jul 30, 2026 at 10:33 AM Gianluca Graziadei <
>>> [email protected]> wrote:
>>>
>>>>
>>>> +1 in principle (non-binding).
>>>>
>>>> A few things I'd like to see settled before the flip, mostly around
>>>> commit latency variance rather than the average.
>>>>
>>>> 1. Jitter, not mean latency. Merging append doesn't make every commit
>>>> uniformly more expensive; it makes most commits slightly more expensive and
>>>> a periodic subset much more expensive, whenever the merge threshold trips
>>>> and a real rewrite happens. For a micro-batch pipeline the mean is close to
>>>> irrelevant. What matters is whether a merge spike can exceed the trigger
>>>> interval, because once a batch overruns you get queueing, and the pipeline
>>>> recovers slowly rather than absorbing the cost locally. I agree with
>>>> Russel, do we have p99 and max commit latency for a representative workload
>>>> at, say, 10-30s triggers under both modes, and how far the worst spike sits
>>>> above the median?
>>>>
>>>> 2. Correlated spikes. Related concern at the fleet level: independent
>>>> pipelines that started around the same time and commit at similar intervals
>>>> will tend to trip their merges in phase, concentrating load on the catalog
>>>> and object store. Is it worth jittering the threshold slightly per writer
>>>> (see in literature how the decorrelation jitter is implemented), or
>>>> bounding the work admitted into any single commit, so the cost amortizes
>>>> instead of clustering?
>>>>
>>>>
>>>> Cheers,
>>>> Gianluca
>>>> ggraziadei
>>>>
>>>> On Thu, 30 Jul 2026, 19:11 Russell Spitzer, <[email protected]>
>>>> wrote:
>>>>
>>>>> I think this makes sense to me. Do we have latency difference numbers
>>>>> for some typical cases between the modes? Or is there a way we can trigger
>>>>> merging heuristically based on the number of small manifests? I'm 
>>>>> wondering
>>>>> if setting that as the default is right vs just switching behaviors
>>>>> automatically when it makes sense.
>>>>>
>>>>> On Thu, Jul 30, 2026 at 12:02 PM Steve <[email protected]>
>>>>> wrote:
>>>>>
>>>>>> Thanks Amogh for the changes and discussion.
>>>>>>
>>>>>> +1 to change default in 1.13 release so Spark streaming can switch to
>>>>>> use merge append by default like Flink and Kafka connect.
>>>>>>
>>>>>> Thanks,
>>>>>> Hongyue Zhang
>>>>>>
>>>>>> On Thu, Jul 30, 2026 at 8:14 AM Amogh Jahagirdar <[email protected]>
>>>>>> wrote:
>>>>>>
>>>>>>> Hey folks,
>>>>>>>
>>>>>>> For context, Spark Streaming has been using the Iceberg Java
>>>>>>> library's fast append, which produces a single new manifest pointing to 
>>>>>>> all
>>>>>>> the new files, without binpacking manifests. The current format's fast
>>>>>>> append accumulates many tiny manifests in streaming cases , and require
>>>>>>> aggressive maintenance (manifest rewriting) so that reads don't 
>>>>>>> degrade. In
>>>>>>> V4, we're improving this so that we achieve true low latency small 
>>>>>>> commits
>>>>>>> without the metadata bloat.
>>>>>>>
>>>>>>> Flink/Kafka Connect use the merging append on their path, because
>>>>>>> for the streaming ingestion workloads that the format can reasonably 
>>>>>>> handle
>>>>>>> today, it makes sense to periodically merge the manifests on write 
>>>>>>> (which
>>>>>>> isn't *that* expensive). One can also reason about this as the
>>>>>>> metadata analog for why it makes sense to merge DVs on write, rather 
>>>>>>> than
>>>>>>> have potentially many position deletes paid at read time.
>>>>>>>
>>>>>>> We recently made changes
>>>>>>> <https://github.com/apache/iceberg/pull/17403> to add a spark write
>>>>>>> conf to change this append option for streaming, but it defaults to the
>>>>>>> existing fast append behavior just to make sure in 1.12 when folks 
>>>>>>> upgrade
>>>>>>> there's no unexpected regression or if for whatever reason clients rely 
>>>>>>> on
>>>>>>> this single manifest per commit assumption, they have time to adjust 
>>>>>>> their
>>>>>>> logic.
>>>>>>>
>>>>>>> It's really not a good default option, so I propose that in 1.13 we
>>>>>>> flip this flag so that new writers start using the merging append.
>>>>>>>
>>>>>>> I would also recommend that any other client libraries that are
>>>>>>> using fast appends as a default re-evaluate their usage of it.
>>>>>>>
>>>>>>> Thanks,
>>>>>>> Amogh Jahagirdar
>>>>>>>
>>>>>>

Reply via email to