It took a while to test on all the various table types, but now the PR
is ready for review: https://github.com/apache/paimon/pull/10109

On Tue, 8 Sept 2026 at 12:07, Jingsong Li <[email protected]> wrote:
>
> Hi Andreas,
>
> 3.0 Yes, we can change some options.
>
> Best,
> Jingsong
>
> On Tue, Sep 8, 2026 at 3:34 PM Andreas Bube <[email protected]> wrote:
> >
> > Hi Jingsong,
> >
> > Thanks for the feedback. That sounds like a solid approach for the
> > initial bug fix. I'll start working on this later this week.
> >
> > Looking ahead to a future major release like Paimon 3.0, could we
> > consider enabling stable UIDs by default? Speaking from experience, we
> > initially missed the operator-uid.suffix settings and had to migrate
> > state across several jobs. Enabling this by default would help prevent
> > new users from running into similar issues with misconfigured Flink
> > jobs.
> >
> > Best regards,
> > Andreas
> >
> >
> > On Tue, 8 Sept 2026 at 05:30, Jingsong Li <[email protected]> wrote:
> > >
> > > Hi Andreas,
> > >
> > > Thanks for the investigation and the reproducer. I would prefer an
> > > opt-in option, with the current behavior preserved by default.
> > >
> > > I checked the current master and confirmed the missing UIDs. I also
> > > ran minimal checks against Flink 1.20.4 and 2.2.0: restoring through
> > > an explicit snapshot path skips unmatched empty operator entries,
> > > while HA checkpoint recovery rejects them. An empty byte[] coordinator
> > > snapshot still produces a non-null state handle, so Collect Statistics
> > > is rejected by the former path as well.
> > >
> > > For the patch, I suggest:
> > > - Keep all existing explicit UIDs unchanged.
> > > - When the new option and the corresponding suffix are set, assign
> > > stable UIDs to the remaining Paimon operators, including the
> > > DataStream row conversions.
> > > - Add coverage for UID stability after upstream topology changes,
> > > recovery from a retained checkpoint store, and migration from the old
> > > UID layout, including PARTITION_DYNAMIC.
> > >
> > > The option lets existing jobs choose when to migrate, but enabling it
> > > still needs a documented migration procedure. Release notes alone
> > > would leave users exposed to an unexpected recovery failure on
> > > upgrade. For cases requiring allowNonRestoredState, we should first
> > > verify exactly which entries are being dropped and that the source,
> > > writer and committer state remains mapped; we should not recommend
> > > ignoring unmatched state unconditionally.
> > >
> > > The current FlinkJobRecoveryITCase restores through
> > > execution.state-recovery.path, which explains why its coverage does
> > > not catch the last-state failure you reported.
> > >
> > > A patch along these lines would be welcome.
> > >
> > > Best,
> > > Jingsong

Reply via email to