Hi Sean,

On Mon, Aug 17 2026, Sean Christopherson wrote:

> On Sat, Aug 15, 2026, Pratyush Yadav wrote:
>> On Wed, Aug 12 2026, Sean Christopherson wrote:
>> > On Wed, Aug 12, 2026, Pratyush Yadav wrote:
>> >> This is ABI between kernels. It needs to be stable-ish so you can move
>> >> from one kernel version to another. At the same time, unlike userspace
>> >> ABI, it can change.
>> >
>> > Uh, yeah, so KVM has been managing such immutable ABI for practically its 
>> > entire
>> > existence.  KVM's save/restore uAPI has exactly what you're describing: 
>> > serialization
>> > ABI that needs to be backwards and forwards compatible between different 
>> > kernels
>> > in order to support both upgrade and rollback scenarios via live migration.
>> 
>> I think there is a slight difference between KVM's save/resture uAPI and
>> live update's ABI. With KVM's uAPI, you need to maintain strict
>> backwards compatibility because userspace reads what you output. So if
>> you change the layout, userspace will interpret it wrong and might
>> break.
>> 
>> With live update, the ABI never gets to userspace. It is used to talk
>> between kernels directly. So if you do break that, your userspace keeps
>> working fine, you just might not be able to live update to the
>> incompatible kernel.
>> 
>> So with live update, we don't need to keep backwards compatibility in
>> the ABI forever. Of course, it is good to minimize changes, but we have
>> more freedom to change it.
>
> Ah.  So this is the heart of the disconnect.  I very strongly disagree with 
> the
> statement that live update doesn't need to support backwards compatibility.  I
> can totally believe that the folks working on live update are ok breaking 
> backwards
> compatibility because their use cases are "fine" with such breakage.  And I 
> can
> also believe live update as an upstream kernel feature being developed by 
> those
> same folks is also ok with breaking backwards compatibility.
>
> But with my upstream KVM maintainer hat on, I am not ok with that.  I did not 
> agree
> to support a world where KVM is allowed to break backwards compatibility, so 
> long
> as it's done "carefully" or whatever.  If y'all want to deal with the 
> resulting
> complexity, that's fine by me, but you'll be doing it without KVM.

Let's step back a bit. I don't think backwards compatibility in the
_ABI_ is all that important in live update's context. What is important
is that our users are able to live update from kernel version X to Y.
The line format the kernel uses to describe its state is an
implementation detail. Our users never see it.

The high level idea is that when you need to make a change to the ABI so
the kernel can better describe the objects/resources/files it is
passing, you create a new version of the ABI and allow transitions from
older versions.

Say you make some changes and need an ABI version v5. You don't get rid
of v4. You keep it around. LUO core will facilitate picking the right
version for the next kernel. So it will be possible to seamlessly
upgrade your kernel that speaks v4 to a new kernel that speaks v4 and
v5. Now this kernel can start speaking v5 if its successor speaks v5
too. And so on for going to v6 and v7, etc.

After a "reasonable" time given for upgrades, you deprecate v4. v4 has
been around long enough and our users have had a chance to go to kernels
speaking newer versions. That's when you remove the code for v4 from the
kernel.

We can argue what "reasonable" means, but the core idea stays. 

It will still be possible to go back to v4 or earlier, but you'd need an
extra stop along the way. And of course, vendors can keep a wider
support matrix downstream if they see the need for it.

Does this idea of "backwards compatibility" sound acceptable to you, at
least at a high level?

Now coming back to today's reality. We don't yet support this version
transition. I proposed the idea at LPC 2025 [0]. Logan Odell has taken
over the work since I have other things on my plate, but it is something
being actively developed. Logan recently sent some RFCs [1] for this
too, though TBH I haven't yet gotten to those patches. We also have
proposed a talk about this at LPC 2026, it would be great if you could
attend (if the talk gets accepted).

If you say that we should figure this version transition out and land it
upstream before we land the KVM code, I think that's a fair ask. I can
work with that.

But I think we first need to agree on the principles of the
compatibility model I described above.

[0] https://lpc.events/event/19/contributions/2049/
[1] 
https://lore.kernel.org/kexec/[email protected]/T/#u

>
> I totally understand that exploratory work and initial development is best 
> done
> in private and/or in a small working groups.  But decisions that will 
> significantly
> impact multiple subsystems need to be made *with* those subystems.  I mean,
> obviously it's possible to make a decision in a small group and then "publish"
> the result later, but then as is happening here, the community and subsystem
> maintainers like me may refuse to play ball if they disagree.

The design of live update was discussed and agreed upon in the open.
David Rientjes (and now Pasha) has been hosting the live update
biweeklies [2] from the very start of the project. The series is open to
everyone and notes shared for those who couldn't join. This meeting was
attended by a lot of people from different companies and different
subsystem backgrounds.

There have also been a microconference at LPC 2025 where topics around
live update in PCI, MM, VFIO, IOMMU, etc. were discussed, along with
presentations at subsystem tracks in other conferences.

Of course, it is impossible to get all maintainers of all subsystems to
attend these sessions or conferences. That's fine.

But I think it is quite unfair to say that live update is developed in
private or in small groups.

[2] https://lore.kernel.org/all/?q=s%3A%22%5BHypervisor+Live+Update%5D%22

-- 
Regards,
Pratyush Yadav

Reply via email to