Hi Sean, On Mon, Aug 17 2026, Sean Christopherson wrote:
> On Sat, Aug 15, 2026, Pratyush Yadav wrote: >> On Wed, Aug 12 2026, Sean Christopherson wrote: >> > On Wed, Aug 12, 2026, Pratyush Yadav wrote: >> >> This is ABI between kernels. It needs to be stable-ish so you can move >> >> from one kernel version to another. At the same time, unlike userspace >> >> ABI, it can change. >> > >> > Uh, yeah, so KVM has been managing such immutable ABI for practically its >> > entire >> > existence. KVM's save/restore uAPI has exactly what you're describing: >> > serialization >> > ABI that needs to be backwards and forwards compatible between different >> > kernels >> > in order to support both upgrade and rollback scenarios via live migration. >> >> I think there is a slight difference between KVM's save/resture uAPI and >> live update's ABI. With KVM's uAPI, you need to maintain strict >> backwards compatibility because userspace reads what you output. So if >> you change the layout, userspace will interpret it wrong and might >> break. >> >> With live update, the ABI never gets to userspace. It is used to talk >> between kernels directly. So if you do break that, your userspace keeps >> working fine, you just might not be able to live update to the >> incompatible kernel. >> >> So with live update, we don't need to keep backwards compatibility in >> the ABI forever. Of course, it is good to minimize changes, but we have >> more freedom to change it. > > Ah. So this is the heart of the disconnect. I very strongly disagree with > the > statement that live update doesn't need to support backwards compatibility. I > can totally believe that the folks working on live update are ok breaking > backwards > compatibility because their use cases are "fine" with such breakage. And I > can > also believe live update as an upstream kernel feature being developed by > those > same folks is also ok with breaking backwards compatibility. > > But with my upstream KVM maintainer hat on, I am not ok with that. I did not > agree > to support a world where KVM is allowed to break backwards compatibility, so > long > as it's done "carefully" or whatever. If y'all want to deal with the > resulting > complexity, that's fine by me, but you'll be doing it without KVM. Let's step back a bit. I don't think backwards compatibility in the _ABI_ is all that important in live update's context. What is important is that our users are able to live update from kernel version X to Y. The line format the kernel uses to describe its state is an implementation detail. Our users never see it. The high level idea is that when you need to make a change to the ABI so the kernel can better describe the objects/resources/files it is passing, you create a new version of the ABI and allow transitions from older versions. Say you make some changes and need an ABI version v5. You don't get rid of v4. You keep it around. LUO core will facilitate picking the right version for the next kernel. So it will be possible to seamlessly upgrade your kernel that speaks v4 to a new kernel that speaks v4 and v5. Now this kernel can start speaking v5 if its successor speaks v5 too. And so on for going to v6 and v7, etc. After a "reasonable" time given for upgrades, you deprecate v4. v4 has been around long enough and our users have had a chance to go to kernels speaking newer versions. That's when you remove the code for v4 from the kernel. We can argue what "reasonable" means, but the core idea stays. It will still be possible to go back to v4 or earlier, but you'd need an extra stop along the way. And of course, vendors can keep a wider support matrix downstream if they see the need for it. Does this idea of "backwards compatibility" sound acceptable to you, at least at a high level? Now coming back to today's reality. We don't yet support this version transition. I proposed the idea at LPC 2025 [0]. Logan Odell has taken over the work since I have other things on my plate, but it is something being actively developed. Logan recently sent some RFCs [1] for this too, though TBH I haven't yet gotten to those patches. We also have proposed a talk about this at LPC 2026, it would be great if you could attend (if the talk gets accepted). If you say that we should figure this version transition out and land it upstream before we land the KVM code, I think that's a fair ask. I can work with that. But I think we first need to agree on the principles of the compatibility model I described above. [0] https://lpc.events/event/19/contributions/2049/ [1] https://lore.kernel.org/kexec/[email protected]/T/#u > > I totally understand that exploratory work and initial development is best > done > in private and/or in a small working groups. But decisions that will > significantly > impact multiple subsystems need to be made *with* those subystems. I mean, > obviously it's possible to make a decision in a small group and then "publish" > the result later, but then as is happening here, the community and subsystem > maintainers like me may refuse to play ball if they disagree. The design of live update was discussed and agreed upon in the open. David Rientjes (and now Pasha) has been hosting the live update biweeklies [2] from the very start of the project. The series is open to everyone and notes shared for those who couldn't join. This meeting was attended by a lot of people from different companies and different subsystem backgrounds. There have also been a microconference at LPC 2025 where topics around live update in PCI, MM, VFIO, IOMMU, etc. were discussed, along with presentations at subsystem tracks in other conferences. Of course, it is impossible to get all maintainers of all subsystems to attend these sessions or conferences. That's fine. But I think it is quite unfair to say that live update is developed in private or in small groups. [2] https://lore.kernel.org/all/?q=s%3A%22%5BHypervisor+Live+Update%5D%22 -- Regards, Pratyush Yadav

