On 10/5/26 5:40 AM, David Vrabel wrote:


On 05/10/2026 05:18, Laine Stump via Devel wrote:

       * 3) If at all possible, we want to avoid needing to change the busNr        *    of a bus in the future, as that changes the guest's device ABI,        *    which could potentially lead to issues with a guest OS that is
       *    picky about such things.

(Side note: I was the author of the above comment, made 10 years ago, and I was mainly concerned about runtime changes in ABI, for example during a migration from one host to another where the guest OS doesn't reboot. That would never happen in this case, since the busNr is stored with other attributes of the controller device whether it was set explicitly by the user or auto-set to a default value by libvirt when the domain was defined (or when the controller was added, if that happened later). A change in the default busNr would only have an effect on a newly defined domain (or newly added pcie-expander-bus).

This change violates this constraint.

When a new domain is created and no busNr is given, the default value decided on by libvirt at that time is written into the XML, so subsequent starts of that particular domain using that definition will use the same busNr - only newly created domains will get a different default busNr than they used to. So in that case, this change in default setting has 0 effect.

As for migration (and domain save/restore) the change here will never have any effect on runtime ABI at all, since whatever the setting of busNr (whether it was the default value set by libvirt, or explicitly set by the user in the XML config, when the domain was defined) that will be written in the XML that is sent along with the migration data and so the destination end of the migration won't ever be using just "the current default busNr setting"; it will always be using "the exact same busNr that was used when the guest was started".

The place where this change *would* have any effect is when the management software doesn't re-use the defined domain for subsequent restarts, but instead just redefines the domain from its original internal format each time it is started. However, note that even in that case the change in busNr will only happen when the guest is completely stopped, shutdown, and booted from scratch (new QEMU process and all). So there will be no case of a running OS having the hardware "change out from under it", there will only be cases where the OS is shut down, and then when it is booted the next time the busNr will be different. Also note that a similar "change in ABI" can happen if, for example, a new device is added to the "raw" XML config of the management software (which is missing PCI addresses), and then that XML is fed back into libvirt - since libvirt internally determines the ordering of the devices (and thus which device is assigned to which PCI address) you can easily end up with, e.g. a network card suddenly appearing at a different PCI address after adding a disk device to the config. If that (changing the PCI address of an endpoint device between OS boots) is acceptable behavior, then I'd say changing the bus number of a PCI controller between guest OS boots is also acceptable.

(another example of guest ABI quietly changing between reboots is when a domain is defined using a machinetype of simply "q35". That generic name will be canonicalized into, e.g. pc-q35-10.1 and the guest OS will be booted with that exact machinetype and any migration will convey that exact machinetype to the destination of the migration. But when the guest is stopped, if the management software re-defines the domain from scratch (rather than continuing to use the same domain definition) still using "q35", then the next time it is booted it might come up with a machinetype of pc-q35-11.2 (if the host's QEMU package has been updated).)



As a non-backward compatible change to the guest ABI I don't think you can change the default bus allocation method.

While I agree that the busNr shouldn't be changed during the lifetime of a guest OS (e.g. during migration), this change in default at domain definition time doesn't lead to a busNr changing during migration. If you are aware of any guest OS that cares whether such a change occurs between reboots, then that's very interesting and is something we need to consider.

You may find it more useful for your management plane to evenly distribute bus numbers for the (presumably) per-vNUMA node pcie-expander buses and explicitly specify the busNr for each.

If someone wants to explicitly specify the busNr they're free to do that (as suggested in the comments) just as they are free to manually specify the PCI address of each device. This is just setting a default value at domain definition time that is more likely to be usable in the general case than the original default; the current default has limited value, since it doesn't work if you have more than a single hotpluggable endpoint device on one pcie-expander-bus anyway. In any case libvirt has always saved the auto-assigned default value of busNr (and other obscure things normally auto-set by libvirt, like chassis number) in the domain config so that they would be maintained during migration, or at a subsequent start of the domain using the same domain definition.

(Again, if you're aware of an example of busNr changing from one boot to the next of a guest OS, then that would require more caution. Or maybe I'm being too laissez faire about it, but my current opinion is that I'm not :-)).

Reply via email to