From: Wei Liu <[email protected]> Sent: Tuesday, August 11, 2026 10:37 AM
> 
> On Tue, Aug 11, 2026 at 04:10:33PM +0000, Michael Kelley wrote:
> > From: [email protected] <[email protected]> Sent: Friday, July 24, 2026 
> > 4:09 PM
> > >
> > > Hyper-V vPCI protocol version 1.5 adds a RESET_DEVICE request for 
> > > projected
> > > PCI functions. Negotiate the new protocol version and issue the request
> > > through the vPCI VMBus channel from the PCI controller reset callback.
> > >
> > > Use the existing VMBus response path. Return -ENOTTY when the host reports
> > > STATUS_NOT_SUPPORTED so PCI core may try another reset method.
> > >
> > > Signed-off-by: Wei Liu <[email protected]>
> > > ---
> > >  Documentation/virt/hyperv/vpci.rst  |  2 +-
> > >  drivers/pci/controller/pci-hyperv.c | 61 +++++++++++++++++++++++++++++
> > >  2 files changed, 62 insertions(+), 1 deletion(-)
> > >
> > > diff --git a/Documentation/virt/hyperv/vpci.rst 
> > > b/Documentation/virt/hyperv/vpci.rst
> > > index b65b2126ede3..6bfee7225c14 100644
> > > --- a/Documentation/virt/hyperv/vpci.rst
> > > +++ b/Documentation/virt/hyperv/vpci.rst
> > > @@ -65,7 +65,7 @@ exchange messages with the vPCI VSP for the purpose of 
> > > setting
> > >  up and configuring the vPCI device in Linux.  Once the device
> > >  is fully configured in Linux as a PCI device, the VMBus
> > >  channel is used only if Linux changes the vCPU to be interrupted
> > > -in the guest, or if the vPCI device is removed from
> > > +in the guest, or if the vPCI device is reset or removed from
> > >  the VM while the VM is running.  The ongoing operation of the
> > >  device happens directly between the Linux device driver for
> > >  the device and the hardware, with VMBus and the VMBus channel
> > > diff --git a/drivers/pci/controller/pci-hyperv.c 
> > > b/drivers/pci/controller/pci-hyperv.c
> > > index cfc8fa403dad..d3c0fd5e1d8e 100644
> > > --- a/drivers/pci/controller/pci-hyperv.c
> > > +++ b/drivers/pci/controller/pci-hyperv.c
> > > @@ -68,6 +68,7 @@ enum pci_protocol_version_t {
> > >   PCI_PROTOCOL_VERSION_1_2 = PCI_MAKE_VERSION(1, 2),      /* RS1 */
> > >   PCI_PROTOCOL_VERSION_1_3 = PCI_MAKE_VERSION(1, 3),      /* Vibranium */
> > >   PCI_PROTOCOL_VERSION_1_4 = PCI_MAKE_VERSION(1, 4),      /* WS2022 */
> > > + PCI_PROTOCOL_VERSION_1_5 = PCI_MAKE_VERSION(1, 5),      /* GE, device 
> > > reset */
> >
> > I wish there were a better way to identify the Hyper-V version than internal
> > code names that mean nothing to the Linux community. "GE" refers to
> > Germanium, which would be the Build 26100 series, right?
> >
> 
> Yes, you're right about Germanium. I see the numbers also match, though
> I don't have a clear idea that the numbering will stay the same.
> 
> > >  };
> > >
> > >  #define CPU_AFFINITY_ALL -1ULL
> > > @@ -77,6 +78,7 @@ enum pci_protocol_version_t {
> > >   * first.
> > >   */
> > >  static enum pci_protocol_version_t pci_protocol_versions[] = {
> > > + PCI_PROTOCOL_VERSION_1_5,
> > >   PCI_PROTOCOL_VERSION_1_4,
> > >   PCI_PROTOCOL_VERSION_1_3,
> > >   PCI_PROTOCOL_VERSION_1_2,
> > > @@ -90,6 +92,7 @@ static enum pci_protocol_version_t 
> > > pci_protocol_versions[] = {
> > >  #define MAX_SUPPORTED_MSI_MESSAGES 0x400
> > >
> > >  #define STATUS_REVISION_MISMATCH 0xC0000059
> > > +#define STATUS_NOT_SUPPORTED     0xC00000BB
> > >
> > >  /* space for 32bit serial number as string */
> > >  #define SLOT_NAME_SIZE 11
> > > @@ -136,6 +139,7 @@ enum pci_message_type {
> > >   PCI_BUS_RELATIONS2              = PCI_MESSAGE_BASE + 0x19,
> > >   PCI_RESOURCES_ASSIGNED3         = PCI_MESSAGE_BASE + 0x1A,
> > >   PCI_CREATE_INTERRUPT_MESSAGE3   = PCI_MESSAGE_BASE + 0x1B,
> > > + PCI_RESET_DEVICE                = PCI_MESSAGE_BASE + 0x1C,
> > >   PCI_MESSAGE_MAXIMUM
> > >  };
> > >
> > > @@ -1397,10 +1401,66 @@ static int hv_pcifront_write_config(struct 
> > > pci_bus *bus, unsigned int devfn,
> > >   return PCIBIOS_SUCCESSFUL;
> > >  }
> > >
> > > +static int hv_pcifront_reset(struct pci_dev *pdev, bool probe)
> > > +{
> > > + struct hv_pcibus_device *hbus =
> > > +         container_of(pdev->bus->sysdata, struct hv_pcibus_device, 
> > > sysdata);
> > > + struct pci_child_message reset = {};
> > > + struct hv_pci_compl comp_pkt;
> > > + struct pci_packet pkt = {
> > > +         .completion_func = hv_pci_generic_compl,
> > > +         .compl_ctxt = &comp_pkt,
> > > + };
> > > + enum hv_pcibus_state state;
> > > + int ret;
> > > +
> > > + /* Device reset was added in vPCI protocol version 1.5. */
> > > + if (hbus->protocol_version < PCI_PROTOCOL_VERSION_1_5)
> > > +         return -ENOTTY;
> > > +
> > > + /* Hyper-V exposes projected functions directly on the root bus. */
> > > + if (!pci_is_root_bus(pdev->bus))
> > > +         return -ENOTTY;
> > > +
> > > + if (probe)
> > > +         return 0;
> > > +
> > > + /* Do not take state_lock: eject holds it while removing/locking pdev. 
> > > */
> > > + state = READ_ONCE(hbus->state);
> > > + if (state != hv_pcibus_probed && state != hv_pcibus_installed)
> > > +         return -ENODEV;
> > > +
> > > + init_completion(&comp_pkt.host_event);
> > > + reset.message_type.type = PCI_RESET_DEVICE;
> > > + reset.wslot.slot = devfn_to_wslot(pdev->devfn);
> > > +
> > > + ret = vmbus_sendpacket(hbus->hdev->channel, &reset, sizeof(reset),
> > > +                        (unsigned long)&pkt, VM_PKT_DATA_INBAND,
> > > +                        VMBUS_DATA_PACKET_FLAG_COMPLETION_REQUESTED);
> > > + if (ret)
> > > +         return ret;
> > > +
> > > + ret = wait_for_response(hbus->hdev, &comp_pkt.host_event);
> > > + if (ret)
> > > +         return ret;
> > > +
> > > + if (comp_pkt.completion_status == STATUS_NOT_SUPPORTED)
> > > +         return -ENOTTY;
> >
> > I tried this patch series in a linux-next20260726 build, and running on a 
> > D16lds v6
> > VM in Azure. The host hypervisor version is 10.0.26100.1652-1-0, and the VM
> > has a paravisor with HvLite.
> >
> > This VM has an NVMe OS disk, two NVMe temp disks, and a MANA network 
> > controller.
> > Absent this patch set, the temp disks and MANA report the "reset_method" as 
> > "flr",
> > while the NVMe OS disk reports no reset methods. With this patch set, 
> > "controller"
> > is added as a reset method for all. The NVMe disks and MANA network 
> > controller
> > are probed with PCI protocol 1.5. That's all good and as expected (though 
> > I'm not
> > sure why the NVMe OS disk doesn't support flr).
> >
> > I then did "echo 1 >reset" for the NVMe OS disk. This returns a -ENOTTY 
> > error
> > from the above line of code. So the host hypervisor (or paravisor?) is 
> > saying that
> > the new PCI_RESET_DEVICE message isn't supported. Presumably this is new
> > functionality that hasn’t been rolled out to where I'm running this Azure 
> > VM.
> > That's fine too.
> >
> 
> That's right. It is not yet rolled out.
> 
> > But interestingly, the NVMe OS disk did a Linux-side reset anyway. That's
> > because of this stack trace from the "echo 1 >reset" command:
> >
> > [   99.642483]  nvme_try_sched_reset+0x25/0x60 [nvme_core]
> > [   99.642495]  nvme_reset_done+0x1c/0x40 [nvme]
> > [   99.642499]  pci_dev_restore+0x35/0x70
> > [   99.642503]  pci_reset_function+0x100/0x140
> > [   99.642506]  reset_store+0x5a/0xa0
> > [   99.642508]  dev_attr_store+0x16/0x30
> > [   99.642512]  sysfs_kf_write+0x71/0x80
> > [   99.642516]  kernfs_fop_write_iter+0x140/0x1d0
> > [   99.642518]  vfs_write+0x313/0x420
> > [   99.642522]  ksys_write+0x68/0xe0
> > [   99.642525]  __x64_sys_write+0x18/0x20
> > [   99.642527]  x64_sys_call+0x1700/0x21c0
> > [   99.642530]  do_syscall_64+0x8d/0x440
> > [   99.642532]  entry_SYSCALL_64_after_hwframe+0x76/0x7e
> >
> > Even though __pci_reset_function_locked() failed, the subsequent
> > code in pci_reset_function() calls nvme_try_sched_reset(), which puts
> > nvme_reset_work() on a workqueue.  nvme_reset_work() tears
> > things down on the Linux side and rebuilds, and the message
> >
> >     nvme nvme0: 16/0/0 default/read/poll queues
> >
> > is output.
> >
> > It's not immediately clear to me how to resolve this issue, so
> > I'm just pointing it out. :-(
> >
> 
> Is resetting the OS disk a real use case?

Perhaps not. However, the same behavior occurs with the NVMe
temp disks if "flr" is removed from "reset_modes" so that the
"controller" method is used. I was just trying to see how this new
controller method behaves, and the result seemed anomalous.
Using the "controller" method against the MANA device (again
presumably with a "not supported" reply from Hyper-V) produced
a hung MANA that I couldn't recover without rebooting the VM.
I did not try to sort out what when wrong there.

> 
> I'm not sure what magic is happening behind this.
> 

FWIW, I don't think anything is happening on the Hyper-V host side.
Hyper-V says PCI_RESET_DEVICE isn't supported and then just goes
on normally. But the Linux path through pci_reset_function() and
the NVMe driver does its reset handling even if none of the available
reset methods succeed. Arguably that's independent of your patch
set, though the existence of Hyper-V versions that support vPCI
protocol 1.5 but not PCI_RESET_DEVICE could make the problem
more likely.

Michael

Reply via email to