On 01/10/26 05:26 PM, Salvatore Bonaccorso wrote:

Hi,

> Do you have  non-production machine with the same setup which you
> couls spare to bisect the changes? If you can that would be the most
> helpful step to isolate the breaking change.
> 
> Or if you can spare the machine for a couple of cycles then bisecting
> the issue would be most efficient forward in the following way:

The machine is semi-productive so we could do it, but so far we've
failed.

>     git clone --single-branch -b linux-6.12.y 
> https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux-stable.git
>     cd linux-stable
>     git checkout v6.12.107
>     cp /boot/config-$(uname -r) .config
>     yes '' | make localmodconfig
>     make savedefconfig
>     mv defconfig arch/x86/configs/my_defconfig
> 
>     # test 6.12.107 to ensure this is "good"
>     make my_defconfig
>     make -j $(nproc) bindeb-pkg
>     ... install the resulting .deb package and confirm problem does not exist
> 
>     # test 6.12.111 to ensure this is "bad"
>     git checkout v6.12.111
>     make my_defconfig
>     make -j $(nproc) bindeb-pkg
>     ... install the resulting .deb package and confirm problem exists

That 6.12.111 is running fine for the last 4,5h, where it usually failed
within a few minutes (up to 30). So right now we don't have a known bad.

Also interesting, while the compiled kernel works fine the binary sizes
are way different from the ones in the Debian packages

-rw-r--r--  1 root root   171116 Oct  2 01:00 config-6.12.107
-rw-r--r--  1 root root   283337 Aug 29 06:19 config-6.12.107+deb13-amd64
-rw-r--r--  1 root root   171310 Oct  2 01:05 config-6.12.111
-rw-r--r--  1 root root   283542 Sep 28 14:18 config-6.12.111+deb13-amd64
-rw-r--r--  1 root root   283542 Sep 28 14:18 config-6.12.111+deb13r1-amd64
-rw-r--r--  1 root root 29809382 Oct  2 01:05 initrd.img-6.12.107
-rw-r--r--  1 root root 53517010 Sep 12 12:11 initrd.img-6.12.107+deb13-amd64
-rw-r--r--  1 root root 29814131 Oct  2 01:09 initrd.img-6.12.111
-rw-r--r--  1 root root 53535601 Sep 30 03:00 initrd.img-6.12.111+deb13-amd64
-rw-r--r--  1 root root 53543460 Oct  2 08:42 initrd.img-6.12.111+deb13r1-amd64
-rw-r--r--  1 root root 11944448 Oct  2 01:00 vmlinuz-6.12.107
-rw-r--r--  1 root root 12142528 Aug 29 06:19 vmlinuz-6.12.107+deb13-amd64
-rw-r--r--  1 root root 11964928 Oct  2 01:05 vmlinuz-6.12.111
-rw-r--r--  1 root root 12154816 Sep 28 14:18 vmlinuz-6.12.111+deb13-amd64
-rw-r--r--  1 root root 12153344 Sep 28 14:18 vmlinuz-6.12.111+deb13r1-amd64

.107 and .111 are ones created by bindeb-pkg from the upstream source,
+deb13 are the ones from the official debian packages and deb13r1 is one
that was created using the manual at 
https://kernel-team.pages.debian.net/kernel-handbook/ch-common-tasks.html#s-common-official

We did do a half-hearted attempt to revert a couple of bnxt_en patches
or pull in additional patches from the 6.12-stable patch queue, but so
far it went nowhere.

However, we do have some more information.

As I said before (but forgot a digit in the kernel version) a friend saw
these errors in Proxmox upgrading from 6.17.13-21-pve to 7.0.14-12-pve.
They used the iommu=pt workaround and have not followed up.

Our own proxmox clusters were throwing the same errors after upgrading
from 7.0.14-19-pve to 7.0.14-20-pve. iommu=pt fixes it. In contrary to
the Dell R740 we have been looking at these are Lenovo AMD systems, so a
different IOMMU implementation?

This is their changelog

---
proxmox-kernel-signed-7.0 (7.0.14+20) trixie; urgency=medium

  * update submodules and patches to Ubuntu-7.0.0-39.39
    - upstream stable changes 6.18.50-6.18.51 and 7.2.2-7.2.5

  * cherry-picks for the following CVEs from stable 7.2 and mainline 7.1:
    - CVE-2026-90063: virtio-net: Ensure that TCP packets don't overflow
      gso_segs
    - CVE-2026-90105: vxlan: fix reading neigh ha
    - CVE-2026-90106: net: bridge: arp/nd proxy: fix reading neigh ha
    - CVE-2026-90232: amt: Don't support cross-netns setup
    - CVE-2026-90061: netfilter: nf_tables: skip double clone set expressions
      on element insert
    - CVE-2026-90071: net/sched: sch_teql: restore skb->dev on the slave
      failure path
    - CVE-2026-90073: net/sched: hhf: clamp quantum before hhf_change() to
      avoid overflow
    - CVE-2026-90074: net/sched: fq_pie: clamp default quantum to avoid signed
      overflow
    - CVE-2026-90075: net/sched: fq_codel: clamp default quantum and mtu
    - CVE-2026-90076: net/sched: fq: add overflow bounds to quantum and initial
      quantum
    - CVE-2026-90050: net/sched: fq: clamp quantum and initial_quantum in
      change path
    - CVE-2026-90078: net/sched: act_skbmod: fix length calculations and avoid
      invalid header warnings
    - CVE-2026-90109: net: sched: fix 32-bit backlog wrap in gred, bfifo and
      plug enqueue
    - CVE-2026-90248: net/sched: cls_api: fix teardown of an adopted proto on
      insert-race loss
    - CVE-2026-90111: ip6mr: do not clone dst in ip6mr_cache_report()
    - CVE-2026-90110: inetpeer: randomize RB-tree node comparison using SipHash
    - CVE-2026-90201: net: page_pool: fix UAF in __page_pool_release_netmem_dma
      on xa_cmpxchg race
    - CVE-2026-64190: net: team: fix NULL pointer dereference in team_xmit
      during mode change
    - CVE-2026-89789: gtp: add synchronize_net() in gtp_newlink() error path to
      prevent use-after-free
    - CVE-2026-90067: libceph: validate banner payload length
    - CVE-2026-90227: nvme/ioctl: check SUBMIT_IO with nvme_cmd_allowed()
    - CVE-2026-90148: NFSv4: Fix incorrect argument passed to
      nfs4_delete_lease() in nfs4_add_lease()
    - CVE-2026-90234: NFS: Return a delegation the client fails to record
    - CVE-2026-90400: md: recheck spare changes before starting sync
    - CVE-2026-90325: blk-cgroup: skip dying blkg in blkcg_activate_policy()
    - CVE-2026-90326: blk-cgroup: fix race between policy activation and blkg
      destruction
    - CVE-2026-90240: iommu/vt-d: Flush context cache with correct SID when
      tearing down aliases
    - CVE-2026-90435: RDMA/mlx5: Fix integer overflow of user QP buffer size
    - CVE-2026-92495: RDMA/bnxt_re: Clear VM_MAYWRITE on DBR/toggle page mmap
    - CVE-2026-90101: bnxt_en: Fix call to hardware monitoring event handler
    - CVE-2026-93182: sched/fair: Fix overflow in update_tg_cfs_runnable()
    - CVE-2026-90246: apparmor: fix integer overflow in verify_tags() bounds
      check
    - CVE-2026-93146: time/namespace: Validate nanosecond field in
      proc_timens_set_offset()
    - CVE-2026-93066: x86/mm/pat: Take cpa_lock around large-page collapse

  * cherry-pick fixes from stable 7.2 that have no CVE assigned yet:
    - af_packet: Don't cast tpacket_hdr.tp_len to int in tpacket_parse_header()
    - af_unix: Update last skb marker in manage_oob()
    - af_unix: Return immediately when manage_oob() returns NULL for 0-length
      buffer
    - ipv6: fix fib6 walker UAF on seq stop
    - ipv6: mcast: fix RCU list diversion in ip6_mc_del1_src()
    - ipv6: mcast: use copy-on-write RCU updates in ip6_mc_source()
    - ipv6: mcast: use rcu_assign_pointer() for __rcu list updates
    - ipv6: null-check fib6_node before accessing in __ip6_del_rt_siblings()
    - ipv6: flowlabel: cap duplicate leases per socket
    - ipv4: fib: bound automatic table ID allocation
    - inet: frags: invalidate queues before flushing them
    - net: iptunnel: fix stale transport header during tunnel decapsulation
    - net: cap tx_queue_len at S16_MAX to prevent oversized ring allocations
    - net/sched: defer qdisc freeing after failed creation
    - net/sched: cls_route: free emptied bucket on filter move
    - net/sched: cls_api: Don't replay RTM_GETCHAIN in tc_ctl_chain()
    - net/sched: cls_u32: fix duplicate handle when node ID pool is exhausted
    - net/sched: drr: clamp quantum in change class
    - net/sched: ets: clamp quantum in parse and fallback paths
    - net/sched: fq_pie: clamp quantum in change path
    - net/sched: hhf: clamp quantum in change and init paths
    - net/sched: sfq: clamp quantum in change path
    - vxlan: mdb: Fix use-after-free in vxlan_mdb_flush()
    - vxlan: mdb: Fix use-after-free in vxlan_mdb_remote_src_del()
    - vxlan: initialize _md in vxlan_xmit_one()
    - net: bridge: mcast: properly convert mglist to rcu
    - netfilter: nfnetlink_log: cope with concurrent instance destruction
    - bonding: alb: fix uninitialized transport header access in
      alb_determine_nd()
    - scsi: mpt3sas: Avoid out-of-bounds cpumask_of_node() call in
      _base_assign_reply_queues()
    - x86/mm/pat: Acquire init_mm write lock on collapse to avoid UAF

  * fix user-space data loss with MADV_FREE on transparent huge pages, where
    reclaim under memory pressure could discard data written after the
    MADV_FREE call if the page protection got changed in between

 -- Proxmox Support Team <[email protected]>  Thu, 24 Sep 2026 12:43:24 +0200
---

There are only two bnxt_en related commits, as far as I can see those
were also included to 6.12.108 and .110. Those would be the prime
suspect (but no, I don't have an explaination why it would fail for my
friend on 7.0.14-12 which did not have those patches).

There are also a few reports of Proxmox breakage being broken with the
recent updates with other network cards, i.e.

https://forum.proxmox.com/threads/broadcom-bcm57504-100g-bnxt_en-tx-timeout-and-nic-reset-on-proxmox-8-1-5-%E2%80%94-while-bcm57414-25g-works-fine-on-same-host.179837/#post-871482
for RTL8125 

so maybe we're completely barking up the wrong tree and the issues lies
somewhere else.

Bernhard

Reply via email to