Public bug reported:

[Impact]

* In a customer's environment which uses a system configured with an
Intel Corporation Ethernet Controller E810-XXV for SFP (rev 02) NIC
attached to an arm64 CPU running the 7.0.0-15-generic kernel on Resolute
we have observed kernel crashes caused by the ICE driver triggering a
DQL Bug

* When this happens, the machine is only recoverable with a reboot

* The issue occurs during their mirroring and high availability database
continuous testing involving hundreds of processes and communication(s)
across a private network.

* The customer has an otherwise identical working environment on Jammy
with the same hardware configuration and the issue was first discovered
during their validation of the workload on Resolute on arm64

* We have not reproduced it internally, but it is reliably hit in their
environment. The machine crashes and produces the same trace each time
and crashes as frequently as a few times a day and as infrequently as
once every two days

* Below is the backtrace of the crash.

kernel BUG at lib/dynamic_queue_limits.c:99
Oops#1 Part2
<2>[781539.214963] kernel BUG at lib/dynamic_queue_limits.c:99!
<0>[781539.215013] Internal error: Oops - BUG: 00000000f2000800 [#1] SMP
<4>[781539.231380] Modules linked in: nf_tables nfsv3 nfs_acl nfs lockd grace 
netfs qrtr cfg80211 nls_iso8859_1 xfs irdma idpf ib_uverbs joydev input_leds 
cppc_cpufreq ast arm_cmn ib_core dax_hmem cxl_acpi cxl_port cxl_pmem acpi_ipmi 
cxl_core ipmi_ssif fwctl ipmi_devintf einj ipmi_msghandler ampere_cspmu 
arm_cspmu_module arm_spe_pmu acpi_tad sch_fq_codel binfmt_misc auth_rpcgss 
efi_pstore dm_multipath sunrpc nfnetlink hid_generic usbhid hid qla2xxx 
onboard_usb_dev nvme_fc nvme ice nvme_fabrics libeth_xdp ghash_ce sm4_ce_gcm 
sm4_ce_ccm sm4_ce nvme_core gnss nvme_keyring sm4_ce_cipher i40e libeth 
xgene_hwmon sm4 nvme_auth libie scsi_transport_fc xhci_pci_renesas igb sm3_ce 
libie_adminq libie_fwlog ahci hkdf i2c_algo_bit dm_mirror dm_region_hash dm_log 
dmi_sysfs autofs4 aes_neon_bs aes_neon_blk aes_ce_blk
<4>[781539.302276] CPU: 53 UID: 0 PID: 0 Comm: swapper/53 Not tainted 
7.0.0-15-generic #15-Ubuntu PREEMPT(lazy)
<4>[781539.311953] Hardware name: Giga Computing 
R183-P92-AAJ1-000/MP93-FS0-000, BIOS F12 09/30/2025
<4>[781539.320572] pstate: 83400009 (Nzcv daif +PAN -UAO +TCO +DIT -SSBS 
BTYPE=--)
<4>[781539.327629] pc : dql_completed+0x268/0x2a0
<4>[781539.331823] lr : ice_clean_tx_irq+0x1d4/0x620 [ice]
<4>[781539.337214] sp : ffff8000801abc80
<4>[781539.340616] x29: ffff8000801abc80 x28: 00000000ffffff7b x27: 
ffff8000cf5e57b0
<4>[781539.347853] x26: 0000000000000000 x25: ffffb96bcb76e000 x24: 
0000000000000001
<4>[781539.355088] x23: 0000000000000042 x22: ffff20015f2fd3b8 x21: 
0000000000000001
Oops#1 Part1
<4>[781539.362323] x20: ffff20011d821e00 x19: ffff200148d44240 x18: 
ffff8000801ad038
<4>[781539.369559] x17: 0000000000000000 x16: ffffb96bc85c9980 x15: 
0000000000000000
<4>[781539.376793] x14: 0000000000000000 x13: 0000000000000000 x12: 
0000000000000000
<4>[781539.384028] x11: 0000000000000000 x10: 0000000077afe1d4 x9 : 
ffffb96b79286bf8
<4>[781539.391263] x8 : 0000000000000000 x7 : ffff200148d442c0 x6 : 
0000000000000000
<4>[781539.398500] x5 : 0000000000000000 x4 : 0000000000000000 x3 : 
0000000000000000
<4>[781539.405734] x2 : 0000000077afe1d4 x1 : 0000000000000042 x0 : 
0000000000000000
<4>[781539.412970] Call trace:
<4>[781539.415503] dql_completed+0x268/0x2a0 (P)
<4>[781539.419691] ice_napi_poll+0x94/0x520 [ice]
<4>[781539.424081] __napi_poll+0x48/0x3f0
<4>[781539.427662] net_rx_action+0x194/0x420
<4>[781539.431499] handle_softirqs+0x128/0x538
<4>[781539.435519] __do_softirq+0x20/0x3c
<4>[781539.439098] ____do_softirq+0x1c/0x40
<4>[781539.442858] call_on_irq_stack+0x48/0x68
<4>[781539.446873] do_softirq_own_stack+0x28/0x60
<4>[781539.451148] __irq_exit_rcu+0x184/0x1c8
<4>[781539.455073] irq_exit_rcu+0x1c/0x40
<4>[781539.458651] el1_interrupt+0x50/0xd8
<4>[781539.462324] el1h_64_irq_handler+0x1c/0x40
<4>[781539.466512] el1h_64_irq+0x84/0x88
<4>[781539.470001] tick_nohz_idle_exit+0x60/0x1d0 (P)
<4>[781539.474630] do_idle+0xe8/0x150
<4>[781539.477861] cpu_startup_entry+0x44/0x50
<4>[781539.481873] secondary_start_kernel+0xe8/0x128
<4>[781539.486413] __secondary_switched+0xc8/0xd0
<0>[781539.491146] Code: 7a401800 5400008b 2a0b03e1 17ffff84 (d4210000)
<4>[781539.497797] ---[ end trace 0000000000000000 ]---

* Affected supported releases: Resolute and Stonking
* Affected kernel versions: confirmed 7.0.0+, but presumably 6.18+ as that's 
when [1] was merged
* Affected environments: ice NICs on arm64

[Investigation]

Kernel version: 7.0.0-15-generic

See below the dql_completed code that's throwing the bug, which
indicates that the machine is crashing because the driver believes the
number of bytes the NIC has sent is greater than the amount of bytes
that were queued for work. As this ordinarily shouldn't be possible, the
fact that this occurs indicate that something went wrong, potentially in
the accounting of packets or a race that causes values to be updated
inconsistently.

void dql_completed(struct dql *dql, unsigned int count)
{
...
 num_queued = READ_ONCE(dql->num_queued);
...
 /* Can't complete more than what's in queue */
 BUG_ON(count > num_queued - dql->num_completed);
 ...
}

Specifically the crash is invoked from ice_clean_tx_irq's inlined call
to netdev_tx_completed_queue which calls dql_completed

eu-addr2line -ifae 
/usr/lib/debug/lib/modules/7.0.0-15-generic/kernel/drivers/net/ethernet/intel/ice/ice.ko.zst
 ice_clean_tx_irq+0x1d4
0x0000000000059fbc
netdev_tx_completed_queue inlined at 
/build/linux-IpI1TW/linux-7.0.0/drivers/net/ethernet/intel/ice/ice_txrx.c:370:2 
in ice_clean_tx_irq
/build/linux-IpI1TW/linux-7.0.0/include/linux/netdevice.h:3907:2 
ice_clean_tx_irq
/build/linux-IpI1TW/linux-7.0.0/drivers/net/ethernet/intel/ice/ice_txrx.c:370:2

static bool ice_clean_tx_irq(struct ice_tx_ring *tx_ring, int napi_budget)
{
 // Iterate through end of packet descriptors and unmap memory associated with 
processed packets
 ...
 ice_update_tx_ring_stats(tx_ring, total_pkts, total_bytes);
 netdev_tx_completed_queue(txring_txq(tx_ring), total_pkts, total_bytes);
}

Note that from the perspective of dql_completed: count =
txring_txq(tx_ring), whose stats are accumulated in the driver's hot
path which has been stable for several years. Note that the below line
references correspond to the 7.0.0-15-generic kernel present in Resolute
as this matches the kernel on which the crash was observed.

As for the other variables in the BUG_ON, dql->num_queued and 
dql->num_completed members are updated in two paths
ice_txrx.c: ice_start_xmit #L2288 -> ice_txrx.c: ice_xmit_frame_ring #L2258 -> 
ice_txrx.c: ice_tx_map #L1518 -> netdevice.h: __netdev_tx_sent_queue #L3853 -> 
dynamic_queue_limits.h: dql_queued which updates num_queued on L139
ice_txrx.c: ice_clean_tx_irq #L370 -> netdevice.h: netdev_tx_completed_queue 
#L3900 -> dynamic_queue_limits.c: dql_completed which updates num_completed on 
L183

The observed crash is in the second path before we get to the point
where we would have updated num_completed (we're crashing on L99 in
dql_completed). This makes sense because we're still currently in the
process of determining the packets that were transmitted (completed), so
num_completed should reflect the value of the previous cycle at this
point. It's less likely that this path is the problem because it would
mean that dql->num_completed was not updated correctly at the end of the
previous accumulation and if this were the case it would probably mask
the issue since num_completed would be a smaller value and, therefore,
less likely to precipitate this issue. This is because num_queued -
dql->num_completed would be larger than it should be and so count would
less likely be greater than that

If we look at the other path, there have been recent changes.
Particularly in the ice_tx_map function we decide whether we need to
notify the hardware via kick = __netdev_tx_sent_queue(...) and then we
proceed to use writel_relaxed() to inform the NIC about the index up to
which we have produced timestamp descriptors. Using writel_relaxed()
does not have the same ordering guarantees as writel() across
architectures with regards to visibility of memory accesses across CPUs
and peripherals. On arm writel_relaxed() allows for reordering of writes
to optimize performance and this potentially introduces a race between
__netdev_tx_sent_queue, which updates the accumulator (dql->num_queued),
and the write to the NIC's MMIO to kick off work. In other words, the
NIC may start transmitting descriptors before the updated
dql->num_queued value is globally visible and so the CPU crashes on the
NIC's return because it thinks the NIC has sent packets that haven't
been queued.

kick = __netdev_tx_sent_queue(txring_txq(tx_ring), first->bytecount,
          netdev_xmit_more());
if (!kick)
 return;

if (ice_is_txtime_cfg(tx_ring)) {
...
 tstamp_ring->next_to_use = j;
 writel_relaxed(j, tstamp_ring->tail);
} else {
 writel_relaxed(i, tx_ring->tail);
}

This logic is a recent addition from the commit at [1]. Previously the
logic was much simpler (see below) and notably used writel() which
guarantees all previously memory writes are visible across CPUs and
peripherals.

if (kick)
    /* notify HW of packet */
    writel(i, tx_ring->tail);

* I have provided the customer with a test kernel that replaces both
instances of writel_relaxed() with writel() and the issue did not recur
with new kernel after 2 weeks of testing where both NICs were under
heavy and continuous load for multiple days without issue

[1]
https://github.com/torvalds/linux/commit/ccde82e909467abdf098a8ee6f63e1ecf9a47ce5

** Affects: linux (Ubuntu)
     Importance: Undecided
     Assignee: Bryan Fraschetti (bryanfraschetti)
         Status: New

** Affects: linux (Ubuntu Resolute)
     Importance: Undecided
         Status: New

** Affects: linux (Ubuntu Stonking)
     Importance: Undecided
     Assignee: Bryan Fraschetti (bryanfraschetti)
         Status: New

** Also affects: linux (Ubuntu Resolute)
   Importance: Undecided
       Status: New

** Also affects: linux (Ubuntu Stonking)
   Importance: Undecided
     Assignee: Bryan Fraschetti (bryanfraschetti)
       Status: New

** Description changed:

  [Impact]
  
  * In a customer's environment which uses a system configured with an
  Intel Corporation Ethernet Controller E810-XXV for SFP (rev 02) NIC
  attached to an arm64 CPU running the 7.0.0-15-generic kernel on Resolute
  we have observed kernel crashes caused by the ICE driver triggering a
  DQL Bug
  
  * When this happens, the machine is only recoverable with a reboot
  
  * The issue occurs during their mirroring and high availability database
  continuous testing involving hundreds of processes and communication(s)
  across a private network.
  
  * The customer has an otherwise identical working environment on Jammy
  with the same hardware configuration and the issue was first discovered
  during their validation of the workload on Resolute on arm64
  
  * We have not reproduced it internally, but it is reliably hit in their
  environment. The machine crashes and produces the same trace each time
  and crashes as frequently as a few times a day and as infrequently as
  once every two days
  
  * Below is the backtrace of the crash.
  
  kernel BUG at lib/dynamic_queue_limits.c:99
  Oops#1 Part2
  <2>[781539.214963] kernel BUG at lib/dynamic_queue_limits.c:99!
  <0>[781539.215013] Internal error: Oops - BUG: 00000000f2000800 [#1] SMP
  <4>[781539.231380] Modules linked in: nf_tables nfsv3 nfs_acl nfs lockd grace 
netfs qrtr cfg80211 nls_iso8859_1 xfs irdma idpf ib_uverbs joydev input_leds 
cppc_cpufreq ast arm_cmn ib_core dax_hmem cxl_acpi cxl_port cxl_pmem acpi_ipmi 
cxl_core ipmi_ssif fwctl ipmi_devintf einj ipmi_msghandler ampere_cspmu 
arm_cspmu_module arm_spe_pmu acpi_tad sch_fq_codel binfmt_misc auth_rpcgss 
efi_pstore dm_multipath sunrpc nfnetlink hid_generic usbhid hid qla2xxx 
onboard_usb_dev nvme_fc nvme ice nvme_fabrics libeth_xdp ghash_ce sm4_ce_gcm 
sm4_ce_ccm sm4_ce nvme_core gnss nvme_keyring sm4_ce_cipher i40e libeth 
xgene_hwmon sm4 nvme_auth libie scsi_transport_fc xhci_pci_renesas igb sm3_ce 
libie_adminq libie_fwlog ahci hkdf i2c_algo_bit dm_mirror dm_region_hash dm_log 
dmi_sysfs autofs4 aes_neon_bs aes_neon_blk aes_ce_blk
  <4>[781539.302276] CPU: 53 UID: 0 PID: 0 Comm: swapper/53 Not tainted 
7.0.0-15-generic #15-Ubuntu PREEMPT(lazy)
  <4>[781539.311953] Hardware name: Giga Computing 
R183-P92-AAJ1-000/MP93-FS0-000, BIOS F12 09/30/2025
  <4>[781539.320572] pstate: 83400009 (Nzcv daif +PAN -UAO +TCO +DIT -SSBS 
BTYPE=--)
  <4>[781539.327629] pc : dql_completed+0x268/0x2a0
  <4>[781539.331823] lr : ice_clean_tx_irq+0x1d4/0x620 [ice]
  <4>[781539.337214] sp : ffff8000801abc80
  <4>[781539.340616] x29: ffff8000801abc80 x28: 00000000ffffff7b x27: 
ffff8000cf5e57b0
  <4>[781539.347853] x26: 0000000000000000 x25: ffffb96bcb76e000 x24: 
0000000000000001
  <4>[781539.355088] x23: 0000000000000042 x22: ffff20015f2fd3b8 x21: 
0000000000000001
  Oops#1 Part1
  <4>[781539.362323] x20: ffff20011d821e00 x19: ffff200148d44240 x18: 
ffff8000801ad038
  <4>[781539.369559] x17: 0000000000000000 x16: ffffb96bc85c9980 x15: 
0000000000000000
  <4>[781539.376793] x14: 0000000000000000 x13: 0000000000000000 x12: 
0000000000000000
  <4>[781539.384028] x11: 0000000000000000 x10: 0000000077afe1d4 x9 : 
ffffb96b79286bf8
  <4>[781539.391263] x8 : 0000000000000000 x7 : ffff200148d442c0 x6 : 
0000000000000000
  <4>[781539.398500] x5 : 0000000000000000 x4 : 0000000000000000 x3 : 
0000000000000000
  <4>[781539.405734] x2 : 0000000077afe1d4 x1 : 0000000000000042 x0 : 
0000000000000000
  <4>[781539.412970] Call trace:
  <4>[781539.415503] dql_completed+0x268/0x2a0 (P)
  <4>[781539.419691] ice_napi_poll+0x94/0x520 [ice]
  <4>[781539.424081] __napi_poll+0x48/0x3f0
  <4>[781539.427662] net_rx_action+0x194/0x420
  <4>[781539.431499] handle_softirqs+0x128/0x538
  <4>[781539.435519] __do_softirq+0x20/0x3c
  <4>[781539.439098] ____do_softirq+0x1c/0x40
  <4>[781539.442858] call_on_irq_stack+0x48/0x68
  <4>[781539.446873] do_softirq_own_stack+0x28/0x60
  <4>[781539.451148] __irq_exit_rcu+0x184/0x1c8
  <4>[781539.455073] irq_exit_rcu+0x1c/0x40
  <4>[781539.458651] el1_interrupt+0x50/0xd8
  <4>[781539.462324] el1h_64_irq_handler+0x1c/0x40
  <4>[781539.466512] el1h_64_irq+0x84/0x88
  <4>[781539.470001] tick_nohz_idle_exit+0x60/0x1d0 (P)
  <4>[781539.474630] do_idle+0xe8/0x150
  <4>[781539.477861] cpu_startup_entry+0x44/0x50
  <4>[781539.481873] secondary_start_kernel+0xe8/0x128
  <4>[781539.486413] __secondary_switched+0xc8/0xd0
  <0>[781539.491146] Code: 7a401800 5400008b 2a0b03e1 17ffff84 (d4210000)
  <4>[781539.497797] ---[ end trace 0000000000000000 ]---
  
+ * Affected supported releases: Resolute and Stonking
+ * Affected kernel versions: confirmed 7.0.0+, but presumably 6.18+ as that's 
when [1] was merged
+ * Affected environments: ice NICs on arm64
+ 
  [Investigation]
  
  Kernel version: 7.0.0-15-generic
  
  See below the dql_completed code that's throwing the bug, which
  indicates that the machine is crashing because the driver believes the
  number of bytes the NIC has sent is greater than the amount of bytes
  that were queued for work. As this ordinarily shouldn't be possible, the
  fact that this occurs indicate that something went wrong, potentially in
  the accounting of packets or a race that causes values to be updated
  inconsistently.
  
  void dql_completed(struct dql *dql, unsigned int count)
  {
  ...
-       num_queued = READ_ONCE(dql->num_queued);
+  num_queued = READ_ONCE(dql->num_queued);
  ...
-       /* Can't complete more than what's in queue */
-       BUG_ON(count > num_queued - dql->num_completed);
-       ...
+  /* Can't complete more than what's in queue */
+  BUG_ON(count > num_queued - dql->num_completed);
+  ...
  }
  
  Specifically the crash is invoked from ice_clean_tx_irq's inlined call
  to netdev_tx_completed_queue which calls dql_completed
  
  eu-addr2line -ifae 
/usr/lib/debug/lib/modules/7.0.0-15-generic/kernel/drivers/net/ethernet/intel/ice/ice.ko.zst
 ice_clean_tx_irq+0x1d4
  0x0000000000059fbc
  netdev_tx_completed_queue inlined at 
/build/linux-IpI1TW/linux-7.0.0/drivers/net/ethernet/intel/ice/ice_txrx.c:370:2 
in ice_clean_tx_irq
  /build/linux-IpI1TW/linux-7.0.0/include/linux/netdevice.h:3907:2 
ice_clean_tx_irq
  
/build/linux-IpI1TW/linux-7.0.0/drivers/net/ethernet/intel/ice/ice_txrx.c:370:2
  
  static bool ice_clean_tx_irq(struct ice_tx_ring *tx_ring, int napi_budget)
  {
-       // Iterate through end of packet descriptors and unmap memory 
associated with processed packets
-       ...
-       ice_update_tx_ring_stats(tx_ring, total_pkts, total_bytes);
-       netdev_tx_completed_queue(txring_txq(tx_ring), total_pkts, total_bytes);
+  // Iterate through end of packet descriptors and unmap memory associated 
with processed packets
+  ...
+  ice_update_tx_ring_stats(tx_ring, total_pkts, total_bytes);
+  netdev_tx_completed_queue(txring_txq(tx_ring), total_pkts, total_bytes);
  }
  
  Note that from the perspective of dql_completed: count =
  txring_txq(tx_ring), whose stats are accumulated in the driver's hot
  path which has been stable for several years. Note that the below line
  references correspond to the 7.0.0-15-generic kernel present in Resolute
  as this matches the kernel on which the crash was observed.
  
  As for the other variables in the BUG_ON, dql->num_queued and 
dql->num_completed members are updated in two paths
  ice_txrx.c: ice_start_xmit #L2288 -> ice_txrx.c: ice_xmit_frame_ring #L2258 
-> ice_txrx.c: ice_tx_map #L1518 -> netdevice.h: __netdev_tx_sent_queue #L3853 
-> dynamic_queue_limits.h: dql_queued which updates num_queued on L139
  ice_txrx.c: ice_clean_tx_irq #L370 -> netdevice.h: netdev_tx_completed_queue 
#L3900 -> dynamic_queue_limits.c: dql_completed which updates num_completed on 
L183
  
  The observed crash is in the second path before we get to the point
  where we would have updated num_completed (we're crashing on L99 in
  dql_completed). This makes sense because we're still currently in the
  process of determining the packets that were transmitted (completed), so
  num_completed should reflect the value of the previous cycle at this
  point. It's less likely that this path is the problem because it would
  mean that dql->num_completed was not updated correctly at the end of the
  previous accumulation and if this were the case it would probably mask
  the issue since num_completed would be a smaller value and, therefore,
  less likely to precipitate this issue. This is because num_queued -
  dql->num_completed would be larger than it should be and so count would
  less likely be greater than that
  
  If we look at the other path, there have been recent changes.
  Particularly in the ice_tx_map function we decide whether we need to
  notify the hardware via kick = __netdev_tx_sent_queue(...) and then we
  proceed to use writel_relaxed() to inform the NIC about the index up to
  which we have produced timestamp descriptors. Using writel_relaxed()
  does not have the same ordering guarantees as writel() across
  architectures with regards to visibility of memory accesses across CPUs
  and peripherals. On arm writel_relaxed() allows for reordering of writes
  to optimize performance and this potentially introduces a race between
  __netdev_tx_sent_queue, which updates the accumulator (dql->num_queued),
  and the write to the NIC's MMIO to kick off work. In other words, the
  NIC may start transmitting descriptors before the updated
  dql->num_queued value is globally visible and so the CPU crashes on the
  NIC's return because it thinks the NIC has sent packets that haven't
  been queued.
  
  kick = __netdev_tx_sent_queue(txring_txq(tx_ring), first->bytecount,
-                                     netdev_xmit_more());
+           netdev_xmit_more());
  if (!kick)
-       return;
-  
+  return;
+ 
  if (ice_is_txtime_cfg(tx_ring)) {
  ...
-       tstamp_ring->next_to_use = j;
-       writel_relaxed(j, tstamp_ring->tail);
+  tstamp_ring->next_to_use = j;
+  writel_relaxed(j, tstamp_ring->tail);
  } else {
-       writel_relaxed(i, tx_ring->tail);
+  writel_relaxed(i, tx_ring->tail);
  }
  
  This logic is a recent addition from the commit at [1]. Previously the
  logic was much simpler (see below) and notably used writel() which
  guarantees all previously memory writes are visible across CPUs and
  peripherals.
  
  if (kick)
-     /* notify HW of packet */
-     writel(i, tx_ring->tail);
+     /* notify HW of packet */
+     writel(i, tx_ring->tail);
  
  * I have provided the customer with a test kernel that replaces both
  instances of writel_relaxed() with writel() and the issue did not recur
  with new kernel after 2 weeks of testing where both NICs were under
  heavy and continuous load for multiple days without issue
  
  [1]
  
https://github.com/torvalds/linux/commit/ccde82e909467abdf098a8ee6f63e1ecf9a47ce5

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2161572

Title:
  ICE Driver triggers DQL BUG on arm64 Under Sustained Transmit Load

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2161572/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to