** Description changed: [ Impact ] - * When Intel's E810 driver (ice) initializes, one of the critical steps - is to read the size of the NVM. This process is as follows: - 1) driver acquires the NVM lock (a shared resource across all PFs - and FW components) - 2) driver keeps reading the NVM through AdminQueue commands until FW - signals the end of the read - 3) driver releases the NVM lock + * When Intel's E810 driver (ice) initializes, one of the critical steps + is to read the size of the NVM. This process is as follows: + 1) driver acquires the NVM lock (a shared resource across all PFs + and FW components) + 2) driver keeps reading the NVM through AdminQueue commands until FW + signals the end of the read + 3) driver releases the NVM lock - It is possible that while the driver is reading the NVM (step 2) FW - returns an EBUSY status. This status is not handled by the driver and - leads to a critical error that stops further initialization. As a - result the interface is not brought up. + It is possible that while the driver is reading the NVM (step 2) FW + returns an EBUSY status. This status is not handled by the driver and + leads to a critical error that stops further initialization. As a + result the interface is not brought up. - * The root cause is in how the driver acquires the NVM lock. One of the - required arguments is a timeout specifying how long the lock is valid. - When that timeout expires, FW frees the lock internally and allows a - different component to grab it - a different driver, a different PF, or - a different FW component. In that case, if the driver that originally - acquired the lock keeps reading the NVM (step 2), it will eventually - receive an EBUSY status from FW. + * The root cause is in how the driver acquires the NVM lock. One of the + required arguments is a timeout specifying how long the lock is valid. + When that timeout expires, FW frees the lock internally and allows a + different component to grab it - a different driver, a different PF, or + a different FW component. In that case, if the driver that originally + acquired the lock keeps reading the NVM (step 2), it will eventually + receive an EBUSY status from FW. - * The maximum timeout for this lock is 3000ms, and the ice driver - currently does not respect it. It is rare but possible that reading a - large NVM sector takes more than 3s. This is especially visible with - some Dell-branded E810 NVMs. + * The maximum timeout for this lock is 3000ms, and the ice driver + currently does not respect it. It is rare but possible that reading a + large NVM sector takes more than 3s. This is especially visible with + some Dell-branded E810 NVMs. - * The fix changes how the ice driver acquires the NVM lock: instead of - acquiring it once at the beginning of the bulk read and releasing it at - the end, the driver now requests the lock before each and every NVM - read. This adds a small, measurable overhead to driver initialization - time - under 50ms per PF initialization with full debug logs enabled. + * The fix changes how the ice driver acquires the NVM lock: instead of + acquiring it once at the beginning of the bulk read and releasing it at + the end, the driver now requests the lock before each and every NVM + read. This adds a small, measurable overhead to driver initialization + time - under 50ms per PF initialization with full debug logs enabled. - * The fix is approved by Intel and is currently in net-dev review. - [TODO: update once the patch reaches linux-next] - https://lore.kernel.org/netdev/[email protected]/ + * Upstream Fix: + https://github.com/torvalds/linux/commit/df88d6f1ed653993bd5c8647aef0e6498f4b1647 + https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=df88d6f1ed653993bd5c8647aef0e6498f4b1647 [ Test Plan ] - * The issue can be reproduced when the ice driver is initializing and - reading the NVM takes more than 3s. I have not found a way to reproduce - it locally, nor to force the FW to slow down its responses. + * The issue can be reproduced when the ice driver is initializing and + reading the NVM takes more than 3s. I have not found a way to reproduce + it locally, nor to force the FW to slow down its responses. - * The issue was reported on a system using a Dell-branded 4.60 0x8001f4f3 - NVM. The driver responds with the following error: + * The issue was reported on a system using a Dell-branded 4.60 0x8001f4f3 + NVM. The driver responds with the following error: - [Sun Apr 19 21:37:37 2026] ice 0000:41:00.0: ice_init_hw failed: -5 - [Sun Apr 19 21:37:37 2026] ice: probe of 0000:41:00.0 failed with error -5 + [Sun Apr 19 21:37:37 2026] ice 0000:41:00.0: ice_init_hw failed: -5 + [Sun Apr 19 21:37:37 2026] ice: probe of 0000:41:00.0 failed with error -5 - After this, the 0000:41:00.0 interface is not initialized. + After this, the 0000:41:00.0 interface is not initialized. - * With debug logs enabled (ice.dyndbg=+p), the following pattern can be - observed: + * With debug logs enabled (ice.dyndbg=+p), the following pattern can be + observed: - [ 102.699610] ice 0000:05:00.0: CQ CMD: opcode 0x0008, flags 0x2003, datalen 0x0000, retval 0x0000 <- Lock acquired, timeout 3s - [ 102.699634] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3000, datalen 0x0001, retval 0x0000 <- Read operation starts - [ 103.020213] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3003, datalen 0x0001, retval 0x0000 <- Read operation ends - [ 104.648070] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3000, datalen 0x0001, retval 0x0000 <- Read operation starts - [ 105.434810] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3007, datalen 0x0001, retval 0x000E <- Read operation ends - [ 106.820932] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3000, datalen 0x0001, retval 0x0000 <- This read is executed after the lock has already expired - [ 109.603077] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3007, datalen 0x0001, retval 0x000C <- We receive EBUSY + [ 102.699610] ice 0000:05:00.0: CQ CMD: opcode 0x0008, flags 0x2003, datalen 0x0000, retval 0x0000 <- Lock acquired, timeout 3s + [ 102.699634] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3000, datalen 0x0001, retval 0x0000 <- Read operation starts + [ 103.020213] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3003, datalen 0x0001, retval 0x0000 <- Read operation ends + [ 104.648070] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3000, datalen 0x0001, retval 0x0000 <- Read operation starts + [ 105.434810] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3007, datalen 0x0001, retval 0x000E <- Read operation ends + [ 106.820932] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3000, datalen 0x0001, retval 0x0000 <- This read is executed after the lock has already expired + [ 109.603077] ice 0000:05:00.0: CQ CMD: opcode 0x0701, flags 0x3007, datalen 0x0001, retval 0x000C <- We receive EBUSY - * Regression testing was based on rebooting the system and verifying that - all interfaces initialized properly. + * Regression testing was based on rebooting the system and verifying that + all interfaces initialized properly. [ Where problems could occur ] - * This change alters the NVM locking granularity in the ice driver: the - lock is now acquired and released per firmware read command rather than - once for the whole bulk read. A regression would be confined to the ice - driver's NVM/flash access paths (flash-size discovery during init, - devlink NVM snapshots, ethtool/devlink flash logging). + * This change alters the NVM locking granularity in the ice driver: the + lock is now acquired and released per firmware read command rather than + once for the whole bulk read. A regression would be confined to the ice + driver's NVM/flash access paths (flash-size discovery during init, + devlink NVM snapshots, ethtool/devlink flash logging). - * Acquiring and releasing the lock more frequently increases the number - of AdminQueue commands issued during initialization. If FW is slow to - grant the lock, or another PF/FW component holds it, initialization - could take longer or, in the worst case, fail to acquire the lock. This - would show up as slower probe times or, in the unlikely event of a - locking mistake, contention with concurrent NVM operations. + * Acquiring and releasing the lock more frequently increases the number + of AdminQueue commands issued during initialization. If FW is slow to + grant the lock, or another PF/FW component holds it, initialization + could take longer or, in the worst case, fail to acquire the lock. This + would show up as slower probe times or, in the unlikely event of a + locking mistake, contention with concurrent NVM operations. - * A regression would most likely manifest as NVM read failures during - driver probe (interface failing to initialize), or as failures reading - the flash via devlink/ethtool. It does not affect the data path. + * A regression would most likely manifest as NVM read failures during + driver probe (interface failing to initialize), or as failures reading + the flash via devlink/ethtool. It does not affect the data path. [ Other Info ] - * SRU for: - Ubuntu 22.04 LTS (Jammy) - linux 5.15 - Ubuntu 24.04 LTS (Noble) - linux 6.8 - Ubuntu 26.04 (Resolute) - linux 7.0 - Ubuntu 26.10 (Stonking) - linux 7.1 + * SRU for: + Ubuntu 22.04 LTS (Jammy) - linux 5.15 + Ubuntu 24.04 LTS (Noble) - linux 6.8 + Ubuntu 26.04 (Resolute) - linux 7.0 + Ubuntu 26.10 (Stonking) - linux 7.1 - * Fix details: - - Move the NVM lock acquire/release into ice_read_flat_nvm() so the - lock brackets each individual ice_aq_read_nvm() command instead of - the whole multi-sector read. This guarantees the lock is never held - across more than one firmware read, so firmware never reclaims it - mid-read. - - Because ice_release_nvm() issues its own AQ command and overwrites - hw->adminq.sq_last_status, an optional read_aq_err output parameter - is added to ice_read_flat_nvm() to capture the failing read's AQ - error before the release. Callers that need it - (ice_discover_flash_size() and the ethtool/devlink log paths) use it - instead of sq_last_status; others pass NULL. Callers that previously - took the lock around ice_read_flat_nvm(), ice_read_sr_word() or - ice_read_flash_module() now call them without it. - - Fixes: e94509906d6b ("ice: create function to read a section of the - NVM and Shadow RAM") + * Fix details: + - Move the NVM lock acquire/release into ice_read_flat_nvm() so the + lock brackets each individual ice_aq_read_nvm() command instead of + the whole multi-sector read. This guarantees the lock is never held + across more than one firmware read, so firmware never reclaims it + mid-read. + - Because ice_release_nvm() issues its own AQ command and overwrites + hw->adminq.sq_last_status, an optional read_aq_err output parameter + is added to ice_read_flat_nvm() to capture the failing read's AQ + error before the release. Callers that need it + (ice_discover_flash_size() and the ethtool/devlink log paths) use it + instead of sq_last_status; others pass NULL. Callers that previously + took the lock around ice_read_flat_nvm(), ice_read_sr_word() or + ice_read_flash_module() now call them without it. + - Fixes: e94509906d6b ("ice: create function to read a section of the + NVM and Shadow RAM") - * Stonking backport notes: the commit applies as a clean cherry-pick. The - 7.1 tree already carries the upstream libie_aq_* API (enum libie_aq_err, - LIBIE_AQ_RC_*, libie_aq_str) and the devlink/devlink.c layout that the - upstream patch targets, so no adaptations were required. The four - modified files (devlink/devlink.c, ice_ethtool.c, ice_nvm.c, ice_nvm.h) - patched without offset or fuzz, and the affected objects build cleanly. + * Stonking backport notes: the commit applies as a clean cherry-pick. The + 7.1 tree already carries the upstream libie_aq_* API (enum libie_aq_err, + LIBIE_AQ_RC_*, libie_aq_str) and the devlink/devlink.c layout that the + upstream patch targets, so no adaptations were required. The four + modified files (devlink/devlink.c, ice_ethtool.c, ice_nvm.c, ice_nvm.h) + patched without offset or fuzz, and the affected objects build cleanly. - * Resolute backport notes: the commit applies as a clean cherry-pick. The - 7.0 tree already carries the upstream libie_aq_* API (enum libie_aq_err, - LIBIE_AQ_RC_*, libie_aq_str) and the devlink/devlink.c layout that the - upstream patch targets, so no adaptations were required. The four - modified files (devlink/devlink.c, ice_ethtool.c, ice_nvm.c, ice_nvm.h) - patched without offset or fuzz, and the affected objects build cleanly. + * Resolute backport notes: the commit applies as a clean cherry-pick. The + 7.0 tree already carries the upstream libie_aq_* API (enum libie_aq_err, + LIBIE_AQ_RC_*, libie_aq_str) and the devlink/devlink.c layout that the + upstream patch targets, so no adaptations were required. The four + modified files (devlink/devlink.c, ice_ethtool.c, ice_nvm.c, ice_nvm.h) + patched without offset or fuzz, and the affected objects build cleanly. - * Noble backport notes: the commit is a backport (not a clean - cherry-pick). The following adaptations were required for the 6.8 tree: - - 6.8 still uses the ice_aq_* API rather than upstream's libie_aq_* types - (enum ice_aq_err, ICE_AQ_RC_*, ice_aq_str); - - devlink changes are in ice_devlink.c instead of devlink/devlink.c; - - the ice_get_eeprom() aq_err is logged via ice_aq_str(); - - in the new lock-acquire failure path of ice_read_flat_nvm() the aq_err is - logged with %d on hw->adminq.sq_last_status + * Noble backport notes: the commit is a backport (not a clean + cherry-pick). The following adaptations were required for the 6.8 tree: + - 6.8 still uses the ice_aq_* API rather than upstream's libie_aq_* types + (enum ice_aq_err, ICE_AQ_RC_*, ice_aq_str); + - devlink changes are in ice_devlink.c instead of devlink/devlink.c; + - the ice_get_eeprom() aq_err is logged via ice_aq_str(); + - in the new lock-acquire failure path of ice_read_flat_nvm() the aq_err is + logged with %d on hw->adminq.sq_last_status - * Jammy backport notes: the commit is a backport (not a clean - cherry-pick). The following adaptations were required for the 5.15 tree: - - 5.15 uses enum ice_status, ICE_ERR_AQ_ERROR and enum ice_aq_err rather - than upstream's int and libie_aq_* types; - - devlink changes are in ice_devlink.c instead of devlink/devlink.c; - - the ice_get_eeprom() aq_err is logged via ice_aq_str(); - - the ice_devlink_nvm_snapshot() is the single-read (non-block-loop) form, - so different changes were required; - - in the new lock-acquire failure path of ice_read_flat_nvm() the aq_err is - logged with %d on hw->adminq.sq_last_status + * Jammy backport notes: the commit is a backport (not a clean + cherry-pick). The following adaptations were required for the 5.15 tree: + - 5.15 uses enum ice_status, ICE_ERR_AQ_ERROR and enum ice_aq_err rather + than upstream's int and libie_aq_* types; + - devlink changes are in ice_devlink.c instead of devlink/devlink.c; + - the ice_get_eeprom() aq_err is logged via ice_aq_str(); + - the ice_devlink_nvm_snapshot() is the single-read (non-block-loop) form, + so different changes were required; + - in the new lock-acquire failure path of ice_read_flat_nvm() the aq_err is + logged with %d on hw->adminq.sq_last_status - * There was a different attempt to fix the issue focusing on retrying the - read instead of acquiring the lock on every read. It was rejected - because of the race conditions it could introduce: - https://patchwork.ozlabs.org/project/intel-wired-lan/patch/[email protected]/ + * There was a different attempt to fix the issue focusing on retrying the + read instead of acquiring the lock on every read. It was rejected + because of the race conditions it could introduce: + https://patchwork.ozlabs.org/project/intel-wired-lan/patch/[email protected]/ - * The patch is currently in net-dev tree review: - https://lore.kernel.org/netdev/[email protected]/ + * The patch was merged upstream: + https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=df88d6f1ed653993bd5c8647aef0e6498f4b1647
-- You received this bug notification because you are a member of Ubuntu Bugs, which is subscribed to Ubuntu. https://bugs.launchpad.net/bugs/2163508 Title: ice: E810 interface fails to initialize (ice_init_hw failed: -5) during NVM read To manage notifications about this bug go to: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2163508/+subscriptions -- ubuntu-bugs mailing list [email protected] https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs
