Thank you all for your time and comments. Revisiting my patch after my vacation, I now think none of the barriers are actually necessary.

On 21/08/2026 11:40, Richard Lyu wrote:
memory reordering between the wake up and clearing the flag is allowed.

When exactly can this reordering happen?

diff --git a/drivers/char/tpm/tpm-chip.c b/drivers/char/tpm/tpm-chip.c
index 12b7394b34bd..7f500797b7a7 100644 --- a/drivers/char/tpm/tpm- chip.c
+++ b/drivers/char/tpm/tpm-chip.c @@ -173,6 +173,9 @@ int
tpm_try_get_ops(struct tpm_chip *chip) if (chip->flags &
TPM_CHIP_FLAG_SUSPENDED) goto out_lock;

+    /* Ensure that device is fully resumed */
+    rmb();
+
     rc = tpm_chip_start(chip);
     if (rc)
         goto out_lock;

Where inside tpm_chip_start do we actually need to avoid loading a stale state or flag?
tpm_chip_start() reads chip->locality for example. I'm not sure it can be written to concurrently but some drivers write to it. It might not be a problem as it would just trigger a call to tpm_request_locality().

Since speculative writes are not possible, so rmb() is sufficient; in any case, no need for a read-to-write memory barrier.

diff --git a/drivers/char/tpm/tpm-interface.c
b/drivers/char/tpm/tpm-interface.c index f745a098908b..2de12b02f62b 100644
--- a/drivers/char/tpm/tpm-interface.c +++
b/drivers/char/tpm/tpm-interface.c @@ -474,13 +474,12 @@ int
tpm_pm_resume(struct device *dev) if (chip == NULL) return -ENODEV;

-    chip->flags &= ~TPM_CHIP_FLAG_SUSPENDED;
-
     /*
      * Guarantee that SUSPENDED is written last, so that hwrng does not
      * activate before the chip has been fully resumed.
      */
     wmb();
+    chip->flags &= ~TPM_CHIP_FLAG_SUSPENDED;

Can you rationalize this change?
This change is mostly based on your comment: tpm_pm_resume() is called after the hardware-specific operations have been performed, and the flag must be set after these operations are complete.

I assumed hwrng corresponds to tpm_hwrng_read() which calls tpm_get_random() which itself calls tpm_try_get_ops(). After tpm_chip_start() is called, tpm_try_get_ops() can return a valid pointer, and tpm_hwrng_read() can work.

Individual drivers implement the .resume() method by doing hardware-specific operations then calling tpm_pm_resume(), so ordering needs to happen between those hardware-specific operations and tpm_pm_resume(). This ensures that the order of operations is:
  - hardware-specific resume operations,
  - tpm_pm_resume(), including clearing the TPM_CHIP_FLAG_SUSPENDED flag,
  - check that TPM_CHIP_FLAG_SUSPENDED is clear,
  - run of tpm_chip_start().

If tpm_try_get_ops() succeeds, then tpm_chip_start() is run so the hardware-specific resume operations are complete.


I agree the clearing of the flag should be moved after wmb() to guarantee that SUSPENDED is written last, the flag has to be cleared after the barrier.
That part makes sense.

My remaining question is whether we actually need the wmb() barrier here?
Revisiting the patch, I agree that this point is not fully clear. Looking at the implementations of tpm drivers, most use the tpm_pm_resume() function directly as the .resume() method.

Three drivers implement specific .resume() methods:
- tpm_infineon.c: which uses writeb/outb that adds the necessary barriers, - st33zp24/st33zp24.c: which calls tpm_pm_resume() without driver-specific operation depending on the config,
  - tpm_tis_core.c: which makes a self-test before calling tpm_pm_resume().

None of those cases seem to require a barrier at all.



Reply via email to