Thanks for the review. All three new findings are legitimate; v5
(following as a separate thread) addresses them:

- [Medium] mana_gf_stats_work_handler(): agreed. When
  mana_rdma_probe() outlives the 2s stats period, the gate dropped the
  reset request with no latch and no re-arm, so a probe that then
  succeeded never recovered the stats work. v5 takes the re-arm
  option: the gated branch re-arms the delayed work, retrying once
  the probe completes; a failing probe still cancels it in the
  mana_remove() unwind. The latch alternative would roll back a
  healthy probe on a single transient timeout, so I left it out.

- [Low] system_wq: agreed, and the work also escaped
  flush_scheduled_work(), since system_wq is a separate compatibility
  queue. v5 returns to schedule_work(), the pre-patch
  system_percpu_wq placement. Whether a long queue suits the 10s+
  service cycles better can be a follow-up of its own.

- [Low] service_quiesce comment: agreed. It now states that only the
  rollback path can have admitted a cycle, because it runs
  mana_service_probe_complete() before jumping there.

On the [High]: the PM suspend/resume path is outside this series'
remove/free closure, and applying quiesce there directly would close
admission one-way across suspend/resume and would deadlock a service
cycle that is itself calling mana_gd_suspend(). A reversible wait
plus a resume-side reopen is needed, which will follow as a separate
patch.

pw-bot: cr


Reply via email to