https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=298416

            Bug ID: 298416
           Summary: [nfs] NFSv4 slot loss on aborted compound operations
           Product: Base System
           Version: 14.5-STABLE
          Hardware: amd64
                OS: Any
            Status: New
          Severity: Affects Some People
          Priority: ---
         Component: kern
          Assignee: [email protected]
          Reporter: [email protected]

Created attachment 274654
  --> https://bugs.freebsd.org/bugzilla/attachment.cgi?id=274654&action=edit
Patch against 14-STABLE.

We have identified some cases where silent slot loss can occur when operations
on NFS mounts are aborted. We experience this when using NFSv4.2, but it likely
also occurs with NFSv4.1.

A slot is acquired for compound operations by nfsv4_setsequence() and freed by
newnfs_request(). Any call path that abandons the compound before reaching
newnfs_request() loses the slot permanently.

We identified four call sites where this happens, one of which where it
actually does happen for us in a semi-reproducible way, which allowed us to
develop a candidate patch, attached.

The patch adds one function, nfsv4_freeunsentslot(), to nfs_commonsubs.c. It is
called from each of the four call sites: nfsrpc_writerpc(), nfsrpc_writeds(),
and two in nfsrpc_setextattr().

nfsv4_freeunsentslot() calls nfsv4_freeslot() with resetseq set to true because
the aborted RPC will never be sent and therefore the slot sequence ID should
not advance.

There are a handful of other call sites inside newnfs_request() where the RPC
dies after sending, which this patch does not address because it's no longer
certain that the server won't see them, making it unsafe to roll back the slot
sequence ID. Perhaps the slot should be explicitly marked bad in those cases,
matching the existing behavior when the stat is one of RPC_CANTSEND,
RPC_CANTRECV, RPC_SYSTEMERROR, or RPC_INTR.

The patch is against 14-STABLE, but it's pretty straightforward.

With this patch and other diagnostics, we have confirmed that slot loss does
not occur under the conditions that otherwise cause it in our environment. (We
are still trying to understand the underlying cause of the aborted RPCs in our
environment, which is a separate and less straightforward issue.)

-- 
You are receiving this mail because:
You are the assignee for the bug.

Reply via email to