On 9/8/26 12:34 PM, Mike Snitzer wrote:
> nfsd_dio_iter_is_aligned() approves a write iterator against the
> file's STATX_DIOALIGN attributes (a whole-iterator iov_iter_alignment()
> test against dio_mem_align), but the block stack applies stricter
> geometry tests at bio split time: bio_split_io_at() checks each bvec's
> offset and length against the queue's dma_alignment and may find no
> valid block-size-aligned split at all. An ITER_BVEC WRITE payload can
> pass the former and fail the latter: bio_iov_bvec_set() hands nfsd's
> bvec array to the queue as-is, and the payload's first fragment starts
> mid-page (the RPC header precedes it in the receive buffer), so the
> iterator's interior page boundaries need not be logical-block aligned
> and a bio the queue must split may have no valid split point. When
> that happens, nfsd_direct_write() returned the -EINVAL to the client
> as a failed WRITE (NFS4ERR_INVAL) -- for a perfectly valid request.
>
> Observed against a brd-backed nvme-loop XFS export (dio_mem_align=4)
> with 1 MiB WRITEs, e.g. arriving as 65 bvecs with bv0=(408,15976):
> the gate admits the iterator, the block layer rejects it, and every
> large write on the affected connection errors out (dd: Invalid
> argument).
Thanks for chasing this down. The bv0 numbers make the gate defect
clear: 15976 is not a multiple of the logical block size, so the
direct segment's first interior bvec boundary lands mid-sector. The
boundaries after that are page boundaries, which are fine.
A small correction for the commit message: nfsd_dio_iter_is_aligned()
doesn't exist. The gate is the first-bvec offset test in
nfsd_write_dio_iters_init(), and it checks only that one offset
against nf_dio_mem_align. Likewise bio_iov_bvec_set() is now
bio_iov_iter_set().
> Treat -EINVAL from the direct attempt as "not direct-able": restore the
> segment's iterator and retry it as (uncached when FOP_DONTCACHE)
> buffered I/O, the same fallback nfsd_write_dio_iters_init() picks for
> geometries it rejects itself.
[ ... ]
> @@ -1467,6 +1469,33 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct
> svc_fh *fhp,
> expected = iov_iter_count(&segments[i].iter);
>
> host_err = vfs_iocb_iter_write(file, kiocb, &segments[i].iter);
> + if (unlikely(host_err == -EINVAL &&
> + (kiocb->ki_flags & IOCB_DIRECT))) {
[ ... ]
> + segments[i].iter = saved_iter;
> + kiocb->ki_flags &= ~IOCB_DIRECT;
> + if (file->f_op->fop_flags & FOP_DONTCACHE)
> + kiocb->ki_flags |= IOCB_DONTCACHE;
> + trace_nfsd_write_vector(rqstp, fhp, kiocb->ki_pos,
> + segments[i].iter.count);
> + host_err = vfs_iocb_iter_write(file, kiocb,
> + &segments[i].iter);
> + }
Per our discussion last October:
https://lore.kernel.org/linux-nfs/[email protected]/
The conclusion then was that -EINVAL from ->write_iter can come from
a number of conditions in the filesystem, so NFSD can't treat it as
meaning only that the I/O was misaligned. That still holds, so I'd
rather not use -EINVAL to signal a retry. An -EINVAL that really is
the filesystem rejecting the request would now cost a second full
write attempt before surfacing anyway.
nfsd_write_dio_iters_init() already has the segment start and
nf_dio_offset_align, and after the first bvec every boundary is
page-aligned. If it also requires the first bvec's remaining length
(from the segment start) to be a multiple of offset_align and takes
the no_dio path otherwise, that rejects bv0=(408,15976) up front
using only data NFSD already has.
What would help me understand the failure even better:
- Which -EINVAL in bio_split_io_at() fired: the per-bvec dma_alignment
test, or the zero-length result after ALIGN_DOWN()?
- On the reproducer, how does stx_dio_offset_align compare with the
queue's logical_block_size?
If there turn out to be cases the gate can't predict from the statx
data, that seems like a question for the block and fs folks about
what error the filesystem should surface, rather than something to
work around in NFSD.
--
Chuck Lever (Come to NFS bake-a-thon! https://nfsv4bat.org)