Apologies for the git-send-email misfire.. still not sure what
happened, but will sort it out for future.

On Tue, Sep 08, 2026 at 12:34:35PM -0400, Mike Snitzer wrote:
> While qualifying NFSD's NFSD_IO_DIRECT write path with byte-level data
> verification, it was found that the ITER_BVEC payloads nfsd submits 
> expose silent data corruption in two bio-based block drivers and
> two defects in nfsd itself. Example problematic payloads is the first
> fragment starts mid-page because the RPC header precedes it in the
> receive buffer, and fragment lengths need not be sector multiples.
> bio_iov_bvec_set() passes such an array to the queue as-is; nothing
> below it validates per-bvec sector alignment.
> 
> Patches 1-2 fix silent corruption (write completes successfully, data
> lands wrong) and are stable candidates:
> 
> - brd re-derives each segment's device position from
>   bio->bi_iter.bi_sector, which bio_advance_iter_single() advances by
>   whole sectors only, so a sub-sector segment length skews everything
>   that follows.
> 
> - zram hardwires is_partial_io() to false on 4K-page kernels, sending
>   sub-page bvecs down a whole-page path that ignores bv_offset/bv_len
>   entirely, and has the same sector-cursor skew.
> 
> Both are verified with a synthetic-bio reproducer (stamped pattern,
> write, read back, compare) across mid-page and page-aligned
> geometries.
> 
> Patches 3-4 fix nfsd: the filecache never fetches DIO alignment
> attributes on the supplied-file acquire branch, so every WRITE to a
> file created via NFSv4 OPEN(CREATE) is refused direct I/O for the
> file's cached lifetime; and nfsd's statx-based DIO gate is weaker than
> bio_split_io_at()'s split-time checks, so an admitted iterator can
> still be rejected by the block layer -- retry the segment buffered
> instead of failing a valid WRITE with NFS4ERR_INVAL.
> 
> One open question for the iomap/block maintainers: should ITER_BVEC
> direct I/O with sub-sector bvec boundaries be validated or bounced
> centrally rather than trusted to every driver's iteration?  An audit
> of in-tree bio-based drivers found the same bi_sector-derived position
> pattern in dm-io, dm-log-writes, dm-writecache (pmem path) and
> dm-integrity -- unreachable through nfsd today only because dm queues
> advertise dma_alignment >= 511, which nfsd's alignment gate refuses.
> 
> Tested with the reproducer matrix on brd, zram and nvme-loop at 4K and
> 16K page size (aarch64) and 4K (x86_64), plus 30-connection NFS write
> rigs comparing source against export byte-for-byte: clean with the
> fixes, corrupting or erroring without them.
> 
> Mike Snitzer (3):
>   brd: iterate the bio by byte position, not bi_sector
>   zram: handle sub-page bvec segments without corrupting data
>   nfsd: fall back to buffered I/O when a direct write gets -EINVAL
> 
> David Flynn (1):
>   nfsd: fetch direct I/O alignment for files handed to the filecache
> 
>  drivers/block/brd.c           | 29 +++++++++++++++++++++++------
>  drivers/block/zram/zram_drv.c | 36 +++++++++++++++++------------------
>  fs/nfsd/filecache.c           |  4 ++--
>  fs/nfsd/vfs.c                 | 29 +++++++++++++++++++++++++++++
>  4 files changed, 77 insertions(+), 31 deletions(-)
> 
> --
> 2.52.0

Reply via email to