While qualifying NFSD's NFSD_IO_DIRECT write path with byte-level data
verification, it was found that the ITER_BVEC payloads nfsd submits 
expose silent data corruption in two bio-based block drivers and
two defects in nfsd itself. Example problematic payloads is the first
fragment starts mid-page because the RPC header precedes it in the
receive buffer, and fragment lengths need not be sector multiples.
bio_iov_bvec_set() passes such an array to the queue as-is; nothing
below it validates per-bvec sector alignment.

Patches 1-2 fix silent corruption (write completes successfully, data
lands wrong) and are stable candidates:

- brd re-derives each segment's device position from
  bio->bi_iter.bi_sector, which bio_advance_iter_single() advances by
  whole sectors only, so a sub-sector segment length skews everything
  that follows.

- zram hardwires is_partial_io() to false on 4K-page kernels, sending
  sub-page bvecs down a whole-page path that ignores bv_offset/bv_len
  entirely, and has the same sector-cursor skew.

Both are verified with a synthetic-bio reproducer (stamped pattern,
write, read back, compare) across mid-page and page-aligned
geometries.

Patches 3-4 fix nfsd: the filecache never fetches DIO alignment
attributes on the supplied-file acquire branch, so every WRITE to a
file created via NFSv4 OPEN(CREATE) is refused direct I/O for the
file's cached lifetime; and nfsd's statx-based DIO gate is weaker than
bio_split_io_at()'s split-time checks, so an admitted iterator can
still be rejected by the block layer -- retry the segment buffered
instead of failing a valid WRITE with NFS4ERR_INVAL.

One open question for the iomap/block maintainers: should ITER_BVEC
direct I/O with sub-sector bvec boundaries be validated or bounced
centrally rather than trusted to every driver's iteration?  An audit
of in-tree bio-based drivers found the same bi_sector-derived position
pattern in dm-io, dm-log-writes, dm-writecache (pmem path) and
dm-integrity -- unreachable through nfsd today only because dm queues
advertise dma_alignment >= 511, which nfsd's alignment gate refuses.

Tested with the reproducer matrix on brd, zram and nvme-loop at 4K and
16K page size (aarch64) and 4K (x86_64), plus 30-connection NFS write
rigs comparing source against export byte-for-byte: clean with the
fixes, corrupting or erroring without them.

Mike Snitzer (3):
  brd: iterate the bio by byte position, not bi_sector
  zram: handle sub-page bvec segments without corrupting data
  nfsd: fall back to buffered I/O when a direct write gets -EINVAL

David Flynn (1):
  nfsd: fetch direct I/O alignment for files handed to the filecache

 drivers/block/brd.c           | 29 +++++++++++++++++++++++------
 drivers/block/zram/zram_drv.c | 36 +++++++++++++++++------------------
 fs/nfsd/filecache.c           |  4 ++--
 fs/nfsd/vfs.c                 | 29 +++++++++++++++++++++++++++++
 4 files changed, 77 insertions(+), 31 deletions(-)

--
2.52.0

Reply via email to