On Tue, Aug 04, 2026 at 11:26:57AM +0200, Paolo Abeni wrote:
On 7/30/26 10:18 AM, [email protected] wrote:
From: Nguyen Dinh Phi <[email protected]>
Syzbot report an issue which can be reproduced with these steps:
r0 = socket(AF_VSOCK, SOCK_STREAM, 0)
bind(r0, {VMADDR_CID_ANY, PORT})
connect(r0, {VMADDR_CID_LOCAL, PORT}) -> -1, EPROTO (self-connect)
listen(r0, backlog) -> 0
r1 = socket(AF_VSOCK, SOCK_STREAM, 0)
connect(r1, {VMADDR_CID_LOCAL, PORT}) -> 0
accept(r0) -> -1, EPROTO (stale sk_err)
Basically, it creates a socket (r0) and triggers a self-connect after
binding it. This self-connect fails with EPROTO because it loops back to
r0 while the socket is still in the TCP_SYN_SENT state, causing it to be
incorrectly dispatched to the connecting-client path. The unexpected
packet type encountered there sets sk_err to EPROTO.
After that, it invokes a listen() call on the same socket. This listen()
call succeeds because the kernel's listening path never inspects or
clears sk_err. Then, a new socket (r1) is created as a normal client and
connects to r0. However, vsock_accept() rejects this incoming connection
because the listener's sk_err still holds the EPROTO error from the
earlier failed self-connect.
This rejection causes the child socket created for r1's connection to
never be freed on virtio or hyperv transports; only the VMCI transport
implements pending_work to revisit and clean up a rejected socket
Fix the issue by using sock_error() to read the sk_err to prevent the
rejection branch from occurring in this scenario.
sock_error() atomically reads and clears sk_err, ensuring the error is
consumed when vsock_connect() returns and cannot affect subsequent
operations on the same socket. This matches the established pattern
used by other protocol connect() implementations in the network
stack like __inet_stream_connect(), tipc_wait_for_connect()...
Reported-by: [email protected]
Closes: https://syzkaller.appspot.com/bug?extid=1b2c9c4a0f8708082678
Fixes: d021c344051af ("VSOCK: Introduce VM Sockets")
Signed-off-by: Nguyen Dinh Phi <[email protected]>
Tested-by: Wupeng Ma <[email protected]>
Sashiko nipa points out that the race still exits:
https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260730081843.287563-1-phind.uet%40gmail.com
Yeah, it seems the same conclusion we reached with Michal on v1 and Phi
agreed on: https://lore.kernel.org/netdev/[email protected]/
Not sure why sk_err check was not removed in vsock_accept.
Phi can you check?
Thanks,
Stefano