On 8/3/26 12:36 AM, Alexis Lothoré (eBPF Foundation) wrote:
> Some tests, like tc_tunnel or tc_edt, sporadically fail in CI with the
> following logs:
>
> (network_helpers.c:309: errno: Operation now in progress) \
> Failed to connect to server
> send_and_test_data:FAIL:connect to server unexpected error: -115
>
> This is due to SO_RCVTIMEO and SO_SNDTIMEO being set on the client
> socket (see settimeo() in client_socket()), allowing connect() to return
> an error and to set errno to EINPROGRESS instead of blocking until
> connection result is known. Increasing the timeout value for those tests
> is likely not a good solution (and it has already been done by commit
> 2790db208b44 ("selftests/bpf: Improve tc_tunnel test reliability")):
> they involve subtests that expect the connection to fail, and so
> increasing the timeout value would increase overall test execution
> duration again (not only the connection, but any socket operation).
>
> Another solution, as documented in man 2 connect, is to poll the socket
> for POLLOUT once connect has returned EINPROGRESS, and to get the actual
> connection result through getsockopt: this allows to keep the overall
> timeout values low for the general traffic, while letting a chance to
> the connection to succeed even if CI runners are loaded.
>
> When connect() returns EINPROGRESS, poll the socket for POLLOUT and
> check the connection result via getsockopt(SO_ERROR).
>
> Fixes: 99126abec5e5 ("bpf: selftests: A few improvements to
> network_helpers.c")
> Signed-off-by: Alexis Lothoré (eBPF Foundation) <[email protected]>
> ---
> Changes in v2:
> - drop unneeded initialization
> - add back error message for immediate connection failure, and slightly
> reword the async connection failure error message
> - Link to v1:
> https://patch.msgid.link/[email protected]
> ---
> I manage to reproduce the issue locally by running `./test_progs -a
> tc_tunnel` in a qemu machine, while making all my CPUs busy with
> stress-ng on host side; the issue happens pretty quickly. I have not
> been able to reproduce the issue anymore with this fix.
> ---
> tools/testing/selftests/bpf/network_helpers.c | 41
> ++++++++++++++++++++++++---
> 1 file changed, 37 insertions(+), 4 deletions(-)
>
> diff --git a/tools/testing/selftests/bpf/network_helpers.c
> b/tools/testing/selftests/bpf/network_helpers.c
> index db935a9d9fc1..52e7b72f8a77 100644
> --- a/tools/testing/selftests/bpf/network_helpers.c
> +++ b/tools/testing/selftests/bpf/network_helpers.c
> @@ -14,6 +14,7 @@
> #include <sys/types.h>
> #include <sys/un.h>
> #include <sys/eventfd.h>
> +#include <sys/poll.h>
>
> #include <linux/err.h>
> #include <linux/in.h>
> @@ -40,6 +41,8 @@
> #define IPPROTO_MPTCP 262
> #endif
>
> +#define CONNECTION_IN_PROGRESS_TIMEOUT_MS 3000
> +
> #define clean_errno() (errno == 0 ? "None" : strerror(errno))
> #define log_err(MSG, ...) ({ \
> int __save = errno; \
> @@ -294,7 +297,8 @@ int client_socket(int family, int type,
> int connect_to_addr(int type, const struct sockaddr_storage *addr, socklen_t
> addrlen,
> const struct network_helper_opts *opts)
> {
> - int fd;
> + socklen_t errlen;
> + int fd, err;
>
> if (!opts)
> opts = &default_opts;
> @@ -305,13 +309,42 @@ int connect_to_addr(int type, const struct
> sockaddr_storage *addr, socklen_t add
> return -1;
> }
>
> - if (connect(fd, (const struct sockaddr *)addr, addrlen)) {
> + err = connect(fd, (const struct sockaddr *)addr, addrlen);
> + if (err && errno == EINPROGRESS) {
> + struct pollfd pfd = { .fd = fd, .events = POLLOUT };
> +
> + err = poll(&pfd, 1, CONNECTION_IN_PROGRESS_TIMEOUT_MS);
Hi Alexis, thanks for the patch.
When I ran tc_* selftests in parallel it looked like they hanged.
I think what's happening is that this 3s timeout is additive, making
some tests to wait for too long. For example tc_tunnel has 1s timeout,
but with this change it actually becomes 4s. AI says a successful
tc_tunnel run makes 50+ connections, so it adds up to minutes.
We should probably be using opts->timeout_ms as an absolute budget set
by the caller, and pass it (or remainder?) to the poll().
> +
> + if (err <= 0) {
> + if (err == 0) {
> + log_err("Connection timeout");
> + errno = ETIMEDOUT;
> + } else {
> + log_err("Failed to poll connection status");
> + }
> + goto close;
> + }
Also poll() can return EINTR here. I think we should retry EINTR,
taking into account the absolute deadline.
pw-bot: cr
> +
> + errlen = sizeof(err);
> + if (getsockopt(fd, SOL_SOCKET, SO_ERROR, &err, &errlen) < 0) {
> + log_err("Failed to getsockopt");
> + goto close;
> + }
> +
> + if (err) {
> + log_err("Eventually failed to connect to server");
> + errno = err;
> + goto close;
> + }
> + } else if (err) {
> log_err("Failed to connect to server");
> - save_errno_close(fd);
> - return -1;
> + goto close;
> }
>
> return fd;
> +close:
> + save_errno_close(fd);
> + return -1;
> }
>
> int connect_to_addr_str(int family, int type, const char *addr_str, __u16
> port,
>
> ---
> base-commit: 2efc18d4bb9ad28d240b69bd324937f1ec12e93d
> change-id: 20260710-tc_tunnel_flaky-27e9a191bd03
>
> Best regards,
> --
> Alexis Lothoré (eBPF Foundation) <[email protected]>
>