Subject: [Bug Report] Persistent CLOSE_WAIT socket leak on smtpDestination
(backend) connections - ASSP 2.8.2(24261)
Hi all,
I'm seeing a reproducible file descriptor leak on our ASSP installation that
eventually leads to "Too many open files" errors and worker crashes if left
unaddressed. I've isolated it fairly precisely and wanted to share the
evidence in case it's a known issue or others are hitting it silently.
## Environment
- ASSP version: 2.8.2(24261)
- OS: Debian/Ubuntu (Linux), Perl 5.26.1
- Run mode: systemd-managed daemon (Type=forking, PIDFile), ulimit raised to
65536 (soft and hard) to rule out an undersized limit as the cause
- smtpDestination:=192.168.1.54:25|SSL:192.168.1.54:465
(backend is Zimbra, both plaintext :25 and SSL :465 destinations
configured)
## Symptom
Open file descriptor count on the assp.pl process climbs steadily over
several hours, with no plateau, until it approaches the ulimit ceiling and
worker threads start failing with:
Worker_X accept to client failed IO::Socket::INET=GLOB(...) (timeout: 2
s) : Too many open files
main exception: error: AsspSelfLoader is unable to load code from file
.../sl-cache/main-resetFH.sl - Too many open files
This has now happened on multiple separate occasions (Aug 28, Sep 1, Sep 4,
and again this morning, Sep 15), each following the same growth pattern:
flat baseline FD count (~40-50) for hours, then a sustained climb beginning
at an arbitrary point in the day and continuing until intervention.
## Isolating the leak
Using `lsof -p <pid>`, I confirmed the growth is entirely in socket file
descriptors, not regular files, disk cache handles, or DNS-related opens
(REG/CHR/DIR counts stay flat throughout):
12:13:47 86 REG 46 IPv4 7 CHR 5 sock 2 DIR
12:28:47 86 REG 115 IPv4 7 CHR 11 sock 2 DIR
12:43:48 87 REG 293 IPv4 7 CHR 50 sock 2 DIR
12:58:48 86 REG 361 IPv4 92 sock 7 CHR 2 DIR
13:13:49 88 REG 536 IPv4 126 sock 7 CHR 2 DIR
13:43:56 86 REG 668 IPv4 255 sock 7 CHR 2 DIR
Breaking down socket state with `lsof -p <pid> -i` during an active leak
event this morning:
335 CLOSE_WAIT
28 ESTABLISHED
12 LISTEN
CLOSE_WAIT dominates - meaning the remote peer has already closed its end
(sent FIN) but ASSP has not called close() on its side.
Breaking the CLOSE_WAIT sockets down by remote host:
246 our.backend.mailserver <- our smtpDestination backend
72 mailsrv.graphic.com.gh
2 mail-qv2-f12.google.com
1 (various other single-connection external senders - normal
background noise)
The overwhelming majority (73%+) are connections to our own backend/
destination server (Zimbra), not inbound client connections. Inspecting the
port on these:
perl 7243 root 46u IPv4 ... TCP
assp.eomega.org:37038->our.backend.mailserver:smtp
(CLOSE_WAIT)
perl 7243 root 51u IPv4 ... TCP
assp.eomega.org:43076->our.backend.mailserver:smtp
(CLOSE_WAIT)
... (continues, ~295 similar entries)
Nearly all of these are the plaintext `smtp` (port 25) destination, not the
SSL (:465) destination - so this does not appear to be an SSL-shutdown
issue, but rather plain-TCP backend connections that are not being closed
after the delivery/proxy transaction completes.
## What this looks like to me
It appears that when ASSP proxies a message through to smtpDestination and
the backend (Zimbra) closes its end of the connection after completing the
SMTP transaction, ASSP does not follow up with its own close() on that
socket - leaving it in CLOSE_WAIT indefinitely. Over hours, under normal
mail volume, this accumulates into the thousands and exhausts the process's
file descriptor limit.
## What I've ruled out
- Undersized ulimit: raised to 65536, leak still occurs, just takes longer
to become fatal
- DNS resolver file-open failures: present in logs but unrelated; these are
a side-effect of FD exhaustion (resolver can't open /etc/protocols), not
a cause
- SSL socket shutdown handling: leak is concentrated on the plaintext :25
backend connection, not :465
- Spam-flood/PenaltyBox tempfail handling: this was our working theory
after an earlier incident (leak did correlate with a spam sender getting
tempfailed repeatedly), but this most recent capture shows the leak is
overwhelmingly backend-connection-related regardless of that theory -
possibly a contributing trigger, but not the core mechanism
- tcp_fin_timeout / kernel-level socket reclaim: not applicable, since
CLOSE_WAIT is governed entirely by the application, not any kernel
timeout
## Questions for the list
1. Is there a known issue with smtpDestination connection handling/closing
in this version or fixed in a later release?
2. Is there a config option to disable backend connection reuse/pooling (if
that's what's happening here), or to force ASSP to close backend
connections immediately after each transaction rather than potentially
holding them open for reuse?
3. Has anyone else seen CLOSE_WAIT accumulation specifically on the
smtpDestination leg rather than the client-facing leg?
Happy to provide additional lsof/tcpdump captures, debug logging, or test
config changes if it helps narrow this down. Currently mitigating with a
raised ulimit + systemd auto-restart to avoid crashes while we get to the
bottom of the actual leak.
Thanks,
Rob
_______________________________________________
Assp-user mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/assp-user