Hope this is the correct list for this question.
 
I'm running nessus 1.0.10 as "nessus -D" on Linux 7.1
 
It appears the main server is not cleaning up the connection servers that are spawned when a client connects. There are numerous zombies lying around. I looked at the process group management code and it looks ... confused. My guess is that an the connection server has changed it's pgrpID so that it is relatively easy to kill all it's children when it decides to terminate. It just issues a kill(0, ...) and waits until there is no other process in it's process group. Kinda makes sense. Unfortunately the server "threads" and the main server use the same handler for the death-of-child signal. The process selection variable in the waitpid() call must be different for the main server and the connection servers. The connection server wants to wait for a child in it's group (waitpid(0,...)) and the main server wants to wait for any of its children (waitpid(-1,...)) or so it would seem.
 
This is incorrect thinking. Creating process groups is reasonable since this allows a member of the group to blast the whole group into oblivion. Waiting for children from "my" group is unnecessary in the case of the connection servers and incorrect in the case of the main server. The connection server will only receive DoC signals for it's children so selecting it's own process group doesn't do anything useful. Waiting for any child is just as effective. In the case of the main server, since it's children always create a new process group waiting for a process from "my" group will alway fail.
 
In addition, since pluggins may also decide to change groups for it's own reasons, waiting for "my" process group is probably not a good idea either although the zombies left around in this case will be cleaned up by init when the connection server terminates.
 
I also don't believe the loop in the DoC signal handler is correct. The child is dead (and therefore "wait" able) by the time the signal is posted so waitpid() should always find at least one child. If waitpid returns 0 or -1(and errno == ECHILD) then the loop can terminate. Unconditionally checking n times is unnecessary.
 
Dave Braun
 
BTW - changing the waitpid() call to waitpid(-1, ...) in nessusd/sighand.c seems to have solved my zombie problem.
 

Reply via email to