Slurm already has connect retry logic (10 times with 0.1 sec between retries). DRMAA should need no changes unless it directly accesses munge.

Has anyone else seen this problem?

Quoting E V <[email protected]>:

Had 2 jobs die yesterday morning with a slurm_load_jobs error:
Protocol authentication error from inside DRMAA, and this interesting
message in the log:

If munged is up, restart with --num-threads=10
error: Munge encode failed: Unable to access
"/var/run/munge/munge.socket.2": No such file or directory
error: authentication: Munged communication error

The slurmctl.log has this
error: slurm_receive_msg: Zero Bytes were transmitted or received

right about the same time.

Digging deeper it appears that the jobs state's were changing in the
slurmctl just as the munge daemon got restarted for a logrotate. I
changed logrotate to rotate munge.log based on size instead of daily,
which may fix the problem, but feels more like a work around. Any
other suggestions? I'd be nice to have some sort of retry in the code,
but not really sure if it'd be in the slurmctl or the DRMAA code.


--
Morris "Moe" Jette
CTO, SchedMD LLC

Reply via email to