Dear developers,

a few minutes after a complete restart of all the slurmd/slurmctld in the cluster (slurm.conf as homogeneously as possible) the communication between master and clients on vm-SL6-[008-010] gets unstable:

tail slurmctld.log | egrep "acct_gather|SL6"

[2015-09-02T23:45:35.280] Updating acct_gather data for taurusi[2045-2079,2081-2108,4001-4073,4091-4180,4199-4216,4219,4225,4230-4232,5001-5408,5410-5432,6091-6233,6235-6270,6451-6477,6479-6522,6524-6536,6541-6612],vm-SL6-[008-010] [2015-09-02T23:45:35.284] debug3: Tree sending to taurusi6567 along with taurusi[6568-6612],vm-SL6-[008-010] [2015-09-02T23:47:05.387] agent/is_node_resp: node:vm-SL6-009 rpc:1017 : Can't find an address, check slurm.conf [2015-09-02T23:47:05.387] agent/is_node_resp: node:vm-SL6-010 rpc:1017 : Can't find an address, check slurm.conf [2015-09-02T23:47:05.387] agent/is_node_resp: node:vm-SL6-010 rpc:1017 : Can't find an address, check slurm.conf [2015-09-02T23:47:05.387] agent/is_node_resp: node:vm-SL6-008 rpc:1017 : Can't find an address, check slurm.conf [2015-09-02T23:47:05.652] error: Nodes taurusi2073,vm-SL6-[008-010] not responding

and a little later

[2015-09-02T23:47:43.443] debug3: Tree sending to vm-SL6-008
[2015-09-02T23:47:43.444] debug3: Tree sending to vm-SL6-010
[2015-09-02T23:47:43.445] debug3: Tree sending to vm-SL6-009
[2015-09-02T23:47:43.591] Node vm-SL6-010 now responding
[2015-09-02T23:47:43.591] debug2: node_did_resp vm-SL6-010
[2015-09-02T23:47:43.591] Node vm-SL6-008 now responding
[2015-09-02T23:47:43.591] debug2: node_did_resp vm-SL6-008
[2015-09-02T23:47:43.591] Node vm-SL6-009 now responding
[2015-09-02T23:47:43.591] debug2: node_did_resp vm-SL6-008
[2015-09-02T23:47:43.591] Node vm-SL6-009 now responding
[2015-09-02T23:47:43.591] debug2: node_did_resp vm-SL6-009

and again...

[2015-09-02T23:48:05.798] Updating acct_gather data for taurusi[2045-2079,2081-2108,4001-4073,4091-4180,4199-4216,4219,4225,4230-4232,5001-5408,5410-5432,6091-6233,6235-6270,6451-6477,6479-6522,6524-6536,6541-6612],vm-SL6-[008-010] [2015-09-02T23:48:05.823] debug3: Tree sending to taurusi6567 along with taurusi[6568-6612],vm-SL6-[008-010] [2015-09-02T23:49:35.926] agent/is_node_resp: node:vm-SL6-008 rpc:1017 : Can't find an address, check slurm.conf [2015-09-02T23:49:35.926] agent/is_node_resp: node:vm-SL6-010 rpc:1017 : Can't find an address, check slurm.conf [2015-09-02T23:49:35.926] agent/is_node_resp: node:vm-SL6-009 rpc:1017 : Can't find an address, check slurm.conf [2015-09-02T23:49:36.130] error: Nodes taurusi2073,vm-SL6-[008-010] not responding

All addresses are propagated via DNS but sporadically, these nodes are marked as unavail, and re-appear.

The slurm.confs are just provisioned throughout the system.
I simply do not understand what "Can't find an address" means.

Thank you,
Ulf

--
___________________________________________________________________
Dr. Ulf Markwardt

Technische Universität Dresden
Center for Information Services and High Performance Computing (ZIH)
01062 Dresden, Germany

Phone: (+49) 351/463-33640      WWW:  http://www.tu-dresden.de/zih

Attachment: smime.p7s
Description: S/MIME Cryptographic Signature

Reply via email to