Our problem has disappear when the number of jobs has dicreased.
We are not sure about what is happening in our cluster. Everything was
working without any problem. Some users started a new research working
with short executions, with hundreds of jobs each one, and our slurmdbd
started to die.
We have decreased the number of jobs in our system and now Slurm works
as previously.
Our version is 2.4.3, the same Marcin Stolarek uses as reported in the
"Problems with slurm during many sacctmgr" post.
It could be the same problem?
Thanks in advance,
Pablo Sanz
--
Pablo Sanz Mercado
Director Técnico
Centro de Computación Científica
Universidad Autónoma de Madrid