Hi all,
it's been 5 days now without any crash, I just wanted to report what I did. I restarted the working nodes (not enough), upgraded to 2.6.5, and set the "defer" parameter to SchedulerParameters.

I was very happy to see that the upgraded "slurmctld" was able to recover all the jobs, and to pass almost every test in the testsuite.

About the possible causes that lead to the crash:
- I can't think of an inconsistency in gres.conf as it is a shared file for all the working nodes. - A user was identified ssh'ing here and there. There may have been consequences, specially because of being a GPU cluster. We're thinking about using the PAM plugin. - Submitting a large amount of jobs through scripts. I still have to quantify the amount of jobs, etc. but will follow the advise to configure Slurm with the high throughput guidelines.

Thanks to all,
Albert

On 24/01/14 14:04, Albert Solernou wrote:

Hi,
while this could have something to do with my error, I don't think it is
the main problem:
  - slurmctld keeps freezing when bringing jobs from "wait" to "run" in
the queue.
  - the error  [2014-01-24T13:55:47.030] gres_gresid_to_gresname--The
gres_conf_list is NULL!!!, reappeared.
  - the last freeze was so hard that I'll need to "startclean".

I am thinking seriously about upgrading to 2.6.5.

Best regards,
Albert

On 24/01/14 10:13, Dominik Friedrich wrote:
Hi Albert,

have a look at the defer scheduler parameter from the high throughput
guide. This should help in your case.

Best regards
Dominik
________________________________________
Von: Albert Solernou [[email protected]]
Gesendet: Freitag, 24. Januar 2014 10:58
An: slurm-dev
Betreff: [slurm-dev] Re: unresponsive slurmctld

Hi,
It's not that we have too many jobs in the queue, but that some users
submit many quickly through shell scripts.

I'll have a look and see if this sorts the problem out.

Thank you,
Albert

On 23/01/14 16:56, Moe Jette wrote:

There is advise about high-throughput computing here:
http://slurm.schedmd.com/high_throughput.html

While circumstances at each site vary, I would generally recommend a
default_queue_depth value lower than 100 for systems with high rates of
job submissions.

Quoting Ulf Markwardt <[email protected]>:

Hello Albert,

* what is the numer of jobs in the queue?
* did you see some heavy submitting ?

-

Just yesterday, we had 15.000 Jobs in the queue and a user submitting
~15.000 more in a shell script. To cope with that (surmctld
scheduling, but not responding)  we had to modify our slurm.conf. We
now have something like this:

SchedulerType=sched/backfill
SchedulerParameters=default_queue_depth=100,max_job_bf=600,bf_interval=30,bf_max_job_user=10,bf_window=1440,bf_continue



This wave is over now, but I know our biologists are preparing their
next high-throughput experiment...

Regards,
Ulf

--
___________________________________________________________________
Dr. Ulf Markwardt

Dresden University of Technology
Center for Information Services and High Performance Computing (ZIH)
01062 Dresden, Germany

Phone: (+49) 351/463-33640      WWW:  http://www.tu-dresden.de/zih




--
---------------------------------
    Dr. Albert Solernou
    Research Associate
    Oxford Supercomputing Centre,
    University of Oxford
    Tel: +44 (0)1865 610631
---------------------------------
Bull GmbH
Sitz Köln, Amtsgericht Köln, HR B 8173
Ust-Id-Nr.: DE 121965133, WEEE-Reg.-Nr. DE 64193985
Geschäftsführer: Gerd-Lothar Leonhart, Michael Heinrichs, Philippe Miltin
Zentrale:
51149 Köln, Von-der-Wettern-Strasse 27
Telefon: +49 (0) 2203 305-0
Telefax: +49 (0) 2203 305-1699
http://www.bull.de

Bull, Architect of an Open World TM
** Folgen Sie uns auf Twitter: http://twitter.com/bull_de
** Bull Firmenprofil bei XING: https://www.xing.com/companies/bullgmbh



--
---------------------------------
  Dr. Albert Solernou
  Research Associate
  Oxford Supercomputing Centre,
  University of Oxford
  Tel: +44 (0)1865 610631
---------------------------------

Reply via email to