You've configured the system to allocate whole nodes to jobs rather
than sockets or cores, but to oversubscribe each node 4x. You probably
want something like this
SelectType=select/cons_res
SelectTypeParameters=CR_SOCKET
NodeName=... CPUs=12 Sockets=4 CoresPerSocket=12
See:
http://slurm.schedmd.com/cons_res.html
Quoting Peter Van Buren <[email protected]>:
Hello,
I am new to slurm and just started running it on a single node with
4 x 12 core CPU's (48 physical cores total) and 1TB of memory. I set
the following in slurm.conf:
Under scheduling I have:
SchedulerType=sched/backfill
SelectType=select/linear
and under the NodeName section:
CPUs=4 Sockets=4 CoresPerSocket=12 Shared=FORCE
When I run sinfo --format "%z %h"
the output is:
S:C:T SHARE
4:12:1 FORCE:4
I am able to submit and run jobs, but no more than 4 will run at one
time. The 5th job is queued waiting for resources.
It looks like something is setting a maximum of 4 jobs to run
concurrently (maybe one per CPU). I'm not sure where this limitation
is set. There should be plenty of free cores
and memory available. Should I be setting CR_ONE_TASK_PER_CORE?
Is there a way to specify the number of jobs that can be run
concurrently on a single shared node (assuming available resources)?
Thanks,
Peter.
--
Morris "Moe" Jette
CTO, SchedMD LLC
Commercial Slurm Development and Support