Check your SlurmdLogFile.
You might also increase SlurmdDebug for more info.

Quoting Chris Read <[email protected]>:

> Greetings all...
>
> Due to some bugs in our user code we've not managed to track down yet, we
> often get bursts of nodes starting to drain due to "batch job complete
> failure".
>
> I've not found a way to disable this behaviour, and there's no trigger on
> draining to allow us to automatically resume the node. At the moment when
> this happens we simply resume the node by hand and they're good for days.
>
> I'm also not 100% sure on what our code is doing that could trigger this.
>
> Anyone able to shed some light on the best way to fix this while we try get
> to the bottom of the root cause?
>
> Thanks in advance...
>
> Chris
>

Reply via email to