Check your SlurmdLogFile. You might also increase SlurmdDebug for more info.
Quoting Chris Read <[email protected]>: > Greetings all... > > Due to some bugs in our user code we've not managed to track down yet, we > often get bursts of nodes starting to drain due to "batch job complete > failure". > > I've not found a way to disable this behaviour, and there's no trigger on > draining to allow us to automatically resume the node. At the moment when > this happens we simply resume the node by hand and they're good for days. > > I'm also not 100% sure on what our code is doing that could trigger this. > > Anyone able to shed some light on the best way to fix this while we try get > to the bottom of the root cause? > > Thanks in advance... > > Chris >
