Hello Sebastian, thank you for the detailed report. I have opened https://phabricator.wikimedia.org/T427204 to follow up and investigate
HTH, Filippo On Wed, May 20, 2026 at 5:36 PM Sebastian Berlin via Cloud < [email protected]> wrote: > Quick update on this: I restarted the job with 2 CPUs and again it started > right away. > > *Sebastian Berlin* > Utvecklare/*Developer* > Wikimedia Sverige (WMSE) > > E-post/*E-Mail*: [email protected] > Telefon/*Phone*: (+46) 0707 - 92 03 84 > > > On Wed, 20 May 2026 at 16:50, Sebastian Berlin < > [email protected]> wrote: > >> It looks like I ran into this again. A job has now been pending for 20 >> minutes. I tried another job with 3 CPUs earlier today and that worked >> fine. >> >> When I ran `kubectl describe pods` I got a lot of info I don't >> understand. There is however a warning: >> >>> Warning FailedScheduling 3m58s (x62 over 14m) default-scheduler >>> 0/81 nodes are available: 11 node(s) didn't match Pod's node >>> affinity/selector, 3 node(s) had untolerated taint { >>> node-role.kubernetes.io/control-plane: }, 3 node(s) had untolerated >>> taint {toolforge.org/gateway: true}, 64 Insufficient cpu. preemption: >>> 0/81 nodes are available: 17 Preemption is not helpful for scheduling, 64 >>> No preemption victims found for incoming pod. >> >> >> I attached the full output. The first job is the one not wanting to start. >> >> *Sebastian Berlin* >> Utvecklare/*Developer* >> Wikimedia Sverige (WMSE) >> >> E-post/*E-Mail*: [email protected] >> Telefon/*Phone*: (+46) 0707 - 92 03 84 >> >> >> On Mon, 23 Mar 2026 at 15:27, David Caro <[email protected]> wrote: >> >>> >>> The workers currently have 8 CPU each, and from what I see there's more >>> than >>> enough cpu free to run a 2 cpu job (just did). >>> >>> It is true though that the more resources you request, the harder it >>> will be >>> to allocate (it might even fail to allocate), but I'm not seeing that the >>> cluster was out of capacity at that moment. >>> >>> Next time it happens, you can also try running `kubectl describe pods` >>> to show >>> the events and some more details on why the job was stuck. >>> >>> We are working on a patch to add also more information to the state of >>> the jobs >>> that will help debug quota issues and similar things too, though not yet >>> there. >>> >>> On 03/23 14:34, Sebastian Berlin via Cloud wrote: >>> > Thanks. The job started when I lowered to 2 CPUs. >>> > >>> > This will probably have some impact on the runtime, which is a factor. >>> I'd >>> > still like to know if it's as simple as the extra CPU preventing it >>> from >>> > starting altogether. If so, does that mean the limit is set too high >>> to in >>> > practice? >>> > >>> > *Sebastian Berlin* >>> > Utvecklare/*Developer* >>> > Wikimedia Sverige (WMSE) >>> > >>> > E-post/*E-Mail*: [email protected] >>> > Telefon/*Phone*: (+46) 0707 - 92 03 84 >>> > >>> > >>> > On Mon, 23 Mar 2026 at 12:44, Nicolas Vervelle <[email protected]> >>> wrote: >>> > >>> > > Hi Sebastian, >>> > > >>> > > I also had this kind of problem, it was with jobs requesting a lot >>> of cpu >>> > > resources, 3 like you (which is the default maximum). When I reduced >>> the >>> > > CPU request, it started a lot faster. >>> > > >>> > > So I guess the cluster is too busy and never find a node with enough >>> CPU >>> > > available. >>> > > >>> > > Nicolas >>> > > >>> > > Le lun. 23 mars 2026 à 12:15, Sebastian Berlin via Cloud < >>> > > [email protected]> a écrit : >>> > > >>> > >> A Toolforge job I'm running is stuck. I started it last Friday after >>> > >> having stopped once before for now apparent reason. When I checked >>> today I >>> > >> got: >>> > >> >>> > >>> tools.commonsdb-registry@tools-bastion-15:~$ toolforge jobs show >>> > >>> make-declarations-wlm-2013 >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Job name: | make-declarations-wlm-2013 >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Command: | src/make_declaration.py -t mem:1.0Gi -t cpu:3.0 >>> -t >>> > >>> version:0.1.11 --verbose batch:category-26152515 | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Job type: | one-off >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Image: | >>> tool-commonsdb-registry/tool-commonsdb-registry:latest >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Port: | none >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | File log: | yes >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Output log: | >>> > >>> /data/project/commonsdb-registry/make-declarations-wlm-2013.out >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Error log: | >>> > >>> /data/project/commonsdb-registry/make-declarations-wlm-2013.err >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Emails: | all >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Resources: | mem: 1.0Gi, cpu: 3.0 >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Replicas: | >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Mounts: | all >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Retry: | no >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Timeout: | no >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Health check: | none >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Status: | Running for 3d23m >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >>> | Hints: | Run not attempted yet. Pod in 'Pending' phase. >>> > >>> | >>> > >>> >>> > >>> >>> +---------------+-----------------------------------------------------------------------------------------------------+ >>> > >> >>> > >> >>> > >> As far as I can tell that means it hasn't actually started. It's >>> still >>> > >> waiting for something, but it's unclear what. I restarted it a >>> couple of >>> > >> times today so far, but it never gets past that stage. >>> > >> >>> > >> When I've run similar jobs in the past they sometimes take a while >>> to >>> > >> start. Eventually they have though, so I though maybe they were >>> waiting for >>> > >> resources or something. >>> > >> >>> > >> Can anyone help me get this started or explain if I've misunderstood >>> > >> something? >>> > >> >>> > >> *Sebastian Berlin* >>> > >> Utvecklare/*Developer* >>> > >> Wikimedia Sverige (WMSE) >>> > >> >>> > >> E-post/*E-Mail*: [email protected] >>> > >> Telefon/*Phone*: (+46) 0707 - 92 03 84 >>> > >> _______________________________________________ >>> > >> Cloud mailing list -- [email protected] >>> > >> List information: >>> > >> >>> https://lists.wikimedia.org/postorius/lists/cloud.lists.wikimedia.org/ >>> > >> >>> > > >>> >>> > _______________________________________________ >>> > Cloud mailing list -- [email protected] >>> > List information: >>> https://lists.wikimedia.org/postorius/lists/cloud.lists.wikimedia.org/ >>> >>> >>> -- >>> David Caro >>> SRE - Cloud Services >>> Wikimedia Foundation <https://wikimediafoundation.org/> >>> PGP Signature: 7180 83A2 AC8B 314F B4CE 1171 4071 C7E1 D262 69C3 >>> >>> "Imagine a world in which every single human being can freely share in >>> the >>> sum of all knowledge. That's our commitment." >>> >> _______________________________________________ > Cloud mailing list -- [email protected] > List information: > https://lists.wikimedia.org/postorius/lists/cloud.lists.wikimedia.org/ > -- *Filippo Giunchedi* Staff Site Reliability Engineer Wikimedia Foundation <https://wikimediafoundation.org/>
_______________________________________________ Cloud mailing list -- [email protected] List information: https://lists.wikimedia.org/postorius/lists/cloud.lists.wikimedia.org/
