lahirujayathilake opened a new pull request, #544: URL: https://github.com/apache/airavata-custos/pull/544
When a user was added to an allocation, their SLURM association was created almost immediately, but their actual cluster account is provisioned seconds to minutes later. If the association landed first, the scheduler cached a failed user lookup and rejected every job that user submitted, until an admin manually refreshed it. Nothing reported the problem, the association looked correct while every job bounced. ## What changed - An association is now written only once the cluster account exists. - A background reconciler in the Association Mapper continuously checks that every active member has the association they should have, and fills in anything missing. This runs on a configurable interval. - The previous approach waited in-line for provisioning and gave up after a fixed deadline, which meant a member could be left permanently without an association and no retry. That wait has been removed. The reconciler is what makes this correct rather than best-effort. Events are delivered in memory with no persistence or retry. The reconciler keeps checking, so a lost event costs a short delay instead of a broken account. It reads the cluster's current state each pass and writes only what differs. ## Also fixed - **Per-member limits were being erased.** Two code paths wrote the same association differently, one of them without limits. Because the scheduler keeps only the last write, adding a member could silently wipe limits that had been set for them. All paths now build the same record. - **A crash on allocations with no resources.** The code warned about the empty case and then used the missing value anyway. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
