lahirujayathilake opened a new pull request, #544:
URL: https://github.com/apache/airavata-custos/pull/544

   When a user was added to an allocation, their SLURM association was created 
almost immediately, but their actual cluster account is provisioned seconds to 
minutes later. If the association landed first, the scheduler cached a failed 
user lookup and rejected every job that user submitted, until an admin manually 
refreshed it. Nothing reported the problem, the association looked correct 
while every job bounced.
   
   ## What changed
   
   - An association is now written only once the cluster account exists.
   - A background reconciler in the Association Mapper continuously checks that 
every active member has the association they should have, and fills in anything 
missing. This runs on a configurable interval.
   - The previous approach waited in-line for provisioning and gave up after a 
fixed deadline, which meant a member could be left permanently without an 
association and no retry. That wait has been removed.
   
   The reconciler is what makes this correct rather than best-effort. Events 
are delivered in memory with no persistence or retry. The reconciler keeps 
checking, so a lost event costs a short delay instead of a broken account. It 
reads the cluster's current state each pass and writes only what differs.
   
   ## Also fixed
   
   - **Per-member limits were being erased.** Two code paths wrote the same 
association differently, one of them without limits. Because the scheduler 
keeps only the last write, adding a member could silently wipe limits that had 
been set for them. All paths now build the same record.
   - **A crash on allocations with no resources.** The code warned about the 
empty case and then used the missing value anyway.
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to