On Fri, Jul 17, 2026 at 11:01:14AM +0000, [email protected] wrote:
> > +   worker = kthread_run_worker_on_cpu(cpu, 0, "cla-dev-worker/%u");
> 
> [Severity: Medium]
> Is it possible this leaks uninitialized memory to user-space? The string
> "cla-dev-worker/%u" expects an integer argument, but the wrapper macro
> kthread_run_worker_on_cpu() only takes the name format and does not forward
> arguments to the underlying kthread_create_worker_on_cpu() function.

kthread_create_worker_on_cpu() adds the cpu number


> [Severity: High]
> When a CPU is hotplugged back online, will this setup path leave the hardware
> out of sync with the domain?

Yes save/restore on CPUHP will be added later


> > +   for (i = 0; i < domain->nr_devs; i++) {
> > +           dev = domain->devs[i];
> > +           call = &dev->call;
> > +
> > +           call->sched.prev_ctx = prev_ctx;
> > +           call->sched.next_ctx = next_ctx;
> > +           call->sched.ctx_id = i;
> > +           call->ret = 0;
> > +
> > +           mutex_lock(&dev->lock);
> > +           if (dev->worker && !dev->broken)
> > +                   kthread_queue_work(dev->worker, &call->switch_ctx);
> > +           mutex_unlock(&dev->lock);
> 
> [Severity: High]
> If dev->worker is NULL because the CPU is offline, does this silently bypass
> the context switch for that device? 

Context cannot be switched on an offline CPU


> > +static void cla_domain_zap(struct cla_domain *domain, struct cla_ctx *ctx)
> > +{
> > +   struct vm_area_struct *vma;
> > +   struct vma_iterator vmi;
> > +
> > +   /*
> > +    * If the context is dying it has already been unmapped, no need to zap
> > +    * it here.
> > +    */
> > +   if (!ctx || cla_ctx_is_dying(ctx))
> > +           return;
> > +
> > +   /*
> > +    * Iterate over vmas in prev_ctx's mm, looking for vmas that map
> > +    * prev_ctx's file. Zap any portions of those vmas that map the domain.
> > +    */
> > +   mmap_read_lock(ctx->key.mm);
> > +   vma_iter_init(&vmi, ctx->key.mm, 0);
> > +   for_each_vma(vmi, vma) {
> 
> [Severity: High]
> Is it safe to acquire the mmap_read_lock() and iterate the maple tree here
> without an active mm reference?
> 
> The context only holds a passive mmgrab() reference. If a concurrent
> unprivileged userspace process exits, mm_users could drop to 0, and
> exit_mmap() might tear down the VMA maple tree. Should this code use
> mmget_not_zero() to prevent a use-after-free while walking the tree?

Yes probably


> > +static void cla_domain_reclaim(struct kthread_work *work)
> > +{
> > +   struct cla_domain *domain;
> > +   struct cla_ctx *ctx, *next;
> > +
> > +   domain = container_of(work, struct cla_domain, reclaim.work);
> > +
> > +   mutex_lock(&domain->lock);
> > +   list_for_each_entry_safe(ctx, next, &domain->dying_ctxs, queue_node) {
> > +           list_del_init(&ctx->queue_node);
> > +           WARN_ON(domain->assigned_ctx == ctx);
> > +           cla_ctx_put(ctx);
> 
> [Severity: Critical]
> If a domain becomes broken, is it possible for the active hardware context to
> be freed while still programmed in the device?
> 
> In cla_domain_schedule_reassignment(), the reassign work is skipped if the
> domain is broken, meaning the context is never uninstalled from the hardware.
> Then here in cla_domain_reclaim(), the dying context is unconditionally freed
> by calling cla_ctx_put(), which drops the PGD and software state.

No, a reference to the mm is kept

Reply via email to