On Fri, Jul 17, 2026 at 11:01:14AM +0000, [email protected] wrote:
> > + worker = kthread_run_worker_on_cpu(cpu, 0, "cla-dev-worker/%u");
>
> [Severity: Medium]
> Is it possible this leaks uninitialized memory to user-space? The string
> "cla-dev-worker/%u" expects an integer argument, but the wrapper macro
> kthread_run_worker_on_cpu() only takes the name format and does not forward
> arguments to the underlying kthread_create_worker_on_cpu() function.
kthread_create_worker_on_cpu() adds the cpu number
> [Severity: High]
> When a CPU is hotplugged back online, will this setup path leave the hardware
> out of sync with the domain?
Yes save/restore on CPUHP will be added later
> > + for (i = 0; i < domain->nr_devs; i++) {
> > + dev = domain->devs[i];
> > + call = &dev->call;
> > +
> > + call->sched.prev_ctx = prev_ctx;
> > + call->sched.next_ctx = next_ctx;
> > + call->sched.ctx_id = i;
> > + call->ret = 0;
> > +
> > + mutex_lock(&dev->lock);
> > + if (dev->worker && !dev->broken)
> > + kthread_queue_work(dev->worker, &call->switch_ctx);
> > + mutex_unlock(&dev->lock);
>
> [Severity: High]
> If dev->worker is NULL because the CPU is offline, does this silently bypass
> the context switch for that device?
Context cannot be switched on an offline CPU
> > +static void cla_domain_zap(struct cla_domain *domain, struct cla_ctx *ctx)
> > +{
> > + struct vm_area_struct *vma;
> > + struct vma_iterator vmi;
> > +
> > + /*
> > + * If the context is dying it has already been unmapped, no need to zap
> > + * it here.
> > + */
> > + if (!ctx || cla_ctx_is_dying(ctx))
> > + return;
> > +
> > + /*
> > + * Iterate over vmas in prev_ctx's mm, looking for vmas that map
> > + * prev_ctx's file. Zap any portions of those vmas that map the domain.
> > + */
> > + mmap_read_lock(ctx->key.mm);
> > + vma_iter_init(&vmi, ctx->key.mm, 0);
> > + for_each_vma(vmi, vma) {
>
> [Severity: High]
> Is it safe to acquire the mmap_read_lock() and iterate the maple tree here
> without an active mm reference?
>
> The context only holds a passive mmgrab() reference. If a concurrent
> unprivileged userspace process exits, mm_users could drop to 0, and
> exit_mmap() might tear down the VMA maple tree. Should this code use
> mmget_not_zero() to prevent a use-after-free while walking the tree?
Yes probably
> > +static void cla_domain_reclaim(struct kthread_work *work)
> > +{
> > + struct cla_domain *domain;
> > + struct cla_ctx *ctx, *next;
> > +
> > + domain = container_of(work, struct cla_domain, reclaim.work);
> > +
> > + mutex_lock(&domain->lock);
> > + list_for_each_entry_safe(ctx, next, &domain->dying_ctxs, queue_node) {
> > + list_del_init(&ctx->queue_node);
> > + WARN_ON(domain->assigned_ctx == ctx);
> > + cla_ctx_put(ctx);
>
> [Severity: Critical]
> If a domain becomes broken, is it possible for the active hardware context to
> be freed while still programmed in the device?
>
> In cla_domain_schedule_reassignment(), the reassign work is skipped if the
> domain is broken, meaning the context is never uninstalled from the hardware.
> Then here in cla_domain_reclaim(), the dying context is unconditionally freed
> by calling cla_ctx_put(), which drops the PGD and software state.
No, a reference to the mm is kept