On Fri, Aug 28, 2026 at 09:52:51PM +0000, Etienne Perot wrote:
> Since commit b69bb476dee9 ("cgroup: fix race between fork and
> cgroup.kill"), the fork path snapshots the kill_seq of the child's
> future cgroup into kargs->kill_seq, and cgroup_post_fork() SIGKILLs
> the child if that cgroup's kill_seq has changed in the meantime, to
> catch forks racing with a cgroup.kill sweep.
> 
> For CLONE_INTO_CGROUP, however, the snapshot in cgroup_css_set_fork()
> is taken before the target cgroup has been resolved: kargs->cgrp is
> always NULL at this point (it is only set at the end of the function).
> So the "if (kargs->cgrp)" branch is dead code and the snapshot always
> records the kill_seq of the parent's cgroup. cgroup_post_fork() then
> compares it with the kill_seq of the target cgroup, so the child gets
> SIGKILLed whenever the two cgroups have been killed a different number
> of times.
> 
> As a result, once cgroup.kill has been written to a cgroup, every
> child subsequently cloned into it with clone3(CLONE_INTO_CGROUP) is
> killed on the spot, for as long as the cgroup exists: kill_seq is not
> exposed to userspace and never resets.
> 
> Re-snapshot kill_seq from the target cgroup once it has been resolved,
> and drop the dead branch at the early snapshot site.
> 
> This does not reopen the race fixed by b69bb476dee9. For
> CLONE_INTO_CGROUP, everything from the snapshot to the check in
> cgroup_post_fork() runs with cgroup_mutex held, and kill_seq is
> only ever incremented under cgroup_mutex.

Thanks for catching this. Overall looks good. Can you please fix the comment
where kill_seq is defined in the header. Currently it says kill_seq is
serialized by css_set_lock. After your change, it should for normal fork it is
serialized by css_set_lock but for clone3(CLONE_INTO_CGROUP), it is serialized
by cgroup_mutex.

Orthogonally, we have plans to remove cgroup_mutex dependency from cgroup.kill,
so we will need to reevaluate this at that time.

> 
> Fixes: b69bb476dee9 ("cgroup: fix race between fork and cgroup.kill")
> Cc: [email protected]
> Cc: Shakeel Butt <[email protected]>
> Assisted-by: LLM
> Signed-off-by: Etienne Perot <[email protected]>
> ---
>  kernel/cgroup/cgroup.c | 6 ++----
>  1 file changed, 2 insertions(+), 4 deletions(-)
> 
> diff --git a/kernel/cgroup/cgroup.c b/kernel/cgroup/cgroup.c
> index c3a12fee7528..2d532bf2c0c7 100644
> --- a/kernel/cgroup/cgroup.c
> +++ b/kernel/cgroup/cgroup.c
> @@ -6873,10 +6873,7 @@ static int cgroup_css_set_fork(struct 
> kernel_clone_args *kargs)
>       spin_lock_irq(&css_set_lock);
>       cset = task_css_set(current);
>       get_css_set(cset);
> -     if (kargs->cgrp)
> -             kargs->kill_seq = kargs->cgrp->kill_seq;
> -     else
> -             kargs->kill_seq = cset->dfl_cgrp->kill_seq;
> +     kargs->kill_seq = cset->dfl_cgrp->kill_seq;
>       spin_unlock_irq(&css_set_lock);
>  
>       if (!(kargs->flags & CLONE_INTO_CGROUP)) {
> @@ -6940,6 +6937,7 @@ static int cgroup_css_set_fork(struct kernel_clone_args 
> *kargs)
>  
>       put_css_set(cset);
>       kargs->cgrp = dst_cgrp;
> +     kargs->kill_seq = dst_cgrp->kill_seq;
>       return ret;
>  
>  err:
> -- 
> 2.55.0.897.gb25b4bd76c-goog
> 

Reply via email to