On 2026-10-01 12:59:09 [+0200], Peter Zijlstra wrote:
> On Thu, Sep 24, 2026 at 10:28:10AM -1000, Tejun Heo wrote:
> > Hello, Shakeel.
> > 
> > On Thu, Sep 24, 2026 at 11:47:04AM -0700, Shakeel Butt wrote:
> > > This series lets a kernel thread say which cgroup it is working for.
> > > That cgroup then sees the CPU time in its cpu.stat and the stalls in
> > > its memory.pressure, and the CPU time comes out of its cpu.max quota.
> > > The first user is the memcg reclaim that runs from high_work.
> > 
> > This doesn't translate to net rx, which is another major source of
> > displaced CPU usage. Switching membership on each packet isn't going to
> > work there. Attribution can't happen that way. We'd much rather count
> > per-cgroup received packets and prorate the CPU consumption. If at all
> > possible, I think it'd be better to adopt an approach which can cover
> > both use cases.
> 
> Ideally RX would be split for each network queue, rather than lumped
> into the one giant softirq that nobody owns.
> 
> I know PREEMPT_RT has been wanting something like that for ages. Not all
> queues are created equal. Some might want RT priority while others
> should definitely not.

What currently kind of works is threaded NAPI. The interrupt wakes the
NAPI thread rather than adding NET_RX to the global flag softirq flags.

I was thinking about making the softirq flags per-thread rather than
per-CPU. This avoid the "catch up" of other raised but unrelated
softirqs.

For now a painless setup is to have "bulk queues" and "real-time" queues
and what gets where is configured via hardware filters.

> Furthermore, without ingress throttling, your RX back charge could
> completely deplete the actual cgroup time quota.
> 
> Anyway, if you get per queue RX processing threads, then you can move
> them into cgroups where so desired.

Sebastian

Reply via email to