Hi Gabriele, On Fri, Jul 24, 2026 at 09:20:44AM +0800, Gabriele Monaco wrote: > On Wed, 2026-07-22 at 12:57 +0800, Chao Liu wrote: > > The rtapp sleep monitor accepts epoll_wait() as a valid sleeping > > reason, but the prose only discusses clock_nanosleep and futexes. Add > > epoll_wait to the list of valid sleeping reasons. > > > > ABORT_SLEEP represents a task restoring TASK_RUNNING before entering > > the scheduler, but its meaning is not described. Explain why the > > aborted sleep attempt is valid. > > > > This RFC is based on Nam Cao's pending "rv: rtapp monitor update" v2 > > series: > > > > https://lore.kernel.org/r/[email protected] > > > > Signed-off-by: Chao Liu <[email protected]> > > --- > > Thanks for the contribution! I have a couple of comments, we want to be > assertive in this docs: in general, let's not say something like "this is > valid > because the monitor thinks it's valid" but "this is valid because it is > rt-safe > for reason X". Thanks for the review. :)
Agreed. In v2, I will update the documentation and the commit message body to say directly why each case is safe for real-time use. > > > Documentation/trace/rv/monitor_rtapp.rst | 6 ++++++ > > 1 file changed, 6 insertions(+) > > > > diff --git a/Documentation/trace/rv/monitor_rtapp.rst > > b/Documentation/trace/rv/monitor_rtapp.rst > > index 238b59395ff5..234ed4d581ac 100644 > > --- a/Documentation/trace/rv/monitor_rtapp.rst > > +++ b/Documentation/trace/rv/monitor_rtapp.rst > > @@ -67,6 +67,8 @@ thread to sleep for one of the following reasons: > > variables as safe for real-time. As an alternative, the librtpi library > > exists to provide a conditional variable implementation that is correct > > for > > real-time applications in Linux. > > + - Real-time thread waiting for events using `epoll_wait`, > > > + which the monitor accepts as a valid sleeping reason for real-time tasks. > > Not adding much value, try: > > which is a real-time-safe syscall for sleeping as it uses PI-aware locking. > Agreed. I will use this wording in v2. > > > > Beside the reason for sleeping, the eventual waker should also be > > real-time-safe. Namely, one of: > > @@ -114,6 +116,10 @@ The monitor's specification is:: > > ALLOWLIST = BLOCK_ON_RT_MUTEX > > or FUTEX_LOCK_PI > > > > +`ABORT_SLEEP` represents a task restoring its state to `TASK_RUNNING` > > before > > +entering the scheduler. In this case, the task does not actually block, > > > > + so the monitor treats the aborted sleep attempt as valid. > > This sentence is a bit vague. Indeed an /aborted/ sleep is a valid wakeup, as > if > the task woke up itself without even sleeping, saying "the monitor treats the > attempt as valid" adds no value. > > You could rephrase it to (full paragraph for clarity): > > `ABORT_SLEEP` represents a task restoring its state to `TASK_RUNNING` before > entering the scheduler. In this case, the task does not actually block, so it > is > back to runnable without any wakeup sequence unsafe for real-time. > > What do you think? > Yes, this is clearer. I will use the proposed paragraph in v2. To check my understanding, I also put together a small eventfd/epoll test and ran it on arm64 and x86_64. The test code, configurations, patches, serial traces, and reproduction steps are available here: https://github.com/zevorn/rtapp-sleep-monitor-lab Test environments ----------------- +---------+--------------------------+---------------+------+------------+ | Guest | Host CPU | QEMU/accel | vCPU | Kernel | +---------+--------------------------+---------------+------+------------+ | arm64 | Apple M5 Pro | 11.0.0/HVF | 2 | PREEMPT_RT | | x86_64 | Intel Xeon Platinum 8163 | 11.0.0/KVM | 2 | PREEMPT_RT | +---------+--------------------------+---------------+------+------------+ All three guest kernels used the same v7.2-rc4-based source baseline. Test design ----------- Phase 1 (P1): P1 runs one epoll wait per iteration, for 2048 iterations, on two vCPUs: - vCPU 0 runs the SCHED_FIFO/50 waiter (`PID_WAITER`). It starts each iteration, waits on epoll, reads the eventfd when epoll returns, and acknowledges completion. - vCPU 1 runs the SCHED_FIFO/60 writer (`PID_WRITER`). It waits for the start signal, uses a different short delay in each iteration, writes the eventfd, and waits for the acknowledgement. The per-iteration flow is: vCPU 0: waiter (PID_WAITER) vCPU 1: writer (PID_WRITER) -------------------------------- --------------------------------- set go = i |------------------------> wait for go == i enter epoll wait apply delay(i) |<------------------------ write(eventfd) epoll_wait() returns read(eventfd) set ack = i |------------------------> wait for ack == i Because the unmodified arm64 monitor does not recognize the `epoll_pwait` entry, each wait that reaches the monitored sleep path produces a P1 error. The test makes 2048 wait attempts, but an iteration that sees readiness before setting `TASK_INTERRUPTIBLE` does not produce a `SLEEP` event. The PIDs are assigned dynamically. During P1, the `event_sleep` and `error_sleep` filters use `pid == PID_WAITER`; `PID_WRITER` only supplies the event. Separate vCPUs let the event arrive while the waiter is preparing to sleep, covering both normal blocking and ABORT_SLEEP. The writer has the higher priority, so the normal wakeup is RT-safe. Phase 2 (P2): P2 is the negative control. A fresh task (`PID_NEGATIVE`) on vCPU 0 performs one relative clock_nanosleep() without TIMER_ABSTIME. The trace filters are switched to `pid == PID_NEGATIVE`. The expected error_sleep confirms that the monitor is active. The RV monitor tracks tasks, not vCPUs. PID 1 enables the sleep monitor and the `event_sleep` and `error_sleep` events. After each phase, PID 1 stops tracing and prints the trace buffer; `scripts/summarize.sh` counts the records in the serial log. For example, the tests can be run and summarized with: $ make initramfs-arm64 $ KERNEL_IMAGE=/path/to/Image ACCEL=hvf \ ./scripts/run-qemu.sh arm64 $ ./scripts/summarize.sh serial-arm64.log $ make initramfs-x86_64 $ KERNEL_IMAGE=/path/to/bzImage ACCEL=kvm \ ./scripts/run-qemu.sh x86_64 $ ./scripts/summarize.sh serial-x86_64.log The complete build steps, including the arm64 mapping variant, are in docs/build-and-run.md. Results ------- The first two rows compare arm64 without and with the experimental mapping. The x86_64 row is the unmodified cross-check using its native `__NR_epoll_wait`. The arm64 rows enter through `epoll_pwait`. Each case was run five times. `M [L-H]` means median `[minimum-maximum]`; a single value means all five runs had that result. +----------------+------------------+------------+---------------+--------+ | Case | P1 err | Ep abort | PI block | P2 err | +----------------+------------------+------------+---------------+--------+ | arm64 baseline | 2043 [2031-2047] | 0 | 0 | 1 | | arm64 mapping | 0 | 26 [20-28] | 355 [308-399] | 1 | | x86_64 native | 0 | 31 [25-92] | 201 [147-204] | 1 | +----------------+------------------+------------+---------------+--------+ The columns mean: - `P1 err` counts `error_sleep` records from the Phase 1 epoll test. - `Ep abort` counts records containing both `ep_wa` (`EPOLL_WAIT`) and `ab_sl` (`ABORT_SLEEP`). - `PI block` counts records containing both `ep_wa` and `bl_on_rt_mu` (`BLOCK_ON_RT_MUTEX`). - `P2 err` counts `error_sleep` records from the Phase 2 negative control. These are records retained in the trace buffer, not syscall totals or rates. The P1/P2 error outcome was the same in all five runs. The `Ep abort` and `PI block` counts vary with timing. The per-run summaries and validation checks are in results/revalidation-2026-07-24.md in the repository. >From the traces, I see the following path for a correctly recognized epoll wait: sys_enter(epoll_wait or epoll_pwait) | | native x86_64 entry, or arm64 experiment mapping v EPOLL_WAIT = true | v ep_poll() | v TASK_INTERRUPTIBLE -> SLEEP | +-- no event | | | v | schedule() -> RT-safe wakeup -> SCHEDULE_IN | `-- event found by the locked recheck | v skip schedule() | v TASK_RUNNING -> ABORT_SLEEP In the abort case, `ep_wa + sle` (`EPOLL_WAIT + SLEEP`) is followed by `ep_wa + ab_sl` (`EPOLL_WAIT + ABORT_SLEEP`), with no `sch_in` (`SCHEDULE_IN`) in between. This matches your wording: the task becomes runnable again without blocking or going through a wakeup sequence. I also observed `ep_wa + bl_on_rt_mu` while epoll wait was active, so the test did reach the PI-aware locking path on PREEMPT_RT. arm64 syscall-coverage gap -------------------------- The arm64 run also turned up what looks like a syscall-coverage gap. The current epoll support checks only `__NR_epoll_wait`. arm64 does not define that syscall; its epoll wait enters the common implementation through `__NR_epoll_pwait`. I went back to the review of the original epoll support: https://lore.kernel.org/r/[email protected] That reply discusses the `ltl_atom_update()` detail. I could not find any discussion of cross-architecture syscall entry coverage there. In all five unmodified arm64 runs, every retained `SLEEP` record has a matching P1 error. Mapping `__NR_epoll_pwait` to the same atom removes those errors, while P2 still produces the expected error. This points to an instrumentation coverage issue rather than a difference in epoll's real-time-safety. I kept the mapping out of this documentation patch. Once the intended scope for `epoll_pwait`, `epoll_pwait2`, and compat syscalls is clear, I can send a separate functional patch. For reference, this is the experiment-only change: diff --git a/kernel/trace/rv/monitors/sleep/sleep.c b/kernel/trace/rv/monitors/sleep/sleep.c index aa5a984853b5..bf9ca4b5414e 100644 --- a/kernel/trace/rv/monitors/sleep/sleep.c +++ b/kernel/trace/rv/monitors/sleep/sleep.c @@ -136,8 +136,11 @@ static void handle_sys_enter(void *data, struct pt_regs *regs, long id) break; #ifdef __NR_epoll_wait case __NR_epoll_wait: +#endif +#ifdef __NR_epoll_pwait + case __NR_epoll_pwait: +#endif ltl_atom_update(current, LTL_EPOLL_WAIT, true); break; -#endif } } Thanks, Chao
