Hi Vincent,
Thanks for sending this! On Mon, 5 Aug 2024 18:32:23 +0100 Vincent Donnefort <[email protected]> wrote: > The growing set of features supported by the hypervisor in protected > mode necessitates debugging and profiling tools. Tracefs is the > ideal candidate for this task: > > * It is simple to use and to script. > > * It is supported by various tools, from the trace-cmd CLI to the > Android web-based perfetto. > > * The ring-buffer, where are stored trace events consists of linked > pages, making it an ideal structure for sharing between kernel and > hypervisor. > > This series introduces a method to create events and to generate them > from the hypervisor (hyp_enter/hyp_exit given as an example) as well as > a Tracefs user-space interface to read them. > > A presentation was given on this matter during the tracing summit in > 2022. [1] > > 1. ring-buffer > -------------- > > To setup the per-cpu ring-buffers, a new interface is created: > > ring_buffer_writer: Describes what the kernel needs to know about the > writer, that is, the set of pages forming the > ring-buffer and a callback for the reader/head > swapping (enables consuming read) > > ring_buffer_reader(): Creates a read-only ring-buffer from a > ring_buffer_writer. > > To keep the internals of `struct ring_buffer` in sync with the writer, > the meta-page is used. It was originally introduced to enable user-space > mapping of the ring-buffer [1]. In this case, the kernel is not the > producer anymore but the reader. The function to read that meta-page is: > > ring_buffer_poll_writer(): > Update `struct ring_buffer` based on the writer > meta-page. Wake-up readers if necessary. > > The kernel has to poll the meta-page to be notified of newly written > events. > > 2. Tracefs interface > -------------------- > > The interface is a hyp/ folder at the root of the tracefs mount point. > This folder is like an instance and you'll find there a subset of the > regular Tracefs user-space interface: > > hyp/ Hmm, do we really need to shorten it? Why not just call it "hypervisor". I mean tab completion helps with the typing. > buffer_size_kb > trace_pipe > trace_pipe_raw > trace > per_cpu/ > cpuX/ > trace_pipe > trace_pipe_raw > events/ > hyp/ > hyp_enter/ > enable > id > > Behind the scenes, kvm/hyp_trace.c must rebuild the tracing hierarchy > without relying on kernel/trace/trace.c. This is due to fundamental > differences: > > * Hypervisor tracing doesn't support trace_array's system-specific > features (snapshots, tracers, etc.). > > * Logged event formats differ (e.g., no PID in hypervisor > events). > > * Buffer operations require specific hypervisor interactions. > > 3. Events > --------- > > In the hypervisor, "hyp events" can be generated with trace_<event_name> > in a similar fashion to what the kernel does. They're also created with > similar macros than the kernel (see kvm_hypevents.h) > > HYP_EVENT("foboar", > HE_PROTO(void), > HE_STRUCT(), > HE_ASSIGN(), > HE_PRINTK(" ") > ) > > Despite the apparent similarities with TRACE_EVENT(), those macros > internally differs: they must be used in parallel between the hypervisor > (for the writing part) and the kernel (for the reading part) which makes > it difficult to share anything with their kernel counterpart. > > Also, events directory isn't using eventfs. > > 4. Few limitations: > ------------------- > > Non consuming reading of the buffer isn't supported (i.e. cat trace) due > to the lack of support in the ring-buffer meta-page. Hmm, I don't think it should be hard to support that. I've been looking into it for the user mapping. But that can be added later. For now, perhaps "cat trace" just returns -EPERM? > > Tracing must be stopped for the buffer to be reset. i.e. (echo 0 > > tracing_on; echo 0 > trace) Hmm, why this? I haven't looked at the patches yet, but why can't the write to trace just stop tracing and re-enable it after the reset? > -- Steve
