On Thu, Aug 13, 2026 at 07:25:20PM +0100, Vincent Donnefort wrote:
> On Fri, Aug 07, 2026 at 05:54:27PM +0200, Thierry Reding wrote:
> > From: Thierry Reding <[email protected]>
> > 
> > NVIDIA Tegra SoCs commonly define a Video-Protection-Region, which is a
> > region of memory dedicated to content-protected video decode and
> > playback. This memory cannot be accessed by the CPU and only certain
> > hardware devices have access to it.
> > 
> > Expose the VPR as a DMA heap so that applications and drivers can
> > allocate buffers from this region for use-cases that require this kind
> > of protected memory.
> > 
> > VPR has a few very critical peculiarities. First, it must be a single
> > contiguous region of memory (there is a single pair of registers that
> > set the base address and size of the region), which is configured by
> > calling back into the secure monitor. The memory region also needs to
> > quite large for some use-cases because it needs to fit multiple video
> > frames (8K video should be supported), so VPR sizes of ~2 GiB are
> > expected. However, some devices cannot afford to reserve this amount
> > of memory for a particular use-case, and therefore the VPR must be
> > resizable.
> > 
> > Unfortunately, resizing the VPR is slightly tricky because the GPU found
> > on Tegra SoCs must be in reset during the VPR resize operation. This is
> > currently implemented by freezing all userspace processes and calling
> > invoking the GPU's freeze() implementation, resizing and the thawing the
> > GPU and userspace processes. This is quite heavy-handed, so eventually
> > it might be better to implement thawing/freezing in the GPU driver in
> > such a way that they block accesses to the GPU so that the VPR resize
> > operation can happen without suspending all userspace.
> > 
> > In order to balance the memory usage versus the amount of resizing that
> > needs to happen, the VPR is divided into multiple chunks. Each chunk is
> > implemented as a CMA area that is completely allocated on first use to
> > guarantee the contiguity of the VPR. Once all buffers from a chunk have
> > been freed, the CMA area is deallocated and the memory returned to the
> > system.
> 
> Hi,
> 
> I believe we (the Android team) are trying to solve similar issue to yours: 
> Arm
> CPUs can still speculatively read memory after it has been transitioned to the
> Secure state, as long as they retain a cacheable mapping to it.
> 
> As modifying the direct mapping is difficult, the workaround ended up in the
> hypervisor which unmaps the pages from the host stage-2 on intercepted FF-A 
> Lend
> invocations.
> 
> If convenient, this is nonetheless the wrong place to do it. Scattering the 
> host
> stage-2 is really terrible for performance and we would like to move it where 
> it
> should be, directly into the kernel...
> 
> This is what I thought was the attempt in v3, but I am now a bit confused
> because I see you are using set_direct_map_invalid_noflush()
> set_direct_map_default_noflush(), but I am not sure that works if rodata=full 
> is
> not set? So does the memory for the NVIDIA IP still need to be unmapped?

Yeah, this currently relies on the circumstances being such that
can_set_direct_map() returns true, and in the case where we want to use
the resizable VPR functionality, we're going to have page-granular
mappings anyway.

> If so, how about having an option where the CMA allocation is backed by a 
> direct
> map with the same page granularity, or at least a smaller and aligned granule?
> 
> With that, we know that for whatever CMA allocation we do, we can safely unmap
> from the direct map without risking splitting blocks. On CMA free, we can 
> safely
> remap into the direct map as I do not believe we coalesce yet.

This sounds intriguing. For VPR we could possibly make the size a
multiple of the memblock size (or whatever might be appropriate). If we
can create a page-granular linear mapping specifically for that region,
that'd be ideal. I don't know if the linear mapping can be subdivided in
this fashion, though.

> (This alignment is not necessary for upcoming systems with BBML3)
> 
> However, to implement this, we'd need to split memblocks and I believe the 
> only
> way at the moment is temporarily mark it as "nomap" which didn't seem very
> popular in the comments on v3.

Could this be simplified if this type of allocation is always memblock
aligned?

Looking at map_mem(), it seems like we could add some sort of special-
casing in the for_each_mem_range() block to check if the memory is VPR
(or generic, page-granular carveout, or whatever we want to call it) and
set NO_BLOCK_MAPPINGS | NO_CONT_MAPPINGS in that case.

If that works, set_direct_map_*() could be enhanced to detect such cases
and always work. Or perhaps a more specific API could be introduced.

Thierry

Attachment: signature.asc
Description: PGP signature

Reply via email to