Hi
Am 29.09.26 um 23:50 schrieb Maíra Canal:
Hi Thomas,
On 29/09/26 06:04, Thomas Zimmermann wrote:
(cc'ing v3d maintainers Maira and Melissa)
Hi
Am 23.09.26 um 18:00 schrieb Fabio Piparo:
Hi Thomas,
Thanks for taking a look.
On 14.09.26 10:43, Thomas Zimmermann wrote:
I assume you write these frames quickly one after the other? As i2c is
really slow, you might write to buffers that are still being
transferred
in the background.
Do you read back the status of the page flips? DRM should tell you
when
it has completed transferring a frame. See [1]
The test renders with the GPU: two dumb buffers allocated on the
ssd130x device and shared with v3d over PRIME. It runs at 10 fps.
Before each flip it calls glFinish(), then flips with
DRM_MODE_PAGE_FLIP_EVENT and waits for the event before drawing into
the other buffer. The same test drawing with the CPU instead comes out
clean.
A number of things come to my mind.
- Dumb buffers (from ssd130x) are not meant for HW rendering. They
are generally for software rendering only.
This reminded me of an issue I recently saw involving GPU rendering to
dumb buffers when combining simpledrm and Panthor. It had the same
symptom: this trail of glitches.
After investigating that issue, I noticed that, on platforms where the
GPU isn't DMA-coherent, when a shmem GEM buffer is exported via dma-
buf to a GPU for rendering, the CPU can read stale data from its cache
when later accessing the buffer, which manifests as rendering artifacts.
Exactly what we're seeing here. Thanks for the analysis.
So, the GPU writes reach memory, but the CPU blit reads stale cache
lines. drm_gem_fb_begin_cpu_access() only syncs imported buffers, and
GEM dma_buf_ops have no begin/end_cpu_access, so nothing invalidates the
CPU cache before the blit and the CPU reads stale cached data.
I have a patch implementing begin/end_cpu_access to
drm_gem_prime_dmabuf_ops, which fixed the issue. However, I'm not sure
we would like to support this use case upstream because, as you
mentioned, dumb buffers are usually used for software-rendering only.
Before we fix anything, I think we should talk to someone with
dma-buf/PRIME credentials. As I outlined, the ideomatic pattern is a
producer-consumer relationship and HW rendering into dumb-buffers is not
supported. Those buffers should have been allocated on the v3d side.
IMHO we should write down these rules in the PRIME documentation and
(soft-)enforce them in the implementation.
- Hence, the ideomatic use for sharing graphics buffers is to model a
producer-consumer relationship. Export the buffer from the device
that generates the frame and import it to the buffer that displays it.
- Since you're compositor's sharing happens in the opposite
direction, ssd130x might not sync correctly. (I'm not sure of v3d
requires a dedicated flush.)
- Or maybe v3d needs to do a final cache flush after drawing to
imported buffers.
I don't think so... Once the job completes, the GPU writes are in
memory. I believe the stale copy is in the CPU cache, so the
invalidation should happen on the CPU side.
OTOH it's not the exporter's (ssd130x here, but really _any_ driver) job
to take care of it either. :/
Best regards
Thomas
Best regards,
- Maíra
If nothing else helpers, we could use map_wc for shared buffers in
ssd130x, but it seems like papering over something else.
Best regards
Thomas
I'm happy to send the test program if it would help.
Best regards,
Fabio
--
--
Thomas Zimmermann
Graphics Driver Developer
SUSE Software Solutions Germany GmbH
Frankenstr. 146, 90461 Nürnberg, Germany, www.suse.com
GF: Stefan Gaiser, Jochen Jaser, Abhinav Puri, (HRB 36809, AG Nürnberg)