Hi Thomas,
On 29/09/26 06:04, Thomas Zimmermann wrote:
(cc'ing v3d maintainers Maira and Melissa)
Hi
Am 23.09.26 um 18:00 schrieb Fabio Piparo:
Hi Thomas,
Thanks for taking a look.
On 14.09.26 10:43, Thomas Zimmermann wrote:
I assume you write these frames quickly one after the other? As i2c is
really slow, you might write to buffers that are still being transferred
in the background.
Do you read back the status of the page flips? DRM should tell you when
it has completed transferring a frame. See [1]
The test renders with the GPU: two dumb buffers allocated on the
ssd130x device and shared with v3d over PRIME. It runs at 10 fps.
Before each flip it calls glFinish(), then flips with
DRM_MODE_PAGE_FLIP_EVENT and waits for the event before drawing into
the other buffer. The same test drawing with the CPU instead comes out
clean.
A number of things come to my mind.
- Dumb buffers (from ssd130x) are not meant for HW rendering. They are
generally for software rendering only.
This reminded me of an issue I recently saw involving GPU rendering to
dumb buffers when combining simpledrm and Panthor. It had the same
symptom: this trail of glitches.
After investigating that issue, I noticed that, on platforms where the
GPU isn't DMA-coherent, when a shmem GEM buffer is exported via dma-
buf to a GPU for rendering, the CPU can read stale data from its cache
when later accessing the buffer, which manifests as rendering artifacts.
So, the GPU writes reach memory, but the CPU blit reads stale cache
lines. drm_gem_fb_begin_cpu_access() only syncs imported buffers, and
GEM dma_buf_ops have no begin/end_cpu_access, so nothing invalidates the
CPU cache before the blit and the CPU reads stale cached data.
I have a patch implementing begin/end_cpu_access to
drm_gem_prime_dmabuf_ops, which fixed the issue. However, I'm not sure
we would like to support this use case upstream because, as you
mentioned, dumb buffers are usually used for software-rendering only.
- Hence, the ideomatic use for sharing graphics buffers is to model a
producer-consumer relationship. Export the buffer from the device that
generates the frame and import it to the buffer that displays it.
- Since you're compositor's sharing happens in the opposite direction,
ssd130x might not sync correctly. (I'm not sure of v3d requires a
dedicated flush.)
- Or maybe v3d needs to do a final cache flush after drawing to imported
buffers.
I don't think so... Once the job completes, the GPU writes are in
memory. I believe the stale copy is in the CPU cache, so the
invalidation should happen on the CPU side.
Best regards,
- Maíra
If nothing else helpers, we could use map_wc for shared buffers in
ssd130x, but it seems like papering over something else.
Best regards
Thomas
I'm happy to send the test program if it would help.
Best regards,
Fabio