On Wed, 2026-09-16 at 20:54 -0400, Vivi, Rodrigo wrote: > On Wed, Sep 16, 2026 at 07:10:11PM +0000, Teres Alexis, Alan Previn wrote: > > On Tue, 2026-09-15 at 18:51 -0400, Vivi, Rodrigo wrote: > > > On Tue, Sep 15, 2026 at 04:00:54PM +0000, Teres Alexis, Alan Previn wrote: > > > > On Tue, 2026-09-15 at 10:34 +0300, Nikula, Jani wrote: > > > > > On Mon, 14 Sep 2026, Alan Previn <[email protected]> > > > > > wrote: > > > > alan:snip > > > > > > int xe_mmio_wait32(struct xe_mmio *mmio, struct xe_reg reg, u32 > > > > > > mask, u32 val, u32 timeout_us, > > > > > > - u32 *out_val, bool atomic) > > > > > > + u32 *out_val) > > > > > > { > > > > > > - return __xe_mmio_wait32(mmio, reg, mask, val, timeout_us, > > > > > > out_val, atomic, true); > > > > > > + u32 read; > > > > > > + int ret; > > > > > > + > > > > > > + ret = poll_timeout_us(read = xe_mmio_read32(mmio, reg), (read & > > > > > > mask) == val, > > > > > > + 10, timeout_us, false); > > > > > > > > > > You probably do need to let the callers pass in the wait too. 10 us > > > > > wait > > > > > with a long timeout is going to be pretty bad. > > > > > > agreed > > > > > > > > > > > > > > > > > alan: okay - perhaps i can make every caller pass in a polling-wait > > > > thats a fraction of their wait time. > > > > (as a starting point since i dont know what's the expected behavior of > > > > every caller). > > > > so perhaps something like "timeout_us << 4" (i.e. 1/16th) but pass in > > > > 10 us if its anything smaller than that > > > > (i.e. smaller than 16 usec). > > > > > > I think we might be complicating this too much.. > > > > > > what about something simpler like: > > > > > > #define XE_MMIO_WAIT_MAX_BACKOFF_US 1000 > > > > > > ... > > > - wait <<= 1; > > > + wait = min_t(s64, wait << 1, XE_MMIO_WAIT_MAX_BACKOFF_US); > > > > > > > > > > alan: i dont understand your this comment on the increasing the "wait by > > x2" in the loop after agreeing with Jani on the earlier statement. > > What I agreed with Jani was that a hardcoded 10us poll interval against > a 2s timeout is bad. That was all. > > > Some historical context: > > > > 1. current baseline code it stands to day IS in violation of linux rules > > for how to use those sleep/delay functions > > right, and I believe that this single line change is enough to fix this > violation. > > The problem is that we double wait every loop 10, 20, 40, 80 … 20480, 40960 > and > never stops doubling. udelay() is only legal up to ~5000us (perhaps 1000?!), > so once the doubling passes that, we're breaking the rule. > > Same on the sleeping side: it ends up asking usleep_range() to sleep 1.3 s, > which is not what that function is for. > > The one line just stops the doubling at 1000us: > > wait = min_t(s64, wait << 1, 1000); > > After that, the wait can be 10, 20, 40 … up to 1000, and then it stays at > 1000 forever. It can never reach 20480 or 1300000. > > So the illegal value never gets passed to udelay() or usleep_range() — ever. > > > 2. my initial revs on fixing this was to minimize the changes so its not > > complicated by simply fixing the code in place. > > I believe this simple line alings with your v1, but just simpler. > > > 3. Jani said we really should use the proper linux kernel helpers: > > poll_timeout_us / poll_timeout_us_atomic. > > 4. Those helpers have the "timeout_us" and the "intra-loop-wait-us" > > period. I hardcoded to 10 usec. > > - Jani said i should not hardcode and ensure all up-the-stack callers > > of the xe_mmio_wait32 function passes in the intra-wait-loop > > value. > > - Then you (Rodrigo) agreed with his request but go on to propose going > > back to exponential 2x intra-wait-loop. > > - but that contradicts Jani's request u agreed to and also that > > means we implement the intra-wait-loop? (i.e. dont use the > > proper linux helper?) > > sorry for not being clear on my previous response. Perhaps we have a cleaner > end code with the poll_timeout_us and poll_timeout_us_atomic indeed. > > Then perhaps we have this single fixes patch and do other attempts in get > cleaner loops? > > Thanks, > Rodrigo. > alan: I think i understand what u mean. But i believe a single line patch will NOT suffice to maintain existing behavior for the non-atomic case. In the old code, if caller was requesting for non-atomic, 'wait' would continue doubling in usleep_range (which has no overflow limit for its params) so we actually CAN keep doubling. Thus callers that have very large timeouts, we'll be saving on CPU cycles as it increases. But with a 1 ms change as you proposed, we'd be increasing CPU cycles to get us to the end. Ofc I agree that doubling forever is not ideal even for the non-atomic usleep_range case, (kernel documentation stating that its good for both short and long timeouts), however kernel documentation DOES state it's an expensive operation with the hrttimer, which is why i wanted to allow the doubling of 'wait' for usleep_range case (like original code) but cap it to 25% of the caller provided final timeout. (while capping the atomic version to MAX_UDELAY_MS for short delays or the same 25% for long delays). This will ensure the overall behavior in terms of cpu cyles and expensive hrttimer calls are similiar as before AND also fixing in accordance to the rules. Actually if u see rev3, this is what i mean Its not 1 line but its much simpler fix and does I believe it does align with your vision.
> > > > ...alan > > > > > > > Also the Fixes tag is not the right one... the bug was there before... > > > > > > > > > > > ...alan > >
