jerpelea opened a new pull request, #20078:
URL: https://github.com/apache/nuttx/pull/20078

   ## Summary
   
   CONFIG_ESP32S3_TICKLESS hangs forever the first time a task calls a 
sleep/timeout with a fractional-second component of roughly 134ms or more (e.g. 
usleep(500000)) while another task is also pending a timeout.
   
   Root cause: NSEC_2_CTICK() computes ((nsec) * CTICK_PER_USEC) / 
NSEC_PER_USEC. `nsec` (struct timespec's tv_nsec) is a 32-bit `long`, and 
CTICK_PER_USEC is 16 (the S3's systimer runs at 16MHz), so the multiplication 
overflows a 32-bit signed int for any tv_nsec at or above INT32_MAX / 16 
(~134,217,728 ns). The overflowed (negative) result then gets added into 
up_timer_start()'s `uint64_t cpu_ticks`, wrapping around to a value near 
UINT64_MAX. tickless_setcounter() then programs the systimer alarm that many 
ticks in the future -- effectively never -- so nxsched_process_timer() is never 
called and the waiting task sleeps forever.
   
   Reproduced on real esp32s3-xiao hardware: apps/testing/ostest hung 
indefinitely right after starting user_main(), whose first statement is 
usleep(500000). Instrumented up_timer_start() to print its inputs and observed 
exactly the described overflow (cpu_ticks close to UINT64_MAX for 
tv_nsec=510000000). Confirmed root cause is the concurrent-timeout case 
specifically: user_main's usleep() alone works, and ostest_main's own usleep() 
alone works, but the two together (matching ostest's actual task_create() + 
concurrent usleep() pattern) reproduce the hang every time.
   
   Fix: cast to uint64_t before multiplying in all three *_2_CTICK macros, 
forcing 64-bit arithmetic throughout, matching how the CTICK_2_* (division) 
macros are already overflow-safe.
   
   Validated on esp32s3-xiao: with the fix, the full ostest suite (built with 
CONFIG_ESP32S3_TICKLESS=y) runs past the point it used to hang and completes 
end to end.
   
   Note: while testing, ostest's own round-robin test (rr_test) failed near the 
end of the run -- the two same-priority SCHED_RR threads did not appear to 
interleave under tickless. That looks like a separate, likely more 
architectural issue (time-slice preemption needs its own periodic re-arm, 
independent of one-shot sleep timeouts) and is not addressed by this fix; 
filing separately.
   
   ## Impact
   
   RELEASE
   
   ## Testing
   
   CI


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to