qq619618919 opened a new pull request, #8758:
URL: https://github.com/apache/hadoop/pull/8758

   ## Summary
   
   `EntityGroupFSTimelineStore` (Timeline Service v1.5) keeps its LRU 
entity-group cache in a `Collections.synchronizedMap(LinkedHashMap)`. On 
eviction, `removeEldestEntry()` calls `EntityCacheItem.forceRelease()` — a 
`synchronized` method — while still holding the global `SynchronizedMap` 
monitor. If another thread is inside `refreshCache()` (also `synchronized`, 
doing HDFS reads + JSON parsing + LevelDB writes that can run for minutes), the 
eviction thread blocks on the `EntityCacheItem` lock while holding the global 
map lock, starving every HTTP request thread that needs the cache.
   
   In production this hung the service: 202 threads BLOCKED for 30+ minutes, 
~20,000 CLOSE_WAIT connections, and file-descriptor exhaustion (32,768 limit).
   
   ## Fix
   
   Evict asynchronously: `removeEldestEntry()` enqueues the evicted item onto a 
`ConcurrentLinkedQueue` and returns immediately, so `put()` no longer holds the 
global map lock while waiting on an `EntityCacheItem` lock. A scheduled 
`drainEvictionQueue()` (1s fixed rate) calls `forceRelease()` on the executor 
thread, outside the global map lock. `serviceStop()` drains any remaining 
queued items.
   
   ## JIRA
   
   https://issues.apache.org/jira/browse/YARN-11994
   
   ## Test
   
   - Added `TestEntityGroupFSTimelineStore#testAsyncEvictionReleasesStore`.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to