Metastarx opened a new issue, #11305:
URL: https://github.com/apache/rocketmq/issues/11305
### Background
`LmqPrefixIndex.forEachLmqByPrefix()`
(broker/src/main/java/org/apache/rocketmq/broker/lite/LmqPrefixIndex.java:78-96)
takes `rwLock.readLock()` at line 85 and then runs the caller's visitor at
line 90 while the lock is still held; the lock is only released in the
`finally` block after the whole traversal finishes.
The visitor is caller code that can do unbounded work.
`AbstractLiteLifecycleManager.forEachLiteTopicByPrefix` (154-163) passes a
lambda that calls `getMaxOffsetInQueue(lmqName)` for every entry, and
`LiteEventDispatcher.doFullDispatchForWildcardGroup` (272-284) runs
`ConsumerOffsetManager.queryOffset()`, subscription lookups, event-queue offers
and long-poll wakeups for every matched lmq. A wildcard group matches every lmq
of a parent topic, so the read lock is held for the entire scan instead of a
short copy of a few keys.
Two consequences:
1. `onLmqCreate` (94-95) and `onLmqDelete` (101-102) need the write lock,
and `onLmqCreate` is reached from `NotifyMessageArrivingListener.arriving` on
the store's `ReputMessageService` thread. While a wildcard dispatch is
iterating, the reput thread blocks on the write lock, so consume-queue
dispatch, long-poll notification and pop triggers stop advancing for every
topic on the broker. `ReentrantReadWriteLock` is non-fair by default, so a
stream of prefix scans can additionally starve the writer.
2. The lock cannot be upgraded, so any callback that transitively calls
`add()`/`remove()` on the same thread requests the write lock while holding the
read lock and hangs forever. `cleanByParentTopic` (224-230) already carries a
collect-then-delete workaround with a "nesting causes deadlock" comment, and
the javadoc of `AbstractLiteLifecycleManager` (139, 152) documents the same
restriction, so the invariant is enforced only by convention.
### How to reproduce
Unit level: build an `LmqPrefixIndex`, add two lmq names under one parent
topic, and call `forEachLmqByPrefix(prefix, name -> { index.remove(name);
return true; })` from a single thread. The call never returns.
Load level: start a broker with a lite parent topic that holds many lmqs,
register a wildcard lite group, and trigger a full dispatch. `jstack` on the
reput thread shows it parked in `LmqPrefixIndex.add` waiting for the write
lock, and `dispatchBehindBytes` keeps growing.
### Expected
A traversal should work on a snapshot so that caller code never runs while
the index lock is held.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]