[
https://issues.apache.org/jira/browse/CXF-9255?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Freeman Yue Fang updated CXF-9255:
----------------------------------
Fix Version/s: 3.6.14
> WS-RM: retransmissions of messages cached in the same millisecond may run out
> of order and stall in-order delivery until the receive timeout
> --------------------------------------------------------------------------------------------------------------------------------------------
>
> Key: CXF-9255
> URL: https://issues.apache.org/jira/browse/CXF-9255
> Project: CXF
> Issue Type: Improvement
> Components: WS-* Components
> Reporter: Freeman Yue Fang
> Assignee: Freeman Yue Fang
> Priority: Major
> Fix For: 3.6.14, 4.2.5, 4.1.10
>
>
> {panel:title=Symptom}
> The in-order one-way tests in {{systests/ws-rm}} time out sometimes, and more
> frequently on a fast machine :
> - {{DeliveryAssuranceOnewayTest.testExactlyOnceInOrder}},
> {{testExactlyOnceInOrderAsyncExecutor}}, {{testInOrder}},
> {{testInOrderAsyncExecutor}}
> - {{MessageCallbackOnewayTest.testExactlyOnceInOrder}}
> {noformat}
> DeliveryAssuranceOnewayTest.testExactlyOnceInOrder:327->testOnewayExactlyOnceInOrder:343
> ยป TestTimedOut test timed out after 60000 milliseconds
> {noformat}
> Each test passes when run on its own. Whether it fails depends on timing, not
> on test order.
> {panel}
> h3. Root cause
> {{RetransmissionQueueImpl.ResendCandidate}} schedules its first resend at
> {{System.currentTimeMillis() + baseRetransmissionInterval}}. Each candidate
> gets its own {{TimerTask}} on the single shared {{java.util.Timer}} of the
> {{RMManager}}. {{java.util.Timer}} gives no ordering guarantee for tasks due
> at the same time.
> When two messages of the same sequence are cached in the same millisecond,
> which is common on fast hardware, the resend of the later message can run
> first. Using the tests' scenario (message 2 lost by {{MessageLossSimulator}},
> in-order delivery):
> Message 3 is held by the destination until message 2 arrives, so the client
> call for message 3 waits for its HTTP response.
> The timer fires the resend of message 3 before the resend of message 2. The
> resend runs synchronously on the timer thread ({{SynchronousExecutor}}), or
> on the client's single-thread executor in the {{*AsyncExecutor}} variants,
> and blocks for the same reason.
> Message 2 is never resent while that thread is blocked. Everything waits for
> the HTTP receive timeout (60 s by default) before message 2 is finally resent.
> In production this means the same thing can happen outside the tests: with
> in-order delivery, a later message resent ahead of an earlier one stalls
> retransmission until the receive timeout.
> h3. Evidence
> Thread dumps of a hung test show the client blocked in
> {{HttpClientHTTPConduit.getResponse()}} while all of the server's Jetty
> threads are idle. RM FINE logs show the resend order in every run:
> ||Run||Messages 2 and 3 cached||Resend order||Result||
> |failed (4 of 4 runs)|same millisecond (e.g. .641 / .641)|1, 3 (2 only ~60 s
> later)|timeout|
> |passed|different milliseconds (e.g. .396 / .397)|1, 2, 3|~4 s|
> Switching to the {{URLConnection}} conduit, or turning off HttpClient
> sharing, doesn't change the result.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)