Freeman Yue Fang created CXF-9255:
-------------------------------------

             Summary: WS-RM: retransmissions of messages cached in the same 
millisecond may run out of order and stall in-order delivery until the receive 
timeout
                 Key: CXF-9255
                 URL: https://issues.apache.org/jira/browse/CXF-9255
             Project: CXF
          Issue Type: Improvement
          Components: WS-* Components
            Reporter: Freeman Yue Fang


{panel:title=Symptom}
The in-order one-way tests in {{systests/ws-rm}} time out sometimes, and more 
frequently on a fast machine :
- {{DeliveryAssuranceOnewayTest.testExactlyOnceInOrder}}, 
{{testExactlyOnceInOrderAsyncExecutor}}, {{testInOrder}}, 
{{testInOrderAsyncExecutor}}
- {{MessageCallbackOnewayTest.testExactlyOnceInOrder}}

{noformat}
DeliveryAssuranceOnewayTest.testExactlyOnceInOrder:327->testOnewayExactlyOnceInOrder:343
 ยป TestTimedOut test timed out after 60000 milliseconds
{noformat}
Each test passes when run on its own. Whether it fails depends on timing, not 
on test order.
{panel}

h3. Root cause

{{RetransmissionQueueImpl.ResendCandidate}} schedules its first resend at 
{{System.currentTimeMillis() + baseRetransmissionInterval}}. Each candidate 
gets its own {{TimerTask}} on the single shared {{java.util.Timer}} of the 
{{RMManager}}. {{java.util.Timer}} gives no ordering guarantee for tasks due at 
the same time.

When two messages of the same sequence are cached in the same millisecond, 
which is common on fast hardware, the resend of the later message can run 
first. Using the tests' scenario (message 2 lost by {{MessageLossSimulator}}, 
in-order delivery):
Message 3 is held by the destination until message 2 arrives, so the client 
call for message 3 waits for its HTTP response.

The timer fires the resend of message 3 before the resend of message 2. The 
resend runs synchronously on the timer thread ({{SynchronousExecutor}}), or on 
the client's single-thread executor in the {{*AsyncExecutor}} variants, and 
blocks for the same reason.

Message 2 is never resent while that thread is blocked. Everything waits for 
the HTTP receive timeout (60 s by default) before message 2 is finally resent.

In production this means the same thing can happen outside the tests: with 
in-order delivery, a later message resent ahead of an earlier one stalls 
retransmission until the receive timeout.

h3. Evidence

Thread dumps of a hung test show the client blocked in 
{{HttpClientHTTPConduit.getResponse()}} while all of the server's Jetty threads 
are idle. RM FINE logs show the resend order in every run:

||Run||Messages 2 and 3 cached||Resend order||Result||
|failed (4 of 4 runs)|same millisecond (e.g. .641 / .641)|1, 3 (2 only ~60 s 
later)|timeout|
|passed|different milliseconds (e.g. .396 / .397)|1, 2, 3|~4 s|

Switching to the {{URLConnection}} conduit, or turning off HttpClient sharing, 
doesn't change the result.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to