[ 
https://issues.apache.org/jira/browse/THRIFT-6392?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=18119882#comment-18119882
 ] 

Sylwester Lachiewicz commented on THRIFT-6392:
----------------------------------------------

Root cause: {{FastestPoolJob}} in 
[lib/d/src/thrift/codegen/async_client_pool.d|https://github.com/apache/thrift/blob/master/lib/d/src/thrift/codegen/async_client_pool.d]
 stops registering completion callbacks at the first child future that has 
already completed, but a child that failed with an RPC fault does not set the 
pool's result, so the remaining children are never counted and the pool's 
future never completes. {{client_pool_test}} hits it in the all-clients-fail 
case when the 1 ms server answers before the pool is constructed. Fix and a 
unit test in [PR #3972|https://github.com/apache/thrift/pull/3972].

> D client_pool_test occasionally hangs until the lib-d job times out
> -------------------------------------------------------------------
>
>                 Key: THRIFT-6392
>                 URL: https://issues.apache.org/jira/browse/THRIFT-6392
>             Project: Thrift
>          Issue Type: Bug
>          Components: D - Library
>            Reporter: Sylwester Lachiewicz
>            Priority: Minor
>          Time Spent: 10m
>  Remaining Estimate: 0h
>
> {{lib/d/test/client_pool_test}} occasionally never finishes, so the lib-d job 
> runs into its 60-minute limit. It happened twice in a row on [PR 
> #3971|https://github.com/apache/thrift/pull/3971], which does not touch lib/d 
> ([job 1|https://github.com/apache/thrift/actions/runs/36332377379], [job 
> 2|https://github.com/apache/thrift/actions/runs/36337045757/job/108670208530]),
>  while {{make -C lib/d check}} took about 2.5 minutes on the other PRs of the 
> same day and on master.
> In both hung runs all 104 unit tests pass, then each of the six server 
> threads logs, three seconds later:
> {noformat}
> src/thrift/server/simple.d:144: Client died unexpectedly: 
> thrift.transport.base.TTransportException@src/thrift/transport/socket.d(340): 
> Timed out
>   thrift.transport.socket.TSocket.read(ubyte[])
>   thrift.transport.buffered.TBufferedTransport.peek()
>   
> thrift.server.simple.TSimpleServer.serve(thrift.util.cancellation.TCancellation)
>   client_pool_test.ServerThread.run()
> {noformat}
> and nothing else is printed until the job is cancelled; the runner then kills 
> an orphaned {{client_pool_test}}. A passing run prints none of these 
> messages. So the test's own client side stopped while holding a connection to 
> every server, and the servers' 3-second {{recvTimeout}} is the only timeout 
> involved.
> The client side of the test has no bound anywhere, so any lost reply waits 
> forever:
> * the synchronous clients are {{TSocket}}s without {{recvTimeout}};
> * the asynchronous tests use {{waitGet()}}, and the implicit {{waitGet}} of 
> {{TFuture}}, instead of {{waitGet(Duration)}};
> * {{main}} ignores the result of {{sem.wait(dur!"seconds"(1))}}, so it goes 
> on even if a server is not listening yet.
> What sets it off is not known yet. Bounding those waits would turn the hang 
> into a test failure that names the call, and a step-level {{timeout-minutes}} 
> on "Run make check for d" in build.yml would stop a hang well before the job 
> limit. Related: THRIFT-4155.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to