Hi Willy,

Thank you for such useful info! I've checked the worst HAProxy nodes and on
every such node all outgoing peers connections are run on the same thread:

echo "show fd" | socat unix-connect:/var/run/haproxy.sock stdio | grep "px="
430 : st=0x010121(cL heopI W:sRa R:srA) tmask=0x40 umask=0x0
owner=0x7f0aa4039250 iocb=0x564c0df95f40(sock_conn_iocb) back=1
cflg=0x10002300 fam=ipv4 lport=22850 rport=1024 px=hap. mux=PASS
ctx=0x7f0aa4039450 xprt=RAW
500 : st=0x010121(cL heopI W:sRa R:srA) tmask=0x40 umask=0x0
owner=0x7f0aa403d4e0 iocb=0x564c0df95f40(sock_conn_iocb) back=1
cflg=0x10002300 fam=ipv4 lport=58906 rport=1024 px=hap mux=PASS
ctx=0x7f0aa403d6e0 xprt=RAW
501 : st=0x010121(cL heopI W:sRa R:srA) tmask=0x40 umask=0x0
owner=0x7f0aa4041770 iocb=0x564c0df95f40(sock_conn_iocb) back=1
cflg=0x10002300 fam=ipv4 lport=54470 rport=1024 px=hap mux=PASS
ctx=0x7f0aa4041970 xprt=RAW
502 : st=0x010121(cL heopI W:sRa R:srA) tmask=0x40 umask=0x0
owner=0x7f0aa4045a00 iocb=0x564c0df95f40(sock_conn_iocb) back=1
cflg=0x10002300 fam=ipv4 lport=48926 rport=1024 px=hap mux=PASS
ctx=0x7f0aa4045c00 xprt=RAW
503 : st=0x010121(cL heopI W:sRa R:srA) tmask=0x40 umask=0x0
owner=0x7f0aa4049c90 iocb=0x564c0df95f40(sock_conn_iocb) back=1
cflg=0x10002300 fam=ipv4 lport=21468 rport=1024 px=hap mux=PASS
ctx=0x7f0aa4049e90 xprt=RAW

echo "show sess" | socat unix-connect:/var/run/haproxy.sock stdio | grep
"proto=?" | grep "epoch=0 " | grep 14h
0x7f0aa402e2c0: proto=? ts=00 epoch=0 age=14h53m calls=10724 rate=1 cpu=0
lat=0 rq[f=848202h,i=0,an=00h,rx=5s,wx=,ax=]
rp[f=80448202h,i=0,an=00h,rx=,wx=,ax=] s0=[8,204048h,fd=-1,ex=]
s1=[8,2000d8h,fd=430,ex=] exp=3s
0x7f0aa402ece0: proto=? ts=00 epoch=0 age=14h53m calls=10724 rate=1 cpu=0
lat=0 rq[f=848202h,i=0,an=00h,rx=5s,wx=,ax=]
rp[f=80448202h,i=0,an=00h,rx=,wx=,ax=] s0=[8,204048h,fd=-1,ex=]
s1=[8,2000d8h,fd=500,ex=] exp=3s
0x7f0aa402f700: proto=? ts=00 epoch=0 age=14h53m calls=10724 rate=1 cpu=0
lat=0 rq[f=848202h,i=0,an=00h,rx=5s,wx=,ax=]
rp[f=80448202h,i=0,an=00h,rx=,wx=,ax=] s0=[8,204048h,fd=-1,ex=]
s1=[8,2000d8h,fd=501,ex=] exp=3s
0x7f0aa4030120: proto=? ts=00 epoch=0 age=14h53m calls=10722 rate=0 cpu=0
lat=0 rq[f=848202h,i=0,an=00h,rx=5s,wx=,ax=]
rp[f=80448202h,i=0,an=00h,rx=,wx=,ax=] s0=[8,204048h,fd=-1,ex=]
s1=[8,2000d8h,fd=502,ex=] exp=2s
0x7f0aa4030b40: proto=? ts=00 epoch=0 age=14h53m calls=10724 rate=1 cpu=0
lat=0 rq[f=848202h,i=0,an=00h,rx=5s,wx=,ax=]
rp[f=80448202h,i=0,an=00h,rx=,wx=,ax=] s0=[8,204048h,fd=-1,ex=]
s1=[8,2000d8h,fd=503,ex=] exp=3s

On one node I was able to rebalance it, but on the node above (and other
nodes) I'm not able to shutdown the sessions:
echo "show sess 0x7f0aa402e2c0" | socat unix-connect:/var/run/haproxy.sock
stdio
0x7f0aa402e2c0: [11/Mar/2022:07:17:02.313221] id=0 proto=?
  flags=0x6, conn_retries=3, srv_conn=(nil), pend_pos=(nil) waiting=0
epoch=0
  frontend=x (id=4294967295 mode=http), listener=? (id=0)
  backend=x (id=4294967295 mode=http) addr=x:22850
  server=<none> (id=0) addr=y:1024
  task=0x7f0aa402e7d0 (state=0x00 nice=0 calls=10741 rate=0 exp=1s
tmask=0x40 age=14h54m)
  si[0]=0x7f0aa402e600 (state=EST flags=0x204048
endp0=APPCTX:0x7f0aa402df90 exp=<NEVER> et=0x000 sub=0)
  si[1]=0x7f0aa402e658 (state=EST flags=0x2000d8 endp1=CS:0x7f0aa4039400
exp=<NEVER> et=0x000 sub=1)
  app0=0x7f0aa402df90 st0=7 st1=0 st2=0 applet=<PEER> tmask=0x40 nice=0
calls=206051832 rate=2860 cpu=0 lat=0
  co1=0x7f0aa4039250 ctrl=tcpv4 xprt=RAW mux=PASS data=STRM
target=PROXY:0x564c0ff53760
      flags=0x10003300 fd=430 fd.state=10121 updt=0 fd.tmask=0x40
      cs=0x7f0aa4039400 csf=0x00008200 ctx=(nil)
  req=0x7f0aa402e2d0 (f=0x848202 an=0x0 pipe=0 tofwd=-1 total=9489764648)
      an_exp=<NEVER> rex=4s wex=<NEVER>
      buf=0x7f0aa402e2d8 data=(nil) o=0 p=0 i=0 size=0
  res=0x7f0aa402e330 (f=0x80448202 an=0x0 pipe=0 tofwd=-1 total=9320288685)
      an_exp=<NEVER> rex=<NEVER> wex=<NEVER>
      buf=0x7f0aa402e338 data=(nil) o=0 p=0 i=0 size=0

echo "shutdown session 0x7f0aa402e2c0" | socat
unix-connect:/var/run/haproxy.sock stdio

echo "show sess 0x7f0aa402e2c0" | socat unix-connect:/var/run/haproxy.sock
stdio
0x7f0aa402e2c0: [11/Mar/2022:07:17:02.313221] id=0 proto=?
  flags=0x6, conn_retries=3, srv_conn=(nil), pend_pos=(nil) waiting=0
epoch=0
  frontend=x (id=4294967295 mode=http), listener=? (id=0)
  backend=x (id=4294967295 mode=http) addr=x:22850
  server=<none> (id=0) addr=x:1024
  task=0x7f0aa402e7d0 (state=0x00 nice=0 calls=10745 rate=1 exp=4s
tmask=0x40 age=14h55m)
  si[0]=0x7f0aa402e600 (state=EST flags=0x204048
endp0=APPCTX:0x7f0aa402df90 exp=<NEVER> et=0x000 sub=0)
  si[1]=0x7f0aa402e658 (state=EST flags=0x2000d8 endp1=CS:0x7f0aa4039400
exp=<NEVER> et=0x000 sub=1)
  app0=0x7f0aa402df90 st0=7 st1=0 st2=0 applet=<PEER> tmask=0x40 nice=0
calls=206100614 rate=3160 cpu=0 lat=0
  co1=0x7f0aa4039250 ctrl=tcpv4 xprt=RAW mux=PASS data=STRM
target=PROXY:0x564c0ff53760
      flags=0x10003300 fd=430 fd.state=10121 updt=0 fd.tmask=0x40
      cs=0x7f0aa4039400 csf=0x00008200 ctx=(nil)
  req=0x7f0aa402e2d0 (f=0x848202 an=0x0 pipe=0 tofwd=-1 total=9491961149)
      an_exp=<NEVER> rex=5s wex=<NEVER>
      buf=0x7f0aa402e2d8 data=(nil) o=0 p=0 i=0 size=0
  res=0x7f0aa402e330 (f=0x80448202 an=0x0 pipe=0 tofwd=-1 total=9322481446)
      an_exp=<NEVER> rex=<NEVER> wex=<NEVER>
      buf=0x7f0aa402e338 data=(nil) o=0 p=0 i=0 size=0

pt., 11 mar 2022 o 18:09 Willy Tarreau <[email protected]> napisaƂ(a):

> Hi Maciej,
>
> On Fri, Mar 04, 2022 at 01:10:37PM +0100, Maciej Zdeb wrote:
> > Hi,
> >
> > I'm experiencing high CPU usage on a single core, idle drops below 40%
> > while other cores are at 80% idle. Peers cluster is quite big (12
> servers,
> > each server running 12 threads) and is used for synchronization of 3
> > stick-tables of 1 million entries size.
> >
> > Is peers protocol single threaded and high usage on single core is
> expected
> > or am I experiencing some kind of bug? I'll keep digging to be sure
> nothing
> > else from my configuration is causing the  issue.
>
> Almost. More precisely, each peers connection runs on a single thread
> at once (like any connection). Some such connections may experience
> heavy protocol parsing so it may be possible that sometimes you end up
> with an unbalance number of connections on threads. It's tricky though,
> because we could imagine a mechanism to try to evenly spread the outgoing
> peers connection on threads but the incoming ones will land where they
> land.
>
> That's something you can check with "show sess" and/or "show fd", looking
> for those related to your peers and checking their thread_mask. If you
> find two on the same thread (same thread_mask), you can shut one of them
> down using "shutdown session <id>" and it will reconnect, possibly on
> another thread. That could confirm that it's the root cause of the
> problem you're experiencing.
>
> I'm wondering if we shouldn't introduce a notion of "heavy connection"
> like we already have heavy tasks for the scheduler. These ones would
> be balanced differently from others. Usually they're long-lived and
> can eat massive amounts of CPU so it would make sense. The only ones
> I'm thinking about are the peers ones but the concept would be more
> portable than focusing on peers. Usually such long connections are not
> refreshed often so we probably prefer to spend time carefully picking
> the best thread rather than saving 200ns of processing and having to
> pay it for the whole connection's life time.
>
> Willy
>

Reply via email to