On 5/1/2018 10:35 PM, Neelakantam Gaddam wrote:
Hi All,

The issue here is we are trying to send the packets on the same device in a loop. While sending on a device, a spinlock for the tx queue has to be acquired in dev_queue_xmit function. This is where we are trying to acquire the same lock again, which is leading to the kernel crash. This issue becomes worse if internal ports are involved in the configuration.

I think, we should avoid these loops in the vport send functions. But in case of tunneling, the tunnel send function should take care of these checks.

Please share your thoughts on this issue.

Hello Neelakantam,

Please provide the full kernel panic backtrace and I'll have a look at the problem.  I'm rather busy dealing with some other critical issues but will try to get back to you next week.

Thanks,

- Greg


On Mon, Apr 30, 2018 at 11:17 AM, Neelakantam Gaddam <[email protected] <mailto:[email protected]>> wrote:

    Hi All,

    OVS misconfiguration leading to spinlock recursion in dev_queue_xmit.

    We are running ovs-2.8.1 with openvswitch kernel modules on two
    hosts connected back to back. We are running OVS on MIPS64 platform.



    We are using the below configuration.

    ovs-vsctl add-br br0

    ovs-vsctl add-bond br0 bond0 p1p1 p1p2

    ovs-vsctl set port bond0 lacp=active bond_mode=balance-tcp

    ifconfig br0 100.0.0.1 up

    ovs-vsctl add-port br0 veth0

    ovs-vsctl add-port br0 vx0 -- set interface vx0 type=vxlan
    options:local_ip=100.0.0.1 options:remote_ip=100.0.0.2 option:key=flow

    ovs-ofctl add-flow br0 "table=0, priority=1, cookie=100,
    tun_id=100, in_port=4, action=output:3"

    ovs-ofctl add-flow br0 "table=0, priority=1, cookie=100,
    in_port=3, actions=set_field:100->tun_id output:4"

    When this configuration is applied on both hosts, we are seeing
    the below spinlock recursion bug.

    [<ffffffff80864cd4>] show_stack+0x6c/0xf8

    [<ffffffff80ad1628>] do_raw_spin_lock+0x168/0x170

    [<ffffffff80bf7b1c>] dev_queue_xmit+0x43c/0x470

    [<ffffffff80c32c08>] ip_finish_output+0x250/0x490

    [<ffffffffc0115664>] rpl_iptunnel_xmit+0x134/0x218 [openvswitch]

    [<ffffffffc0120f28>] rpl_vxlan_xmit+0x430/0x538 [openvswitch]

    [<ffffffffc00f9de0>] do_execute_actions+0x18f8/0x19e8 [openvswitch]

    [<ffffffffc00fa2b0>] ovs_execute_actions+0x90/0x208 [openvswitch]

    [<ffffffffc0101860>] ovs_dp_process_packet+0xb0/0x1a8 [openvswitch]

    [<ffffffffc010c5d8>] ovs_vport_receive+0x78/0x130 [openvswitch]

    [<ffffffffc010ce6c>] internal_dev_xmit+0x34/0x98 [openvswitch]

    [<ffffffff80bf74d0>] dev_hard_start_xmit+0x2e8/0x4f8

    [<ffffffff80c10e48>] sch_direct_xmit+0xf0/0x238

    [<ffffffff80bf78b8>] dev_queue_xmit+0x1d8/0x470

    [<ffffffff80c5ffe4>] arp_process+0x614/0x628

    [<ffffffff80bf0cb0>] __netif_receive_skb_core+0x2e8/0x5d8

    [<ffffffff80bf4770>] process_backlog+0xc0/0x1b0

    [<ffffffff80bf501c>] net_rx_action+0x154/0x240

    [<ffffffff8088d130>] __do_softirq+0x1d0/0x218

    [<ffffffff8088d240>] do_softirq+0x68/0x70

    [<ffffffff8088d3a0>] local_bh_enable+0xa8/0xb0

    [<ffffffff80bf5c88>] netif_rx_ni+0x20/0x30

    The packet path traced is : netif_rx->arp->dev_queue_xmit(internal
    port)->vxlan_xmit->dev_queue_xmit(internal port). According to the
    configuration, this packet path is valid. But we should not hit
    the crash.


    Questions:


      * Is it a kernel bug or ovs bug ?
      * How OVS handles these kind of misconfigurations especially
        packet loops involved?

    Any suggestion or help is greatly appreciated.



    Thanks



-- Thanks & Regards
    Neelakantam Gaddam




--
Thanks & Regards
Neelakantam Gaddam


_______________________________________________
discuss mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-discuss

_______________________________________________
discuss mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-discuss

Reply via email to