[ 
https://issues.apache.org/jira/browse/MESOS-9031?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16530017#comment-16530017
 ] 

Kirill Plyashkevich commented on MESOS-9031:
--------------------------------------------

[~qianzhang], yes, of course:
2 standalone containers/tasks are launched on the same slave, joining the same 
`mesos-cni0` bridge network.

tasks themselves are services of an akka cluster, so they communicate with 
other services.nodes of the cluster and between each other using host's ips and 
ports.

both tasks fail to cummunicate between each other due to timeout (with 
`excludeDevices` set to one including `mesos-cni0` connection will just get 
refused).

unfortunatelly striped json won't be a lot of help here, short interaction can 
be discribed as e.g.: [email protected]:2552 tries to reach 
[email protected]:31303 (host ip and other service port, which is effectively 
[email protected]:2552)

 

I've been digging into `bridge` recently as well and ACCEPT rule is added 
[here|[https://github.com/containernetworking/plugins/blob/master/pkg/ip/ipmasq_linux.go#L63].]
 that said it's related to `cni/bridge` plugin.

> Mesos CNI portmap plugins' iptables rules doesn't allow connections via host 
> ip and port from the same bridge container network
> -------------------------------------------------------------------------------------------------------------------------------
>
>                 Key: MESOS-9031
>                 URL: https://issues.apache.org/jira/browse/MESOS-9031
>             Project: Mesos
>          Issue Type: Bug
>          Components: cni, containerization
>    Affects Versions: 1.6.0
>            Reporter: Kirill Plyashkevich
>            Priority: Major
>
> using `mesos-cni-port-mapper` with folllowing config:
> {noformat}
> { 
>    "name" : "dcos", 
>    "type" : "mesos-cni-port-mapper", 
>    "excludeDevices" : [], 
>    "chain": "MESOS-CNI0-PORT-MAPPER", 
>    "delegate": { 
>        "type": "bridge", 
>        "bridge": "mesos-cni0", 
>        "isGateway": true, 
>        "ipMasq": true, 
>        "hairpinMode": true, 
>        "ipam": { 
>            "type": "host-local", 
>            "ranges": [ 
>                [{"subnet": "172.26.0.0/16"}] 
>            ], 
>            "routes": [ 
>                {"dst": "0.0.0.0/0"} 
>            ] 
>        } 
>    } 
> }
> {noformat}
>  - 2 services running on the same mesos-slave using unified containerizer in 
> different tasks and communicating via host ip and host port
>  - connection timeouts due to iptables rules per container CNI-XXX chain
>  - actually timeouts are caused by
> {noformat}
> Chain CNI-XXX (1 references)
> num  target     prot opt source               destination         
> 1    ACCEPT     all  --  anywhere             172.26.0.0/16        /* name: 
> "dcos" id: "YYYY" */
> 2    MASQUERADE  all  --  anywhere            !base-address.mcast.net/4  /* 
> name: "dcos" id: "YYYY" */
> {noformat}
> rule #1 is executed and no masquerading happens.
> there are multiple solutions:
>  - simpliest and fastest one is not to add that ACCEPT
>  - perhaps, there's a better change in iptables rules that can fix it
>  - proper one (imho) is to finally implement cni spec 0.3.x in order to be 
> able to use chaining of plugins and use cni's `bridge` and `portmap` plugins 
> in chain (and get rid of mesos-cni-port-mapper completely eventually).



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Reply via email to