[
https://issues.apache.org/jira/browse/MESOS-8125?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16345186#comment-16345186
]
Qian Zhang commented on MESOS-8125:
-----------------------------------
Had a discussion with [~vinodkone], basically we should not allow user to
launch a Docker container with `–restart=always`, because it may cause the
Docker container is running even after the task completes, I have created a
separate [ticket|https://issues.apache.org/jira/browse/MESOS-8509] to trace it.
And there is still another case that the above solution can not handle: Agent
is restarted and recovers when Docker executor is trying to launch Docker
container but it has NOT been launched yet, in this case, we will not find the
Docker container in `DockerContainerizerProcess::_recover` since it has not
been launched yet, and then set `container->status` to `None()` which will make
agent shutdown the executor, and the task will fail. I think this is not wha we
want.
> Agent should properly handle recovering an executor when its pid is reused
> --------------------------------------------------------------------------
>
> Key: MESOS-8125
> URL: https://issues.apache.org/jira/browse/MESOS-8125
> Project: Mesos
> Issue Type: Bug
> Reporter: Gastón Kleiman
> Assignee: Qian Zhang
> Priority: Critical
>
> Here's how to reproduce this issue:
> # Start a task using the Docker containerizer (the same will probably happen
> with the command executor).
> # Stop the corresponding Mesos agent while the task is running.
> # Change the executor's checkpointed forked pid, which is located in the meta
> directory, e.g.,
> {{/var/lib/mesos/slave/meta/slaves/latest/frameworks/19faf6e0-3917-48ab-8b8e-97ec4f9ed41e-0001/executors/foo.13faee90-b5f0-11e7-8032-e607d2b4348c/runs/latest/pids/forked.pid}}.
> I used pid 2, which is normally used by {{kthreadd}}.
> # Reboot the host
--
This message was sent by Atlassian JIRA
(v7.6.3#76005)