Repository navigation
Implement best-effort mount clean up when using host mount namespace - #3118
SteeleDesmond wants to merge 2 commits into
Conversation
0651a1a to
c3cf57f
Compare
|
@SteeleDesmond this calls for an integration test (see |
|
@SteeleDesmond please rebase |
cfa1a71 to
389d9f7
Compare
|
@SteeleDesmond ^^^ |
| config := c.Config() | ||
|
|
||
| // Unmount recursive | ||
| err := unix.Unmount(config.Rootfs, unix.MNT_DETACH) |
There was a problem hiding this comment.
NB: this relies on the assumption that Rootfs is a mount point, and that assumption is correct.
|
|
||
| // If recursive unmount fails, try best-effort unmount | ||
| for i := len(config.Mounts) - 1; i >= 0; i-- { | ||
| mountpoint := config.Rootfs + config.Mounts[i].Destination |
kolyshkin
left a comment
There was a problem hiding this comment.
Thank you for your work, @SteeleDesmond!
Aside from a nit about using unmount vs unix.Unmount, this definitely needs a test case (e.g. an integration test, look into tests/integration/mounts.bats).
|
This still needs a rebase (to pick up latest and greatest CI). |
|
@SteeleDesmond do you intend to keep working on this? If yes, I can help with a test case. |
bd6623a to
e592a34
Compare
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
This comment was marked as outdated.
There was a problem hiding this comment.
Tests need to be changed to the new style (see #5429), otherwise they will fail.
And it still needs a rebase @SteeleDesmond
| } | ||
|
|
||
| // If recursive unmount fails, try best-effort unmount | ||
| // We iterate in reverse to unmount children before parents |
There was a problem hiding this comment.
Hmm, should this be sorted by length instead? Or do we actually rely on the order of mounts? (we may)
Signed-off-by: Steele Ray Desmond <steele@desmond.sh>
Signed-off-by: Steele Ray Desmond <steele@desmond.sh>
7517e72 to
879ab44
Compare
|
Rebased (mostly to see how the old-style int tests fail) |
|
Also, there is # https://github.com/opencontainers/runc/pull/3118
@test "runc delete [host mount ns, mounts are cleaned up]" {
update_config ' .linux.namespaces -= [{"type": "mount"}]
| .linux.maskedPaths = []
| .linux.readonlyPaths = []
| .root.readonly = false'
run -0 runc run -d --console-socket "$CONSOLE_SOCKET" test_host_mntns
testcontainer test_host_mntns running
# Container's rootfs is a mount point on the host.
mountpoint -q rootfs
run -0 runc delete -f test_host_mntns
# Check the mount is gone.
run ! mountpoint -q rootfs
}With this PR, the workaround in that file's function teardown() {
[ ! -v ROOT ] && return 0 # nothing to teardown
# In case a test failed before the container was deleted.
if mountpoint -q "$ROOT"/bundle/rootfs; then
umount -R --lazy "$ROOT"/bundle/rootfs
fi
teardown_bundle
} |
When the container does not have its own mount namespace, the container rootfs (bind mounted onto itself by runc) and all the container mounts (proc, dev etc.) are created in the host mount namespace, and are never unmounted, so they are left behind after the container is deleted. To fix, once the rootfs is mounted, remember its mount ID (and save it to the container state), and on container destroy, unmount it (together with all the container mounts under it), provided it is still the same mount. This way, a mount which was not created by runc is never unmounted. One exception is when the rootfs is a mount which is a peer of another mount in the same mount namespace (e.g. rootfs is bind mounted onto itself on a shared mount). In this case, runc rootfs mount is propagated to that peer, and the copy ends up tucked under the user mount, so unmounting runc mount results in the user mount being unmounted, too. Do not unmount anything in such case. Add integration tests, and remove the workaround from host-mntns.bats teardown. This is a carry of PR opencontainers#3118 (with the implementation redone). Fixes opencontainers#2095. Co-authored-by: Steele Ray Desmond <steele@desmond.sh> Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
|
Happy to see this contribution make it in! A half decade slow burn. Thanks for getting it over the finish line. Cheers @kolyshkin. |
When the container does not have its own mount namespace, the container rootfs (bind mounted onto itself by runc) and all the container mounts (proc, dev etc.) are created in the host mount namespace, and are never unmounted, so they are left behind after the container is deleted. To fix, once the rootfs is mounted, remember its mount ID (and save it to the container state), and on container destroy, unmount it (together with all the container mounts under it), provided it is still the same mount. This way, a mount which was not created by runc is never unmounted. If rootfs is itself a shared mount, runc rootfs mount is propagated to its peers. If one of these peers is the mount rootfs is mounted on (e.g. rootfs is bind mounted onto itself on a shared mount), the copy gets tucked under the rootfs mount, and unmounting runc mount on destroy unmounts the rootfs mount, too. To prevent that, make rootfs a slave before creating runc rootfs mount. Add integration tests, and remove the workaround from host-mntns.bats teardown. This is a carry of PR opencontainers#3118 (with the implementation redone). Fixes opencontainers#2095. Co-authored-by: Steele Ray Desmond <steele@desmond.sh> Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
When the container does not have its own mount namespace, the container rootfs (bind mounted onto itself by runc) and all the container mounts (proc, dev etc.) are created in the host mount namespace, and are never unmounted, so they are left behind after the container is deleted. To fix, once the rootfs is mounted, remember its mount ID (and save it to the container state), and on container destroy, unmount it (together with all the container mounts under it), provided it is still the same mount. This way, a mount which was not created by runc is never unmounted. If rootfs is itself a shared mount, runc rootfs mount is propagated to its peers. If one of these peers is the mount rootfs is mounted on (e.g. rootfs is bind mounted onto itself on a shared mount), the copy gets tucked under the rootfs mount, and unmounting runc mount on destroy unmounts the rootfs mount, too. To prevent that, make rootfs a slave before creating runc rootfs mount. Add integration tests, and remove the workaround from host-mntns.bats teardown. This is a carry of PR opencontainers#3118 (with the implementation redone). Fixes opencontainers#2095. Co-authored-by: Steele Ray Desmond <steele@desmond.sh> Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
When the container does not have its own mount namespace, the container rootfs (bind mounted onto itself by runc) and all the container mounts (proc, dev etc.) are created in the host mount namespace, and are never unmounted, so they are left behind after the container is deleted. To fix, once the rootfs is mounted, remember its mount ID (and save it to the container state), and on container destroy, unmount it (together with all the container mounts under it), provided it is still the same mount. This way, a mount which was not created by runc is never unmounted. If rootfs is itself a shared mount, runc rootfs mount is propagated to its peers. If one of these peers is the mount rootfs is mounted on (e.g. rootfs is bind mounted onto itself on a shared mount), the copy gets tucked under the rootfs mount, and unmounting runc mount on destroy unmounts the rootfs mount, too. To prevent that, make rootfs a slave before creating runc rootfs mount. Add integration tests, and remove the workaround from host-mntns.bats teardown. This is a carry of PR opencontainers#3118 (with the implementation redone). Fixes opencontainers#2095. Co-authored-by: Steele Ray Desmond <steele@desmond.sh> Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
|
Carried in #5538 (with the implementation redone: runc now remembers the rootfs mount it creates in the host mount namespace, and unmounts it on container destroy only if it is still the same mount). It is based on #5537, which fixes host mounts propagation being changed when running a container without a mount namespace. Thank you @SteeleDesmond for your work on this! Closing. |
The goal in the runtime spec is to unmount container mounts created during container create processing.
Below shows an example case where mounts persist on the host after container delete. With this commit, a best-effort clean up is made to remove mounts in LIFO order during container delete when the host mount namespace is used. This is done by matching the container rootfs and mount destination paths given in the container config.
See the references below for similar discussions around this issue.
References:
Issue #2095
Issue #1909
libcontainer: containers with host fs root
Signed-off-by: Steele Ray Desmond steele.desmond@ibm.com