Skip to content

libct: unmount container rootfs in host mount ns on destroy - #5538

Draft
kolyshkin wants to merge 3 commits into
opencontainers:mainfrom
kolyshkin:kir/carry-3118
Draft

kolyshkin wants to merge 3 commits into
opencontainers:mainfrom
kolyshkin:kir/carry-3118

Conversation

@kolyshkin

@kolyshkin kolyshkin commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Based on (and includes) #5537; please review the last two commits only.

When the container does not have its own mount namespace, the container rootfs (bind mounted onto itself by runc) and all the container mounts (proc, dev etc.) are created in the host mount namespace, and are never unmounted, so they are left behind after the container is deleted.

To fix, once the rootfs is mounted, remember its mount ID (and save it to the container state), and on container destroy, unmount it (together with all the container mounts under it), provided it is still the same mount. This way, a mount which was not created by runc is never unmounted.

If rootfs is itself a shared mount, runc rootfs mount is propagated to its peers. If one of these peers is the mount rootfs is mounted on (e.g. rootfs is bind mounted onto itself on a shared mount), the copy gets tucked under the rootfs mount, and unmounting runc mount on destroy unmounts the rootfs mount, too. To prevent that, make rootfs a slave before creating runc rootfs mount.

Add integration tests, and remove the workaround from host-mntns.bats teardown.

This is a carry of #3118 (with the implementation redone). Thanks to @SteeleDesmond for the original work.

Fixes #2095.
Closes #3118.

When the container does not have its own mount namespace, prepareRoot
is run in the host mount namespace. Since commit 91ca331 ("chroot when
no mount namespaces is provided"), this results in host's / being
recursively made a slave (which, in the host mount namespace, means
private), and the parent mount of the container rootfs being made
private (or slave).

In the host mount namespace, do not change the propagation of the
existing mounts. Instead, make the new rootfs mount a slave, so that
container mounts do not propagate to the host and other mount
namespaces (otherwise, if the rootfs parent mount is shared, every
container mount appears twice in the host mount namespace, and is
propagated to all its slaves).

Add an integration test.

Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
@kolyshkin
kolyshkin force-pushed the kir/carry-3118 branch 3 times, most recently from cda9440 to 90b7509 Compare October 9, 2026 20:46
kolyshkin and others added 2 commits October 9, 2026 13:48
All the tests in this file need the same config modifications to run
a container in the host mount namespace, so move those to setup.

No functional change.

Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
When the container does not have its own mount namespace, the container
rootfs (bind mounted onto itself by runc) and all the container mounts
(proc, dev etc.) are created in the host mount namespace, and are never
unmounted, so they are left behind after the container is deleted.

To fix, once the rootfs is mounted, remember its mount ID (and save it
to the container state), and on container destroy, unmount it (together
with all the container mounts under it), provided it is still the same
mount. This way, a mount which was not created by runc is never
unmounted.

If rootfs is itself a shared mount, runc rootfs mount is propagated to
its peers. If one of these peers is the mount rootfs is mounted on (e.g.
rootfs is bind mounted onto itself on a shared mount), the copy gets
tucked under the rootfs mount, and unmounting runc mount on destroy
unmounts the rootfs mount, too. To prevent that, make rootfs a slave
before creating runc rootfs mount.

Add integration tests, and remove the workaround from host-mntns.bats
teardown.

This is a carry of PR opencontainers#3118 (with the implementation redone).

Fixes opencontainers#2095.

Co-authored-by: Steele Ray Desmond <steele@desmond.sh>
Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Mountpoints not being cleaned if the mount namespace is shared

1 participant