Skip to content

Implement best-effort mount clean up when using host mount namespace - #3118

Closed
SteeleDesmond wants to merge 2 commits into
opencontainers:mainfrom
SteeleDesmond:clean-mounts-no-ns
Closed

SteeleDesmond wants to merge 2 commits into
opencontainers:mainfrom
SteeleDesmond:clean-mounts-no-ns

Conversation

@SteeleDesmond

Copy link
Copy Markdown

The goal in the runtime spec is to unmount container mounts created during container create processing.

Below shows an example case where mounts persist on the host after container delete. With this commit, a best-effort clean up is made to remove mounts in LIFO order during container delete when the host mount namespace is used. This is done by matching the container rootfs and mount destination paths given in the container config.

root@srd-server:/# tree runc-test-bundles/ -L 2
runc-test-bundles/
└── busybox-mount-bundle
    ├── config.json
    └── rootfs
root@srd-server:/# cat runc-test-bundles/busybox-mount-bundle/config.json | grep namespaces -A 13
		"namespaces": [
			{
				"type": "pid"
			},
			{
				"type": "network"
			},
			{
				"type": "ipc"
			},
			{
				"type": "uts"
			}
		]
root@srd-server:/# cat runc-test-bundles/busybox-mount-bundle/config.json | grep mounts -A 75
	"mounts": [
		{
			"destination": "/proc",
			"type": "proc",
			"source": "proc"
		},
		{
			"destination": "/dev",
			"type": "tmpfs",
			"source": "tmpfs",
			"options": [
				"nosuid",
				"strictatime",
				"mode=755",
				"size=65536k"
			]
		},
		{
			"destination": "/dev/pts",
			"type": "devpts",
			"source": "devpts",
			"options": [
				"nosuid",
				"noexec",
				"newinstance",
				"ptmxmode=0666",
				"mode=0620",
				"gid=5"
			]
		},
		{
			"destination": "/dev/shm",
			"type": "tmpfs",
			"source": "shm",
			"options": [
				"nosuid",
				"noexec",
				"nodev",
				"mode=1777",
				"size=65536k"
			]
		},
		{
			"destination": "/dev/mqueue",
			"type": "mqueue",
			"source": "mqueue",
			"options": [
				"nosuid",
				"noexec",
				"nodev"
			]
		},
		{
			"destination": "/sys",
			"type": "sysfs",
			"source": "sysfs",
			"options": [
				"nosuid",
				"noexec",
				"nodev",
				"ro"
			]
		},
		{
			"destination": "/sys/fs/cgroup",
			"type": "cgroup",
			"source": "cgroup",
			"options": [
				"nosuid",
				"noexec",
				"nodev",
				"relatime",
				"ro"
			]
		}
	],
root@srd-server:/# ./runc create -b /runc-test-bundles/busybox-mount-bundle/ testid1
root@srd-server:/# mount | grep /runc-test-bundles/
/dev/mapper/ubuntu--vg-ubuntu--lv on /runc-test-bundles/busybox-mount-bundle/rootfs type ext4 (rw,relatime)
proc on /runc-test-bundles/busybox-mount-bundle/rootfs/proc type proc (rw,relatime)
tmpfs on /runc-test-bundles/busybox-mount-bundle/rootfs/dev type tmpfs (rw,nosuid,size=65536k,mode=755)
devpts on /runc-test-bundles/busybox-mount-bundle/rootfs/dev/pts type devpts (rw,nosuid,noexec,relatime,gid=5,mode=620,ptmxmode=666)
shm on /runc-test-bundles/busybox-mount-bundle/rootfs/dev/shm type tmpfs (rw,nosuid,nodev,noexec,relatime,size=65536k)
mqueue on /runc-test-bundles/busybox-mount-bundle/rootfs/dev/mqueue type mqueue (rw,nosuid,nodev,noexec,relatime)
sysfs on /runc-test-bundles/busybox-mount-bundle/rootfs/sys type sysfs (ro,nosuid,nodev,noexec,relatime)
tmpfs on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup type tmpfs (rw,nosuid,nodev,noexec,relatime,mode=755)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/systemd type cgroup (ro,nosuid,nodev,noexec,relatime,xattr,name=systemd)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/blkio type cgroup (ro,nosuid,nodev,noexec,relatime,blkio)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/freezer type cgroup (ro,nosuid,nodev,noexec,relatime,freezer)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/cpu,cpuacct type cgroup (ro,nosuid,nodev,noexec,relatime,cpu,cpuacct)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/cpuset type cgroup (ro,nosuid,nodev,noexec,relatime,cpuset)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/perf_event type cgroup (ro,nosuid,nodev,noexec,relatime,perf_event)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/memory type cgroup (ro,nosuid,nodev,noexec,relatime,memory)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/net_cls,net_prio type cgroup (ro,nosuid,nodev,noexec,relatime,net_cls,net_prio)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/devices type cgroup (ro,nosuid,nodev,noexec,relatime,devices)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/hugetlb type cgroup (ro,nosuid,nodev,noexec,relatime,hugetlb)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/rdma type cgroup (ro,nosuid,nodev,noexec,relatime,rdma)
cgroup on /runc-test-bundles/busybox-mount-bundle/rootfs/sys/fs/cgroup/pids type cgroup (ro,nosuid,nodev,noexec,relatime,pids)
root@srd-server:/# ./runc delete testid1
# The above container mounts will persist after container delete. With the change they are unmounted during delete.

See the references below for similar discussions around this issue.
References:
Issue #2095
Issue #1909
libcontainer: containers with host fs root

Signed-off-by: Steele Ray Desmond steele.desmond@ibm.com

@kolyshkin

Copy link
Copy Markdown
Contributor

@SteeleDesmond this calls for an integration test (see tests/integration/ for examples).

@kolyshkin

Copy link
Copy Markdown
Contributor

@SteeleDesmond please rebase

@SteeleDesmond
SteeleDesmond force-pushed the clean-mounts-no-ns branch 2 times, most recently from cfa1a71 to 389d9f7 Compare August 3, 2021 18:50
@kolyshkin

Copy link
Copy Markdown
Contributor

@SteeleDesmond ^^^

Comment thread libcontainer/state_linux.go Outdated
config := c.Config()

// Unmount recursive
err := unix.Unmount(config.Rootfs, unix.MNT_DETACH)

@kolyshkin kolyshkin Aug 31, 2021 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NB: this relies on the assumption that Rootfs is a mount point, and that assumption is correct.

Comment thread libcontainer/state_linux.go Outdated

// If recursive unmount fails, try best-effort unmount
for i := len(config.Mounts) - 1; i >= 0; i-- {
mountpoint := config.Rootfs + config.Mounts[i].Destination

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we use securejoin (or, better, WithProcfd()) here?

Probably not, since we're in a host mountns and we can't have a situation where the mount point is an absolute symlink which only make sense inside the mountns (something like #3047).

@cyphar PTAL 🙏🏻

Comment thread libcontainer/state_linux.go
Comment thread libcontainer/state_linux.go Outdated

@kolyshkin kolyshkin left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your work, @SteeleDesmond!

Aside from a nit about using unmount vs unix.Unmount, this definitely needs a test case (e.g. an integration test, look into tests/integration/mounts.bats).

@kolyshkin

Copy link
Copy Markdown
Contributor

This still needs a rebase (to pick up latest and greatest CI).

Comment thread libcontainer/state_linux.go Outdated
@kolyshkin

Copy link
Copy Markdown
Contributor

@SteeleDesmond do you intend to keep working on this? If yes, I can help with a test case.

@SteeleDesmond
SteeleDesmond force-pushed the clean-mounts-no-ns branch 5 times, most recently from bd6623a to e592a34 Compare December 20, 2025 01:12
@kolyshkin

This comment was marked as outdated.

@kolyshkin

This comment was marked as outdated.

@kolyshkin kolyshkin left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tests need to be changed to the new style (see #5429), otherwise they will fail.

And it still needs a rebase @SteeleDesmond

}

// If recursive unmount fails, try best-effort unmount
// We iterate in reverse to unmount children before parents

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, should this be sorted by length instead? Or do we actually rely on the order of mounts? (we may)

Signed-off-by: Steele Ray Desmond <steele@desmond.sh>
Signed-off-by: Steele Ray Desmond <steele@desmond.sh>
@kolyshkin

Copy link
Copy Markdown
Contributor

Rebased (mostly to see how the old-style int tests fail)

@kolyshkin

Copy link
Copy Markdown
Contributor

Also, there is tests/integration/host-mntns.bats which already has the proper setup/teardown for host mount namespace tests, so I suggest to move the test there. Something like this:

# https://github.com/opencontainers/runc/pull/3118
@test "runc delete [host mount ns, mounts are cleaned up]" {
      update_config '   .linux.namespaces -= [{"type": "mount"}]
                      | .linux.maskedPaths = []
                      | .linux.readonlyPaths = []
                      | .root.readonly = false'
      run -0 runc run -d --console-socket "$CONSOLE_SOCKET" test_host_mntns
      testcontainer test_host_mntns running

      # Container's rootfs is a mount point on the host.
      mountpoint -q rootfs

      run -0 runc delete -f test_host_mntns
      # Check the mount is gone.
      run ! mountpoint -q rootfs
}

With this PR, the workaround in that file's teardown is no longer needed (and umount will fail as rootfs is no longer mounted), so it should be changed to something like:

function teardown() {
      [ ! -v ROOT ] && return 0 # nothing to teardown

      # In case a test failed before the container was deleted.
      if mountpoint -q "$ROOT"/bundle/rootfs; then
              umount -R --lazy "$ROOT"/bundle/rootfs
      fi

      teardown_bundle
}

kolyshkin added a commit to kolyshkin/runc that referenced this pull request Oct 9, 2026
When the container does not have its own mount namespace, the container
rootfs (bind mounted onto itself by runc) and all the container mounts
(proc, dev etc.) are created in the host mount namespace, and are never
unmounted, so they are left behind after the container is deleted.

To fix, once the rootfs is mounted, remember its mount ID (and save it
to the container state), and on container destroy, unmount it (together
with all the container mounts under it), provided it is still the same
mount. This way, a mount which was not created by runc is never
unmounted.

One exception is when the rootfs is a mount which is a peer of another
mount in the same mount namespace (e.g. rootfs is bind mounted onto
itself on a shared mount). In this case, runc rootfs mount is propagated
to that peer, and the copy ends up tucked under the user mount, so
unmounting runc mount results in the user mount being unmounted, too.
Do not unmount anything in such case.

Add integration tests, and remove the workaround from host-mntns.bats
teardown.

This is a carry of PR opencontainers#3118 (with the implementation redone).

Fixes opencontainers#2095.

Co-authored-by: Steele Ray Desmond <steele@desmond.sh>
Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
@SteeleDesmond

Copy link
Copy Markdown
Author

Happy to see this contribution make it in! A half decade slow burn. Thanks for getting it over the finish line. Cheers @kolyshkin.

kolyshkin added a commit to kolyshkin/runc that referenced this pull request Oct 9, 2026
When the container does not have its own mount namespace, the container
rootfs (bind mounted onto itself by runc) and all the container mounts
(proc, dev etc.) are created in the host mount namespace, and are never
unmounted, so they are left behind after the container is deleted.

To fix, once the rootfs is mounted, remember its mount ID (and save it
to the container state), and on container destroy, unmount it (together
with all the container mounts under it), provided it is still the same
mount. This way, a mount which was not created by runc is never
unmounted.

If rootfs is itself a shared mount, runc rootfs mount is propagated to
its peers. If one of these peers is the mount rootfs is mounted on (e.g.
rootfs is bind mounted onto itself on a shared mount), the copy gets
tucked under the rootfs mount, and unmounting runc mount on destroy
unmounts the rootfs mount, too. To prevent that, make rootfs a slave
before creating runc rootfs mount.

Add integration tests, and remove the workaround from host-mntns.bats
teardown.

This is a carry of PR opencontainers#3118 (with the implementation redone).

Fixes opencontainers#2095.

Co-authored-by: Steele Ray Desmond <steele@desmond.sh>
Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
kolyshkin added a commit to kolyshkin/runc that referenced this pull request Oct 9, 2026
When the container does not have its own mount namespace, the container
rootfs (bind mounted onto itself by runc) and all the container mounts
(proc, dev etc.) are created in the host mount namespace, and are never
unmounted, so they are left behind after the container is deleted.

To fix, once the rootfs is mounted, remember its mount ID (and save it
to the container state), and on container destroy, unmount it (together
with all the container mounts under it), provided it is still the same
mount. This way, a mount which was not created by runc is never
unmounted.

If rootfs is itself a shared mount, runc rootfs mount is propagated to
its peers. If one of these peers is the mount rootfs is mounted on (e.g.
rootfs is bind mounted onto itself on a shared mount), the copy gets
tucked under the rootfs mount, and unmounting runc mount on destroy
unmounts the rootfs mount, too. To prevent that, make rootfs a slave
before creating runc rootfs mount.

Add integration tests, and remove the workaround from host-mntns.bats
teardown.

This is a carry of PR opencontainers#3118 (with the implementation redone).

Fixes opencontainers#2095.

Co-authored-by: Steele Ray Desmond <steele@desmond.sh>
Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
kolyshkin added a commit to kolyshkin/runc that referenced this pull request Oct 9, 2026
When the container does not have its own mount namespace, the container
rootfs (bind mounted onto itself by runc) and all the container mounts
(proc, dev etc.) are created in the host mount namespace, and are never
unmounted, so they are left behind after the container is deleted.

To fix, once the rootfs is mounted, remember its mount ID (and save it
to the container state), and on container destroy, unmount it (together
with all the container mounts under it), provided it is still the same
mount. This way, a mount which was not created by runc is never
unmounted.

If rootfs is itself a shared mount, runc rootfs mount is propagated to
its peers. If one of these peers is the mount rootfs is mounted on (e.g.
rootfs is bind mounted onto itself on a shared mount), the copy gets
tucked under the rootfs mount, and unmounting runc mount on destroy
unmounts the rootfs mount, too. To prevent that, make rootfs a slave
before creating runc rootfs mount.

Add integration tests, and remove the workaround from host-mntns.bats
teardown.

This is a carry of PR opencontainers#3118 (with the implementation redone).

Fixes opencontainers#2095.

Co-authored-by: Steele Ray Desmond <steele@desmond.sh>
Signed-off-by: Kir Kolyshkin <kolyshkin@gmail.com>
@kolyshkin

Copy link
Copy Markdown
Contributor

Carried in #5538 (with the implementation redone: runc now remembers the rootfs mount it creates in the host mount namespace, and unmounts it on container destroy only if it is still the same mount). It is based on #5537, which fixes host mounts propagation being changed when running a container without a mount namespace.

Thank you @SteeleDesmond for your work on this! Closing.

@kolyshkin kolyshkin closed this Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants