mirror of
https://github.com/torvalds/linux.git
synced 2026-09-24 06:24:02 +02:00
Document the failfs semantics, the FD_FAILFS_ROOT sentinel, the fchroot() entry requirements, and the ways back out. Link: https://patch.msgid.link/20260724-work-failfs-v2-7-485dabbae185@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
74 lines
3.1 KiB
ReStructuredText
74 lines
3.1 KiB
ReStructuredText
.. SPDX-License-Identifier: GPL-2.0
|
|
|
|
======
|
|
failfs
|
|
======
|
|
|
|
failfs is a kernel-internal filesystem that fails every operation
|
|
reaching it with ``EOPNOTSUPP``. It is the counterpart to nullfs. Where
|
|
nullfs is permanently empty, failfs means "nothing is supported here".
|
|
It cannot be mounted from userspace, nothing can be mounted on top of
|
|
it. It cannot be cloned.
|
|
|
|
The only way into it is the ``FD_FAILFS_ROOT`` file descriptor sentinel which
|
|
is understood by ``fchdir(2)`` and ``fchroot(2)``.
|
|
|
|
Semantics
|
|
=========
|
|
|
|
Every path walk of a component through failfs fails with
|
|
``EOPNOTSUPP`` before that component is parsed, including ``.``.
|
|
|
|
No path lookup can open the root, not even with ``O_PATH``.
|
|
|
|
A process with its working directory in failfs fails every
|
|
``AT_FDCWD``-relative lookup. As with any working directory that is
|
|
unreachable from the process root, the ``getcwd(2)`` system call returns
|
|
a path prefixed with ``(unreachable)``.
|
|
|
|
A process with its root directory in failfs fails every absolute path
|
|
lookup including absolute symlinks and the interpreter of dynamically
|
|
linked binaries. In other words, this fails exec.
|
|
|
|
Lookups anchored at explicit directory file descriptors keep working. It
|
|
is the ``fs_struct`` equivalent of ``RESOLVE_BENEATH``. The process must
|
|
anchor every lookup at a file descriptor it explicitly holds.
|
|
|
|
Entering
|
|
========
|
|
|
|
``fchroot(FD_FAILFS_ROOT, 0)`` requires ``CAP_SYS_CHROOT`` in the
|
|
caller's user namespace, mirroring ``chroot(2)``. Unprivileged callers
|
|
may enter if all of the following hold:
|
|
|
|
* ``no_new_privs`` is set: setuid binaries on regular mounts remain
|
|
reachable via inherited directory file descriptors and executing them
|
|
with an unusable root directory is the classic confused deputy.
|
|
|
|
* The caller is not already chrooted: the root directory is what
|
|
confines ``..`` resolution and the failfs root can never be reached by
|
|
walking up a real mount tree, so moving the root of a chrooted task to
|
|
failfs would allow it to escape its chroot via ``openat(fd, "..")``.
|
|
|
|
* The caller does not share its ``fs_struct``: ``no_new_privs`` is
|
|
checked on the calling thread, but the root lives in the ``fs_struct``.
|
|
A ``CLONE_FS`` sibling without ``no_new_privs`` could otherwise execute
|
|
a setuid binary with the failfs root, so entry requires ``fs->users ==
|
|
1``, the same restriction ``setns(2)`` applies for the mount and user
|
|
namespaces.
|
|
|
|
Leaving
|
|
=======
|
|
|
|
Backing out is currently hard, but this is a property of the current
|
|
implementation, not a guaranteed interface, and may be loosened later.
|
|
For now a process that entered failfs counts as chrooted, so it cannot
|
|
create user namespaces to regain ``CAP_SYS_CHROOT``, and ``chroot(2)``
|
|
or ``fchroot(2)`` back out require ``CAP_SYS_CHROOT``. The remaining way
|
|
out today is ``setns(2)`` with a mount namespace file descriptor, which
|
|
requires ``CAP_SYS_ADMIN`` over the target mount namespace as well as
|
|
``CAP_SYS_CHROOT`` and ``CAP_SYS_ADMIN`` in the caller's user namespace
|
|
and resets both root and working directory. A process that holds no such
|
|
file descriptor and restricts ``*chdir()``/``*chroot()``/``setns()`` via
|
|
seccomp cannot currently get back out.
|