mirror of
https://github.com/torvalds/linux.git
synced 2026-09-23 05:04:02 +02:00
Documentation: add failfs documentation
Document the failfs semantics, the FD_FAILFS_ROOT sentinel, the fchroot() entry requirements, and the ways back out. Link: https://patch.msgid.link/20260724-work-failfs-v2-7-485dabbae185@kernel.org Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
This commit is contained in:
parent
df4b2889ea
commit
a45a6605cd
73
Documentation/filesystems/failfs.rst
Normal file
73
Documentation/filesystems/failfs.rst
Normal file
|
|
@ -0,0 +1,73 @@
|
|||
.. SPDX-License-Identifier: GPL-2.0
|
||||
|
||||
======
|
||||
failfs
|
||||
======
|
||||
|
||||
failfs is a kernel-internal filesystem that fails every operation
|
||||
reaching it with ``EOPNOTSUPP``. It is the counterpart to nullfs. Where
|
||||
nullfs is permanently empty, failfs means "nothing is supported here".
|
||||
It cannot be mounted from userspace, nothing can be mounted on top of
|
||||
it. It cannot be cloned.
|
||||
|
||||
The only way into it is the ``FD_FAILFS_ROOT`` file descriptor sentinel which
|
||||
is understood by ``fchdir(2)`` and ``fchroot(2)``.
|
||||
|
||||
Semantics
|
||||
=========
|
||||
|
||||
Every path walk of a component through failfs fails with
|
||||
``EOPNOTSUPP`` before that component is parsed, including ``.``.
|
||||
|
||||
No path lookup can open the root, not even with ``O_PATH``.
|
||||
|
||||
A process with its working directory in failfs fails every
|
||||
``AT_FDCWD``-relative lookup. As with any working directory that is
|
||||
unreachable from the process root, the ``getcwd(2)`` system call returns
|
||||
a path prefixed with ``(unreachable)``.
|
||||
|
||||
A process with its root directory in failfs fails every absolute path
|
||||
lookup including absolute symlinks and the interpreter of dynamically
|
||||
linked binaries. In other words, this fails exec.
|
||||
|
||||
Lookups anchored at explicit directory file descriptors keep working. It
|
||||
is the ``fs_struct`` equivalent of ``RESOLVE_BENEATH``. The process must
|
||||
anchor every lookup at a file descriptor it explicitly holds.
|
||||
|
||||
Entering
|
||||
========
|
||||
|
||||
``fchroot(FD_FAILFS_ROOT, 0)`` requires ``CAP_SYS_CHROOT`` in the
|
||||
caller's user namespace, mirroring ``chroot(2)``. Unprivileged callers
|
||||
may enter if all of the following hold:
|
||||
|
||||
* ``no_new_privs`` is set: setuid binaries on regular mounts remain
|
||||
reachable via inherited directory file descriptors and executing them
|
||||
with an unusable root directory is the classic confused deputy.
|
||||
|
||||
* The caller is not already chrooted: the root directory is what
|
||||
confines ``..`` resolution and the failfs root can never be reached by
|
||||
walking up a real mount tree, so moving the root of a chrooted task to
|
||||
failfs would allow it to escape its chroot via ``openat(fd, "..")``.
|
||||
|
||||
* The caller does not share its ``fs_struct``: ``no_new_privs`` is
|
||||
checked on the calling thread, but the root lives in the ``fs_struct``.
|
||||
A ``CLONE_FS`` sibling without ``no_new_privs`` could otherwise execute
|
||||
a setuid binary with the failfs root, so entry requires ``fs->users ==
|
||||
1``, the same restriction ``setns(2)`` applies for the mount and user
|
||||
namespaces.
|
||||
|
||||
Leaving
|
||||
=======
|
||||
|
||||
Backing out is currently hard, but this is a property of the current
|
||||
implementation, not a guaranteed interface, and may be loosened later.
|
||||
For now a process that entered failfs counts as chrooted, so it cannot
|
||||
create user namespaces to regain ``CAP_SYS_CHROOT``, and ``chroot(2)``
|
||||
or ``fchroot(2)`` back out require ``CAP_SYS_CHROOT``. The remaining way
|
||||
out today is ``setns(2)`` with a mount namespace file descriptor, which
|
||||
requires ``CAP_SYS_ADMIN`` over the target mount namespace as well as
|
||||
``CAP_SYS_CHROOT`` and ``CAP_SYS_ADMIN`` in the caller's user namespace
|
||||
and resets both root and working directory. A process that holds no such
|
||||
file descriptor and restricts ``*chdir()``/``*chroot()``/``setns()`` via
|
||||
seccomp cannot currently get back out.
|
||||
|
|
@ -91,6 +91,7 @@ Documentation for filesystem implementations.
|
|||
ext3
|
||||
ext4/index
|
||||
f2fs
|
||||
failfs
|
||||
gfs2/index
|
||||
hfs
|
||||
hfsplus
|
||||
|
|
|
|||
Loading…
Reference in New Issue
Block a user