bpf_binprm_set_interp() tests path[0] != '/' on the buffer its load
program passes and then reads the same buffer again to copy it with
kmemdup_nul(). The buffer can be a BPF map value that another CPU
rewrites between the two reads. If byte 0 is overwritten in that
window, the kfunc stages a relative or empty interpreter path. The
staged path is not checked again, so open_exec() resolves a relative
path against the working directory of the task doing the exec.
bpf_binprm_set_interp_arg() has the same pattern for its "!len" test
and can stage an empty argument, which the interpreter then receives
as an empty argv entry.
The verifier checks the path and path__sz pair with BPF_READ |
BPF_WRITE, so a writable array map value is an accepted argument.
bpf(BPF_MAP_UPDATE_ELEM) on an array map copies the new value over the
old one in place and takes no lock. Both kfuncs are KF_SLEEPABLE and
allocate with GFP_KERNEL between the test and the copy, so the task
can sleep inside the window:
load program bpf(BPF_MAP_UPDATE_ELEM)
bpf_binprm_set_interp()
strnlen(path, path__sz)
path[0] != '/' is false
kmemdup_nul(path, len, GFP_KERNEL)
allocation may sleep
array_map_update_elem()
copy_map_value()
rewrites byte 0
copy reads path again
bm_bpf_stage_selection()
The test in the load program's column proves what byte 0 held only at
the moment the test ran. The map update takes no lock, so it can store
to byte 0 right after. kmemdup_nul() then copies the rewritten bytes,
and bm_bpf_stage_selection() publishes them as bprm->bpf_interp.
The staged path is not checked again on its way to open_exec():
load_misc_binary()
entry_select_interpreter() returns bprm->bpf_interp unchanged
build_interp_argv()
copy_string_kernel() copies it as argv[0]
bprm_change_interp()
kstrdup()
entry_open_interpreter()
open_exec() unless a bound file is staged or
the entry is an 'F' entry
None of these functions tests the first byte, and load_misc_binary()
hands the pointer to nothing else.
In bpf_binprm_set_interp_arg(), strnlen() finds a non-zero len, a NUL
is then stored to byte 0, and build_interp_argv() later copies the
empty bprm->bpf_interp_arg with copy_string_kernel().
The handler's own load program has to pass a writable map value, and
something has to store into it while the kfunc runs. The allocation can
sleep inside the window, and with a BPF_F_MMAPABLE array the store is a
plain user space write into the mapped value, so a loop can hit it
without a single bpf() call.
Check the private copy in both kfuncs, so that the string that gets
staged is the string that was checked. bpf_binprm_select_interp()
already looks its name up in a private copy for the same reason. The
remaining tests work on path__sz, arg__sz or the local len, and the
copy length is len, so the copy stays inside the extent the verifier
checked.
Results of bpf_binprm_set_interp() with the check on the copy:
- A NUL stored to byte 0 fails interp[0] != '/' and gets -EINVAL.
- For len == 0, kmemdup_nul() returns an empty string, so an empty path
still gets -EINVAL.
- A NUL stored further into the string only shortens it to another
absolute path, or another non-empty argument, that the program could
have passed anyway.
- A path that both lacks the leading '/' and is PATH_MAX or longer now
gets -ENAMETOOLONG instead of -EINVAL.
- A path that is empty or lacks the leading '/' is now rejected after
the copy rather than before it, so such a call makes an allocation
and returns -ENOMEM instead of -EINVAL if that allocation fails.
bpf_binprm_set_interp_arg() still rejects an empty argument before
allocating, so its results are unchanged apart from the raced case
fixed here.
Both new checks run before the previously staged string is freed or
replaced. A failing call frees only its own allocation and leaves the
earlier selection in place, as the -ENOMEM path already does.
Fixes: b4bfe2f6b0 ("binfmt_misc: add binfmt_misc_ops bpf struct_ops")
Signed-off-by: Chris Mason <mason@kernel.org>
Link: https://patch.msgid.link/20260918-work-binfmt_misc-fixes-v1-2-647b24bc1c46@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
bpf_binprm_select_interp() checks the name its load program passes with
strnlen(name, name__sz) and then hands the same buffer to
binfmt_misc_find_interp(), which compares it with an unbounded strcmp().
The buffer can be a BPF map value that another CPU rewrites between the
two reads. If the terminating NUL is overwritten in that window, strcmp()
reads past the name__sz bytes the verifier checked. That is an
out-of-bounds read of up to 31 bytes of whatever follows the checked
name__sz bytes.
The verifier checks the name and name__sz pair with BPF_READ | BPF_WRITE,
so a writable array map value is an accepted argument.
bpf(BPF_MAP_UPDATE_ELEM) on an array map copies the new value over the
old one in place and takes no lock. The NUL that strnlen() finds can be
overwritten before strcmp() reads the buffer again:
CPU0 CPU1
bpf_binprm_select_interp()
strnlen(name, name__sz)
finds the NUL inside name__sz
bpf(BPF_MAP_UPDATE_ELEM)
array_map_update_elem()
copy_map_value()
overwrites the NUL
binfmt_misc_find_interp()
strcmp(interp->name, name)
reads past name__sz
strnlen() proves that a NUL lies inside name__sz only at the moment it
runs. The map update on CPU1 takes no lock, so it can store over the NUL
right after. The lookup on CPU0 then walks the live buffer again, once
per bound interpreter:
fs/binfmt_misc.c:binfmt_misc_find_interp
list_for_each_entry(interp, interps, list)
if (!strcmp(interp->name, name))
return interp;
strcmp() stops at the first mismatch or at the end of interp->name.
bm_entry_add_interp() caps a bound name at BINFMT_MISC_INTERP_NAME_MAX
(32) bytes, so strcmp() reads at most 33 bytes of name. The smallest
name__sz the kfunc accepts is 2, which leaves up to 31 bytes read beyond
the checked extent. The handler's own load program has to pass a
writable map value, and something has to store into it while the kfunc
runs. The window between strnlen() and strcmp() is short, but with a
BPF_F_MMAPABLE array the store is a plain user space write into the
mapped value, so a loop can hit it without a single bpf() call.
Copy the name into a stack buffer of BINFMT_MISC_INTERP_NAME_MAX + 1
bytes, terminate it, and look up the copy. The memcpy() length is below
name__sz, so the copy stays inside the extent the verifier checked, and
the BPF buffer is not read again afterwards.
Return -ENOENT first for a name longer than BINFMT_MISC_INTERP_NAME_MAX.
bm_entry_add_interp() rejects a longer name, and the only other binding
site attaches the empty name. No entry can bind such a name, so that
lookup already ended in -ENOENT and no result changes.
Check the first byte of the copy and return -EINVAL if it is NUL, as the
existing "!len" test does for an empty name. Only an 'F' entry binds the
empty name and a 'B' entry cannot carry 'F', so without that check a
racing store of NUL to byte 0 would look up a name no entry binds and
end in -ENOENT rather than -EINVAL. A NUL stored further into the name
only shortens it to another name the program could have passed anyway.
binfmt_misc_find_interp() itself is left alone: entry_attach_interpreter()
calls it with a kernel string, and this kfunc now calls it with a private
copy.
Fixes: 6ec7c96bee ("binfmt_misc: let a 'B' entry bind its interpreters")
Signed-off-by: Chris Mason <mason@kernel.org>
Link: https://patch.msgid.link/20260918-work-binfmt_misc-fixes-v1-1-647b24bc1c46@kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
init_mount_tree() mounts the mutable rootfs on top of nullfs via
LOCK_MOUNT_EXACT(). That declares a pinned mountpoint with a cleanup
attribute in the scope of the whole function so the nullfs root inode
lock and namespace_sem are only dropped when init_mount_tree() returns.
This became a problem when the private nullfs instance for kthreads was
added. kern_mount() allocates a new superblock and alloc_super() takes
the new s_umount with SINGLE_DEPTH_NESTING and then shrinker_mutex via
shrinker_alloc(). Doing that with namespace_sem held teaches lockdep the
dependency
namespace_sem -> s_umount/1 -> shrinker_mutex
With CONFIG_SHRINKER_DEBUG shrinker_debugfs_rename() takes the debugfs
directory inode lock under shrinker_mutex every time a block device is
mounted and lock_mount_exact() takes namespace_sem under the inode lock
of the mountpoint for every mount. So mounting anything on debugfs,
e.g. the tracefs automount on /sys/kernel/debug/tracing, closes the
cycle:
WARNING: possible circular locking dependency detected
7.3.0-rc3+ #17 Not tainted
------------------------------------------------------
rasdaemon/4449 is trying to acquire lock:
(namespace_sem){++++}-{4:4}, at: lock_mount_exact+0x4c/0x308
but task is already holding lock:
(&sb->s_type->i_mutex_key#17){++++}-{4:4}, at: lock_mount_exact+0x3c/0x308
which lock already depends on the new lock.
...
Chain exists of:
namespace_sem --> shrinker_mutex --> &sb->s_type->i_mutex_key#17
This can't actually deadlock. init_mount_tree() runs single-threaded
during early boot before any other task exists and nothing allocates a
superblock under namespace_sem after that. But lockdep can't know that
and disables itself for the rest of the boot.
Move mounting the rootfs on top of nullfs into a helper so the locks
are dropped when it returns.
Fixes: 32750c77e8 ("fs: start all kthreads in nullfs")
Reported-by: Zenghui Yu <yuzenghui@huawei.com>
Closes: https://lore.kernel.org/15174353-3f4a-a1ca-5bd1-ea2a4c77828e@huawei.com
Link: https://patch.msgid.link/20260917-atemtechnik-bleichen-befassen-9a57db01baf0@brauner
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
cleanup_offline_cgwb() prepares at most WB_MAX_INODES_PER_ISW inodes
per call and is called again until the dying wb is drained, but every
call walks wb->b_attached and then wb->b_dirty_time from the same end.
Inodes already prepared (they stay on the list with I_WB_SWITCH set
until the switch worker runs) and inodes that cannot be switched
(I_FREEING, I_WILL_FREE, !SB_ACTIVE, DAX, already on the target wb)
stay where they are, so each pass rescans a growing run of them under
wb->list_lock and a full drain is quadratic in the number of inodes on
the list. With ~17M inodes attached to one dying cgwb we saw this end
in soft lockups, with CPUs reported stuck for 21-48s.
Walk both lists from the oldest end and move every scanned inode to
the newest end, so the next pass starts where the previous one stopped
and the drain becomes linear. b_attached is unordered, so nobody sees
the reorder there. b_dirty_time is ordered by dirtied_when, but the
oldest unscanned inode stays at the end move_expired_inodes() picks
from, sync takes the whole list regardless of order, and prepared
inodes leave the list as soon as the switch work runs and get a new
dirtied_time_when on the new wb anyway, so the only inodes left out of
order are the ones that can never switch (DAX), and only on the dying
wb.
Fixes: c22d70a162 ("writeback, cgroup: release dying cgwbs by switching attached inodes")
Cc: stable@vger.kernel.org
Acked-by: Tejun Heo <tj@kernel.org>
Acked-by: Roman Gushchin <roman.gushchin@linux.dev>
Signed-off-by: Patrick Lu (Anthropic) <perf.patrick.lu@gmail.com>
Link: https://patch.msgid.link/20260911-wb-cgwb-rotate-v2-1-a9ab253a1295@gmail.com
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
We observed hung tasks when users attempted to unmount a filesystem
after its disk had been removed while still in use. During device
removal, fs_bdev_mark_dead() calls evict_inodes() while holding s_umount.
Each time evict_inodes() drops s_inode_list_lock to reschedule, it
restarts the walk from the head of s_inodes. With many referenced inodes
at the head of the list, these restarts repeatedly scan the same inodes
without reclaiming them. This can keep s_umount held for a long time,
blocking concurrent umount attempts and triggering hung-task reports.
Keep the current inode, already marked I_FREEING, out of the disposal
batch until s_inode_list_lock is reacquired. Resume the walk from this
inode and dispose of it in a later batch or at the end of the walk.
The zero-refcount and state checks under i_lock allow this walker to
claim the inode by setting I_FREEING and removing it from the LRU.
Other reclaimers skip the inode, leaving this walker responsible for
eviction. Only evict() removes it from s_inodes, so keeping it out of
the disposal batch ensures that it remains on the list while the lock
is dropped. After reacquiring the lock, reading its current next pointer
accounts for concurrent removal of following inodes.
The existing inode lifetime rules prohibit acquiring a reference to an
inode marked I_FREEING or I_WILL_FREE. __iget() requires its caller to
hold i_lock and establish that taking a reference is valid. Inode lookup
and igrab() check these flags under i_lock when acquiring a reference
from zero. ihold() requires an existing reference, which would keep
i_count nonzero and prevent this walker from claiming the inode. These
rules already allow iput_final() and the inode shrinker to release
i_lock after setting I_FREEING and before eviction completes.
A temporary __iget() reference would also keep the inode on the list,
but its release must preserve last-reference handling. Another user can
acquire a reference, update lazy timestamps and drop its reference while
the pin is held. If the pin becomes the last reference, dropping it with
atomic_dec_and_test() and evicting directly bypasses iput()'s lazytime
handling and can lose those timestamp updates.
Releasing the pin with iput() preserves that handling, but does not
guarantee eviction. fs_bdev_mark_dead() runs with SB_ACTIVE set, so iput()
may retain the inode in cache, whereas evict_inodes() must evict eligible
zero-reference inodes. The inode may also have been freed when iput()
returns, so the walker cannot then use it to force eviction. Using
I_FREEING preserves the existing eviction behavior without introducing
an additional last-reference transition.
The xfstests auto group passed on ext4 and XFS with known unrelated
failures excluded. No new issues were observed, and the previously
reproducible hung task no longer occurs with this patch.
Fixes: ac05fbb400 ("inode: don't softlockup when evicting inodes")
Signed-off-by: Julian Sun <sunjunchao@bytedance.com>
Link: https://patch.msgid.link/20260915044912.3183440-1-sunjunchao@bytedance.com
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
Fix the reading of symlinks from the cache in afs by making netfslib trim
the amount read down to i_size. The problem is that afs sets the size of
the iterator to the size of the buffer (PAGE_SIZE) so that the cache can
round the read size up to the cache's DIO size.
Note that this also impacts the reading of AFS mountpoints as they're just
stored as symlinks with an odd file mode.
Link: https://patch.msgid.link/3912795.1789489319@warthog.procyon.org.uk
Fixes: c0410adf3d ("afs: Fix the locking used by afs_get_link()")
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Paulo Alcantara <pc@manguebit.org>
cc: Marc Dionne <marc.dionne@auristor.com>
cc: linux-afs@lists.infradead.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
When an abnormal SquashFS image (COMP_OPTS flag is 1 but dictionary size
is 0) is mounted, and performs shift operations using dictionarysize, the
shift exponent is -1, causing a shift-out-of-bounds.
Detail as below:
squashfs_comp_opts(msblk, buffer, length)
squashfs_xz_comp_opts()
if (comp_opts)
n = ffs(opts->dict_size) - 1;<----opts->dict_size=0, n=-1
if (opts->dict_size != (1 << n) && opts->dict_size !=
(1 << n) + (1 << (n + 1))) <----shift-out-of-bounds
Fix it by adding a dictionary size range check before the shift operation.
Fixes: ff750311d3 ("Squashfs: add compression options support to xz decompressor")
Signed-off-by: Ran Hongyun <ranhongyun1@huawei.com>
Link: https://patch.msgid.link/20260713115525.2661734-1-ranhongyun1@huawei.com
Reviewed-by: Phillip Lougher <phillip@squashfs.org.uk>
Reviewed-by: Zhihao Cheng <chengzhihao1@huawei.com>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
ntfs_create_inode() creates a new inode via ntfs_new_inode(). It hashes
it with insert_inode_locked() and so it's marked as I_NEW until
unlock_new_inode().
ntfs 3 calls d_instantiate() in between though... Since the dentry was
already hashed by the lookup before the create any path walk finds it
without touching the parent's i_rwsem and so can lock the inode.
If the inode is a directory unlock_new_inode() calls
lockdep_annotate_inode_mutex_key() and marks i_rwsem with the
i_mutex_dir_key class.
That resets the count and the owner of a lock somebody else may already
hold by now...
syzbot has been spamming us with the same godforsaken bug
"WARNING in do_new_mount"
since 2023. I can't take it anymore so I went looking. Afaict, syzbot's
executor chdirs into a freshly mounted ntfs3 image, creates a
directory and then mounts some pseudofs on it. Everytime the mkdir()
takes longer than syzbot waits mount() runs concurrently:
mkdir("./sys") mount(NULL, "./sys", "sysfs")
ntfs_create_inode()
d_instantiate()
user_path_at() finds the dentry
do_lock_mount()
inode_lock(inode)
namespace_lock()
unlock_new_inode()
lockdep_annotate_inode_mutex_key()
init_rwsem(&inode->i_rwsem)
unlock_mount()
inode_unlock(inode)
The mount side then releases a lock that according to the rwsem nobody
holds:
DEBUG_RWSEMS_WARN_ON((rwsem_owner(sem) != current) && ...):
count = 0x0, magic = 0xffff888043a854e8, owner = 0x0,
curr 0xffff888000244880, list empty
WARNING: CPU: 0 PID: 5346 at kernel/locking/rwsem.c:1368 __up_write
Call Trace:
inode_unlock include/linux/fs.h:877 [inline]
unlock_mount fs/namespace.c:2892 [inline]
do_new_mount_fc fs/namespace.c:3828 [inline]
do_new_mount+0x777/0xa40 fs/namespace.c:3887
On PREEMPT_RT the same thing shows up as
DEBUG_LOCKS_WARN_ON(rt_mutex_owner(lock) != current)
WARNING: kernel/locking/rtmutex_common.h:193 at rt_mutex_slowunlock
The up_write() underflows the reset count. A following inode_lock() on
that directory then never returns. A path walk into the new directory
racing with the mkdir() corrupts the lock the same way via
inode_lock_shared() in lookup_slow().
Switch to d_instantiate_new() and drop the trailing unlock_new_inode().
All error paths bail out before that point with I_NEW still set and
keep using discard_new_inode().
May we never see this fscking bug report again.
Link: https://patch.msgid.link/20260909-work-ntfs3-d_instantiate_new-v1-1-2db697162ce8@kernel.org
Fixes: 82cae269cf ("fs/ntfs3: Add initialization of super block")
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: stable@vger.kernel.org # v5.15+
Reported-by: syzbot+2a13ad6914e6fcec716c@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/6a9beced.a5e650b3.26d8a.000b.GAE@google.com
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
filesystems/eventfd, filesystems/open_tree_ns and filesystems/xattr were
never added to TARGETS when introduced. filesystems/openat2 was moved
from selftests/openat2/ but the TARGETS entry was never updated, leaving a
stale entry pointing at a directory that no longer exists.
Fix this by adding the four missing subdirectories to TARGETS and
removing the stale openat2 entry.
Link: https://lore.kernel.org/20260703150742.58991-1-disgoel@linux.ibm.com
Fixes: 7c37857fc2 ("selftests: add eventfd selftests")
Fixes: b8f7622aa6 ("selftests/open_tree: add OPEN_TREE_NAMESPACE tests")
Fixes: 7e28fef5d4 ("selftests/xattr: path-based AF_UNIX socket xattr tests")
Fixes: fe08792704 ("selftests: move openat2 tests to selftests/filesystems/")
Signed-off-by: Disha Goel <disgoel@linux.ibm.com>
Reviewed-by: Christian Brauner (Amutable) <brauner@kernel.org>
Cc: "Darrick J. Wong" <djwong@kernel.org>
Cc: Jan Kara <jack@suse.cz>
Cc: Jeff Layton <jlayton@kernel.org>
Cc: Shuah Khan <shuah@kernel.org>
Cc: Wen Yang <wenyang.linux@foxmail.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Link: https://patch.msgid.link/20260904183659.B81CD1F00A3D@smtp.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
Commit 14b007e178 added an address check using iter_iov_addr() and a
length check using iter_iov_len() to iov_iter_extract_bvecs(), but these
cannot be used so and are unsafe in this circumstance as the functions have
hardwired assumptions about the iterator type. They should only be used
with ITER_UBUF or ITER_IOVEC-type iterators; they shouldn't be used with
ITER_BVEC, ITER_KVEC, ITER_FOLIOQ, ITER_XARRAY or ITER_DISCARD iterators.
This proves to be a problem for cachefiles as an iterator of type
ITER_FOLIOQ is passed and iter_iov_addr() and iter_iov_len() both
malfunction because iter->__iov in iter_iov() is not pointing to an iovec
array.
Fix this by using iov_iter_alignment() instead.
Fixes: 14b007e178 ("block: validate user space vectors during extraction")
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/1667275.1788941191@warthog.procyon.org.uk
Reviewed-by: Keith Busch <kbusch@kernel.org>
Reviewed-by: Christoph Hellwig <hch@lst.de>
cc: Hannes Reinecke <hare@kernel.org>
cc: Christoph Hellwig <hch@infradead.org>
cc: Jens Axboe <axboe@kernel.dk>
cc: Alexander Viro <viro@zeniv.linux.org.uk>
cc: Paulo Alcantara <pc@manguebit.org>
cc: netfs@lists.linux.dev
cc: linux-block@vger.kernel.org
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
-----BEGIN PGP SIGNATURE-----
iIYEABYKAC4WIQSVyBthFV4iTW/VU1/l49DojIL20gUCaqF0GxAcbWljQGRpZ2lr
b2QubmV0AAoJEOXj0OiMgvbSs+YBALj3Ttl+T8cnEmxExfOYnPt6eL+oIsZFo6HU
zSXUqyiNAQDxtpucp/JgwBNbuk0XA+BfLSVWuw94jdqbPpCrjUW0BA==
=egGv
-----END PGP SIGNATURE-----
Merge tag 'landlock-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux
Pull Landlock fixes from Mickaël Salaün:
"This fixes a use-after-free and a lockdep assert NULL dereferencing,
and properly truncates too-long strings printed by a Landlock
tracepoint. Most of the changes are brought by new tests"
* tag 'landlock-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux:
landlock: Test trace path output boundaries
landlock: Bound escaped trace path output
landlock: Clean up ruleset validation checks
selftests/landlock: Test abstract socket trace name limits
landlock: Fix use-after-free of the source's parent directory
Please consider pulling these changes from the signed vfs-7.3-rc3.fixes tag.
Thanks!
Christian
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQRAhzRXHqcMeLMyaSiRxhvAZXjcogUCaqFWOQAKCRCRxhvAZXjc
otLlAP9X02ybdUt9NndBK8LjslDWwB9hOXzPgYsOKYODEqODjQD/aLpbXVEsA1yy
SLdSDtbtpf+01z4KHorvAakBzk/jrw4=
=OzBX
-----END PGP SIGNATURE-----
Merge tag 'vfs-7.3-rc3.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs
Pull vfs fixes from Christian Brauner:
- netfs:
- Fix an uninitialized return value in netfs_unbuffered_write()
when preparing the first subrequest fails
- For partial unbuffered/DIO writes return the amount transferred
rather than an error
- Update i_size with the amount actually written when a partial
transfer ends in an error
- Fix a subrequest reference leak when the io_iter ends up empty
- Handle netfs_alloc_subrequest() failure during unbuffered writes
- Load all readahead folios into the rolling buffer upfront and
drop the readahead references once the first subrequest is
dispatched
- Mark folios for copy-to-cache while issuing subrequests
- Fix read progress reporting
- afs:
- Add the missing kunmap in the error path of afs_dir_search_bucket()
- Fix a double kunmap in afs_edit_dir_remove()
- Don't free an existing server's endpoint state when cleaning up a
candidate server in afs_lookup_server()
- Unbind peers removed from a server's address list
- ufs:
- Load the cylinder group metadata before creating the root dentry
- Validate the cylinder group index and rotor positions before
caching them
- Treat an unreadable directory block as not empty
- exec:
- Close the close-on-exec files before taking exec_update_lock
Closing a file can block on the filesystem, so a hung filesystem
blocked everything that takes exec_update_lock and a FUSE server
inspecting the calling process could deadlock
- Drop the bprm loader before closing bprm->file in free_bprm()
- exit: Hold a reference to thread_pid across proc_flush_pid()
- reboot: Fix a use-after-free on cad_pid
- nsfs: Keep the namespace tree fields out of the rcu_head used by
kfree_rcu()
- nstree: Check listing permission before taking a namespace
reference in listns()
- super: Return 0 when a nested thaw drops its hold while other
freezers remain
- ext4: Don't set I_METADATA_WRITEBACK during fastcommit replay
- adfs: Free s_fs_info in ->kill_sb()
- autofs: Free the inode info allocated in autofs_fill_super() when
the root inode allocation fails
- ovl: Return EINVAL instead of EIO on a user namespace mismatch now
that it's a plain refusal and not an internal error
- cachefiles: Don't cast the variable-length coherency data to a
__be64 in the coherency tracepoint
* tag 'vfs-7.3-rc3.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (28 commits)
nstree: check listing permission before taking a namespace reference
exec: do_close_on_exec() before taking exec_update_lock
exit: hold a reference to thread_pid across proc_flush_pid
fs: autofs: fix memory leak in autofs_fill_super()
exec: Drop bprm loader before closing bprm->file
afs: Clear stale peer app data after address list changes
afs: Fix incorrect free in candidate cleanup in afs_lookup_server()
afs: Fix double-unmap of directory block
afs: Fix missing kunmap in afs_dir_search_bucket()
ovl: return EINVAL instead of EIO in case of mismatched user_ns
reboot: fix cad_pid use-after-free race
cachefiles: Fix potential UAF/KASAN warning
netfs: Fix read progress reporting
netfs: Mark folios with COPY_TO_CACHE whilst issuing subreqs
netfs: Fix readahead synchronisation issues by loading all folios upfront
netfs: break unbuffered write when netfs_alloc_subrequest() fails
netfs: Fix subreq ref leak
netfs: Fix i_size update for partial transfer
netfs: Fix error vs transferred passed to ->ki_complete()
netfs: Fix unbuffered/DIO write partial transfer error return
...
Just a ton of small fixes all over the place.
Also includes virtio and virtio-rng MAINTAINERS updates.
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
-----BEGIN PGP SIGNATURE-----
iQFDBAABCgAtFiEEXQn9CHHI+FuUyooNKB8NuNKNVGkFAmqhKYcPHG1zdEByZWRo
YXQuY29tAAoJECgfDbjSjVRppBUIAK/QswxhFU0fUQPFQ4YU5xB8/ANGBBpaE1D0
D6g7LYJsB9SguzdiSWOK8BV9/2h8A485yoU98kQBHLCM/Qraclr/t8sNel0Vq3V/
FZmCW21EQZnbcsEbct5WlBlU2veUP2mAhBlRruHEFdMil/W2k4ifF26jFnKAgS8y
ixirBte0LRCo/Ho42D2mZrY40Z1viRKL03Uhl4jiJz+16bx8uRWGd0UELjr7fMT0
095HUYvOdCFLLedhFe9LFN5VFb+gy/iQjv/wOBAcwjjwHNIIAF40gedu8bRh8ovu
ZA17syGe5DjLnp/C1nnEwZhGSDngVEC/aBogKXKJ1MHtVbVZn+A=
=3WXI
-----END PGP SIGNATURE-----
Merge tag 'for_linus' of git://git.kernel.org/pub/scm/linux/kernel/git/mst/vhost
Pull virtio fixes from Michael Tsirkin:
"Just a ton of small fixes all over the place.
Also includes virtio and virtio-rng MAINTAINERS updates"
* tag 'for_linus' of git://git.kernel.org/pub/scm/linux/kernel/git/mst/vhost: (27 commits)
vduse: return compat ioctl results directly
virtio_input: stop callbacks before unregistering input device
virtio_input: reset device if input_register_device() fails
vhost: invalidate vring access on IOTLB transitions
vduse: validate virtqueue alignment
vduse: do not take dev->rwsem in the virtqueue kick path
vhost-scsi: clamp max_io_vqs module parameter
vhost-scsi: use kvzalloc for vq array allocation
virtio-pci: return IRQ_HANDLED after non-zero ISR
virtio: add Eugenio Pérez as Maintainer
vhost: limit outstanding IOTLB misses per virtqueue
MAINTAINERS: Add a section for virtio-rng
vdpa_sim_net: check TX pull result before RX copy
vdpa_sim_blk: reject out-of-range sector starts
virtio-vdpa: Use queue id when setting vq affinity
vdpa: octeon_ep: Check dev_set_name() in dev add
vdpa: ifcvf: Put device on unsupported feature error
vdpa: solidrun: Free IRQs after request failure
vdpa: alibaba: Keep DRIVER_OK clear if IRQ setup fails
vdpa/pds: check virtqueue notify mapping
...
legitimize_ns() takes a reference on the candidate namespace before
may_list_ns() has decided whether the caller may see it. The
__free(ns_put) cleanup on the denied path can drop the last reference to a
mount namespace while we still hold the rcu read lock, and put_mnt_ns()
may sleep there. This is the same problem commit 2ec2aff3c8 ("ns: make
sure reference are dropped outside of rcu lock") fixed for the put_user()
path. Neither ns_requested() nor may_list_ns() needs a reference, both
only look at the namespace type and at the caller's own namespaces, so do
the checks first and take the reference last.
Splat:
Voluntary context switch within RCU read-side critical section!
WARNING: kernel/rcu/tree_plugin.h:332 at rcu_note_context_switch+0x238/0x2a0, CPU#5: a/3442
CPU: 5 UID: 1000 PID: 3442 Comm: a Not tainted 7.0.0-30-generic #30-Ubuntu PREEMPT(lazy)
RIP: 0010:rcu_note_context_switch+0x238/0x2a0
Call Trace:
<TASK>
__schedule+0xcf/0x650
schedule+0x27/0x90
schedule_preempt_disabled+0x15/0x30
__mutex_lock.constprop.0+0x550/0xaf0
__mutex_lock_slowpath+0x13/0x20
mutex_lock+0x3b/0x50
exp_funnel_lock+0xb2/0x260
synchronize_rcu_expedited+0xe7/0x220
namespace_unlock+0x26a/0x320
put_mnt_ns+0xd3/0x120
mntns_put+0xe/0x20
do_listns+0x13e/0x560
__do_sys_listns+0x126/0x2d0
__x64_sys_listns+0x20/0x30
x64_sys_call+0x2366/0x2390
do_syscall_64+0x105/0x5a0
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
Fixes: 76b6f5dfb3 ("nstree: add listns()")
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Link: https://patch.msgid.link/ABA32239-733B-438C-B95A-B13ED69FF0F3@doyensec.com
Reviewed-by: Bradley Morgan <brads@mainlining.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
do_close_on_exec() currently happens while holding the exec_update_lock,
which is used in a lot of places that access process state to
synchronize access checks.
I recently added another such use of exec_update_lock, causing a
regression.
do_close_on_exec() can block waiting for a reply from a filesystem.
That means a hung filesystem can block codepaths that use
exec_update_lock; and it also means that a FUSE filesystem which
attempts to inspect the calling process can deadlock.
To avoid such problems, move do_close_on_exec() before the
exec_update_lock is taken, but after the FD table has been copied if
necessary.
I have looked through all the calls between the old and new position of
the do_close_on_exec() call; there seems to be no file descriptor table
access in between.
Reported-by: Benjamin Peterson <benjamin@locrian.net>
Closes: https://lore.kernel.org/r/f5e8166a-88be-46c5-8939-1e5227ffe4c2@app.fastmail.com
Fixes: 6650527444 ("proc: protect ptrace_may_access() with exec_update_lock (part 1)")
Cc: stable@vger.kernel.org
Signed-off-by: Jann Horn <jannh@google.com>
Link: https://patch.msgid.link/20260907-cloexec-before-exec-update-lock-v1-1-8018c201a7df@google.com
Tested-by: Benjamin Peterson <benjamin@locrian.net>
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEV76QKkVc4xCGURexaDWVMHDJkrAFAmqgZ18ACgkQaDWVMHDJ
krBVhw//cq1xgLpL9e3Y/U21cM0WVN/X02R8Q3TV7sN+x28SJTnN+ZMclwvhTtqZ
F3wdprRbe/KKUZpA1PbxRdlApvHXGJw7QmveiBeO4P+solXVMxsLoEQsdDTRK6i0
dbkUxlEqi+0K4SpzrUa1HKE3fdTFEFDF+bVbm12dw1uOS7Le4qVmPx8xa2tvrWXb
sELQzg5Qzt4VIm9ltx935rXUUVp7fFZDTdnOTqQqj7lPjxp8QdtfNR1kAyy2d1hF
1SdqeYJLLexZqYHSkr49xF4o3pdHoPz55O/isP+3drOReN3a2w9Pj3uADx/WdAH3
+vx1+qSejiSPwneoIdFxDnpayG/TOwpuEjwcnAlp5TVHQNIMzoqJKRsKd1PTaE2e
P62VT3/BNEohlMGPjDZn6d9udUSzckJ9GwFOKOw0M4pqgRbU6Fs5mpJBbjvpyUzA
pI9CRnDSQ2tUWXAr6vypvus5GFCxw7phalqeU0vv0D+u4VedjLpGa6M57QPOOOX1
BpHBb+0mBeFIVdr8NEKvRYPM8wR7dhUGv5AV2m7CsUP8M0uGHU3qNjst6bOzcitB
m7vdKE9mJoxkQ/W1EEYI8xFGDsnKtlhAcA7eyFYbBqCpc83uXwDrV6mDgc+v7dTF
vYzEctEkArJtfQ8zv40uSCivjamCK60SA3At1LyfKIuuQS1lNK4=
=DVVF
-----END PGP SIGNATURE-----
Merge tag 'x86_urgent_for_7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip
Pull x86 fixes from Dave Hansen:
"These are fixes for some older AMD device topology and machine check
issues. But, they are issues that are affecting real users and aren't
just cleaning up AI drive-by reports.
These is coming a wee bit later than the usual Sundays because of a
late breaking issue with one of the patches which is now temporarily
kicked out"
* tag 'x86_urgent_for_7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
x86/MCE/AMD: Fix inverted interrupt enablement during storm handling
x86/amd_node: Fix potential NULL pointer dereference
x86/amd_node: Avoid divide by zero on virtualized systems
Use focused KUnit tests to exercise the renderer's internal boundary and
composition contracts with synthetic scratch states, including both
sibling-helper evaluation orders. Check the exact output and
reservation boundaries, including a four-byte octal escape accepted at
exact capacity and rejected one byte short. Also verify an unchanged
cursor on failure, that bracketed process names and embedded NUL bytes
remain data, and that input ellipsis bytes are escaped rather than
mistaken for the raw truncation marker.
The composition test requires generic trace output helpers. Enable
CONFIG_FTRACE and CONFIG_SCHED_TRACER because the latter selects the
otherwise-hidden CONFIG_TRACING support required by
trace_print_flags_seq().
Use kselftests to exercise the complete tracefs path for both affected
filesystem events. A valid path containing 2640 spaces exceeds the
scratch output budget. Require its escaped prefix to end in the raw
UTF-8 ellipsis while access_rights and blockers remain intact.
This division keeps the exact safety contract compiler-independent while
proving that real tracepoints preserve their surrounding symbolic
fields. The end-to-end assertions fail after a full fix revert with
both GCC and Clang, while the composition KUnit test fails if the
scratch reserve is removed.
Cc: Günther Noack <gnoack@google.com>
Link: https://patch.msgid.link/20260907154401.124362-2-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
Filesystem paths may expand fourfold when trace text escapes spaces and
other untrusted bytes. A sufficiently long representation can exhaust
the shared scratch sequence. A sibling __print_flags() helper may then
return an unterminated one-past pointer because TP_printk() argument
ordering is unspecified.
Use a fixed budget rather than the scratch space available at call time,
so output does not vary with sibling evaluation order. Limit an
untrusted string to three quarters of the trace sequence, leaving the
rest for sibling helpers and final event metadata. Compute and commit
complete escaped output transactionally so an exact fill cannot consume
the terminating NUL or poison the scratch sequence.
For strings that exceed the limit, retain the largest prefix ending at a
complete escape unit, then append a raw UTF-8 ellipsis. Keep the
helper's existing octal fallback so complete values remain unchanged.
Hex fallback would consume the same four bytes per escaped byte without
increasing the prefix or strengthening the marker. ESCAPE_NAP renders
every non-ASCII input byte in octal, so legitimate data cannot reproduce
the marker without being escaped.
Cc: Günther Noack <gnoack@google.com>
Link: https://patch.msgid.link/20260907154401.124362-1-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
landlock_merge_ruleset() checks for a NULL ruleset after dereferencing
it in lockdep_assert_held(). Move the assertion after the check so the
defensive path remains effective.
The mask-validation comment originated in landlock_add_fs_access_mask()
to explain that its WARN_ON_ONCE() checked a caller invariant. It
became self-referential when this helper and its network and scope
counterparts were inlined into landlock_create_ruleset(). Restate the
invariant without naming the caller.
Keep both as defensive callee checks. Moving the assertion preserves
the NULL check's ability to warn and return -EINVAL, while invalid masks
remain warned about and masked.
Reported-by: Günther Noack <gnoack@google.com>
Closes: https://patch.msgid.link/aobYhIt3vcs2xN0b@google.com
Closes: https://patch.msgid.link/aobasxUDQ8b7GYXl@google.com
Reviewed-by: Günther Noack <gnoack@google.com>
Link: https://patch.msgid.link/20260907103609.113325-1-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
The compat handler handles VDUSE_IOTLB_GET_FD and VDUSE_VQ_GET_INFO, but
then calls the native handler. Their different command sizes make native
dispatch return -ENOIOCTLCMD.
For GET_FD, this overwrites receive_fd()'s return value after the
descriptor is installed, leaking one fd per call. Return handled compat
results directly and use native dispatch only for other commands.
Fixes: 455a2a1af9 ("vduse: fix compat handling for VDUSE_IOTLB_GET_FD/VDUSE_VQ_GET_INFO")
Signed-off-by: Linfeng Sun <linfeng.sun.dev@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260908-fix-vduse_dev_compat_ioctl-v1-1-62264d9bfb8d@gmail.com>
mce_amd_handle_storm() currently does the opposite of what storm
handling needs: it enables thresholding interrupts when a storm is
detected and disables them when the storm subsides.
Flip the "on" function argument before passing it to threshold_restart_bank()
as it should have been done.
To clarify: "on" to mce_handle_storm() means, the storm is on now when
"on" is true, and off when "on" is false.
[ bp: Simplify. ]
Fixes: 5c4663ed1e ("x86/mce: Handle AMD threshold interrupt storms")
Signed-off-by: Jasjeet Rangi <jrangi@purestorage.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260812221514.598842-2-jrangi@purestorage.com
amd_smn_read/write() are exported functions around __amd_smn_rw(), so
they are always available even if amd_smn_init() fails. In that case,
'amd_roots' is NULL and __amd_smn_rw() will access uninitialized memory.
Then, commit:
8351845307 ("x86/amd_node: Add SMN offsets to exclusive region access")
added the 'smn_exclusive' flag, which indicated the calls to
pci_request_config_region_exclusive() succeeded, to prevent
concurrent userspace access.
Commit:
0a4b61d9c2 ("x86/amd_node: Fix AMD root device caching")
re-ordered initialization so pci_request_config_region_exclusive() is
called earlier and a failure exits amd_smn_init() before allocating
'amd_roots'. The setting of 'smn_exclusive' moved to the end of
amd_smn_init(), after 'amd_roots' is allocated. It became redundant
and can be removed.
Replace 'smn_exclusive' with directly checking 'amd_roots', to fix a
potential NULL pointer dereference and to simplify the logic.
[ bp: Reorg commit message, touchup comment. ]
[ mingo: Rebase & further touchups. ]
Fixes: 77466b798d ("x86/amd_node: Remove dependency on AMD_NB")
Signed-off-by: Jason Andryuk <jason.andryuk@amd.com>
Signed-off-by: Borislav Petkov (AMD) <bp@alien8.de>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Reviewed-by: Yazen Ghannam <yazen.ghannam@amd.com>
Reviewed-by: Mario Limonciello (AMD) <superm1@kernel.org>
Cc: stable@vger.kernel.org
Link: https://patch.msgid.link/20260825214805.39148-3-jason.andryuk@amd.com
virtinput_remove() unregisters the input device before resetting the
virtio device. virtinput_recv_events() drops vi->lock around input_event(),
so clearing vi->ready does not stop a callback that passed the entry check.
It can still use vi->idev, requeue buffers and kick the queue.
Reset first, as virtinput_freeze() already does. With the preceding core
change, reset waits for callbacks before input_unregister_device() can
free vi->idev. Recheck vi->ready after taking the lock again: keep draining
completed events so an input packet is not truncated, but stop requeueing
buffers and kicking the queue.
With evdev attached, input_unregister_handle() currently waits for an RCU
grace period, which also waits out IRQ callbacks. This masks the lifetime
bug on PCI and MMIO, but does not protect sleepable callbacks on other
transports.
Fixes: 271c865161 ("Add virtio-input driver.")
Assisted-by: LLM
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260905152059.89560-3-kmehltretter@gmail.com>
Probe marks the device DRIVER_OK with virtio_device_ready() before
calling input_register_device(). If registration fails, the error path
cleared vi->ready and called del_vqs() while the device was still live,
so the device could keep DMA to queues that were already torn down.
Match remove/freeze: call virtio_reset_device() on that path before
tearing down the virtqueues.
Fixes: 271c865161 ("Add virtio-input driver.")
Signed-off-by: Xiong Weimin <xiongweimin@kylinos.cn>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260805032931.1606652-1-xiongweimin@kylinos.cn>
When VIRTIO_F_ACCESS_PLATFORM changes, cached vring pointers and IOTLB
metadata are interpreted in a different address space. Keeping them
across the transition can leave stale ring mappings in use.
Clearing d->iotlb before taking the VQ locks also lets a worker observe
a transient NULL d->iotlb and fall back to d->umem while translating a
descriptor.
Add a common vhost_clear_device_iotlb() helper for vhost-net and
vhost-vsock. Take all VQ mutexes in index order before dropping the
device-wide IOTLB, invalidate each VQ's cached ring access and metadata,
clear pending IOTLB messages, and free the old table after the handoff.
This serializes the transition with workers and prevents mixed address
space mappings.
On the first direct-to-IOTLB transition, invalidate the cached vring
addresses. When an existing device IOTLB is replaced, preserve the
GIOVA ring addresses and reset only the metadata cache. After clearing
ACCESS_PLATFORM, userspace must configure the vring addresses for the
new address mode.
vhost_vq_invalidate_access() clears desc, avail, and used together.
Treat the VQ as invalidated only when all three are NULL, since a single
GIOVA address may legitimately be zero.
Fixes: 6b1e6cc785 ("vhost: new device IOTLB API")
Fixes: e13a6915a0 ("vhost/vsock: add IOTLB API support")
Suggested-by: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Jia Jia <physicalmtea@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260828085721.57816-1-physicalmtea@gmail.com>
vduse_validate_config() only checks the upper bound of vq_align. Invalid
values can therefore reach vring_create_virtqueue_map(). The split-ring
helpers use align - 1 as a bit mask, so the alignment must be a non-zero
power of two. A zero value makes vring_size() drop the descriptor and
available-ring part and vring_init() leave the used ring pointer NULL.
The VIRTIO spec requires the used ring to start at an address
aligned to at least 4 bytes. Reject values below VRING_USED_ALIGN_SIZE as
well as non-power-of-two values before they reach the virtio ring helpers.
Opening a virtio-net device created with vq_align=0 triggered:
BUG: KASAN: null-ptr-deref in virtqueue_kick_prepare_split+0xe3/0x100
Read of size 2 at addr 0000000000000000 by task systemd-network/1062
Call Trace (relevant frames):
dump_stack_lvl
print_report
kasan_report
__asan_load2
virtqueue_kick_prepare_split+0xe3/0x100
virtqueue_kick_prepare+0x40/0x60
try_fill_recv+0x857/0x1250
virtnet_open+0x189/0x460
__dev_open+0x225/0x390
__dev_change_flags+0x368/0x3b0
netif_change_flags+0x56/0xc0
do_setlink.isra.0+0x68c/0x1e30
Validate the value before it reaches the virtio ring helpers.
Fixes: c8a6153b6c ("vduse: Introduce VDUSE - vDPA Device in Userspace")
Signed-off-by: Jia Jia <physicalmtea@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260830023354.115333-1-physicalmtea@gmail.com>
vduse_vq_kick() runs in the context of the vdpa .kick_vq callback. With
the virtio_vdpa bus driver that callback is invoked by virtqueue_notify()
from the virtio device driver, which may be an atomic context: virtio-blk
kicks from ->queue_rq(), which blk-mq dispatches under rcu_read_lock()
(the tag set does not use BLK_MQ_F_BLOCKING), and virtio-net kicks from
its xmit path with the tx queue lock held.
Commit b282418bc3 ("vduse: Add suspend") made vduse_vq_kick() take
dev->rwsem for reading in order to check dev->suspended. down_read() may
sleep, so with CONFIG_DEBUG_ATOMIC_SLEEP the first I/O on a VDUSE-backed
virtio-blk device bound to virtio_vdpa now triggers:
BUG: sleeping function called from invalid context at kernel/locking/rwsem.c:1573
in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 27, name: kworker/1:0H
preempt_count: 0, expected: 0
RCU nest depth: 1, expected: 0
3 locks held by kworker/1:0H/27:
#0: ((wq_completion)kblockd){+.+.}-{0:0}, at: process_one_work+0xac7/0xcf0
#1: ((work_completion)(&(&hctx->run_work)->work)){+.+.}-{0:0}, at: process_one_work+0x51f/0xcf0
#2: (rcu_read_lock){....}-{1:3}, at: blk_mq_run_work_fn+0x119/0x220
Workqueue: kblockd blk_mq_run_work_fn
Call Trace:
<TASK>
dump_stack_lvl+0x80/0xa0
__might_resched+0x231/0x370
down_read+0x73/0x330
vduse_vq_kick+0x30/0x120
virtio_vdpa_notify+0x63/0x80
virtqueue_notify+0x45/0x70
virtio_queue_rq+0x19d/0x300
blk_mq_dispatch_rq_list+0x269/0xe20
__blk_mq_sched_dispatch_requests+0x761/0xa60
blk_mq_sched_dispatch_requests+0x6b/0xc0
blk_mq_run_work_fn+0x143/0x220
process_one_work+0x581/0xcf0
worker_thread+0x2fc/0x5a0
kthread+0x1cc/0x210
ret_from_fork+0x3c4/0x540
ret_from_fork_asm+0x1a/0x30
</TASK>
Without CONFIG_DEBUG_ATOMIC_SLEEP, a kick that finds the rwsem
write-locked by vduse_dev_reset() or vduse_vdpa_suspend() blocks inside
an RCU read-side critical section. The vhost_vdpa path kicks from the
vhost worker, i.e. process context, which is why this went unnoticed.
Check dev->suspended under vq->kick_lock instead, which the kick path
already takes, and have vduse_vdpa_suspend() cycle every virtqueue's
kick_lock after setting the flag. A kick that observed suspended == false
has thus finished signalling before suspend returns, which is the
guarantee the rwsem used to provide. The flag is now also read outside
the rwsem, so access it with READ_ONCE()/WRITE_ONCE().
Fixes: b282418bc3 ("vduse: Add suspend")
Signed-off-by: Nikhil <nikhilljatt@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260829225457.1037867-1-nikhilljatt@gmail.com>
max_io_vqs is currently validated only when a vhost-scsi device is opened.
This allows sysfs to show values larger than the driver will actually use,
e.g. writing 2048 succeeds even though vhost_scsi_open() later clamps it to
VHOST_SCSI_MAX_IO_VQ. This makes the sysfs value differ from the value that
will actually be used.
hv# echo 2048 > /sys/module/vhost_scsi/parameters/max_io_vqs
hv# cat /sys/module/vhost_scsi/parameters/max_io_vqs
2048
[ 315.630495] Invalid max_io_vqs of 2048. Using 1024.
Keep accepting out-of-range values for compatibility, but clamp them in the
module parameter setter and store the effective value. This preserves the
existing behavior that invalid values do not make module loading or sysfs
writes fail. It also makes reads report the value that will actually be
used.
With the parameter value kept in range, remove the duplicate validation
from vhost_scsi_open().
Signed-off-by: Dongli Zhang <dongli.zhang@oracle.com>
Reviewed-by: Mike Christie <michael.christie@oracle.com>
Reviewed-by: Stefan Hajnoczi <stefanha@redhat.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260802172534.260047-3-dongli.zhang@oracle.com>
vp_interrupt() reads the ISR before dispatching config-change and
vring handling. Reading the ISR also clears it, so once the read
returns non-zero the interrupt was from this device and has already
been consumed.
Currently vp_interrupt() returns the result of vp_vring_interrupt().
For a config-change interrupt with no vring work, that can return
IRQ_NONE even though the ISR was non-zero and the interrupt was
handled.
Call vp_vring_interrupt() for any queue work, but once the ISR is
non-zero return IRQ_HANDLED.
Tested with QEMU virtio-blk-pci forced to INTx using vectors=0 and
pci=nomsi. On an idle device, 200 config-change interrupts were
generated using QMP block_resize.
Before this change, irq_handler_exit reported ret=unhandled and
/proc/irq/11/spurious increased from 0 to 200 unhandled interrupts.
After this change, irq_handler_exit reported ret=handled and the
unhandled count remained at 0.
The issue was found during an LLM-assisted Quality Playbook review.
Fixes: 77cf524654 ("virtio_pci: split up vp_interrupt")
Suggested-by: Michael S. Tsirkin <mst@redhat.com>
Assisted-by: LLM
Signed-off-by: Andrew Stellman <astellman@stellman-greene.com>
Message-ID: <20260904141318.30278-1-astellman@stellman-greene.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
vhost allocates a message node whenever address translation misses. If
userspace reads these messages without resolving them, repeated virtqueue
kicks can grow the pending message list until the host runs out of memory.
Virtqueue processing stops at the first translation miss and cannot make
progress until userspace installs a mapping. Keep a pointer to that
outstanding message in the virtqueue and suppress additional misses until
the node is resolved or discarded.
The pointer remains set while the message is queued for reading, copied to
userspace, or waiting on the pending list. Clear it under the IOTLB lock
when the owning node is freed. This bounds outstanding miss messages by the
fixed number of virtqueues without introducing an arbitrary queue limit.
Signed-off-by: Linfeng Sun <linfeng.sun.dev@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260903-fix-kernel-panic-in-vhost_iotlb_miss_pending_list-v1-1-39b8cd427978@gmail.com>
At Michael's request, add a MAINTAINERS entry for the virtio-rng driver
and list myself as its maintainer.
I already maintain the corresponding QEMU implementation.
Cc: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Laurent Vivier <lvivier@redhat.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260818133913.162471-1-lvivier@redhat.com>
vringh_iov_pull_iotlb() returns a signed byte count. A failed TX pull is
currently added to the unsigned byte counter and then passed as a size_t
length to receive_filter() and vringh_iov_push_iotlb(). A negative error
can therefore become a large length in the RX path.
Handle non-positive pull results before every length use. Count the TX
error and complete the consumed TX descriptor with zero bytes.
I found this bug myself, though the patch was written with AI assistance.
Fixes: cfe2268929 ("vdpa_sim: filter destination mac address")
Assisted-by: OpenAI-Codex:GPT-5
Signed-off-by: Linfeng Sun <linfeng.sun.dev@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260901094842.25875-1-linfeng.sun.dev@gmail.com>
vdpasim_blk_check_range() logs an invalid start sector but continues
validating the request. The subsequent unsigned capacity subtraction can
underflow and let an out-of-range buffer offset reach the data path.
The invalid offset is used by three request paths. VIRTIO_BLK_T_OUT
copies guest data to blk->buffer + offset through
vringh_iov_pull_iotlb(), causing an out-of-bounds write in
_copy_from_iter() or memcpy(). VIRTIO_BLK_T_IN copies from
blk->buffer + offset to the guest through vringh_iov_push_iotlb(),
causing an out-of-bounds read in _copy_to_iter().
VIRTIO_BLK_T_WRITE_ZEROES passes blk->buffer + offset to memset(),
causing an out-of-bounds write.
Reject starts at or beyond the capacity before the subtraction. Treat the
capacity boundary as invalid because the IN and OUT paths round byte counts
down to sectors for validation but later copy the original byte counts. A
sub-sector request at the capacity boundary would otherwise still access
past the end of the buffer.
I found this bug myself, though the patch was written with AI assistance.
Fixes: 7d189f617f ("vdpa_sim_blk: implement ramdisk behaviour")
Assisted-by: OpenAI-Codex:GPT-5
Signed-off-by: Linfeng Sun <linfeng.sun.dev@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260901094800.25475-1-linfeng.sun.dev@gmail.com>
When optional queues are skipped, pass the compressed vDPA queue id to
set_vq_affinity() so affinity is applied to the queue that was actually
created.
Signed-off-by: Xiong Weimin <xiongweimin@kylinos.cn>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260804092649.1344478-1-xiongweimin@kylinos.cn>
Handle dev_set_name() failures before registering the vDPA device so
allocation is unwound through the existing put_device() path.
Signed-off-by: Xiong Weimin <xiongweimin@kylinos.cn>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260804092636.1344431-1-xiongweimin@kylinos.cn>
Route unsupported provisioned features through the common error path after
vdpa_alloc_device() so the allocated device and adapter pointer are
released consistently.
Fixes: 46fc0917bb ("vDPA/ifcvf: implement features provisioning")
Cc: stable@vger.kernel.org # v6.3+
Signed-off-by: Xiong Weimin <xiongweimin@kylinos.cn>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <178589471294.1556376.4816776800128323034@kylinos.cn>
Unwind IRQs already requested by snet_request_irqs() before returning a
VQ IRQ request error so a later DRIVER_OK retry starts from a clean
state. The IRQs are requested and freed while the PCI device remains
bound, so the driver cannot wait for devres cleanup at detach time.
Fixes: 51a8f9d7f5 ("virtio: vdpa: new SolidNET DPU driver.")
Cc: stable@vger.kernel.org # v6.3+
Signed-off-by: Xiong Weimin <xiongweimin@kylinos.cn>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <178589471328.1556376.15570536900532373521@kylinos.cn>
If requesting MSI-X interrupts fails while DRIVER_OK is being set, leave
the device status unchanged instead of advertising a ready device without
working interrupts.
Signed-off-by: Xiong Weimin <xiongweimin@kylinos.cn>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260804092608.1344269-1-xiongweimin@kylinos.cn>
vp_modern_map_vq_notify() can fail and return NULL. Check the notify
mapping while adding a pds vDPA device and use the existing teardown path
instead of storing a NULL doorbell pointer in the virtqueue state.
Signed-off-by: Xiong Weimin <xiongweimin@kylinos.cn>
Reviewed-by: Brett Creeley <brett.creeley@amd.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260806005809.1875257-1-xiongweimin@kylinos.cn>
When the DT node has "wakeup-source", vm_find_vqs() calls
enable_irq_wake() on the shared IRQ, but vm_del_vqs() freed that IRQ
without a matching disable_irq_wake(). That leaves a wake reference
behind and can warn on later free_irq()/request_irq() cycles.
Record whether enable_irq_wake() succeeded, and disable it in
vm_del_vqs() before free_irq().
Fixes: 02213273f7 ("virtio_mmio: add support to set IRQ of a virtio device as wakeup source")
Cc: stable@vger.kernel.org
Signed-off-by: Xiong Weimin <xiongweimin@kylinos.cn>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260805032937.1606737-1-xiongweimin@kylinos.cn>
vhost_vdpa_config_cb() loads v->config_ctx and signals it without taking
a reference and without holding any lock:
struct eventfd_ctx *config_ctx = v->config_ctx;
if (config_ctx)
eventfd_signal(config_ctx);
VHOST_VDPA_SET_CONFIG_CALL replaces that field and drops what is normally
the last reference to the old context:
swap(ctx, v->config_ctx);
if (ctx)
eventfd_ctx_put(ctx);
eventfd_ctx_put() drops the last kref and frees the context immediately,
with no RCU grace period, so a callback that has already loaded the
pointer goes on to dereference freed memory. The two sides share no
lock: the ioctl runs under vhost_dev.mutex, while the parent invokes the
callback from its own interrupt or workqueue context.
This is not the reopen refcount underflow fixed by commit f6bbf0010b
("vhost-vdpa: fix use-after-free of v->config_ctx"), which was about
vhost_vdpa_config_put() leaving a stale pointer behind. Here the pointer
is maintained correctly and it is the read side that is unprotected.
With VDUSE as the parent this is reachable from userspace with access to
/dev/vduse (root by default). VDUSE_DEV_INJECT_CONFIG_IRQ queues
dev->inject, and vduse_dev_irq_inject() runs the callback under VDUSE's
own dev->irq_lock, which vhost does not hold. vduse_dev_reset() does
flush_work(&dev->inject), but VHOST_VDPA_SET_CONFIG_CALL never goes
through reset, so an inject already in flight is not waited for. A
process that injects config interrupts on the VDUSE fd while another
thread swaps the call fd on the vhost-vdpa fd hits it in seconds:
BUG: KASAN: slab-use-after-free in native_queued_spin_lock_slowpath
Read of size 4 at addr ffff888107d21808 by task kworker/u17:1/2993
Workqueue: vduse-irq vduse_dev_irq_inject
Call Trace:
native_queued_spin_lock_slowpath+0x97/0x5b0
_raw_spin_lock_irqsave+0xd4/0xe0
eventfd_signal_mask+0x69/0x120
vhost_vdpa_config_cb+0x34/0x50
vduse_dev_irq_inject+0x46/0x60
process_one_work+0x468/0x950
Allocated by task 2992:
do_eventfd+0x50/0x200
__x64_sys_eventfd2+0x2e/0x40
Freed by task 2992:
eventfd_ctx_put+0xb9/0xc0
vhost_vdpa_unlocked_ioctl+0x116c/0x2190
Add a spinlock covering every access to config_ctx, so the callback
either signals a context that is still alive or observes NULL, and the
put happens only once no callback can reach the old value.
Clearing the parent's callback before the put would not be enough: of the
in-tree set_config_cb() implementations only VDUSE takes a lock, the rest
store the pointer unlocked, so that would not order against an in-flight
invocation.
Fixes: 776f395004 ("vhost_vdpa: Support config interrupt in vdpa")
Signed-off-by: Yu Zhang <yuz08559@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260807100025.19750-3-yuz08559@gmail.com>
vhost_vdpa_set_config_call() swaps the eventfd_ctx_fdget() return value
into v->config_ctx before checking it, so on failure the field briefly
holds an ERR_PTR:
ctx = fd == VHOST_FILE_UNBIND ? NULL : eventfd_ctx_fdget(fd);
swap(ctx, v->config_ctx);
if (!IS_ERR_OR_NULL(ctx))
eventfd_ctx_put(ctx);
if (IS_ERR(v->config_ctx)) {
long ret = PTR_ERR(v->config_ctx);
v->config_ctx = NULL;
return ret;
}
Commit 0bde59c172 ("vhost-vdpa: set v->config_ctx to NULL if
eventfd_ctx_fdget() fails") added that clearing, and spelled out the
invariant the rest of the file relies on: "we consider 'v->config_ctx'
valid if it is not NULL". The window between the swap and the clearing
still breaks it. vhost_vdpa_config_cb() only tests for NULL, so a config
interrupt delivered inside the window hands the ERR_PTR to
eventfd_signal().
Check the fd before installing it instead. That closes the window and
matches how vhost_vring_ioctl() handles the same failure for the vq call
fd.
It also stops a rejected fd from tearing down a config interrupt that was
working: until now the swap replaced the live context and put it, so
after an EBADF the device silently stopped delivering config interrupts
until userspace installed a new fd.
Fixes: 776f395004 ("vhost_vdpa: Support config interrupt in vdpa")
Signed-off-by: Yu Zhang <yuz08559@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260807100025.19750-2-yuz08559@gmail.com>
vhost_vring_set_num() accepts any non-zero power-of-two queue size that
fits in 16 bits. vhost-vdpa then passes that value to set_vq_num()
without comparing it with get_vq_num_max().
A process with access to /dev/vhost-vdpa-* can therefore configure a
queue larger than the device advertises. With vdpa_sim, the worker can
walk descriptors beyond the mapped descriptor ring. KASAN reports a
16-byte out-of-bounds read, corresponding to one vring_desc, in the
vringh IOTLB path:
BUG: KASAN: out-of-bounds in _copy_from_iter
Read of size 16
copy_from_iotlb
copydesc_iotlb
vringh_getdesc_iotlb
vdpasim_net_work
Cache get_vq_num_max() immediately after reset. Some backends derive
it from writable queue-size state, so querying it after SET_NUM may
return the current size instead of the device capability. Invalidate
the cached value before reset so a failed reset leaves SET_NUM
disabled.
For VHOST_SET_VRING_NUM, copy the complete vring state once and use
the same index and size for validation, vq->num, and set_vq_num().
This ensures that validation and use operate on the same copied values.
Fixes: 4c8cf31885 ("vhost: introduce vDPA-based backend")
Signed-off-by: Jia Jia <physicalmtea@gmail.com>
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Message-ID: <20260810010300.132959-1-physicalmtea@gmail.com>