Commit Graph

1460583 Commits

Author SHA1 Message Date
Guopeng Zhang
1765a153d9 cgroup/pids: Restore pids.events notifications in local mode
A fork rejected by the pids controller increments the counter reported by
pids.events. When local event accounting is selected, however, pids_event()
returns after notifying only events_local_file, leaving pids.events pollers
asleep.

On legacy hierarchies, pids.events.local does not exist. With
pids_localevents, pids.events reports the same local counter. In both
cases, pids.events changes without generating a notification.

This can be reproduced with a pids_localevents mount:

    mkdir /tmp/test
    mount -t cgroup2 -o pids_localevents none /tmp/test
    mkdir /tmp/test/t
    echo 1 > /tmp/test/t/pids.max
    cat /tmp/test/t/pids.events                 # max 0
    timeout 3 inotifywait -e modify /tmp/test/t/pids.events &
    sh -c 'echo $$ > /tmp/test/t/cgroup.procs; (true &)' 2>/dev/null
    wait
    cat /tmp/test/t/pids.events                 # max 1

Without this patch, inotifywait times out without reporting an event.
Notify pids.events before returning from the local event path.

Fixes: 3f26a885a0 ("cgroup/pids: Add pids.events.local")
Cc: stable@vger.kernel.org # v6.11+
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-24 06:22:11 -10:00
Eva Kurchatova
c774ec8f0a selftests: cgroup: give the O_TMPFILE open in get_temp_fd() a mode
O_TMPFILE, like O_CREAT, needs the third argument. Without it glibc
refuses the call at compile time as soon as fortification is on:

  In function 'open',
      inlined from 'get_temp_fd' at test_memcontrol.c:33:9:
  /usr/include/bits/fcntl2.h:52:11: error: call to '__open_missing_mode'
    declared with attribute error: open with O_CREAT or O_TMPFILE in
    second argument needs 3 arguments

The fortify checks take effect only once the compiler optimises, and
cgroup/Makefile builds with "-Wall -pthread" alone, so this goes
unnoticed in a plain build. Building the tests with the flags
distributions commonly use, -O2 -D_FORTIFY_SOURCE=3, loses
test_memcontrol entirely.

Fixes: 84092dbcf9 ("selftests: cgroup: add memory controller self-tests")
Signed-off-by: Eva Kurchatova <eva.kurchatova@virtuozzo.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-16 11:19:14 -10:00
Michal Koutný
057dac23d3 cgroup: Avoid iteration of dying tasks with zero refcount
The commit 260fbcb92b ("cgroup: Move dying_tasks cleanup from
cgroup_task_release() to cgroup_task_free()") extended the lifetime of
tasks on the dying_tasks list.
The iterators have provision to go through dying_tasks because of
dying threadgroup leaders or explicit CSS_TASK_ITER_WITH_DEAD, however,
it was expected that such tasks can obtain a new reference (that is
possible before cgroup_task_release()/put_task_struct_rcu_user()).
The tasks after cgroup_task_release() and before cgroup_task_free()
are subject to race when they may or may not have ->usage count > 0.

The race window is between css_task_iter_next() invocations
when css_set_lock is released and we may arrive at a new ->task_pos.
The iterator should not attempt to resurrect tasks whose ->usage count
dropped to zero. (When that happens, __put_task_struct_rcu_cb() is
already imminent and the returned task_struct would could be used
after free.)

As for the fix, we cannot simply check the signal->live count of a task
on the dying list because that won't distinguish regular zombies waiting
to be reaped from RCU remnant tasks that are going to be free'd.
Therefore add an extra check to rule out ->usage==0 tasks from any
iteration.

The repeat: loop in css_task_iter_advance() doesn't consider ->usage
count, so add a new loop to css_task_iter_next() to skip de-used tasks
on the dying_list.

Rough illustration of the possible race

  R (reader of cgroup.procs)         T (thread)                       L (group leader)
  ---------------------------------  -------------------------------- --------------------------------
                                                                      L exits, signal->live > 0
                                                                      cgroup_task_dead(L)
                                                                        css_set_skip_task_iters() // skips only cset->tasks
                                                                        list_add_tail(&L->cg_list, &cset->dying_tasks)
  css_task_iter_next()
    take css_set_lock
    css_task_iter_advance()
      leader && signal->live != 0
      => it->task_pos = &L->cg_list
    release css_set_lock
                                     T exits
                                     --signal->live == 0
				     cgroup_task_dead(T) // css_set_lock
                                     release_task(T)
                                       cgroup_task_release(T)
                                       release_task(L) // zap_leader
                                         cgroup_task_release(L)
                                         put_task_struct_rcu_user(L)
                                         ...RCU...
                                         put_task_struct(L)
                                           L->usage = 0
                                           /* L still on dying_tasks */
                                           ...RCU...
                                           __put_task_struct(L)
  css_task_iter_next() // another iteration
    take css_set_lock
    it->task_pos = &L->cg_list
    get_task_struct(L)
      => addition on 0
    drop css_set_lock
                                           cgroup_task_free(L)
                                             css_set_skip_task_iters() // dying skip comes too late
                                           free_task(L)
  cgroup_procs_show()
    task_pid_vnr(L)

Fixes: 260fbcb92b ("cgroup: Move dying_tasks cleanup from cgroup_task_release() to cgroup_task_free()")
Cc: stable@vger.kernel.org # v6.19+
Link: https://lists.debian.org/debian-kernel/2026/08/msg00220.html
Reported-by: Noah Elias Feldt <N.Feldt@mittwald.de>
Reported-by: Salvatore Bonaccorso <carnil@debian.org>
Tested-by: Salvatore Bonaccorso <carnil@debian.org>
Signed-off-by: Michal Koutný <mkoutny@suse.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-09-14 12:43:50 -10:00
Etienne Perot
3f4b7d1a49 selftests/cgroup: test clone3() into a previously killed cgroup
Once cgroup.kill had been written to a cgroup, a stale kill_seq
snapshot (taken in cgroup_css_set_fork() before the target cgroup was
resolved) caused every child subsequently cloned into that cgroup with
clone3(CLONE_INTO_CGROUP) to be SIGKILLed on the spot.

Add a regression test: create a cgroup, kill it while it is empty,
then clone a child into it and check that the child runs and exits
cleanly. On a kernel without the fix, the test fails:

  not ok 4 test_cgkill_clone_into_killed

The test is skipped on kernels without clone3() or without
CLONE_INTO_CGROUP.

Cc: Shakeel Butt <shakeel.butt@linux.dev>
Assisted-by: LLM
Signed-off-by: Etienne Perot <eperot@google.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-31 06:19:53 -10:00
Etienne Perot
8e35992021 cgroup: fix spurious SIGKILL of CLONE_INTO_CGROUP children
Since commit b69bb476de ("cgroup: fix race between fork and
cgroup.kill"), the fork path snapshots the kill_seq of the child's
future cgroup into kargs->kill_seq, and cgroup_post_fork() SIGKILLs
the child if that cgroup's kill_seq has changed in the meantime, to
catch forks racing with a cgroup.kill sweep.

For CLONE_INTO_CGROUP, however, the snapshot in cgroup_css_set_fork()
is taken before the target cgroup has been resolved: kargs->cgrp is
always NULL at this point (it is only set at the end of the function).
So the "if (kargs->cgrp)" branch is dead code and the snapshot always
records the kill_seq of the parent's cgroup. cgroup_post_fork() then
compares it with the kill_seq of the target cgroup, so the child gets
SIGKILLed whenever the two cgroups have been killed a different number
of times.

As a result, once cgroup.kill has been written to a cgroup, every
child subsequently cloned into it with clone3(CLONE_INTO_CGROUP) is
killed on the spot, for as long as the cgroup exists: kill_seq is not
exposed to userspace and never resets.

Re-snapshot kill_seq from the target cgroup once it has been resolved,
and drop the dead branch at the early snapshot site.

This does not reopen the race fixed by b69bb476de. For
CLONE_INTO_CGROUP, everything from the snapshot to the check in
cgroup_post_fork() runs with cgroup_mutex held, and kill_seq is
only ever incremented under cgroup_mutex.

tj: Updated the comment above kill_seq to reflect the new serialization
rules as suggested by Shakeel Butt.

Fixes: b69bb476de ("cgroup: fix race between fork and cgroup.kill")
Cc: stable@vger.kernel.org
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Assisted-by: LLM
Signed-off-by: Etienne Perot <eperot@google.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-31 06:19:38 -10:00
Guopeng Zhang
87d347a8c8 selftests/cgroup: Add test for preserving boot-isolated CPUs
Put a CPU isolated at boot into an isolated partition, change the
partition back to member and check that the CPU remains isolated.

Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Reviewed-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-24 07:01:44 -10:00
Guopeng Zhang
6c37d7e074 cgroup/cpuset: Preserve boot-isolated CPUs on partition release
isolated_cpus tracks CPUs isolated with isolcpus= as well as CPUs in
isolated cpuset partitions. When an isolated partition is released,
isolated_cpus_update() removes its whole CPU mask. This also clears CPUs
which were already isolated at boot.

This can be reproduced on a cgroup v2 system booted with
isolcpus=domain,15:

    cd /sys/fs/cgroup
    echo +cpuset > cgroup.subtree_control
    mkdir cpuset-repro
    echo 15 > cpuset-repro/cpuset.cpus
    echo isolated > cpuset-repro/cpuset.cpus.partition
    echo member > cpuset-repro/cpuset.cpus.partition
    cat cpuset.cpus.isolated

CPU 15 is absent before the change. It must remain in
cpuset.cpus.isolated after the partition is released.

Update isolated_cpus one CPU at a time and keep CPUs outside the
boot-time domain housekeeping mask isolated.

Fixes: c188f33c86 ("cgroup/cpuset: Account for boot time isolated CPUs")
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Acked-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-24 07:01:30 -10:00
Guopeng Zhang
2bf404b1bd selftests/cgroup: Drop invalid boot isolation comparison
check_isolcpus() clears ISOLCPUS before rebuilding it from sched domain
data. Comparing that empty value with
/sys/devices/system/cpu/isolated makes the test fail whenever
isolcpus=domain is present.

That sysfs file is generated from HK_TYPE_DOMAIN_BOOT and does not change
when cpuset updates HK_TYPE_DOMAIN. Re-reading it cannot validate dynamic
housekeeping updates. The cpuset.cpus.isolated and sched domain checks
already cover the two dynamic interfaces, so remove the invalid comparison.

This can be reproduced on a kernel booted with isolcpus=domain,15:

    # tools/testing/selftests/cgroup/test_cpuset_prs.sh

The test fails its first state-matrix isolation check before the change and
continues past that check afterward.

Fixes: 6df415aa46 ("cgroup/cpuset: Defer housekeeping_update() calls from CPU hotplug to workqueue")
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Reviewed-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-24 07:01:15 -10:00
Cheng Lingfei
909a3f0e9d docs: cgroup-v2: fix misc.events key format description
In misc cgroup, misc.events does not output a simple "max" key. Instead,
each registered misc resource outputs a separate key suffixed with ".max"
(i.e., "<res>.max").
Update the documentation to clarify that the entry key is "<res>.max".

Suggested-by: Michal Koutný <mkoutny@suse.com>
Signed-off-by: Cheng Lingfei <chenglingfei@foxmail.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-24 06:55:59 -10:00
Hongfu Li
a8c6daab4b selftests/cgroup: Fix cg_run_in_subcgroups ignoring arg parameter
cg_run_in_subcgroups() discards its arg and always passes NULL to cg_run(),
turning the (void *)100 from test_kmem_dead_cgroups() into NULL so no
allocation occurs.

This makes test_kmem_dead_cgroups() falsely pass without exercising the
"dying cgroup with charged slab" scenario it intends to test.

Pass the arg through to cg_run() to fix this.

Fixes: 933dc80ec2 ("kselftests: cgroup: add kernel memory accounting tests")
Signed-off-by: Hongfu Li <lihongfu@kylinos.cn>
Reviewed-by: Michal Koutný <mkoutny@suse.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-20 07:48:01 -10:00
Hemanth Selam
0c893d170f selftests/cgroup: set the test plan after the setup checks
The cgroup tests announce their plan before checking whether cgroup v2 is
available, so on a host without it they promise a number of results and
then skip out after the first one:

	TAP version 13
	1..3
	ok 1 # SKIP cgroup v2 isn't mounted
	# Planned tests != run tests (3 != 1)
	# Totals: pass:0 fail:0 xfail:0 xpass:0 skip:1 error:0

ksft_exit_skip() can only emit a well formed "1..0 # SKIP" line while no
plan has been printed, as the comment above it in kselftest.h points out.

Move ksft_set_plan() below the setup checks that can skip, so that a
skipped run reports:

	TAP version 13
	1..0 # SKIP cgroup v2 isn't mounted

Several of the tests skip more than once while setting up, for a missing
or unwritable controller as well, so the plan goes after the last of
them.  test_core joins its two setup paths at the post_v2_setup label and
sets the plan there.

Reporting each planned test as skipped instead would keep the plan where
it is, but the setup failures here mean the whole test cannot run rather
than its individual cases being skipped, which is what "1..0 # SKIP" is
for.

Fixes: 1dc830ee4c ("selftests/cgroup: conform test to KTAP format output")
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Reviewed-by: Sarthak Sharma <sarthak.sharma@arm.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-19 10:04:43 -10:00
Shaojie Sun
2d19207f3f selftests/cgroup: Remove redundant chown in test_cgcore_lesser_ns_open
test_cgcore_lesser_ns_open runs as root throughout and never changes its
euid, so chowning the two cgroup.procs files to a non-root uid has no
effect on the test.

The ENOENT the test expects comes from the cgroup namespace delegation
check in cgroup_procs_write_permission(): the source and destination
cgroups must both be descendants of the namespace root captured at open
time.  That check does not depend on file ownership.  In addition, the
permission check only examines the common ancestor's cgroup.procs file
(the test root here), which the chown calls do not touch.

Remove the redundant chown calls and the now unused test_euid and
cg_test_a_procs variables.

Signed-off-by: Shaojie Sun <sunshaojie@kylinos.cn>
Reviewed-by: Tao Cui <cuitao@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-14 08:32:58 -10:00
Rui Qi
ddabc5dbd2 selftests/cgroup: Preserve CPU hotplug write errors
The cpuset partition root state selftest checks several CPU hotplug
transitions. If writing to a CPU online file fails, the helper still
runs pause afterwards and returns the status of pause instead of the
failed write.

This hides the real hotplug failure and can make later checks run
against expectations for a transition that never happened. Move the
write before the bookkeeping and return when it fails, so callers can
observe the hotplug error and the test does not record a CPU as offline
unless the offline operation actually succeeded.

Also change the O* command handler in set_ctrl_state() to use
"eval $COMM $REDIRECT" like all other handlers. The previous version
set COMM but still called write_cpu_online directly, bypassing the
redirect that captures stderr for error reporting.

Changes since v1:
 - Use eval $COMM $REDIRECT in the O* handler instead of calling
   write_cpu_online directly (Waiman Long)

Fixes: a8c52eba88 ("kselftest/cgroup: Add cpuset v2 partition root state test")
Signed-off-by: Rui Qi <qirui.001@bytedance.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-13 07:07:34 -10:00
Shaojie Sun
44f57a2b1b cgroup/cpuset: Add test for partition root invalidation returning wrong CPUs
Add a test case to REMOTE_TEST_MATRIX covering the bug fixed by
commit 345f401666 ("cgroup/cpuset: Return only actually allocated
CPUs during partition invalidation"). The test verifies that when a
sibling partition root changes its cpuset.cpus to overlap with another
partition root, only actually allocated CPUs (effective_xcpus) are
returned to the parent, not all CPUs in cpus_allowed.

Signed-off-by: Shaojie Sun <sunshaojie@kylinos.cn>
Reviewed-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-11 15:59:35 -10:00
Guopeng Zhang
6dd5d93f6c cgroup/cpuset: Remove obsolete PFA_SPREAD_SLAB task flag
Commit 16a1d96835 ("mm/slab: remove mm/slab.c and slab_def.h")
removed the SLAB allocator, the only allocator that implemented cpuset
slab spreading. Commit 61a182ab61 ("cgroup/cpuset: Remove
cpuset_do_slab_mem_spread()") then removed the last task_spread_slab()
caller. Commit 3ab67a9ce8 ("cgroup/cpuset: Mark memory_spread_slab as
obsolete") marked the legacy control obsolete.

cpuset still updates PFA_SPREAD_SLAB when tasks attach to a legacy
cpuset and walks all tasks in a cpuset when memory_spread_slab changes.
Remove the unused task flag and its helpers, and make spread task
updates depend only on memory_spread_page.

Keep the memory_spread_slab control and CS_SPREAD_SLAB state so legacy
users retain the existing write, readback and inheritance behavior.
Update the comments and documentation to describe only page-cache
spreading as functional.

Assisted-by: LLM
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Reviewed-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-11 15:43:08 -10:00
Zhe Liu
3345c7248a docs: cgroup-v2: fix stale "io" controller introduction
The introductory paragraph for the IO controller still states that
weight based distribution is "available only if cfq-iosched is in use"
and that "neither scheme is available for blk-mq devices".  This text
dates from when the cgroup v2 documentation was first written (2015)
and was correct at the time, but is no longer accurate:

  * cfq-iosched was removed in v5.0;
  * blk-mq is now the only block I/O path, and both the absolute limit
    scheme (io.max via blk-throttle) and the weight based scheme
    (io.weight via iocost, or io.bfq.weight under BFQ) work on it;
  * latency based protection (iolatency) and I/O priority (ioprio)
    controllers have since been added.

The rest of the section already documents io.weight, io.max,
io.cost.{qos,model}, io.latency and io.prio.class correctly, so the
introduction is the only part that contradicts them.  Rewrite it to
reflect the current state.

Signed-off-by: Zhe Liu <liuzhe1@kylinos.cn>
Reviewed-by: Tao Cui <cuitao@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-10 09:34:22 -10:00
Rui Qi
75e79a2019 selftests/cgroup: Avoid awk -e in cpuset tests
The cpuset selftests use awk -e to parse cgroup mount points. This
works with gawk, but mawk rejects the option. In test_cpuset_prs.sh,
this leaves CGROUP2 empty and causes the test to skip as if cgroup v2
were not mounted. The same non-portable invocation exists in the cpuset
v1 hotplug test.

The scripts only need to pass a single awk program. Use the standard awk
invocation without -e so mount point detection works with awk
implementations that do not support the gawk extension.

Fixes: a8c52eba88 ("kselftest/cgroup: Add cpuset v2 partition root state test")
Fixes: 812c5945bd ("cgroup/cpuset: Add test_cpuset_v1_hp.sh")
Signed-off-by: Rui Qi <qirui.001@bytedance.com>
Acked-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-10 09:28:03 -10:00
Guopeng Zhang
26d3a59e02 cgroup/cpuset: Use WRITE_ONCE() for shared prs_err updates
cpuset_partition_show() reads cs->prs_err without cpuset_mutex using
READ_ONCE(). The field is documented as not lock protected, but several
updates to live cpusets still use plain stores.

Convert the remaining prs_err stores on live cpusets to WRITE_ONCE().

Fixes: 0c7f293efc ("cgroup/cpuset: Add cpuset.cpus.exclusive.effective for v2")
Assisted-by: LLM
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Reviewed-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-10 09:27:38 -10:00
Shaojie Sun
77bc7e952f selftests/cgroup: add user_usec sanity check in test_cpucg_nice
In test_cpucg_nice, after the child process exits, user_usec is
read from cpu.stat but the value is not checked. Add a sanity check
to ensure user_usec > 0, analogous to test_cpucg_stats(), so that
the test fails early if CPU usage wasn't properly accounted.

Signed-off-by: Shaojie Sun <sunshaojie@kylinos.cn>
Reviewed-by: Michal Koutný <mkoutny@suse.com>
Acked-by: Tao Cui <cuitao@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-02 09:50:55 -10:00
Julia Lawall
ae649c9636 cgroup: drop unneeded semicolon
The trailing semicolon belongs at the point of use, not in the macro
definition. All uses have been verified to have their own semicolons.

This was found using the following Coccinelle semantic patch:

@r@
identifier i : script:ocaml() { String.lowercase_ascii i = i };
expression e;
@@

*#define i(...) e;

Signed-off-by: Julia Lawall <Julia.Lawall@inria.fr>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-01 23:42:05 -10:00
Tao Cui
3864c95587 docs: cgroup-v2: mark memory.pressure and io.pressure as read-write
The cgroup-v2 documentation describes memory.pressure and io.pressure as
"read-only nested-keyed file", but both files accept trigger writes
(cgroup_memory_pressure_write / cgroup_io_pressure_write) and are therefore
read-write. cpu.pressure and irq.pressure are already documented as
read-write, so this also resolves an internal inconsistency.

Signed-off-by: Tao Cui <cuitao@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-23 23:16:06 -10:00
Waiman Long
d56281fefb selftests/cgroup: Fix minor defects in test_cpuset
With commit 98149f5425 ("selftests/cgroup: Add test for cpuset affinity
on controller disable"), sashiko [1] reported 3 different issues with
the new test_cpuset_affinity_on_controller_disable() test.

 1) `cpu_set_equal` iterates over mask bytes instead of bits, ignoring
    CPUs >= 8.
 2) Thread synchronization logic allows the main thread to read
    uninitialized stack memory, causing test flakiness.
 3) Test fails instead of skipping gracefully on uniprocessor systems
    or when CPU 1 is unavailable.

Fix the reported issues by:
 1) Iterates over the bit size of the mask.
 2) Test the new ready flag for each thread to end the wait
    on the conditional variable and eliminate the now unneeded
    AFFINITY_THREAD_A_READY and AFFINITY_THREADS_READY test phases.
 3) Return KSFT_SKIP on "cpuset.cpus" setting failure.

[1] https://sashiko.dev/#/patchset/20260712235510.373125-1-longman%40redhat.com

Fixes: 98149f5425 ("selftests/cgroup: Add test for cpuset affinity on controller disable")
Signed-off-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-17 12:12:44 -10:00
Tao Cui
5da800b016 Docs/admin-guide/cgroup-v2: fix delay_nsec unit in io.latency doc
The io.latency doc says the io.stat delay field counts microseconds.  The
field is delay_nsec and is reported in nanoseconds.  Refer to it by its
real name and correct the unit.

Signed-off-by: Tao Cui <cuitao@kylinos.cn>
Acked-by: Michal Koutný <mkoutny@suse.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-17 07:05:59 -10:00
Jian Guo
9133e0e2f7 selftests/cgroup: Remove redundant cg_enter_current() call in test_core
The test_cgcore_no_internal_process_constraint_on_threads test has two
back-to-back cg_enter_current(root) calls in its cleanup path.

A single cg_enter_current() call atomically migrates the entire thread
group to the target cgroup even for multi-threaded processes, and this
test creates no extra threads or child processes that would require a
second migration attempt. The second call is a harmless no-op once the
process is already in the root cgroup, but it is redundant and
inconsistent with the cleanup logic used in all other cgroup core
selftest cases.

Remove the duplicate call to clean up the code. No functional change is
intended.

Signed-off-by: Jian Guo <guojian@kylinos.cn>
Acked-by: Michal Koutný <mkoutny@suse.com>
Tested-by: Tao Cui <cuitao@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-17 07:05:40 -10:00
Michal Koutný
98149f5425 selftests/cgroup: Add test for cpuset affinity on controller disable
Add a new selftest that exposes a bug in cpuset_attach() where thread
CPU affinity is not properly updated when the cpuset controller is
disabled in a threaded cgroup hierarchy.

The test creates a threaded cgroup hierarchy with two child cgroups
(A and B) having different cpuset.cpus constraints:
- Parent: cpuset.cpus=0-1
- Child A: cpuset.cpus=0-1
- Child B: cpuset.cpus=1 (restricted to CPU 1 only)

A multithreaded process is created with threads placed in different
cgroups. When the cpuset controller is disabled on the parent, thread
affinities should be updated to match the parent's cpuset.

Expected behavior:
- thread_a affinity: {0-1} before and after (unchanged)
- thread_b affinity: {1} before, {0-1} after (expanded)

Current buggy behavior:
- thread_b affinity remains {1} after controller disable

Assisted-by: Claude:claude-sonnet-4-5
Signed-off-by: Michal Koutný <mkoutny@suse.com>
Acked-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-16 09:38:22 -10:00
Waiman Long
9637786d38 cgroup/cpuset: Handle the special case of non-moving tasks in cpuset_can_attach()
With cgroup v2 migration of a multithreaded process having threads
in different cgroups of a threaded subtree, it is possible that
cpuset_can_attach() can be called with tasks that are not migrating with
respect to cpuset if cpuset controller is not enabled in some of the
subtree cgroups. IOW, the old cpuset can be the same as the new one. This
can cause problem when we need to track the set of old cpusets and the
new cpusets in singly linked lists as a cpuset cannot be in both lists.

As reported by Tejun, the following is an example threaded subtree with
partial cpuset delegation that can cause this issue to show up.

  P (+cpuset)
  |- R (cpuset)        <- destination
  |  `- C (no cpuset)  -> effective cpuset == R
  `- W (cpuset)

Group leader in R, thread_a in C, thread_b in W; migrate the whole
process into R (echo $PID > R/cgroup.procs). thread_a moves C->R:
its cgroup changes so compare_css_sets() keeps it in the taskset, but
its cpuset css is unchanged (C inherits R's), so task_cs() == cs ==
R. cpuset is in ss_mask because thread_b (W->R) changed. can_attach()
then tags R as a source (thread_a) and the destination (thread_b):

Handle this special case by skipping tasks that are not migrating in
cpuset_can_attach() and avoid calling cpuset_can_attach_check() in this
case. By doing so, the destination cpuset will not be put into source
cpuset linked list.

As the source cpuset cannot be easily determined in cpuset_attach(),
unnecessary work can be performed if a task is not actually
migrating. However, no harm will be done except wasting some
CPU cycles. If it happens that none of the tasks is migrating,
attach_ctx.old_cs will be NULL and task iteration won't be needed.

Reported-by: Tejun Heo <tj@kernel.org>
Closes: https://lore.kernel.org/lkml/e254af713b5345aec3d086771ecf1e71@kernel.org
Signed-off-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-16 09:37:50 -10:00
Waiman Long
7309352a04 cgroup/cpuset: Support multiple destination cpusets for cpuset_*attach()
The only case where the cgroup_taskset structure requires task migration
to multiple cpusets is when enabling a cpuset controller in cgroup v2
where the newly created child cpusets inherits the same effective CPUs
and memory nodes from the parent. In that case, task migration can happen
directly with no update to tasks' CPU and memory nodes assignment and no
further work needed from the cpuset side except updating nr_deadline_tasks
when DL tasks are involved and setting old_mems_allowed in the child
cpusets.

Do that by tracking all the destination cpusets with a new dst_cs_head
singly linked list. The reset_migrate_dl_data() function is integrated
into clear_attach_data() so that it can be used for both source and
destination cpusets.

A warning will be printed if there are multiple destination cpusets but
it is not on default hierarchy or when the CPUs or memory nodes change.

Signed-off-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-16 09:37:42 -10:00
Song Hu
97e7efdda8 selftests/cgroup: fix missing TAP output in test_hugetlb_memcg
main() in test_hugetlb_memcg never calls ksft_print_header(),
ksft_set_plan(), or ksft_finished(), so its output has no TAP plan and is
not valid TAP, unlike the sibling test_memcontrol and test_kmem tests.
Add the header/plan/finished calls following the same pattern.

Signed-off-by: Song Hu <husong@kylinos.cn>
Reviewed-by: Tao Cui <cuitao@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-14 11:42:10 -10:00
Waiman Long
ac1607366c cgroup/cpuset: Support multiple source cpusets for cpuset_*attach()
There are 2 possible scenarios where the cgroup_taskset structure
passed into the cgroup can_attach() and attach() methods can contain
task migration data with multiple source cpusets.

 - A multithread application with threads in different cpusets is
   fully migrated into a new cpuset.
 - Disabling v2 cpuset controller will move all the tasks in child
   cpusets to the parent cpuset.

The current cpuset_can_attach() and cpuset_attach() functions still
expect task migration is from one source cpuset to one destination
cpuset.

Fix that by tracking the set of source (old) cpusets in singly linked
lists. The list will be iterated when necessary to properly update
internal data.

To ensure proper DL tasks accounting, the nr_migrate_dl_tasks in both
the source and destination cpusets are decremented/incremented with
their values added to nr_deadline_tasks when the migration is successful.

The setting of the global attach_ctx.cpus_updated and
attach_ctx.mems_updated flags are also moved from cpuset_attach()
to cpuset_can_attach() as the correct source cpuset can no longer be
determined in cpuset_attach() and cpuset states will not be changed
between cpuset_attach() and cpuset_can_attach() with an earlier patch.

Signed-off-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Waiman Long
65e510cd30 cgroup/cpuset: Move mpol_rebind_mm/cpuset_migrate_mm() calls inside cpuset_attach_task()
The cpuset_attach_task() was introduced in commit 42a11bf5c5
("cgroup/cpuset: Make cpuset_fork() handle CLONE_INTO_CGROUP properly")
to enable the CLONE_INTO_CGROUP flag of clone(2) to behave more like
moving a task from one cpuset into another one. That commits didn't
move the mpol_rebind_mm() and cpuset_migrate_mm() calls for group leader
into cpuset_attach_task().

When the CLONE_INTO_CGROUP flag is used without CLONE_THREAD, the new
task is its own group leader. So it is still not equivalent to moving
task between cpusets in this case. Make CLONE_INTO_CGROUP behaves
more close to cpuset_attach() by moving the mpol_rebind_mm() and
cpuset_migrate_mm() calls inside cpuset_attach_task().

Also move the stack local cpus_updated, mems_updated and queue_task_work
flags into attach_ctx so that these flags can be accessed inside and
outside of cpuset_attach_task(). The cpuset_fork() function is updated
to set up these flags and do memory migration if necessary.

Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Waiman Long
1bc48a502a cgroup/cpuset: Make attach_ctx.old_cs track task group leader
There are two possible ways that migration of tasks from multiple source
cpusets to a target cpuset can happen. Either a multithread application
with threads in different cpusets is wholely migrated to a new cpuset
or disabling of v2 cpuset controller will move all the tasks in child
cpusets to the parent cpuset.

In the former case, it is the mm setting of the group leader that
really matters. So attach_ctx.old_cs should track the oldcs of the
thread leader. In the latter case, effective_mems of child cpusets
must always be a subset of the parent. So no real page migration
will not be necessary no matter which child cpuset is selected as
attach_ctx.old_cs.

IOW, attach_ctx.old_cs should be updated to match the latest task
group leader in cpuset_can_attach(), but fall back to that of the first
task if there is no group leader in the taskset.

Suggested-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Waiman Long
74eda6ea70 cgroup/cpuset: Expand the scope of cpuset_can_attach_check()
Expand the scope of cpuset_can_attach_check() by including the setting
of setsched flag inside cpuset_can_attach_check() with the new @oldcs
and @psetsched argument. As cpuset_can_attach_check() is also called
from cpuset_can_fork(), set the new arguments to NULL from that caller.

Signed-off-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Michal Koutný <mkoutny@suse.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Waiman Long
e165f243fe cgroup/cpuset: Add a cpuset_reserve_dl_bw() helper
Extract the DL bandwidth allocation code in cpuset_attach() to a new
cpuset_reserve_dl_bw() helper to simplify code.

No functional change is expected.

Signed-off-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Reviewed-by: Gregory Price <gourry@gourry.net>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Waiman Long
892b8bb3fb cgroup/cpuset: Put all task attach related variables into attach_ctx
Put the task attach related cpuset_attach_old_cs and
cpuset_attach_nodemask_to static variables into the new attach_ctx
structure to improve readability and ease maintanence.

No functional change is expected.

Suggested-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Waiman Long
75f7a25ec6 cgroup/cpuset: Prevent race between task attach and cpuset state change
Commit e44193d39e ("cpuset: let hotplug propagation work wait for
task attaching") was introduced to let hotplug operation to wait
until the completion of task attach operation. However, it is still
possible that the states of the source or destination cpuset can
be changed between the cpuset_can_attach() call and the subsequent
cpuset_attach()/cpuset_cancel_attach() call.

As a result, data gathered during cpuset_can_attach() cannot be reliably
used in the subsequent cpuset_attach()/cpuset_cancel_attach()
call at all. Make the task attach operation more robust
and allow the sharing of data between cpuset_can_attach() and
cpuset_attach()/cpuset_cancel_attach() by making cpuset_write_resmask()
and cpuset_partition_write() wait for the completion of task attach
as well.

Ideally, an ongoing task attach operation should block any cpuset write
operation that can change its internal state until the operation is
completed. However, the attach_in_progress flag is currently per cpuset
and only the destination cpuset will have this flag set. The flag is not
set in the source cpuset where the tasks will be moved from. Even if we
extend the scope to include the source cpuset, it will not block cpuset
operation that changes the state of one of its ancestor cpuset which may
indirectly impact the state of the source or destination cpuset. It may
be too costly to set the flag for the whole subtree, it is far easier
to just make the flag global and block all the cpuset write operation
whenever a task attach operation is in progress.

Make that change by creating a new cpuset attach context (attach_ctx)
structure to hold the global in_progress flag and use it for blocking
cpuset write operation if a cpuset attach operation is in progress. Also
add a new wait_attach_done_lock() helper to do the waiting for an
ongoing attach operation and acquire the cpuset_mutex.

The comments about validate_change() are no longer valid as it won't
be called at all if an attach operation is in progress. So the comments
can be removed.

The per-cpuset attach_in_progress flag is also currently used in
partition_is_populated() and cpuset_is_populated() to determine if
an empty cpuset will have incoming task. This check will no longer be
needed as this function will not be called when there is a task attach
in progress. So the flag check is now removed.

Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Waiman Long
4d73368514 cgroup/cpuset: Fix node inconsistencies between cpuset_update_tasks_nodemask() and cpuset_attach()
Whenever memory node mask is changed, there are 4 places where the node
mask has to be updated or used.
 1) task's node mask via cpuset_change_task_nodemask()
 2) memory policy binding via mpol_rebind_mm()
 3) if memory migration is enabled, migrate from old_mems_allowed to
    the new node mask via cpuset_migrate_mm().
 4) setting old_mems_allowed

These memory actions are done in cpuset_update_tasks_nodemask() and
cpuset_attach(). However there are inconsistencies in what node masks
are being used in these 2 functions.

In cpuset_update_tasks_nodemask(),
 - cpuset_change_task_nodemask(): guarantee_online_mems()
 - mpol_rebind_mm(): mems_allowed
 - cpuset_migrate_mm(): guarantee_online_mems()
 - old_mems_allowed: guarantee_online_mems()

In cpuset_attach(),
 - cpuset_change_task_nodemask(): guarantee_online_mems()
 - mpol_rebind_mm(): effective_mems
 - cpuset_migrate_mm(): effective_mems
 - old_mems_allowed: effective_mems

These inconsistencies dates back to quite a long time ago and it is
hard to say what should be the correct values.

The guarantee_online_mems() function returns a node mask from current or
an ancestor cpuset that is a subset of node_states[N_MEMORY]. Nodes in
node_states[N_MEMORY] are all online, i.e. in node_states[N_ONLINE].
However, node in node_states[N_ONLINE] may not have memory. So
node_states[N_MEMORY] should be a subset of node_states[N_ONLINE].

The guarantee_online_mems() function should mostly be useful for v1
where mems_allowed is the same as effective_mems. With v2, the memory
nodes in effective_mems should be a subset of node_states[N_MEMORY]
except when a memory hot-unplug operation is in progress and a memory
node is removed from node_states[N_MEMORY] but not yet reflected in
the effective_mems's as cpuset_handle_hotplug() has not been called
from cpuset_track_online_nodes().

Let use the following setup for both of them and make them consistent.
 - cpuset_change_task_nodemask(): guarantee_online_mems()
 - mpol_rebind_mm(): effective_mems
 - cpuset_migrate_mm(): guarantee_online_mems()
 - old_mems_allowed: guarantee_online_mems()

So for v2, it is effectively all effective_mems most of the time. For
v1, mpol_rebind_mm() uses mems_allowed which may differ from what
guarantee_online_mems() returns, but it conforms to what the cpuset v1
documentation says with respect to setting memory policy.

Signed-off-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Reviewed-by: Gregory Price <gourry@gourry.net>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Waiman Long
95220e1f18 cgroup/cpuset: Make nr_deadline_tasks an atomic_t
The nr_deadline_tasks variable in the cpuset structure was introduced by
commit 6c24849f55 ("sched/cpuset: Keep track of SCHED_DEADLINE task
in cpusets"). It is reported by sashiko [1] that nr_deadline_tasks
can currently be modified by inc_dl_tasks_cs() under rq->lock and
by cpuset_attach() under cpuset_mutex. So if both updates happen
simultaneously, the nr_deadline_tasks variable can be corrupted leading
to incorrect operations down the road.

Fix that by changing its type to atomic_t so that nr_deadline_tasks
are always atomically updated. This fix patch is a low hanging fruit.
It can handle some of the races between a concurrent sched_setscheduler()
and cpuset_can_attach()/cpuset_attach() calls, but not all of them like
the other issue raised by sashiko [2]. This will be handled hopefully
in a future follow up patch.

[1] https://sashiko.dev/#/patchset/20260626181923.133658-1-longman%40redhat.com
[2] https://sashiko.dev/#/patchset/20260630033344.352702-1-longman%40redhat.com

Fixes: 6c24849f55 ("sched/cpuset: Keep track of SCHED_DEADLINE task in cpusets")
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Waiman Long <longman@redhat.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 14:51:01 -10:00
Manuel Ebner
e9d189aa4b docs: cgroup: Fix bracket
Remove single ')'.

Signed-off-by: Manuel Ebner <manuelebner@mailbox.org>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-06 09:40:33 -10:00
Sun Shaojie
171569f8ee cgroup/cpu: document cpu.stat.local and clarify cpu.stat behavior
Add documentation for the cpu.stat.local interface file, which reports
the throttled_usec stat -- the actual throttling time incurred by the
cgroup's own runqueues, which may include throttling inherited from
ancestor cgroup bandwidth limits. Unlike cpu.stat's throttled_usec
which only accounts for throttling caused by the cgroup's own CFS
bandwidth limit.

When the controller is not enabled, the stat is not reported.

Also clarify cpu.stat descriptions: note that the three base CPU usage
stats (usage_usec, user_usec, system_usec) include descendant cgroups,
and that the five CFS bandwidth stats are non-hierarchical -- they only
account for throttling caused by the cgroup's own bandwidth limit.

Signed-off-by: Sun Shaojie <sunshaojie@kylinos.cn>
Acked-by: Michal Koutný <mkoutny@suse.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-06-29 08:53:09 -10:00
Joe Simmons-Talbott
a2703c2980 selftests/cgroup: Adjust cpu test duration based on HZ
For lower HZ values a quota of 1000us is much lower than the amount
of microseconds per tick which makes the tests test_cpucg_max and
test_cpugc_max_nested fail. Increase the test duration to accommodate
for lower HZ values.

Link: https://lore.kernel.org/lkml/20260625203307.1114538-1-joest@redhat.com/
Signed-off-by: Joe Simmons-Talbott <joest@redhat.com>
Acked-by: Michal Koutný <mkoutny@suse.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-06-29 08:52:56 -10:00
Guopeng Zhang
da43ea2139 cgroup: Use data_race() for task->flags in task_css_set_check()
task_css_set_check() uses rcu_dereference_check() to verify that
task->cgroups can be dereferenced. One accepted condition is that the
task is already exiting, tested by checking PF_EXITING in task->flags.

This check is only part of the CONFIG_PROVE_RCU lockdep predicate. This
was found by KCSAN during fuzz testing. KCSAN can report a data race
when another task flag bit is updated concurrently. One report shows
pids_release() reading task->flags through task_css_set_check() while
do_task_dead() sets PF_NOFREEZE:

  KCSAN: data-race in task_css() [inline]
  KCSAN: data-race in pids_release()

  task_css()
  pids_release()
  cgroup_release()
  release_task()
  wait_task_zombie()

  value changed: 0x0040004c -> 0x0040804c

The changed bit is PF_NOFREEZE, not PF_EXITING. PF_EXITING remains set
before and after the update, so the task_css_set_check() condition does
not change. This is not a race on task->cgroups and does not indicate
incorrect pids charging or uncharging.

tools/memory-model/Documentation/access-marking.txt recommends
data_race() for data-racy loads used only for diagnostic purposes. Use
data_race() here to mark the intended diagnostic-only access.

No functional change intended.

Suggested-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Guopeng Zhang <zhangguopeng@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-06-25 11:12:19 -10:00
Zenghui Yu (Huawei)
8a564dfdfd cgroup: Fix a typo of the function name in comment
... which was wrongly written as cgroup_threadcgroup_change_begin().

Signed-off-by: Zenghui Yu (Huawei) <zenghui.yu@linux.dev>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-06-24 11:09:22 -10:00
Waiman Long
eda17a3a70 cgroup/cpuset: Rebind/migrate mm only for threadgroup leader in cpuset_update_tasks_nodemask()
As reported by sashiko [1], cpuset_update_tasks_nodemask() will do
mpol_rebind_mm() and possibly cpuset_migrate_mm() for all threads of
a multithreaded process. Since commit 3df9ca0a2b ("cpuset: migrate
memory only for threadgroup leaders"), cpuset_attach() had been updated
to rebind and migrate memory only for threadgroup leaders to mark the
group leader as the owner of the mm_struct.

To be consistent and avoid unnecessary performance overhead for heavily
multithreaded processes, follow the cpuset_attach() example and perform
memory rebind and migration only for threadgroup leaders.

Also add a paragraph in cgroup-v2.rst under cpuset.mems that the
threadgroup leader is the memory owner of that threadgroup. Therefore
the non-leading threads shouldn't be in other cgroups whose "cpuset.mems"
doesn't fully overlap that of the group leader.

[1] https://sashiko.dev/#/patchset/20260621032816.1806773-1-longman%40redhat.com

Signed-off-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-06-24 09:42:04 -10:00
Waiman Long
866f587e9c cgroup/cpuset: Avoid unnecessary cpus & mems update in cpuset_hotplug_update_tasks()
As reported by sashiko [1], cpuset_hotplug_update_tasks() may perform
unnecessary task iteration and updating of tasks' CPU and node masks
when mems_allowed and/or cpus_allowed are not set in cpuset v2. It is
due to the fact that the temporary new_cpus and new_mems masks do not
inherit parent's effective_cpus/mems when they are empty which is the
expected behavior for cpuset v2 since commit 4ec22e9c5a ("cpuset:
Enable cpuset controller in default hierarchy").

Fix that and avoid unnecessary work by enhancing
compute_effective_cpumask() to add the empty cpumask check
and inheriting the parent's versions if empty when in v2. A new
compute_effective_nodemask() helper is also added to perform a similar
function for new effective_mems.

Add new test_cpuset_prs.sh test cases to confirm that effective_cpus
will inherit the parent's version if cpuset.cpus is empty.

[1] https://sashiko.dev/#/patchset/20260621032816.1806773-1-longman%40redhat.com

Suggested-by: Ridong Chen <ridong.chen@linux.dev>
Fixes: 4ec22e9c5a ("cpuset: Enable cpuset controller in default hierarchy")
Signed-off-by: Waiman Long <longman@redhat.com>
Reviewed-by: Ridong Chen <ridong.chen@linux.dev>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-06-24 09:41:37 -10:00
Yousef Alhouseen
a2ba161d17 tools/cgroup: iocost_monitor: parse help before importing drgn
iocost_monitor.py imports drgn before argparse can handle "-h" or report
argument errors. That makes basic usage help fail on systems where drgn is
not installed.

Parse arguments before importing drgn so the help and argument-error paths
work without the runtime debugging dependency. Normal execution still
imports drgn before reading kernel state.

Signed-off-by: Yousef Alhouseen <alhouseenyousef@gmail.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-06-24 08:59:01 -10:00
Linus Torvalds
f0e6f20cb5 Changes for 7.2-rc1
Added:
     depth limit to indx_find_buffer() to prevent stack overflow
     validate split-point offset in indx_insert_into_buffer()
     bounds check to run_get_highest_vcn()
     fileattr_get() and fileattr_set() support
     zero stale pagecache beyond valid data length
     handle delayed allocation overlap in run lookup
     validate lcns_follow in log_replay() conversion
     cap RESTART_TABLE free-chain walker at rt->used
     resize log->one_page_buf when adopting on-disk page size
     reject direct userspace writes to reserved $LX* xattrs
 
 Fixed:
     out-of-bounds read in decompress_lznt()
     avoid -Wmaybe-uninitialized warnings
     hold ni_lock across readdir metadata walk
     preserve non-DOS attribute bits in system.dos_attrib
     validate index entry key bounds
     syncing wrong inode on DIRSYNC cross-directory rename
     validate Dirty Page Table capacity in log_replay() copy_lcns
     wrong LCN in run_remove_range() when splitting a run
     allocate iomap inline_data using alloc_page
     mount failure on 64K page-size kernels
     out-of-bounds read in ntfs_dir_emit() and hdr_find_e()
     bound attr_off in UpdateResidentValue against data_off
     bound DeleteIndexEntryAllocation memmove length
     bound copy_lcns dp->page_lcns[] index in analysis pass
     bound NTFS_DE view.data_off in UpdateRecordData{Root,Allocation}
     prevent potential lcn remains uninitialized
 
 Changed:
     bound to_move in indx_insert_into_root() before hdr_insert_head()
     call _ntfs_bad_inode() when failing to rename
     fold resident writeback into writepages loop
     force waiting for direct I/O completion
     fold file size handling into ntfs_set_size()
     reject SEEK_DATA and SEEK_HOLE past EOF early
     format code, add descriptive comments and remove non-useful
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEh0DEKNP0I9IjwfWEqbAzH4MkB7YFAmoqcPYACgkQqbAzH4Mk
 B7ZQ3A//ZYsz0s0qIZ0ErRuxQqmliZc1hzVGbFdKi046AKeRfhN1nV/1MP75+F4V
 eD3sJ4kiROT4oc1x//uJdCoMrrH7qZs2Rcrzv3azC4F9BEFxcxLtJkyZ5NVU4eCj
 vfKaRWZ8ewKeMm37Laz8DOpsz8193KzAVYK/Fm1KYoKMR0Jt+/sdOkIO/NVczZEk
 gY4EAqKUTORfN0a/iELaA+NIrViTk2Wjzzu74YNl/1RDii1LGFaTfa3cmB6jylTY
 AVjPX/lMMtdhy5k9Thcp4lG4uK6x6fSPYHEqvB8+Q3/JGbJfS92Oz3FR9zNKlKz4
 y8depBmT85aZJ2psKetFJCXcVj3EIC1aVY/1CCgJCnLymANUuHlwFDv+3QY4D2bR
 Me//cob7zLFNuul22Uveb00+34H+Tqf1QQNFNUtam6aeXC3bK9PEMbYIQ2+rU6jr
 8M9MTDqfdc6WjyjXwVTNXqki5ZEZmEbSk0FXJM9JsKAyLeo2aentrQtCj5CuTBaH
 sfIbJjT6g/UCfQudNxUDPtlKbhCqYA8SU23iOUlVe7TQDNKnHaQ+NEJ0prqAt303
 2uVupQSJJTu+qv4s2s7ZdaA8z44WrrZfFFQm0okUIA/4NkrAKMYLTQsmki9xpW0h
 KADepNmqtQ4tva9OjTkdOYiy8FUCGTKdZTJOgihjqGO+Z881xOY=
 =OVdT
 -----END PGP SIGNATURE-----

Merge tag 'ntfs3_for_7.2' of https://github.com/Paragon-Software-Group/linux-ntfs3

Pull ntfs3 updates from Konstantin Komarov:
 "Added:
   - depth limit to indx_find_buffer() to prevent stack overflow
   - validate split-point offset in indx_insert_into_buffer()
   - bounds check to run_get_highest_vcn()
   - fileattr_get() and fileattr_set() support
   - zero stale pagecache beyond valid data length
   - handle delayed allocation overlap in run lookup
   - validate lcns_follow in log_replay() conversion
   - cap RESTART_TABLE free-chain walker at rt->used
   - resize log->one_page_buf when adopting on-disk page size
   - reject direct userspace writes to reserved $LX* xattrs

  Fixed:
   - out-of-bounds read in decompress_lznt()
   - avoid -Wmaybe-uninitialized warnings
   - hold ni_lock across readdir metadata walk
   - preserve non-DOS attribute bits in system.dos_attrib
   - validate index entry key bounds
   - syncing wrong inode on DIRSYNC cross-directory rename
   - validate Dirty Page Table capacity in log_replay() copy_lcns
   - wrong LCN in run_remove_range() when splitting a run
   - allocate iomap inline_data using alloc_page
   - mount failure on 64K page-size kernels
   - out-of-bounds read in ntfs_dir_emit() and hdr_find_e()
   - bound attr_off in UpdateResidentValue against data_off
   - bound DeleteIndexEntryAllocation memmove length
   - bound copy_lcns dp->page_lcns[] index in analysis pass
   - bound NTFS_DE view.data_off in UpdateRecordData{Root,Allocation}
   - prevent potential lcn remains uninitialized

  Changed:
   - bound to_move in indx_insert_into_root() before hdr_insert_head()
   - call _ntfs_bad_inode() when failing to rename
   - fold resident writeback into writepages loop
   - force waiting for direct I/O completion
   - fold file size handling into ntfs_set_size()
   - reject SEEK_DATA and SEEK_HOLE past EOF early
   - format code, add descriptive comments and remove non-useful"

* tag 'ntfs3_for_7.2' of https://github.com/Paragon-Software-Group/linux-ntfs3: (34 commits)
  ntfs3: reject direct userspace writes to reserved $LX* xattrs
  fs/ntfs3: resize log->one_page_buf when adopting on-disk page size
  fs/ntfs3: prevent potential lcn remains uninitialized
  ntfs3: cap RESTART_TABLE free-chain walker at rt->used
  fs/ntfs3: bound NTFS_DE view.data_off in UpdateRecordData{Root,Allocation}
  fs/ntfs3: validate lcns_follow in log_replay conversion
  fs/ntfs3: bound copy_lcns dp->page_lcns[] index in analysis pass
  fs/ntfs3: bound DeleteIndexEntryAllocation memmove length
  fs/ntfs3: bound attr_off in UpdateResidentValue against data_off
  ntfs3: fix out-of-bounds read in ntfs_dir_emit() and hdr_find_e()
  fs/ntfs3: fix mount failure on 64K page-size kernels
  ntfs3: avoid another -Wmaybe-uninitialized warning
  ntfs3: Allocate iomap inline_data using alloc_page
  fs/ntfs3: format code, deal with comments
  fs/ntfs3: reject SEEK_DATA and SEEK_HOLE past EOF early
  fs/ntfs3: fold file size handling into ntfs_set_size()
  fs/ntfs3: force waiting for direct I/O completion
  fs/ntfs3: fold resident writeback into writepages loop
  fs/ntfs3: handle delayed allocation overlap in run lookup
  fs/ntfs3: zero stale pagecache beyond valid data length
  ...
2026-06-24 10:05:53 -07:00
Linus Torvalds
840ef6c78e NFS Client Updates for Linux 7.2
New Features:
  * XPRTRDMA: Decouple req recycling from RPC completion
  * NFS: Expose FMODE_NOWAIT for read-only files
 
 Bugfixes:
  * SUNRPC: Fix sunrpc sysfs error handling
  * SUNRPC: Fix uninitialized xprt_create_args structure
  * XPRTRDMA: Harden connect and reply handling
  * NFS: Fix EOF updates after fallocate/zero-range
  * NFS: Keep PG_UPTODATE clear after read errors in page groups
  * NFS: Use nfsi->rwsem to protect traversal of the file lock list
  * NFS: Prevent resource leak in nfs_alloc_server()
  * NFSv4: Clear exception state on successful mkdir retry
  * NFSv4: Don't skip revalidate when holding a dir delegation and attrs are stale
  * pNFS: Fix use-after-free in pnfs_update_layout()
  * pNFS: Defer return_range callbacks until after inode unlock
  * pNFS: Fix LAYOUTCOMMIT retry loop on OLD_STATEID
  * pNFS: Reject zero-length r_addr in nfs4_decode_mp_ds_addr
  * NFS/flexfiles: Reject zero-length filehandle version arrays
  * NFS/flexfiles: Fix checking if a layout is striped
  * NFS/flexfiles: Fixes for honoring FF_FLAGS_NO_IO_THRU_MDS
 
 Other Cleanups and Improvements:
  * Remove the fileid field from struct nfs_inode
  * Move long-delayed xprtrdma work onto the system_dfl_long_wq
  * Convert xprtrdma send buffer free list to an llist
  * Show "<redacted>" for cert_serial and privkey_serial mount options
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEnZ5MQTpR7cLU7KEp18tUv7ClQOsFAmo64NUACgkQ18tUv7Cl
 QOvVMRAAnto2SAwqPkUf2V6dET141qKhWLRKLUqbYxkzc1PKqJBfJuJBwNWHNtyb
 M9JXpx00WSCjfksP5SyD5YugOzom1/SbMJlZB2FCBW6+LTyP/jwsBmqzWXdiKc/d
 x2pD7dkKVdjQUg8siNRLkJR4cyquySUlV39JNKHtPzhHTyWCVYqpvBcsFZwvPPPp
 TKC2ubpbu3zFlZUIYUEKMpPq44dOOlLzMzjWxMO8yTy/s/+5LsNLFRiSadr2sINp
 EWdPn2rpaQT1KmkHdklwUy8xtS+Zw0LaH0g0bVGJfd2ptiMz2VdFIFzxJkQh8jMT
 x0FkUBWDbTdVyiI0OZDo3uh/pJiKzTQI2SecE9to4rNHlNVDeOT9n8UanSYs71rz
 emXQIgszv2juiUvbSRcgzQ+SFKcxq332eDRWmpPIQox+/NvMFK+aMLS7aTd319Up
 bfVMRp5uAp5r2oVz3ETg7RDqatMJ2S0/J2HB3zVf5ONzaBaA//TUrCiSAt49Ep7a
 SsK7VXJCnxw2S23fa3RqlylZ3Gw29QiRjK7INoe8iNjLTqxAvtwcCTum7Ys+IGEl
 VVyxzBzgeGLlT4mU9BpMRZ9BZUjqgmflL8t4FwiFZQD1nZmJLwulZ8zSjIJ7OK2g
 8G8SWP3K7igEbWGCOwqqZWTtkzQC7OYR27vQuz6aPcgIS/fuMxg=
 =hj8e
 -----END PGP SIGNATURE-----

Merge tag 'nfs-for-7.2-1' of git://git.linux-nfs.org/projects/anna/linux-nfs

Pull NFS client updates from Anna Schumaker:
 "New features:
   - XPRTRDMA: Decouple req recycling from RPC completion
   - NFS: Expose FMODE_NOWAIT for read-only files

  Bugfixes:
   - SUNRPC:
      - Fix sunrpc sysfs error handling
      - Fix uninitialized xprt_create_args structure
   - XPRTRDMA:
      - Harden connect and reply handling
   - NFS:
      - Fix EOF updates after fallocate/zero-range
      - Keep PG_UPTODATE clear after read errors in page groups
      - Use nfsi->rwsem to protect traversal of the file lock list
      - Prevent resource leak in nfs_alloc_server()
   - NFSv4:
      - Clear exception state on successful mkdir retry
      - Don't skip revalidate when holding a dir delegation and attrs are stale
   - pNFS:
      - Fix use-after-free in pnfs_update_layout()
      - Defer return_range callbacks until after inode unlock
      - Fix LAYOUTCOMMIT retry loop on OLD_STATEID
      - Reject zero-length r_addr in nfs4_decode_mp_ds_addr
   - NFS/flexfiles:
      - Reject zero-length filehandle version arrays
      - Fix checking if a layout is striped
      - Fixes for honoring FF_FLAGS_NO_IO_THRU_MDS

  Other cleanups and improvements:
   - Remove the fileid field from struct nfs_inode
   - Move long-delayed xprtrdma work onto the system_dfl_long_wq
   - Convert xprtrdma send buffer free list to an llist
   - Show "<redacted>" for cert_serial and privkey_serial mount options"

* tag 'nfs-for-7.2-1' of git://git.linux-nfs.org/projects/anna/linux-nfs: (42 commits)
  NFS: Use common error handling code in nfs_alloc_server()
  NFS: Prevent resource leak in nfs_alloc_server()
  NFSv4/pNFS: reject zero-length r_addr in nfs4_decode_mp_ds_addr
  nfs: don't skip revalidate on directory delegation when attrs flagged stale
  xprtrdma: Return sendctx slot after Send preparation failure
  xprtrdma: Repost Receive buffers for malformed replies
  xprtrdma: Sanitize the reply credit grant after parsing
  xprtrdma: Fix bcall rep leak and unbounded peek
  xprtrdma: Resize reply buffers before reposting receives
  xprtrdma: Check frwr_wp_create() during connect
  xprtrdma: Initialize re_id before removal registration
  xprtrdma: Fix ep kref imbalance on ADDR_CHANGE
  xprtrdma: Convert send buffer free list to llist
  NFS: correct CONFIG_NFS_V4 macro name in #endif comment
  nfs: use nfsi->rwsem to protect traversal of the file lock list
  NFSv4.1/pNFS: fix LAYOUTCOMMIT retry loop on OLD_STATEID
  nfs: expose FMODE_NOWAIT for read-only files
  nfs: add nowait version of nfs_start_io_direct
  NFSv4/flexfiles: honor FF_FLAGS_NO_IO_THRU_MDS in pg_get_mirror_count_write
  NFSv4/flexfiles: honor FF_FLAGS_NO_IO_THRU_MDS on fatal DS connect errors
  ...
2026-06-23 18:36:41 -07:00
Linus Torvalds
09ca8dc7d6 f2fs-for-7.2-rc1
In this round, the changes primarily focus on filesystem error reporting,
 reducing memory footprint by reverting in-memory data structures used for
 runtime validation, honoring FDP hints, and adding trace and debug logs.
 In addition, there are critical bug fixes resolving out-of-bounds read
 vulnerabilities in inline directory and ACL handling, potential deadlocks
 in balance_fs, use-after-free issues in atomic writes, and false data/node
 type assignments in large sections.
 
 Enhancement:
  - Revert  in-memory sit version and block bitmaps
  - support to report fserror
  - add trace_f2fs_fault_report
  - add iostat latency tracking for direct IO
  - add logs in f2fs_disable_checkpoint()
  - honor per-I/O write streams for direct writes
  - map data writes to FDP streams
  - skip inode folio lookup for cached overwrite
  - skip direct I/O iostat context when disabled
  - revert "check in-memory block bitmap"
  - revert "check in-memory sit version bitmap"
 
 Bug fix:
  - optimize representative type determination in GC
  - fix incorrect FI_NO_EXTENT handling in __destroy_extent_node()
  - fix potential deadlock in f2fs_balance_fs()
  - fix potential deadlock in gc_merge path of f2fs_balance_fs()
  - atomic: fix UAF issue on f2fs_inode_info.atomic_inode
  - fix missing read bio submission on large folio error
  - pass correct iostat type for single node writes
  - fix to do sanity check on f2fs_get_node_folio_ra()
  - validate orphan inode entry count
  - keep atomic write retry from zeroing original data
  - read COW data with the original inode during atomic write
  - validate inline dentry name lengths before conversion
  - validate dentry name length before lookup compares it
  - reject setattr size changes on large folio files
  - revert "remove non-uptodate folio from the page cache in move_data_block"
  - validate ACL entry sizes in f2fs_acl_from_disk()
  - bound i_inline_xattr_size for non-inline-xattr inodes
  - fix listxattr handling of corrupted xattr entries
  - fix to round down start offset of fallocate for pin file
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE00UqedjCtOrGVvQiQBSofoJIUNIFAmo5sQwACgkQQBSofoJI
 UNIB7g/+IMAr2UVWdpZ88Uho58GvkkIuZceoZoPuSfUi1hTy0o4+wxeo2ecO06v6
 pDlxivkWXRDpdW1iXbUqcmk2HKEsM3ysWb6jvsFXz+eC4QeWKQJ2uZTyLpEVpW2j
 phXv3TwENPEkku2Mncv907hqUZG/SBxJ2H7jcxb0jHRoaLUwZHuGF0VU/MEodyuy
 ZJifGI3BMwm7Gu2GXwuliDBjbUHRaBs+8kYQ21NZGv0FuQsCNQ+bLhWz3q4FVfg6
 nt5FStKgfoKPHIhamltP6uc4E4KlNDtFgKxluEfzrVqxqvHvUpBxj718DtZLbpNN
 zD6PUHCI0MU0L7qW+RVJx8TOaceYB5xHVcNi8d+CDQPCJgG0LV0ilykqzQ4LRSob
 JcPIjEVkrIgNSzYh/PcDHkUBZmt3MiZZf6xaxviqxDoPqyY6TFATF27ZIZbc7jSa
 hF4XO6mNtbDLhSIMrFUXBnGfnKvIK42OyM5aFLEMxBm7akYYr64h+r6mR+apjDb1
 4FQ1YIuKfIHm7DuphUiazmyOV7P4kcOGYbqyiOk/HNxf6Cc3/kOXKTRZq00ORNbX
 X0FXUOy94xrrjdqoWTldv2o4I49zf8RAEBNLAxDneV2qXsITQtsHWUXp5BbKdA0Q
 7nyUBUTOdDvUq3qyXXA9BIKWdTr0XeGGp/z8rrVR0aT6ksqiTzM=
 =YvkR
 -----END PGP SIGNATURE-----

Merge tag 'f2fs-for-7.2-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs

Pull f2fs updates from Jaegeuk Kim:
 "The changes primarily focus on filesystem error reporting, reducing
  memory footprint by reverting in-memory data structures used for
  runtime validation, honoring FDP hints, and adding trace and debug
  logs. In addition, there are critical bug fixes resolving
  out-of-bounds read vulnerabilities in inline directory and ACL
  handling, potential deadlocks in balance_fs, use-after-free issues in
  atomic writes, and false data/node type assignments in large sections.

  Enhancements:
   - Revert  in-memory sit version and block bitmaps
   - support to report fserror
   - add trace_f2fs_fault_report
   - add iostat latency tracking for direct IO
   - add logs in f2fs_disable_checkpoint()
   - honor per-I/O write streams for direct writes
   - map data writes to FDP streams
   - skip inode folio lookup for cached overwrite
   - skip direct I/O iostat context when disabled
   - revert "check in-memory block bitmap"
   - revert "check in-memory sit version bitmap"

  Fixes:
   - optimize representative type determination in GC
   - fix incorrect FI_NO_EXTENT handling in __destroy_extent_node()
   - fix potential deadlock in f2fs_balance_fs()
   - fix potential deadlock in gc_merge path of f2fs_balance_fs()
   - atomic: fix UAF issue on f2fs_inode_info.atomic_inode
   - fix missing read bio submission on large folio error
   - pass correct iostat type for single node writes
   - fix to do sanity check on f2fs_get_node_folio_ra()
   - validate orphan inode entry count
   - keep atomic write retry from zeroing original data
   - read COW data with the original inode during atomic write
   - validate inline dentry name lengths before conversion
   - validate dentry name length before lookup compares it
   - reject setattr size changes on large folio files
   - revert "remove non-uptodate folio from the page cache in move_data_block"
   - validate ACL entry sizes in f2fs_acl_from_disk()
   - bound i_inline_xattr_size for non-inline-xattr inodes
   - fix listxattr handling of corrupted xattr entries
   - fix to round down start offset of fallocate for pin file"

* tag 'f2fs-for-7.2-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs: (42 commits)
  f2fs: fix to round down start offset of fallocate for pin file
  f2fs: fix listxattr handling of corrupted xattr entries
  f2fs: skip direct I/O iostat context when disabled
  f2fs: remove unneeded f2fs_is_compressed_page()
  f2fs: avoid unnecessary fscrypt_finalize_bounce_page()
  f2fs: avoid unnecessary sanity check on ckpt_valid_blocks
  f2fs: misc cleanup in f2fs_record_stop_reason()
  f2fs: fix wrong description in printed log
  f2fs: bound i_inline_xattr_size for non-inline-xattr inodes
  f2fs: validate ACL entry sizes in f2fs_acl_from_disk()
  Revert "f2fs: remove non-uptodate folio from the page cache in move_data_block"
  f2fs: Split f2fs_write_end_io()
  f2fs: Rename f2fs_post_read_wq into f2fs_wq
  f2fs: Prepare for supporting delayed bio completion
  f2fs: reject setattr size changes on large folio files
  f2fs: validate dentry name length before lookup compares it
  f2fs: validate inline dentry name lengths before conversion
  f2fs: read COW data with the original inode during atomic write
  f2fs: skip inode folio lookup for cached overwrite
  f2fs: keep atomic write retry from zeroing original data
  ...
2026-06-23 17:59:36 -07:00
Linus Torvalds
bade58eb06 - Prevent NULL dereference on theoretical missing IO bitmap
(Li RongQing)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmo6wFMRHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1hOmBAAsFd4cotcp2OQt0Cn6ZNMt1WwoJc5Qplw
 RCMEuTzWmf5zIjCOYXNHNd4lDKKTvMQr3BKOX7oG8noZgnYXxdMlhvN4j9y0Pc0h
 UXIXamzGRzI+2I4kSLL2iee3/utj1Srs19g5ONtaHkJEiH37mw00BIJVgE51TT5t
 6KDMbdHZ2ZmtDjtD7BEdecVrhpKLVejFHIpljPJ8GGVPC6QmrF91d+Lv7r2B0IjR
 3WlDHHXbRvDNbd6r0wvlIbxOUEwUtYMCDIErvohfN2ZU/HVNKivJ42L4dgG8yx5q
 q8ZLbmI15Qa2vEcMesr8liXqT992INMvn+TjCJ0huY3qoyRvf285Xla3DKaA50rL
 1btj9c0i4O6GfKN568nwx+K5YGWb2EH1s79vs3VmG6L6pmnX8CGawMRBDIXY2+qQ
 kFfWsiFCid3TDTwoA1bkOpakNM77d4BAAvYg7yGxlN2wYWiddN8slx77iunDhw+s
 xCi7rQ5Kb35SfWAg3DiKHd5wqptvpK/EkgwCVfzkLf6VFxqFPdaiP4bDn23daUP7
 XhUo3symJ8KLh1bLlR1DLgyzRp5t5yW3/Rocs7RS1h1bTvd96vF04PEQncH5c6Lw
 bB8AXKnB3am1AmBfWbQ3kg/AQJV1tlgCkCozjr990uSVgwRAy1XigioPn/y/eNPW
 pBqWCO4B7aA=
 =mOJb
 -----END PGP SIGNATURE-----

Merge tag 'x86-urgent-2026-06-23' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull x86 fix from Ingo Molnar:

 - Prevent NULL dereference on theoretical missing IO bitmap (Li
   RongQing)

* tag 'x86-urgent-2026-06-23' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  x86/ioperm: Prevent NULL dereference on theoretical missing IO bitmap
2026-06-23 17:16:31 -07:00
Linus Torvalds
541643982b Miscellaneous timer fixes:
- Fix timekeeping locking order bug in the timekeeping init code
    (Mikhail Gavrilov)
 
  - Fix u64 multiplication bug in the posix-cpu-timers code
    on 32-bit kernels (Zhan Xusheng)
 
  - Fix macro name in comment block (Ethan Nelson-Moore)
 
  - Fix off-by-one bug in the compat settimeofday() usecs
    validation code (Wang Yan)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmo6v6cRHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1iSVQ//RCLIhw+fUmcGrjN0E/MX8cFsg0ED3+PS
 3Dfi6mNY+vU5y2my+Mg8zMRw0m2xgnTs1rFeF8x9y2X6bbitoYePi/WomT4+uX/v
 Jj/EEfdbTBqvpZbfKiLN2GGwQJyVox5oRwEwKY1ik/xUFNxCqRxmELrNCUwU0Lwg
 b9UNjebX5WjtDufc0cw2RiBUNuAcAyE7gSqaf1IMyJTygS6l0iMQMkI6Xd8KHB+g
 LXRpB4V5RvFncCV6b81/9pKJzSM6EYdysKod/wi0kq0muxBlk+iiE7zdaiTKjzp5
 y/ItEStuNIQn/NAUYUX4ui1I9wRqYMrnNzvUFYLYWySpAU7V9GuyflB08RvXZpWb
 Gp6LtsZy192liyDvSUrYpQBnfkjKAPXxy3FrcnnI95U86UloXSMBvI+aQCHdFyt6
 TJpgkZM0fn7kjb9i/CB5Vyvrwu7iN+gm8lFtpu5DRHNzIVPvk/C4pyaM5Za/z68k
 dZ0Wv7pZzhLBjIxERzuMhr4YI6PG/DyFNz17JKiNr5S5sKk8q/pNhH3Ki99aaz5y
 IykkevHVgldp7/Zz91ixvLP3BYyFRx6++Bl1DOZyN/JAzAukuu2vBgaGnvzqfs5z
 7K5nqaouPNvvYbyCmb8bU9FtZUuHForN9LDQ3QGRYhQsVQH+u2FbEeFjfB8NQdOI
 llNA4FZgVUc=
 =dSO3
 -----END PGP SIGNATURE-----

Merge tag 'timers-urgent-2026-06-23' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull misc timer fixes from Ingo Molnar:

 - Fix timekeeping locking order bug in the timekeeping init code
   (Mikhail Gavrilov)

 - Fix u64 multiplication bug in the posix-cpu-timers code on 32-bit
   kernels (Zhan Xusheng)

 - Fix macro name in comment block (Ethan Nelson-Moore)

 - Fix off-by-one bug in the compat settimeofday() usecs validation code
   (Wang Yan)

* tag 'timers-urgent-2026-06-23' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  time: Fix off-by-one in compat settimeofday() usec validation
  hrtimer: Correct CONFIG_NO_HZ_COMMON macro name in comment
  posix-cpu-timers: Use u64 multiplication in update_rlimit_cpu()
  timekeeping: Register default clocksource before taking tk_core.lock
2026-06-23 16:57:39 -07:00