Commit Graph

1460666 Commits

Author SHA1 Message Date
Tejun Heo
5fd501744b sched_ext: Add bandwidth-limited rescue execution for stranded tasks
A local DSQ insert lacking the needed caps is diverted to the reject DSQ and
bounced back through ops.enqueue() so the scheduler can re-decide. That
recovery assumes the scheduler has somewhere legal to send the task. When it
doesn't, e.g. when the task's affinity is restricted to cids delegated away,
the task starves until the stall watchdog ejects the scheduler. An exiting
task is worse - it skips ops.enqueue() and the rejection becomes a
self-requeuing cycle that burns the CPU until the watchdog fires.

Add SCX_ENQ_RESCUE, a fallback modifier on local DSQ inserts. When the
insert would be rejected for missing caps, the kernel takes over and runs
the task on the target CPU without consulting the owning scheduler. The
kernel sets the flag itself when enqueueing an exiting task.

Rescue is a last-resort forward-progress backstop with a persistent
disadvantage, not a way around cap enforcement. A per-CPU token bucket
accrues rescue_bandwidth_ppt (default 2%) of CPU time and rescues run one at
a time in arrival order. Each is granted a slice of the rescue_quantum_us
(default 5ms) quantum divided across the waiters, waits at the tail of the
local DSQ claiming no priority, and rejoins its scheduler as a fresh arrival
once the slice is served.

The schedulers keep their normal control over an admitted rescuee and may
preempt or reslice it. Service is measured on CPU time actually received, so
neither shortens the rescue. Prolonged denial escalates - the remaining
slice turns into protected execution (SCX_TASK_PROTECTED) and the rescuee
preempts the current task. Escalation is paced by the same bucket, and
delivered service converges on the configured bandwidth no matter how
aggressively the schedulers dispatch.

Both knobs are root-only and SCX_RESCUE_DISABLE turns rescue off, making
SCX_ENQ_RESCUE inserts reject as usual.

v2: - Add SCX_OPS_OPEN() fix-ups for the new ops fields so cpu-form
      schedulers setting them still load on older kernels. (Andrea)

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-08-03 11:01:29 -10:00
Tejun Heo
9cfc6ab34a sched_ext: Add SCX_TASK_PROTECTED
A BPF scheduler can displace any of its tasks at will - cut a running one's
slice with an SCX_ENQ_PREEMPT dispatch, an SCX_KICK_PREEMPT kick or a direct
shortening, and jump a queued one with HEAD insertions. Sometimes the kernel
needs a slice and a DSQ position to stick regardless.

Add SCX_TASK_PROTECTED, guarding both:

- The slice becomes immutable. Every scheduler-reachable write is refused
  and counted as SCX_EV_SLICE_DENIED. Higher scheduling classes are
  unaffected. PREEMPT|IMMED can't preempt a running protected task and gets
  reenqueued.

- A protected task that reached the head of its DSQ keeps it - HEAD
  insertions land behind the leading run of protected tasks and reenqueue
  sweeps skip them. Only rq-owned DSQs can hold protected tasks, so the walk
  runs only for them.

The bit lives in p->scx.flags so that both the refusal and the head walk
read it under the rq lock that protects it.

Protection ends when the slice is consumed, when the task leaves the rq
except for a save/restore on the running task, on a yield, when the
scheduler enters bypass, and when the task leaves scx. The flag is
kernel-internal and not used yet.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-08-03 11:01:20 -10:00
Tejun Heo
13f1eae3b6 sched_ext: Synchronize slice and dsq_vtime writes
p->scx.slice and p->scx.dsq_vtime writes have no synchronization rules. The
dsq insert kfuncs write both fields synchronously from whatever context
they're called in - a direct dispatch from ops.select_cpu() writes with only
pi_lock held - and, as the kfuncs are safe to call spuriously with the
invalid dispatch discarded later, a scheduler can modify any task's slice by
spuriously calling them. The latter stands in the way of an upcoming patch
which adds kernel-granted slices that the schedulers must not be able to
modify.

Give both fields explicit rules. While the task is running, sleeping or
queued on an rq-owned DSQ, the rq lock protects them - these are the states
where the kernel consumes the slice. While queued on a user DSQ or on the
BPF side, the kernel neither consumes nor decides on the fields and every
writer acts for the BPF scheduler - synchronizing the writers is the
scheduler's responsibility and whichever write lands last wins.

To conform, an insert kfunc no longer writes the fields when called. The
values travel with the dispatch and take effect when the task is inserted. A
discarded dispatch has no side effects. The rq lock rule is asserted at the
slice store.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-08-03 11:01:14 -10:00
Tejun Heo
78f8d726e6 sched_ext: Make SCX_ENQ_IGNORE_CAPS waive the preemption cap too
SCX_ENQ_IGNORE_CAPS is kernel-internal and marks a placement the kernel
forces. scx_caps_for_enq() waives the enqueue cap for it, but a PREEMPT
insert still picks up the preemption cap requirement from
scx_caps_for_preempt(). Update scx_caps_for_preempt() to take enq_flags and
require nothing when SCX_ENQ_IGNORE_CAPS is set.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-08-03 11:01:07 -10:00
Tejun Heo
1fd50778b1 sched_ext: Reject internal enq_flags in the dsq move kfuncs
The dsq insert kfuncs reject __SCX_ENQ_INTERNAL_MASK bits in
scx_dsq_insert_preamble() instead of scx_vet_enq_flags(). A scheduler can
smuggle internal flags such as SCX_ENQ_CLEAR_OPSS through the dsq move
kfuncs and corrupt the dispatch protocol. Move the rejection into
scx_vet_enq_flags(). The vtime move wrapper OR'd the internal
SCX_ENQ_DSQ_PRIQ bit into enq_flags before the vet; the bit now goes in
inside scx_dsq_move() after the vet.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-08-03 11:00:57 -10:00
Tejun Heo
f82b16b8e8 sched_ext: Factor out __scx_bpf_now()
scx_bpf_now() couples the valid-or-fresh rq clock read to the current rq.
The read is useful for kernel-internal timing against a specific rq,
including a remotely locked one. Factor it out into __scx_bpf_now().

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-08-03 11:00:47 -10:00
Tejun Heo
2d091012a4 sched_ext: Make several ext.c helpers available outside ext.c
set_task_slice(), task_unlink_from_dsq(), move_local_task_to_local_dsq(),
init_dsq() and dump_line() will be used outside ext.c. Add the scx_ prefix
and declare them in internal.h. The scx_sched_all list will also be used
outside ext.c, drop its static. No functional changes.

v2: Declare scx_sched_all outside the CONFIG_EXT_SUB_SCHED block - the
    definition is unconditional. (sashiko AI)

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-08-03 11:00:39 -10:00
Tejun Heo
8b3b8522c9 sched_ext: Rename scx_local_or_reject_dsq() to scx_resolve_local_dsq()
The following rescue execution addition gives the function a third possible
destination, making a name that enumerates the outcomes a poor fit. Rename
to the destination-neutral scx_resolve_local_dsq(). No functional changes.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-08-03 11:00:31 -10:00
Andrea Righi
12da4723b6 selftests/sched_ext: Make allowed_cpus idle validation race-free
A remotely selected CPU can be re-advertised as idle by an idle-to-idle
re-pick before the BPF program validates the selection. Checking that
the selected CPU remains absent from the idle mask is therefore
inherently racy.

Validate a stable local invariant instead: a CPU executing
ops.select_cpu() or ops.enqueue() in a non-idle scheduling context must
not be advertised as idle. Read the idle mask without modifying it and
also validate selected CPUs against the requested domain and task
affinity.

Suggested-by: Kuba Piecuch <jpiecuch@google.com>
Signed-off-by: Andrea Righi <arighi@nvidia.com>
Reviewed-by: Kuba Piecuch <jpiecuch@google.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-03 06:21:11 -10:00
Andrea Righi
ba190ed3f4 sched_ext: Initialize idle masks as busy
The built-in idle masks are reset with all online CPUs marked idle
before sched_ext is enabled. Busy CPUs can therefore be incorrectly
advertised as idle until their next idle transition.

Initialize the masks empty so that the initial state is conservative.
When bypass is lifted, every CPU is rescheduled and idle-to-idle
re-picks populate the masks with CPUs that are actually idle. Later
idle transitions keep the masks up to date.

Suggested-by: Tejun Heo <tj@kernel.org>
Signed-off-by: Andrea Righi <arighi@nvidia.com>
Reviewed-by: Kuba Piecuch <jpiecuch@google.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-03 06:20:51 -10:00
Liang Luo
c5b9316cf3 sched_ext: Set errno on ENABLING -> ENABLED transition failure
If the SCX_ENABLING -> SCX_ENABLED cmpxchg at the tail of
scx_root_enable_workfn() fails, the function jumps to err_disable
without setting ret. At that point ret still holds the return value
of the last successful __scx_init_task() call, which is 0, so the
err_disable fallback reports the meaningless message:

  scx_root_enable() failed (0)

Set ret = -EBUSY, consistent with the other enable-state guards at
the top of the same function, so the fallback always reports a real
errno.

Signed-off-by: Liang Luo <luoliang@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-02 10:16:39 -10:00
Liang Luo
680e0718b9 sched_ext: Fix stale @cgroup_id in sched_ext_ops kernel-doc
The kernel-doc comment for sched_ext_ops::sub_cgroup_id uses the old
@cgroup_id name, which no longer matches the struct member. This
produces two kernel-doc warnings:

  Warning: struct member sub_cgroup_id not described in sched_ext_ops
  Warning: Excess struct member cgroup_id description in sched_ext_ops

Update the @param name to match the actual member.

Signed-off-by: Liang Luo <luoliang@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-08-02 10:16:39 -10:00
Tejun Heo
ee7aece608 sched_ext: Report scx_link_sched() failures inline
scx_link_sched() carries each failure out of the locked section through
err_msg and ret because scx_error() used to take scx_sched_lock and couldn't
be called under it. That restriction is gone, so report each failure at the
site it's detected and return directly. The scx_error() here claims the exit
on the sched being linked, which has no descendants yet, and the locked
propagation walk is deferred, so nothing reacquires scx_sched_lock inline.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-27 11:20:32 -10:00
Tejun Heo
3c4b380649 sched_ext: Abort directly from the hardlockup handler
scx_hardlockup() defers the abort to an irq_work because exit claiming used
to take scx_sched_lock and couldn't run from NMI. The deferral is now
unnecessary - claiming is NMI-safe and asserting ->aborting is exactly what
breaks the live-locks that hard-lock CPUs. Call handle_lockup() directly and
drop the irq_work. This also makes the self-detected case recoverable: the
perf watchdog fires on the hard-locked CPU itself, where a queued irq_work
never runs with IRQs off.

Also fix the return value: %true used to be returned whenever sched_ext was
loaded, suppressing the kernel's hardlockup report even when the abort was
refused. Return %true only when this call initiated the abort.

Fixes: bd2d76455b ("sched_ext: Defer scx_hardlockup() out of NMI")
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-27 11:20:32 -10:00
Tejun Heo
e06ece82d7 sched_ext: Report NMI kicks with scx_error()
The per-cpu kick lists are protected by IRQ masking which doesn't stop NMIs,
so scx_bpf_kick_cpu() from NMI silently drops the kick after a one-time
warning. A dropped kick can leave a CPU idle when the scheduler believes it
was woken, which is a correctness problem for the scheduler even if the
kernel is fine. Now that scx_error() works from NMI, abort the scheduler
instead so that the bug is surfaced deterministically. The warned_nmi_kick
tracking is no longer needed.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-27 11:20:32 -10:00
Tejun Heo
1bf623ebd5 sched_ext: Format bstr exit messages after claiming the exit
The bstr exit kfuncs format the message into a shared static buffer under a
raw spinlock before initiating the exit. The lock can't be taken from NMI
and needlessly serializes all bstr exits system-wide.

Now that exit claiming is lock-free, reverse the order: claim the exit first
and format directly into the exit_info message buffer which the claim winner
owns exclusively. The new scx_exit_bstr() implements the sequence, replacing
scx_bstr_format(), and the shared buffer and lock are deleted; the formatter
itself is what bpf_trace_printk() already runs from NMI. scx_prog_sched()
callers were relying on the lock for RCU protection, which is now provided
explicitly.

A malformed format no longer changes or fails the requested operation:
scx_bpf_exit_bstr() keeps its graceful exit kind and scx_bpf_sub_kill_bstr()
still kills the child, with a fallback message carrying the formatting
errno, while the sched that supplied the bad format is aborted for its bug.

Before this and the previous patch, an "any" category kfunc called from NMI
context could trigger scx_error() and deadlock - e.g. a tracing prog
attached to a function running in NMI calling scx_bpf_dsq_peek() on a
non-existent DSQ would try to grab scx_sched_lock, which may be held by the
interrupted CPU. This and the previous patch fix the deadlock: scx_error()
and the bstr exit kfuncs, and thus scx_bpf_error() and scx_bpf_exit(), are
now safe to call from any context including NMI.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-27 11:20:32 -10:00
Tejun Heo
f883dbb64c sched_ext: Make exit claiming lock-free
scx_claim_exit() claims descendants' exits by walking the subtree under
scx_sched_lock, making exit claiming, and thus scx_error(), unusable from
NMI and from under scx_sched_lock. However, kfuncs raising errors can run
from NMI-attached BPF progs, the hardlockup handler runs in NMI, and
scx_link_sched() wants to report failures under the lock.

The walk does two things with different urgencies: ->aborting must be
asserted synchronously to break IRQs-off dispatch-path live-locks, while the
descendants' exit_kind claims can happen later. Split them: sweep ->aborting
locklessly under RCU to unwedge the system and defer the locked
SCX_EXIT_PARENT walk to a new irq_work, both of which are NMI-safe.

The sweep stores each node's ->aborting and then reads its children list
while scx_link_sched() inserts and then checks the parent's ->aborting, the
two sides paired by full barriers - one side always sees the other. A link
that sees ->aborting undoes its insert and fails. As the undo's
list_del_rcu() leaves ->sibling non-empty, list_empty() can no longer
identify a never-linked sched during teardown - add sch->linked instead.

trace_sched_ext_exit can now fire from NMI and is called after the
->aborting stores so that its callbacks don't hold up live-lock recovery.
The exit backtrace is skipped for NMI exits as stack_trace_save()'s
NMI-safety is arch-dependent and undocumented.

v2: Move trace_sched_ext_exit() after the ->aborting stores (Andrea).

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-27 11:20:23 -10:00
Tejun Heo
3457220620 tools/sched_ext/include: Regenerate enum_defs.autogen.h
Regenerate enum_defs.autogen.h from the current vmlinux.h to pick up the SCX
enum changes accumulated since the last regeneration, including the
SCX_REENQ_LOCAL_MAX_REPEAT to SCX_REENQ_MAX_REPEAT rename.

Reported-by: Andrea Righi <arighi@nvidia.com>
Link: https://lore.kernel.org/all/amZsEbZJdDgjstPF@gpd4/
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-26 11:26:49 -10:00
Tejun Heo
7706d6e4f2 sched_ext: Bound per-task reenqueues and eject the owning scheduler
Unlike local reenqueues, cap rejections have no repeat limit. A
malfunctioning scheduler can keep re-inserting a task to a cid it lacks caps
on, cycling the task through reject and reenqueue. This was assumed safe
because a task that never runs trips the stall watchdog. However, the
reenqueue irq_work re-arms itself and outranks the timer vector, blocking
everything else on the CPU including stall detection and recovery, until the
NMI hardlockup detector fires.

Local reenqueues already have a repeat cap, SCX_REENQ_LOCAL_MAX_REPEAT,
which needs generalizing to cover all reenqueues. It also has an attribution
problem. Counted per-cpu on root, it tears down the whole hierarchy even
when a sub-scheduler caused the repeated reenqueues.

Generalize by bounding every reenqueue with one per-task counter. reenq_cnt
is bumped in scx_do_enqueue_task() on each SCX_ENQ_REENQ, the single path
every reenqueue producer passes through, and cleared in clr_task_runnable()
when the task is picked to run and in scx_disable_task() when it leaves the
scheduler's control. Past SCX_REENQ_MAX_REPEAT the task's owning scheduler
is ejected with a new SCX_EXIT_ERROR_REENQ and the task is left stranded to
be picked up during sched exit.

The SCX_EV_REENQ_LOCAL_REPEAT event becomes SCX_EV_REENQ_REPEAT, counting
repeat reenqueues from all sources.

v2: Count SCX_EV_REENQ_REPEAT only when a reenqueue leads to another
    reenqueue, not on every reenqueue.

v3: - Also clear reenq_cnt in scx_disable_task() so that the count doesn't
      carry over to the next owner across sched class switches, scheduler
      replacement or sub-scheduler rehoming (Andrea Righi).

    - Update the stale SCX_EV_REENQ_LOCAL_REPEAT references in sched-ext.rst
      (Andrea Righi).

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-26 11:12:19 -10:00
Tejun Heo
6ee471b646 sched_ext: Use rcu_access_pointer() for the first_task comparison
dsq->first_task is __rcu for the lockless scx_bpf_dsq_peek(). The task
removal path compares it against the departing task with a plain load, which
sparse flags. The comparison runs under the dsq lock and only tests
identity, so rcu_access_pointer() is the fit.

Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-25 10:33:24 -10:00
Tejun Heo
ce22834301 sched_ext: Resolve most remaining scx_root accesses
scx_root is __rcu and naked accesses were left as transitional markers for
the multi-scheduler transition, to be converted to accesses through the
associated scheduler instances. Most accesses have since been converted to
resolve the sched from the program or task at hand. The remaining naked
sites divide into ones that semantically always want the root sched, which
this patch resolves, and one that is left to a later patch.

The resolved sites:

- The SCX_OPS_TID_TO_TASK validation and the ecaps sync kick already hold a
  sched whose ancestors[] pins the root as entry 0 with plain pointers
  stable for the sched's lifetime. Reach the root through the sched at hand.

- The dispatch entry, class switch, idle notification and fork init paths
  only execute while the scheduler is live and scx_root never changes inside
  the live window, so no update can race them. Add scx_root_protected_live()
  which documents that invariant and resolves with a plain load.

- The hotplug path, including the ecaps reseeds, runs with the hotplug lock
  held, which excludes the scx_root writers. Add scx_root_protected(), which
  accepts either the hotplug lock or scx_enable_mutex.

- Is-root tests use a zero level instead of comparing against the global.

touch_core_sched_dispatch() stays naked, to be resolved by a later patch.

Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-25 10:33:17 -10:00
Tejun Heo
7947442047 sched_ext: Add scx_cgroup_sched() for cgrp->scx_sched reads
cgrp->scx_sched is __rcu and published with rcu_assign_pointer() but every
reader loads it with a plain access, so sparse flags all of them. The reads
are lock-protected: enable/disable paths rewrite the field under all of
scx_enable_mutex, scx_fork_rwsem and cgroup_mutex, and cgroup creation
inherits the parent's sched under cgroup_mutex before the new cgroup is
reachable, so holding any one of the three locks makes the read stable.

Add scx_cgroup_sched() which states the protection with
rcu_dereference_check() and convert the readers. No functional changes.

Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-25 10:33:04 -10:00
Tejun Heo
3a21e34eb2 sched_ext: Gate scx_bpf_cidperf_set() behind a new SCX_CAP_PERF
scx_bpf_cidperf_set() reaches cpufreq with no cap check, so any cid-form
sub-sched can steer the frequency of any cid in its view, including ones it
holds nothing on.

Gate it behind a new SCX_CAP_PERF rather than SCX_CAP_BASE: hardware control
is a separate axis from queue access - a parent may well delegate scheduling
on a cid without handing over its frequency. PERF neither implies nor is
implied by the other caps. The check runs under the target rq's lock, which
ecaps updates are also folded under, so it is authoritative - a write can
never land after a revoke has taken effect. Denials are counted in
SCX_EV_SUB_CIDPERF_DENIED.

The operation is synchronous and the outcome is reported to the caller:
scx_bpf_cidperf_set() now returns 0 or -errno, -EACCES on denial. The
cid-form interface is still under initial development, so the signature is
changed in place without versioning.

scx_qmap grants PERF alongside its existing cid grants so the cpuperf demo
keeps working in sub-scheds.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-24 12:10:22 -10:00
Tejun Heo
49d6247d64 sched_ext: Factor out scx_cpuperf_set()
Factor the cpuperf target write out of scx_bpf_cpuperf_set() into
scx_cpuperf_set() which takes the acting sched and returns 0 or -errno, and
flatten the nested validation into early returns. No functional change.
Prep for gating the write behind a cap and reporting the outcome from the
cid-form kfunc.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-24 12:10:21 -10:00
Tejun Heo
7c80408fa9 sched_ext: Count kicks denied for lacking baseline cid access
kick_one_cpu() silently skips a kick when the kicking sub-sched lacks
SCX_CAP_BASE on the target cid, as does kick_one_cpu_if_idle() for idle
kicks. The skips are sound with the same logic as the reenq gate but are
invisible today, unlike the preempt degradation counted in
SCX_EV_SUB_PREEMPT_DENIED. Count them in a new SCX_EV_SUB_KICK_DENIED event
so every cap denial is observable.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-24 12:10:21 -10:00
Tejun Heo
f879519db8 sched_ext: Gate local DSQ reenq on baseline cid access
scx_bpf_dsq_reenq() with an SCX_DSQ_LOCAL_ON target schedules deferred reenq
work on the cid's cpu, raising an IPI when the target rq isn't the locked
one. Nothing checks caps along the way, so a sub-sched holding no cap at all
on a cid can force its cpu to take IPIs and rq lock cycles at will. The
analogous scx_bpf_kick_cid() path gates delivery on SCX_CAP_BASE in
kick_one_cpu() to prevent exactly this.

Apply the same rule at the reenq scheduling point: if the calling sched
lacks SCX_CAP_BASE on the target cid, drop the reenq and count it in the new
SCX_EV_SUB_REENQ_DENIED event. The check is lockless, which is fine: a reenq
slipping through right after a revoke is harmless, and a wrong denial can't
happen - if the caller has seen its ownership of the cpu, the check sees it
too.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-24 12:10:21 -10:00
Tejun Heo
457daba18e tools/sched_ext: Don't restart over a pending exit request
The tools restart when the kernel exits the scheduler with
SCX_ECODE_ACT_RESTART. The restart decision doesn't consult exit_req, so an
exit request arriving while the restart condition persists is ignored and
the tool reloads in a tight loop. Test exit_req before restarting.

scx_userland needs more: its main loop never watches the kernel-side exit
and exit_req doubles as the stats printer's stop signal, set by the teardown
and reset on each restart. Add the missing UEI_EXITED() test and give the
printer its own stop flag so that exit_req only means an exit request and
stays latched like in the other tools.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-24 12:10:13 -10:00
Liang Luo
c509d81070 sched_ext: Fix incorrect SCX_PICK_IDLE_CPU_* flag prefix in kernel-doc
The flags passed to the pick-idle kfuncs are values from the
scx_pick_idle_cpu_flags enum, whose members are prefixed
SCX_PICK_IDLE_ (SCX_PICK_IDLE_CORE, SCX_PICK_IDLE_IN_NODE).

Three kernel-doc comments in idle.c erroneously used
%SCX_PICK_IDLE_CPU_* which does not correspond to any defined flag
name, while the adjacent scx_bpf_pick_idle_cpu_node() correctly
documents %SCX_PICK_IDLE_*.

Fix the three occurrences to use the correct SCX_PICK_IDLE_* prefix.

Signed-off-by: Liang Luo <luoliang@kylinos.cn>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-23 23:20:52 -10:00
Randy Dunlap
94ca959110 sched_ext: Repair kernel-doc comments
Add missing function parameter descriptions and use the correct
function name in kernel-doc comments to avoid kernel-doc warnings:

Warning: kernel/sched/ext/ext.c:2692 function parameter 'sch' not described in 'finish_dispatch'
Warning: kernel/sched/ext/ext.c:5309 function parameter 'stalled_mask' not described in 'scx_rcu_cpu_stall'
Warning: kernel/sched/ext/ext.c:5405 function parameter 'cpu' not described in 'scx_hardlockup'
Warning: kernel/sched/ext/ext.c:8470 expecting prototype for scx_bpf_dsq_insert(). Prototype was for scx_bpf_dsq_insert___v2() instead
Warning: kernel/sched/ext/ext.c:8784 expecting prototype for scx_bpf_dsq_move_to_local(). Prototype was for scx_bpf_dsq_move_to_local___v2() instead
Warning: kernel/sched/ext/ext.c:9498 expecting prototype for scx_bpf_reenqueue_local(). Prototype was for scx_bpf_reenqueue_local___v2() instead

Signed-off-by: Randy Dunlap <rdunlap@infradead.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-23 08:01:55 -10:00
Cui Jian
00a08ddfb4 sched_ext: Fix stale errno in scx_sub_enable_workfn()
The nesting depth check and the cgroup online check in
scx_sub_enable_workfn() reach err_disable without setting ret, so
the fallback error added by commit db4e9defd2 ("sched_ext: Record
an error on errno-only sub-enable failure") reports
"scx_sub_enable() failed (0)".

This is currently harmless because both paths record their own
scx_error() first and the first error wins, but it leaves the
fallback broken for these paths. Set -EINVAL and -ENODEV there
so the fallback always reports a real errno.

v2: The validate_ops() path from v1 is already fixed in for-7.3
    (sub.c already has ret = scx_validate_ops()), so only the two
    remaining paths are addressed.

Signed-off-by: Cui Jian <cjian720@163.com>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-22 08:31:30 -10:00
Cheng-Yang Chou
2abee44c7d tools/sched_ext: scx_pair: Convert to sched_switch TP
ops.cpu_acquire/release() are deprecated in favor of tracking CPU
preemption from a sched_switch tracepoint, see
commit a3f5d48222 ("sched_ext: Allow scx_bpf_reenqueue_local() to
be called from anywhere"). Loading scx_pair currently emits a
deprecation warning.

Replace the pair_cpu_acquire/release() callbacks with a
tp_btf/sched_switch program that edge-detects the same transitions the
core used to deliver: a release when a running SCX task loses its CPU
to a higher-priority class, and an acquire when the CPU switches back
to an SCX task or idle while marked preempted.

Tasks are classified by effective priority (p->prio) rather than by
policy: rt_mutex_setprio() boosts a PI beneficiary into the rt/dl
classes while leaving its policy untouched, so a policy test would
both miss the release when a boosted task takes the CPU and fire a
spurious acquire when a boosted task replaces a real rt task.

A switch from idle straight to a higher-priority task is deliberately
not treated as a release. The CPU was not running an SCX task, so
there is nothing to drain, and kicking SCX_KICK_PREEMPT |
SCX_KICK_WAIT on every rt wakeup would make the pair CPU wait out rt
bursts it was never coupled to. The old callbacks behaved the same
way, firing ops.cpu_release() only from switch_class() when an SCX
task was put for a higher class.

The tracepoint runs on every context switch in the system, so the
common no-transition case is filtered before taking the pair-shared
lock. This is safe because a CPU's own preempted_mask bit is only ever
written by this tracepoint running on that CPU.

sched_setscheduler() on a running task changes class in place without
a context switch, so such transitions are only observed at the task's
next switch. The old callbacks had the same blind spot in
switch_class(), and try_dispatch() already bounds the resulting wait.

Verified in virtme-ng with the script below. The scheduler must load
without the deprecation warning, stay enabled through the rt churn and
the idle soak (the watchdog would otherwise abort it with "runnable
task stall"), keep its preemption counter advancing, and unregister
cleanly at the end. A PI rt-mutex churn that repeatedly boosts
SCX tasks into the rt class was exercised separately:

  #!/bin/bash
  # vng --verbose --cpus 8 -m 4G --user root -- ./verify.sh
  # FIFO harness: survives even if all SCHED_NORMAL tasks stall
  [ "${RT:-0}" = 1 ] || exec chrt -f 5 env RT=1 "$0"

  chrt -o 0 ./tools/sched_ext/build/bin/scx_pair &
  PAIR=$!
  sleep 3
  for round in $(seq 10); do
    pids=""
    for i in 0 1 2 3; do  # SCHED_FIFO churn
      chrt -f 10 bash -c \
        'e=$((SECONDS+1)); while [ $SECONDS -lt $e ]; do :; done' &
      pids="$pids $!"
    done
    for i in 0 1; do      # SCHED_NORMAL load under scx
      chrt -o 0 bash -c \
        'n=0; while [ $n -lt 200000 ]; do n=$((n+1)); done' &
      pids="$pids $!"
    done
    wait $pids            # explicit pids, not the scx_pair job
  done
  sleep 300               # idle soak
  kill -INT $PAIR         # expect clean unregister in dmesg

Signed-off-by: Cheng-Yang Chou <yphbchou0911@gmail.com>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-22 08:12:00 -10:00
Tejun Heo
3a773220d3 sched_ext: Build the cid tables privately and publish them with RCU
The cid tables are visible to the cid kfuncs while being modified: the
first enable publishes the global pointers before filling them,
ops.init_cids() overrides rewrite them in place, and re-enables rebuild
them in place. A racing TRACING or SYSCALL program can read unfilled
entries, including uninitialized memory in the kmalloc'd tables, or torn
topo updates.

Tie the tables' lifetimes to the root sched instead: each root enable
builds a fresh set privately and publishes the per-table __rcu globals once
the layout is final, and root disable unpublishes and RCU-frees the set. A
non-NULL global is now always a fully built table which stays valid for the
reader's RCU read section, and lookups stay two loads. Kfuncs treat NULL as
no-mapping, also after the scheduler exits instead of reporting the stale
last mapping.

The cid kfuncs are available whether the root scheduler is cid-form or
cpu-form, the latter to allow gradual migration to cids. Every root
therefore builds and publishes a default mapping.

Every reader must either be gated on scheduler liveness or NULL-check
inside an RCU read section. Fix the two kfuncs that were neither:
scx_bpf_this_cid() read the table with no RCU or preemption protection and
scx_bpf_task_cid() relied on KF_RCU, which doesn't put a sleepable program
in an RCU read section. The hotplug callbacks are instead serialized by
retiring the tables inside the cpus_read_lock() section that clears
scx_root.

v2: Document why every root builds the tables (desc + cid.c comment).

Reported-by: Andrea Righi <arighi@nvidia.com>
Closes: https://lore.kernel.org/r/al3tLtPZZkFjMveK@gpd4
Reviewed-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-21 22:27:27 -10:00
Tejun Heo
8cc4909134 sched_ext: Drop unused scx_cpumask_to_cmask()
scx_cpumask_to_cmask() has no callers.

Reviewed-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-21 22:27:17 -10:00
Tejun Heo
5f860e647d sched_ext: Skip the default CPU selection while bypassing
select_task_rq_scx() falls into the default path when the scheduler has no
ops.select_cpu or is bypassing. There it calls scx_select_cpu_dfl() and
direct-dispatches to the picked CPU's local DSQ.

While bypassing, neither does anything: the enqueue path routes the task to
a bypass DSQ before consulting the direct-dispatch target, so the direct
dispatch never happens, and the CPU pick at most shifts which CPU's bypass
DSQ receives the task. Worse, when the scheduler does its own idle tracking,
the built-in idle cpumasks the pick consults are not even updated, so it
doesn't work anyway.

Return prev_cpu without the default selection while bypassing and let the
bypass enqueue place the task.

Reviewed-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-21 22:27:07 -10:00
Tejun Heo
2c91169377 sched_ext: Blame the DSQ's owning scheduler for a runnable stall
check_rq_for_timeouts() blames a runnable stall on the task's owner. Under
a sub-scheduler hierarchy the stalled task can be sitting on a DSQ that a
different scheduler has to drain, e.g. an ancestor's bypass DSQ while the
owner is bypassing. The drainer then escapes blame while the owner is
exited, and when the owner's exit is already claimed, nothing actionable is
reported at all.

Blame the DSQ's owning scheduler instead. The local DSQ is consumed by the
cpu itself and keeps blame on the owner. Detection keeps the owner's timeout
and single-scheduler behavior is unchanged.

Reviewed-by: Andrea Righi <arighi@nvidia.com>
Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-21 22:26:57 -10:00
Tejun Heo
b20dfde5ec sched_ext: Rename the cid-form cgroup ops to cpuctl_*
Two unrelated things go by "cgroup" in the cid form. Sub-schedulers attach
to cgroups, and the cgroup_*() ops deliver cpu controller events. While the
ops names suggest cgroup2 hierarchy, they actually operate on the cpu
controller.

Rename them to cpuctl_* in struct sched_ext_ops_cid, which has no users
outside scx_qmap yet. The cpu form is deployed ABI and keeps the old names.
The layout is unchanged and the kernel keeps calling through the cpu-form
union view.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:11:04 -10:00
Tejun Heo
d327796bd0 tools/sched_ext: Add SCX_OPS_CID_OPEN for cid-form schedulers
SCX_OPS_OPEN() clears compat-gated ops fields which the running kernel
lacks. The clears dereference cpu-form member names and compile for
cid-form skeletons only because both ops structs currently name their
cgroup ops identically, which an upcoming rename will end. No load-time
fix-up can apply to a cid-form scheduler anyway as the cid form postdates
every compat-gated op.

Factor the skeleton open path out of SCX_OPS_OPEN() and add
SCX_OPS_CID_OPEN() which uses only that shared part. Switch scx_qmap, the
only cid-form scheduler, over.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:11:04 -10:00
Tejun Heo
01cad83043 tools/sched_ext: scx_qmap - Add init fault injection modes
Add -J init-fail which makes ops.init_task() fail with -ENOMEM for tasks
whose comm starts with "qmfail", and -J cgrp-init-fail which does the same
in ops.cgroup_init() for cgroups named "qmfail*".

The former exercises the migration veto path: the cgroup.procs write must
fail with the injected errno while the destination sched stays up and the
task stays put. The latter exercises the ownership-return failure path: a
parent failing to re-init a returned cgroup leaves it unowned, and moves and
set_* ops against it must be skipped instead of dereferencing the missing
owner. Matching on "qmfail" names keeps the injecting scheduler's own enable
unaffected.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:10:57 -10:00
Tejun Heo
29ac3ae9ab tools/sched_ext: scx_qmap - Consume cgroup weights through set_weight
With the set_* ops delivered to the parent's sched, a parent qmap instance
now receives cgroup_set_weight for its child subs' attach points. Update
the matching sub_sched_ctx weight and redistribute() in-kernel, and drop
the userspace feed_weights() polling. This exercises the knob routing end
to end.

The self weight is fixed at 100: a cgroup's weight is its parent's knob and
not the scheduler's own business. This drops the self-weight polling and the
repartition PROG_RUN poke with it.

sub_attach seeds the slot with the cgroup's current weight, read through
bpf_cgroup_from_id(), so a weight set before the sub attaches is picked up.
A write racing the attach can still be lost until the next value-changing
cpu.weight write. Acceptable for a demo.

While at it, add a traced ops.cgroup_move() so tests can observe move
delivery.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:10:57 -10:00
Tejun Heo
a6ec0b62c5 sched_ext: Hand over cgroups at sub-scheduler enable/disable
Sub-schedulers don't get cgroups yet: every task_group is inited on the root
sched and the routing added by the previous patches always resolves to it.
Add the handover: an enabling sub-scheduler takes over the cgroups in its
subtree and a disabling one returns them to its parent.

scx_cgroup_claim_subtree() runs while the sub enables, after the subtree's
cgrp->scx_sched's are set and before any task is claimed. It inits each
subtree task_group on the sub, exits it from the parent and updates
tg->scx.sched. A failed ops.cgroup_init() unwinds the sub-side inits and
aborts the enable with the parent untouched.

Disabling reverses it with scx_cgroup_return_subtree(): exit each cgroup
from the sub, then re-init it on the parent with the current tg->scx.*
values, resyncing weight and bandwidth changes made while the sub had it.
When a re-init fails, the parent is failed and the remaining task_groups
still transfer uninited and get no cgroup ops - the same punting done for
tasks. The dying parent's own disable moves them onward.

The handover walks include dying but not yet offlined task_groups, the same
as root's bulk walks: a removed cgroup keeps hosting scheduling events until
its dying tasks finish their final context switches, and its
ops.cgroup_exit() must follow the last of them. tg on/offlining is excluded
through cgroup_lock(), so either ordering against an rmdir of a subtree
cgroup delivers balanced init/exit pairs.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:10:57 -10:00
Tejun Heo
46932bc5fd sched_ext: Deliver cgroup ops to each task_group's sched
With sub-schedulers claiming cgroup subtrees, cgroup ops must be delivered
to each task_group's sched rather than always to root. Add tg->scx.sched to
track which sched initialized the task_group. It is set and cleared together
with SCX_TG_INITED.

Deliver the ops accordingly:

- ops.cgroup_exit() goes to the sched whose ops.cgroup_init() it pairs with.

- ops.cgroup_prep_move/move/cancel_move() go to the task's sched, and only
  for moves that don't re-home the task. A re-homing move is reported
  through the ops.exit_task/init_task() pair instead. The cgroups passed to
  the move ops can be outside the sched's inited set as the cpu controller
  can be coarser than the sub-scheduler topology.

- Knobs of a cgroup belong to the parent, so ops.set_weight/idle/bandwidth()
  go to the parent task_group's sched.

All task_groups currently resolve to the root sched, so no behavior changes
until sub-schedulers start claiming cgroups.

While at it, scx_cgroup_init() is restructured so both paths share the
recording.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:10:56 -10:00
Tejun Heo
bf9dee58ab sched_ext: Re-home tasks on cgroup migration
A task's sched (p->scx.sched) must match its cgroup's owner
(cgrp->scx_sched). cgroup migration breaks the invariant:
scx_cgroup_move_task() only fires root's ops.cgroup_move() and never
re-homes the task, leading to wrong-sched scheduling and, once the stale
sched is freed, a use-after-free.

Hook into the new cgroup task migration events and re-home each task whose
destination cgroup is owned by a different sched. The events map naturally
to the transfer: MIGRATING runs the fallible init for the destination sched,
letting it reject the migration the same way ops.cgroup_prep_move() can,
MIGRATED does the re-home, which can't fail, and CANCELED undoes the init
when the migration falls through.

Pre-commit, the task's task_group still reflects the source, so
__scx_init_task() grows an explicit cgroup argument for the migration path
to hand ops.init_task() the destination cgroup.

Signed-off-by: Tejun Heo <tj@kernel.org>
Closes: https://lore.kernel.org/r/alnxrsexEe_nQwqL@gpd4
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:10:56 -10:00
Tejun Heo
54880aeb87 sched_ext: Relocate scx_cgroup_enabled
scx_cgroup_enabled is in the CONFIG_EXT_GROUP_SCHED block. The upcoming
cgroup migration re-homing needs the gate outside the block. Move the
definition and flag flips outside CONFIG_EXT_GROUP_SCHED. No functional
changes.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:10:56 -10:00
Tejun Heo
0dc90ce1be sched_ext: Factor out scx_rehome_task() and scx_punt_task()
Factor out scx_rehome_task() and scx_punt_task() from the sub-disable
re-home loop and scx_fail_parent(). The upcoming cgroup migration re-homing
also needs scx_rehome_task(). No functional changes.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:10:56 -10:00
Tejun Heo
52478777b3 cgroup: Add cgroup_task_notifier and task migration events
A subsystem can attach to the cgroup hierarchy itself, independent of which
controllers are enabled where - BPF hooks already behave this way and
sched_ext sub-schedulers do too. Controller callbacks can't track task
migrations for them: sched_ext must re-home a task whose migration crosses a
sub-scheduler boundary, but the cpu controller's attach callbacks fire only
when the task_group changes and miss moves whenever the controller topology
is coarser than the sub-scheduler topology.

Add cgroup_task_notifier with per-task migration events mirroring the
can_attach/attach/cancel_attach phases so that a consumer which prepares
per-task state can also veto a migration: CGROUP_TASK_MIGRATING fires
pre-commit, CGROUP_TASK_MIGRATED post-commit and
CGROUP_TASK_MIGRATE_CANCELED unwinds a failed migration. Only migrations
that change a task's dfl cgroup are reported.

Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-19 21:10:56 -10:00
Tejun Heo
7c2cd76770 Merge branch 'for-7.2-fixes' into for-7.3
Pull to receive:

 477869bfaf ("sched_ext: Reject setting disallow from init_task outside the enable path")
 5f8b69642d ("sched_ext: Take cgroup_lock() first in scx_cgroup_lock()")
 8c13364db9 ("sched_ext: Skip sub-disable teardown for never-linked sub-schedulers")
 5cdc928598 ("sched_ext: Don't enable non-ext tasks in the sub-sched task loops")

as dependencies for the upcoming cgroup migration patchset and to
resolve the conflicts with the ext.c/sub.c split on for-7.3.

5f8b69642d comments scx_cgroup_lock() which for-7.3 exported for
sub.c. Resolved by keeping the exported version with the comment.

8c13364db9 and 5cdc928598 patch the pre-split sub-sched enable and
disable paths in ext.c which for-7.3 moved to sub.c. Resolved by
applying the never-linked teardown skip and the class gates to sub.c.

Signed-off-by: Tejun Heo <tj@kernel.org>
2026-07-17 21:50:21 -10:00
Tejun Heo
5cdc928598 sched_ext: Don't enable non-ext tasks in the sub-sched task loops
Root enable and scx_post_fork() enable a task only if it's on the ext class.
Tasks on other classes, possible under an SCX_OPS_SWITCH_PARTIAL root, are
left READY and enabled by switching_to_scx() when they switch over. The sub
enable-commit pass and the sub-disable re-home loop enable unconditionally,
so a fair-class READY task in the subtree becomes ENABLED while not on
sched_ext. A later switch to SCHED_EXT then trips the task state validation
WARN (ENABLED with the previous state not READY) and calls ops.enable() a
second time.

Gate scx_enable_task() on the task's class in both loops.

Fixes: 337ec00b1d ("sched_ext: Implement cgroup sub-sched enabling and disabling")
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-17 21:29:27 -10:00
Tejun Heo
8c13364db9 sched_ext: Skip sub-disable teardown for never-linked sub-schedulers
A sub-scheduler enable can fail before scx_link_sched() links the sched into
the hierarchy, e.g. when the parent is already being disabled, and cleanup
still runs the full scx_sub_disable().

That is racy against root disable: drain_descendants() is the only ordering
between a sub's disable-time task walk and root disable's all-task teardown,
and an unlinked sub is invisible to it. Root's teardown can thus run between
the never-linked sub's drain and its walk, exiting every task to no
scheduler.

The walk then trips the membership WARN and re-homes the exited tasks onto
the dying hierarchy, a use-after-free.

Skip the cgroup ownership reset and the task walk if @sch was never linked,
indicated by the empty ->sibling as unlinking only happens later in the same
function. The membership WARN remains valid: a linked sub is always waited
on by an ancestor's drain.

Fixes: 337ec00b1d ("sched_ext: Implement cgroup sub-sched enabling and disabling")
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-17 21:29:27 -10:00
Tejun Heo
5f8b69642d sched_ext: Take cgroup_lock() first in scx_cgroup_lock()
scx_cgroup_lock() write-locks scx_cgroup_ops_rwsem and then takes
cgroup_lock(), which can deadlock through kernfs:

  scx enable/disable         cgroup rmdir           cpu.weight write
  ------------------         ------------           ----------------
                             cgroup_lock()
  percpu_down_write(rwsem)
  cgroup_lock()
                                                    kernfs_get_active()
                                                    percpu_down_read(rwsem)
                             kernfs_drain()

The enable path waits for the rmdir to release cgroup_mutex. The rmdir,
deactivating the cpu controller's files, waits in kernfs_drain() for the
write's active reference. The write, in scx_group_set_weight(), waits for
the rwsem behind the pending writer.

Take cgroup_lock() first. The set_* paths take no cgroup locks inside the
read side, so a pending write-lock then only waits for read sections that
always run to completion, and no dependency from the rwsem back to
cgroup_mutex remains.

Fixes: a5bd6ba30b ("sched_ext: Use cgroup_lock/unlock() to synchronize against cgroup operations")
Cc: stable@vger.kernel.org # v6.18+
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-17 21:29:12 -10:00
Tejun Heo
477869bfaf sched_ext: Reject setting disallow from init_task outside the enable path
The p->scx.disallow revert assumes the root enable path, where the switching
loop reads the reverted policy right afterwards and leaves the task off SCX.
The sub-scheduler disable path also reaches it when re-initializing the
returned tasks on a root parent. Nothing reads the policy there: the task is
enabled on root anyway and keeps running on the ext class with a silently
rewritten policy.

Kill the sched instead, matching the fork and non-root branches, and update
the disallow documentation, which equated !fork with the load path and
pointed at a stale debugfs path for nr_rejected.

Fixes: 337ec00b1d ("sched_ext: Implement cgroup sub-sched enabling and disabling")
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Andrea Righi <arighi@nvidia.com>
2026-07-17 21:28:57 -10:00