Merge series "mm/slab: introduce kfree_rcu_nolock() and improve
slub_kunit coverage" from Harry Yoo. From the cover letter [1]:
This series improves kmalloc_nolock() and kfree_nolock() coverage in
slub_kunit and introduces kfree_rcu_nolock() for unknown context as
suggested by Alexei Starovoitov.
Unknown context means the caller does not know whether spinning on a
lock is safe (e.g., a BPF program attached to an arbitrary kernel
function or in NMI context).
The slab allocator already supports unknown context via kmalloc_nolock()
and kfree_nolock(), but te slab allocator does not support freeing
objects by RCU in unknown context.
It is not ideal to have completely separate batching for unknown context
because the worst scenario where spinning on a lock would lead to
deadlock is very rare, and in most cases, it is safe to use the existing
mechanism (kfree_rcu_sheaf()).
Since most part of the slab allocator already supports unknown context
and sheaves support batching kvfree_rcu() calls for slab objects,
implement kfree_rcu_nolock() with minimal changes by teaching
kfree_rcu_sheaf() how to support unknown context and making it a little
bit harder to allocate an empty sheaf, instead of making intrusive
changes to the existing kvfree_rcu batching logic.
kfree_rcu_nolock() tries to free the object to the rcu sheaf if trylock
succeeds. Once the rcu sheaf becomes full, it is submitted to RCU via
call_rcu() if spinning is allowed or IRQs are enabled (to avoid calling
call_rcu() in the middle of call_rcu()). Otherwise, call_rcu() is
deferred via irq work.
When there is no sheaf available, kfree_rcu_sheaf() falls back to
defer_kfree_rcu(). It submits the object to kvfree_rcu batching via irq
work. To do this, patch 6 converts kvfree_rcu to use kvfree_rcu_head
without visible changes to the API for now.
Unlike kfree_rcu(), only the 2-argument variant is supported. This is
because the last resort of the 1-arg variant is synchronize_rcu(), which
cannot be used in an unknown context.
As suggested by Alexei Starovoitov, kfree_rcu_nolock() can be used with
struct kvfree_rcu_head (8 bytes), which is smaller than struct rcu_head
(16 bytes).
Link: https://lore.kernel.org/all/20260729-kfree_rcu_nolock-v5-0-a28cdcda9673@kernel.org/ [1]
Merge series "mm/slab, alloc_tag: reduce obj_ext memory waste" from
myself. From the cover letter [1]:
It's been bothering me that the memory usage of struct slabobj_ext
depend only on config options and not whether the fields are actually
used. So with both CONFIG_MEMCG=y and CONFIG_MEM_ALLOC_PROFILING=y there
is always objcg field and codetag_ref field. And thus:
1) Having memory allocation profiling config-enabled but not
boot-enabled means wasted memory on unused codetag_refs. This makes
it less suitable for a general distro config and the page allocator
side doesn't suffer from this, only slab and percpu.
2) Complementary, with memory allocation profiling enabled, there are
caches/slabs that don't need the objcg field, so memory is wasted on
those.
This series should solve the point 1) fully for slab; pcpuobj_ext
handling can be perhaps improved similarly, haven't looked into that.
For 2) it avoids allocating objcg fields for KMALLOC_NORMAL and
KMALLOC_NO_OBJ_EXT caches where we know they are not necessary because
kmalloc() with __GFP_ACCOUNT will pick a KMALLOC_CGROUP type (except
with SLUB_TINY).
The named kmem_caches are tricky. They can be created with SLAB_ACCOUNT
and then we know objcg fields are always needed. But also they can be
created without SLAB_ACCOUNT and then some allocations have
__GFP_ACCOUNT and some not and we don't know that in advance.
This series introduces a SLAB_MAY_ACCOUNT flag that's currently internal
only and is applied to all caches (unless kmem accounting is disabled)
except KMALLOC_NORMAL (unless that aliases KMALLOC_RECLAIM) and
KMALLOC_NO_OBJ_EXT.
As a followup we can make SLAB_MAY_ACCOUNT explicit and add it to to
caches where we know __GFP_ACCOUNT is used. Then we could only honour
__GFP_ACCOUNT for those, while warning for an unexpected usage
elsewhere.
To check for regressions, I forward-ported a microbenchmark hacked into
slub_kunit that was used to evaluate sheaves.
Tried 3 scenarios, MEMCG and KFENCE were always enabled:
- CONFIG_MEM_ALLOC_PROFILING=n
- CONFIG_MEM_ALLOC_PROFILING=y but _ENABLED_BY_DEFAULT=n
- same but booted with sysctl.vm.mem_profiling=1
The results are quite noisy, but no regression was apparent, except
perhaps few percents for the last case. I don't expect it will be
visible in any real workloads.
Link: https://lore.kernel.org/all/20260727-b4-objext_split-v3-0-c29ef0f1f257@kernel.org/ [1]
We have already disabled memory allocation profiling for objects
allocated for KFENCE to avoid complexity. KFENCE allocations are rare
and there can be only CONFIG_KFENCE_NUM_OBJECTS (default to 255)
outstanding ones at any time, so they are among noise in the profiling
stats.
For the same reasons, we can stop memcg_kmem accounting of kfence
objects as their memory usage will be negligible wrt any practical
memcg limits.
This allows us simplifying the code and getting rid of
is_kfence_address() checks in various places, including slab_obj_ext()'s
usage of obj_to_index(). Instead we rely on the fact that slab_obj_exts()
will now always return 0 for a kfence object's fake slab, which makes
those places unreachable.
All we need to do to keep this assumption valid is not to allocate
obj_exts for kfence objects, so the checks need to guard
alloc_slab_obj_exts() where necessary.
Suggested-by: Harry Yoo <harry@kernel.org>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-13-c29ef0f1f257@kernel.org
Reviewed-by: Hao Li <hao.li@linux.dev>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Start using the slab_needs_objcg() helper to calculate slabobj_ext size.
Caches that we know to never need objcg pointers (currently
KMALLOC_NORMAL caches) will thus stop wasting memory on them when memory
allocation profiling is enabled.
For things to work properly, we need to also add slab_needs_objcg()
checks to mem_cgroup_from_obj_slab() and memcg_slab_free_hook(), because
when obj_exts array exists for a slab only due to mem_alloc profiling,
we would otherwise attempt to access a non-existing objcg pointer in
that slab.
In slab_obj_ext_[set_]objcg() add debug warnings if called on a slab
where slab_needs_objcg() is false.
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-12-c29ef0f1f257@kernel.org
Reviewed-by: Harry Yoo <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Slabs of some caches never need the objcg part of struct slabobj_ext.
Introduce helpers to query this for a cache or a slab.
Introduce SLAB_MAY_ACCOUNT flag that is currently only internal and all
caches have it set except:
- KMALLOC_NORMAL caches, as long as KMALLOC_RECLAIM caches are separate
- KMALLOC_NO_OBJ_EXT caches, if they exist
For named caches we currently can't derive SLAB_MAY_ACCOUNT from
SLAB_ACCOUNT because some caches might be created without SLAB_ACCOUNT
and then used both with and without __GFP_ACCOUNT concurrently,
allocating obj_ext arrays on demand. So just add the SLAB_MAY_ACCOUNT
to all kmem caches, unless kmem accounting is disabled.
This can be improved later by finding out all caches used with
__GFP_ACCOUNT, creating them with the SLAB_MAY_ACCOUNT flag explicitly,
and then ignoring __GFP_ACCOUNT for all other caches (possibly with a
warning).
To make the evaluation of slab_needs_objcg() faster in the allocation
and free fast paths, add a obj_exts_needs_objcg flag into slab itself.
This optimization is only available on 64bit architectures where free
bits are available for the flag.
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-11-c29ef0f1f257@kernel.org
Reviewed-by: Harry Yoo <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
No module code calls these functions directly. Seems it was always the
case. Remove the exports.
Acked-by: Paul E. McKenney <paulmck@kernel.org>
Reviewed-by: Harry Yoo <harry@kernel.org>
Link: https://patch.msgid.link/20260730-unexport-barriers-v1-1-852f6641abe9@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
When slub_kunit is not built-in, call kfree_rcu() and kfree_rcu_nolock()
to test kfree_rcu_nolock() in slub_kunit.
Rename the test case as the test covers more _nolock() APIs.
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260729-kfree_rcu_nolock-v5-8-a28cdcda9673@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Currently, k[v]free_rcu() cannot be called in unknown context since
it could lead to a deadlock when called in the middle of k[v]free_rcu().
Make users' lives easier by introducing kfree_rcu_nolock() variant,
now that kfree_rcu_sheaf() is available on PREEMPT_RT and
__kfree_rcu_sheaf() handles unknown context.
When sheaves path fails, kfree_rcu_nolock() falls back to
defer_kfree_rcu() that uses an irq work to free the object via
kvfree_call_rcu(). In most cases, the sheaves path is expected to
succeed and therefore it's unnecessary to introduce additional
complexity to the existing kvfree_rcu batching by teaching it
how to handle unknown context.
Since defer_kfree_rcu() can be called on caches without sheaves, move
deferred_work_barrier() and rcu_barrier() outside the branch in
kvfree_rcu_barrier_on_cache().
Now that deferred kvfree_rcu objects are submitted to kvfree_call_rcu()
after deferred_work_barrier() and may end up in RCU sheaves,
deferred_work_barrier() must be invoked before
flush_rcu_sheaves_on_cache().
Since the RCU sheaf path has not been used on !KVFREE_RCU_BATCHED
kernels, always fall back when kvfree_rcu() is not batched, for
consistency. kvfree_rcu_barrier{,_on_cache()}() on
!KVFREE_RCU_BATCHED are moved to mm/slab_common.c to invoke
deferred_work_barrier() before rcu_barrier().
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260729-kfree_rcu_nolock-v5-7-a28cdcda9673@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
rcu_head is overkill for kvfree_rcu() because the callback
function is always either kfree(), vfree(), or free_large_kmalloc(),
and thus there is no need for a function pointer.
kvfree_rcu batching reuses the field to store the start address
of an object, however, this is not strictly needed because we can
calculate the start address in the slowpath. For the purpose of
kvfree_rcu batching, it is sufficient to implement a linked list using
a single pointer.
Introduce a new struct called kvfree_rcu_head (the name was suggested
by Vlastimil Babka), which is similar to rcu_head but is only a single
pointer to build a linked list, without a function pointer, when
CONFIG_KVFREE_RCU_BATCHED=y.
When kvfree_rcu is not batched, kvfree_rcu_head is the same size
as rcu_head. Note that shrinking struct kvfree_rcu_head on
CONFIG_KVFREE_RCU_BATCHED=n kernels would inevitably require additional
complexity and also some sort of batching (which defeats the purpose of
the config option) because it cannot fall back to call_rcu().
For now there are no user-visible changes to the API. k[v]free_rcu()
simply casts rcu_head to kvfree_rcu_head. While this does not affect
the API, it allows kfree_rcu_nolock() to reuse kvfree_rcu batching
as a fallback when trylock or sheaf allocation fails.
Stop storing the object pointer in rcu_head.func and instead calculate
the object's start address in kvfree_rcu_list(). Factor out the existing
logic to calculate the start address from kvfree_rcu_cb() to
kvmalloc_obj_start_addr(). To avoid losing the KASAN tag, calculate
the offset and subtract it from the address of the kvfree_rcu_head.
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260729-kfree_rcu_nolock-v5-6-a28cdcda9673@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
When memory allocation profiling is compiled in but permanently disabled
on boot with (implicit or explicit) "never" parameter, stop allocating
(thus wasting) memory for the codetag_ref parts of slabobj_ext metadata.
Do this by using the new slab_obj_ext_has_codetag() helper in
cache_obj_ext_size().
Additionally add a slab_obj_ext_has_codetag() check in
handle_failed_objexts_alloc(). The function might get called with memory
allocation profiling disabled, when the obj_ext array is allocated for
objcg pointers only. Setting codetag refs as empty is unnecessary in
that case, and with them not allocated anymore would now result in
memory corruption.
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-10-c29ef0f1f257@kernel.org
Reviewed-by: Hao Li <hao.li@linux.dev>
Reviewed-by: Harry Yoo <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
mem_alloc_profiling_enabled() allows evaluating (with a static key) if
memory profiling is currently enabled. mem_profiling_support is a
variable where false means it's not possible to enable it anymore,
because the system was booted with "never" or it was later shut down.
This is possible to query by mem_alloc_profiling_permanently_disabled().
To make slabobj_ext array size handling dynamic, we need a snapshot of
mem_alloc_profiling_permanently_disabled() early in boot, so that's not
affected by a later shutdown. We also need it to be static key based for
performance. Neither mem_alloc_profiling_enabled() nor
mem_alloc_profiling_permanently_disabled() satisfy this.
Therefore introduce slab_obj_ext_has_codetag() with an underlying static
key for that use case. Its state is made to reflect the result of
mem_alloc_profiling_permanently_disabled() during kmem_cache_init(),
which does happen after setup_early_mem_profiling().
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-9-c29ef0f1f257@kernel.org
Reviewed-by: Hao Li <hao.li@linux.dev>
Reviewed-by: Harry Yoo <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
As suggested by Vlastimil Babka [1], kfree_rcu_sheaf() can be used
on PREEMPT_RT if we always assume spinning is not allowed on PREEMPT_RT.
This is because local_trylock and spinlock_t are safe to use with
trylock and unlock as long as the kernel does not spin and the context
is not NMI and not hardirq.
Now that __kfree_rcu_sheaf() knows how to handle SLAB_FREE_NOLOCK,
relax the limitation and try the sheaves path on PREEMPT_RT as well.
Keep the lockdep map on non RT kernels. However, do not use the lockdep
map on PREEMPT_RT to avoid suppressing valid lockdep warnings.
As pointed by Vlastimil Babka [2], on PREEMPT_RT it is unnecessary to
defer call_rcu() under IRQ-disabled section or raw spinlock. However,
let us avoid adding more complexity as the scenario is not supposed
to be common on PREEMPT_RT, with a hope that call_rcu_nolock() will be
soon supported in RCU.
Link: https://lore.kernel.org/linux-mm/6811cc17-8ee4-48c8-8cbf-6bf4d9f98162@kernel.org [1]
Link: https://lore.kernel.org/linux-mm/40591888-3a87-433e-b3d2-cda1cab543be@kernel.org [2]
Suggested-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Reviewed-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260729-kfree_rcu_nolock-v5-5-a28cdcda9673@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
__kfree_rcu_sheaf() cannot invoke call_rcu() when spinning is not
allowed and IRQs are disabled. To relax the limitation, extend the
deferred free fallback so that a full rcu sheaf can be submitted to
call_rcu() via the existing IRQ work.
Since the deferred mechanism does more than deferred freeing of objects,
rename the struct to deferred_percpu_work and adjust names accordingly.
When a sheaf is queued on an IRQ work, it is detached from
pcs->rcu_free but call_rcu() is not invoked until the irq_work runs.
To keep the kvfree_rcu barrier's promise, call irq_work_sync() on each
CPU before calling rcu_barrier().
In the meantime, remove the TODO item as apparently there is no simple
and effective way to achieve that. This is because, unlike sheaves,
kfree_rcu() batches objects from different caches together.
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Pedro Falcato <pfalcato@suse.de>
Reviewed-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260729-kfree_rcu_nolock-v5-4-a28cdcda9673@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
call_rcu() disables IRQs with local_irq_save() to protect its per-cpu
data structures. Therefore, if IRQs are not disabled, they cannot be
corrupted by reentrance into call_rcu(). So fall back to the deferred
path only when !allow_spin && irqs_disabled().
The RCU subsystem does not guarantee this contractually, and this
optimization relies on RCU's implementation details. Ideally, it should
be removed once call_rcu_nolock() is supported by the RCU subsystem.
Link: https://lore.kernel.org/linux-mm/CAADnVQKRVD5ZSnEKbZZU7w86gHbGHUug2pvzpgZTngNS+fg4rw@mail.gmail.com
Suggested-by: Alexei Starovoitov <ast@kernel.org>
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260729-kfree_rcu_nolock-v5-3-a28cdcda9673@kernel.org
Reviewed-by: Shengming Hu <hu.shengming@zte.com.cn>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Teach kfree_rcu_sheaf() how to handle the !allow_spin case. Try to get
an empty sheaf from pcs->spare or the barn even when spinning is not
allowed. Unlike __pcs_replace_full_main(), try harder to allocate
an empty sheaf because the fallback path will be more expensive than
kfree_nolock().
Now that slab has internal alloc_flags to describe context, introduce
free_flags analogously and convert free_flags to alloc_flags when
allocating memory in the free path.
When trylock fails or the kernel observes non-NULL pcs->rcu_free after
lock acquisition, free the sheaf instead of putting it to the barn.
This is rare and not worth complicating the code.
Since call_rcu() cannot be called in an unknown context,
kfree_rcu_sheaf() fails when the rcu sheaf becomes full.
Link: https://lore.kernel.org/linux-mm/872bd673-3d45-4111-8a41-31185db3ece5@kernel.org
Reviewed-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260729-kfree_rcu_nolock-v5-2-a28cdcda9673@kernel.org
Reviewed-by: Shengming Hu <hu.shengming@zte.com.cn>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Currently, struct slabobj_ext can hold both objcg pointer and
codetag_ref (when both are compile-enabled) and there is an array of as
many slabobj_ext instances as there are objects in a slab.
This makes the layout fixed so even if codetag_ref is unused (because
memory allocation profiling is disabled), the space for them is
allocated and wasted. Similarly, some caches (currently kmalloc_normal)
do not ever need objcg pointers, leading to wasted memory with memory
allocation profiling enabled.
To make this more flexible, change the layout so that struct slabobj_ext
becomes a union of objcg pointer and codetag_ref (to ensure uniform
size; in practice both are the same size anyway). The slabobj_ext array
then can have twice as many elements as before. For cache locality
purposes, the effective memory layout is unchanged, so objcg and codetag
ref for a given object are still adjacent.
cache_obj_ext_size() returns the effective size of (0-2) struct
slabobj_ext's for a cache, slab_obj_ext_size() for a slab. Currently
both return a constant value derived from the config options, but will
be made dynamic later. Replace all sizeof(slabobj_ext) usage with these.
No functional change intended, the layout is still effectively static.
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-8-c29ef0f1f257@kernel.org
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Hao Li <hao.li@linux.dev>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
The stride field is used to convert object index to an slabobj_ext so
both compact arrays (kmalloc() or in-slab-leftover) and spread
in-object-padding obj_ext layouts are supported.
In practice thus the stride is always sizeof(slabobj_ext) or s->size.
This simplifies the calculations, but with the upcoming slabobj_ext
handling changes, it will be easier to stop storing the stride and
instead just have a flag whether obj_ext is in the object padding.
obj_exts_in_object() can then rely on this flag and slab_obj_ext()
can use that to determine the stride.
No functional change intended.
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-7-c29ef0f1f257@kernel.org
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
In preparation for changes to the structure, abstract access to the ref
field with a slab_obj_ext_codetag_ref() function. Rename the field to
_ctref to make an unexpected direct access a compile error.
No functional change intended.
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Hao Li <hao.li@linux.dev>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-6-c29ef0f1f257@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
In preparation for changes to the structure, abstract getting and
setting the objcg field with slab_obj_ext_objcg() and
slab_obj_ext_set_objcg().
Rename the field to _objcg to make an unexpected direct access a compile
error.
The helpers take a slab pointer, which is currently unused, but will be
used by a debug check later.
Since there is no slab pointer easily available in __kfence_free(), just
drop the debug check there. The whole memcg_kmem accounting in kfence is
to be removed later anyway.
Otherwise, no functional change intended.
Reviewed-by: Hao Li <hao.li@linux.dev>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-5-c29ef0f1f257@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
All callers perform the same obj_to_index() calculation to pass the
index. Simplify by passing object pointer instead and determining the
index by slab_obj_ext().
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-4-c29ef0f1f257@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Users of include/linux/memcontrol.h don't need to see this internal
structure. Further changes to the struct will reduce recompiling.
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-3-c29ef0f1f257@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
The function has an unused kmem_cache argument and almost nothing uses
it anyway; doing slab->objects is simpler. Remove it with the last two
users. KUNIT_EXPECT_EQ() needs a cast to avoid "error: ‘typeof’ applied
to a bit-field" but we don't need to keep a wrapper just for that.
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-2-c29ef0f1f257@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
struct kfence_metadata only contains struct slabobj_ext with
CONFIG_MEMCG, which is then used for the "fake" slab's obj_exts field.
If CONFIG_MEMCG is enabled, the struct can also end up used for memory
allocation profiling. If CONFIG_MEMCG is disabled but profiling is
enabled, it will end up allocating its obj_exts via
prepare_slab_obj_exts_hook() and assigning them to the fake struct slab.
These will probably then never be freed.
So things sorta work, but not always in the intended and optimal way.
The upcoming changes to slabobj_ext layout would additionally need a
proper refactoring to keep working.
However, there's little benefit in accounting KFENCE objects. KFENCE
allocations are rare and there can be only CONFIG_KFENCE_NUM_OBJECTS
(default to 255) outstanding ones at any time. For any callsite
prominent enough in the memory allocation profiling stats, allocations
served from KFENCE will be lost in the noise.
Thus let's not complicate things and simply stop accounting KFENCE
objects in allocation profiling and skip them in the related slab hooks.
We also need to skip kfence objects in mark_obj_codetag_empty() in case
a sheaf is allocated from kfence, per earlier sashiko review.
Link: https://patch.msgid.link/20260727-b4-objext_split-v3-1-c29ef0f1f257@kernel.org
Reviewed-by: Hao Li <hao.li@linux.dev>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
slab_debugfs_init() creates the slab debugfs root at device initcall
time, while slab_sysfs_init() moves slab_state to FULL at late initcall
time. SLAB_STORE_USER caches created in this window miss their debugfs
entries because do_kmem_cache_create() skips debugfs_slab_add() when
slab_state <= UP. This was observed with MPTCP's request_sock_subflow_v6
cache, whose slab debugfs directory was missing.
The affected window is:
slab_debugfs_init()
slab_debugfs_root = debugfs_create_dir(...)
list_for_each_entry(s, &slab_caches, list)
debugfs_slab_add(s)
kmem_cache_create(..., SLAB_STORE_USER, ...)
do_kmem_cache_create()
if (slab_state <= UP)
return without debugfs entries
slab_sysfs_init()
slab_state = FULL
Initialize the debugfs root and add debugfs entries while holding
slab_mutex, walking slab_caches exactly once and handling both sysfs
and debugfs entries in the same pass. This gives the sysfs and debugfs
initialization an explicit order and prevents caches from being
created between the debugfs scan and slab_state reaching FULL.
Gate the new slab_late_init() on either sysfs or debugfs being enabled,
with the slab_kset creation and alias_list processing factored into
helpers that have empty no-sysfs variants, as suggested by Vlastimil
Babka. On slab_kset_init() failure, slab_state stays below FULL so
kmem_cache_create() keeps taking the early-boot path, matching prior
behavior.
Guard debugfs_slab_release() against an uninitialized debugfs root,
since the root is now created later and a cache may be released before
it exists.
Fixes: 1a5ad30b89 ("mm: slub: make slab_sysfs_init() a late_initcall")
Cc: stable@vger.kernel.org
Suggested-by: Vlastimil Babka <vbabka@kernel.org>
Signed-off-by: Li Xiasong <lixiasong1@huawei.com>
Link: https://patch.msgid.link/20260729101849.3734287-1-lixiasong1@huawei.com
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Since kmalloc_nolock() always fails in NMI and hardirq contexts on
PREEMPT_RT, slub_kunit cannot properly test _nolock() APIs.
Register a kprobe pre-handler to invoke kmalloc_nolock() and
kfree_nolock() in the middle of the slab allocator. However, do not
register the handler on UP kernels because that use case is not
well supported [1] in the kernel.
To attach the pre-handler while s->cpu_sheaves->lock or n->list_lock
is held, add a wrapper function for lockdep_assert_held() that calls
a no-op function slab_attach_kprobe_locked() on debug builds. The
function is optimized away when neither CONFIG_PROVE_LOCKING nor
CONFIG_DEBUG_VM is selected and register_kprobe() fails.
The function calls barrier() to prevent the compiler from optimizing
away its callsites. Otherwise, the compiler may consider the function
does not have any side effect and remove callsites.
Compared to using plain kprobe, this has two advantages: 1) it avoids
hardcoding function names in the test, and 2) it can trigger those APIs
in the middle of a function, where the lock is expected to be held as
annotated with lockdep.
While it was proposed [2] to use kunit function redirection to test
this, it is currently infeasible as some lock helpers don't have
symbols.
Factor out the nested loop that calls kmalloc and friends to
test_kmalloc_kfree(), and call them in
test_kmalloc_kfree_nolock_{perf,kprobe}(), each being an independent
test case. During the refactoring, drop alloc_fail handling as it
doesn't provide much benefits.
Link: https://lore.kernel.org/linux-mm/20260427-nolock-api-fix-v2-0-a6b83a92d9a4@kernel.org [1]
Link: https://lore.kernel.org/linux-mm/6edebc2b-5f5a-4b9c-9a4c-564310acee1b@kernel.org [2]
Acked-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Reviewed-by: Shengming Hu <hu.shengming@zte.com.cn>
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260729-kfree_rcu_nolock-v5-1-a28cdcda9673@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Sparse does not know __builtin_infer_alloc_token() and complains:
sparse: sparse: undefined identifier '__builtin_infer_alloc_token'
Fix it by using a dummy variant of __kmalloc_token() if __CHECKER__ is
defined.
Fixes: feb662d916 ("slab: support for compiler-assisted type-based slab cache partitioning")
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202607110912.nZTqfCrH-lkp@intel.com/
Signed-off-by: Marco Elver <elver@google.com>
Link: https://patch.msgid.link/20260721092005.1986693-1-elver@google.com
Acked-by: Harry Yoo (Oracle) <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
kmem_cache_return_sheaf() may refill a partially consumed sheaf before
placing it in the barn. Without an explicit restriction, this refill may
draw objects from pfmemalloc slabs and consume emergency reserves.
Add __GFP_NOMEMALLOC so that returned sheaves are refilled only from
non-pfmemalloc slabs. Also add __GFP_NOWARN, as suggested by Hao Li,
because this refill is a best-effort attempt and failure is acceptable.
If the refill fails, flush and free the sheaf instead.
Fixes: 1ce20c28ea ("slab: handle pfmemalloc slabs properly with sheaves")
Cc: stable@vger.kernel.org
Signed-off-by: Shengming Hu <hu.shengming@zte.com.cn>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260721084522552ZPa16p1SRj3PYat3sqxuN@zte.com.cn
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Commit 280ea9c315 ("mm/slab: avoid allocating slabobj_ext array from
its own slab") avoided recursive allocation of obj_exts from kmalloc
caches of the same size, by bumping the obj_exts array's allocation
size whenever the array size equals the size of the object being
allocated.
However, as reported by Danielle Costantino and Shakeel Butt,
even slabs from kmalloc caches of different sizes can form a cycle
by allocating obj_exts arrays from each other [1]:
What happened: a KMALLOC_NORMAL slab's obj_exts array (used by
allocation profiling / memcg accounting) is itself kmalloc()'d from a
KMALLOC_NORMAL cache, so the "slab holds another slab's obj_exts array"
relation can form cycles. With sizeof(struct slabobj_ext) == 16 and
the host's geometry:
- kmalloc-512 has 64 objects/slab -> array is 64*16 == 1024 bytes,
served from kmalloc-1k;
- kmalloc-1k has 32 objects/slab -> array is 32*16 == 512 bytes,
served from kmalloc-512.
A kmalloc-512 slab and a kmalloc-1k slab therefore hold each other's
obj_exts array. Discarding one frees the other's array, which empties
and discards that slab, which frees the first's array, and so on:
__free_slab() -> free_slab_obj_exts() -> kfree() -> discard_slab() ->
__free_slab() recurses along the cycle until the stack is exhausted.
With memory allocation profiling, this allows unbounded recursion
in the free path and led to a stack overflow on a production host in
the Meta fleet [1]:
BUG: TASK stack guard page was hit
Oops: stack guard page
RIP: 0010:kfree+0x8/0x5d0
Call Trace:
__free_slab+0x66/0xc0
kfree+0x3f0/0x5d0
... ( ~125x __free_slab <-> kfree ) ...
<kernel driver freeing a resource>
do_syscall_64
It is proposed [1] to resolve this issue by always serving the obj_exts
array allocation from kmalloc caches (or large kmalloc) of sizes larger
than the object size. However, as pointed out by Vlastimil Babka [2],
this can waste an excessive amount of memory as slabs from large
kmalloc sizes (e.g. kmalloc-8k) generally need obj_exts arrays much
smaller than the object size.
Therefore, rather than bumping the size, let us take a different
approach; disallow formation of cycles between kmalloc types when
allocating obj_exts arrays. Currently, all obj_exts arrays are served
from normal kmalloc caches. Cycles cannot be created if obj_exts arrays
of normal kmalloc caches are served from a special kmalloc type that can
never have obj_exts arrays.
To achieve this, create a new kmalloc type called KMALLOC_NO_OBJ_EXT.
KMALLOC_NO_OBJ_EXT caches are created with SLAB_NO_OBJ_EXT flag when
either 1) memory allocation profiling is not permanently disabled,
or 2) kmalloc types with a priority higher than KMALLOC_CGROUP are
aliased with KMALLOC_NORMAL.
Sheaf bootstrapping for KMALLOC_NO_OBJ_EXT caches now must be deferred
because allocation of a barn can trigger obj_exts array allocation of
normal kmalloc caches when the KMALLOC_NO_OBJ_EXT cache for that size
is not ready yet. For simplicity, perform bootstrapping of sheaves for
all kmalloc caches later.
Introduce a new slab alloc flag, SLAB_ALLOC_NO_OBJ_EXT, to prevent
allocation of obj_exts arrays, and let kmalloc_slab() override the type
to KMALLOC_NO_OBJ_EXT when specified. Note that kmalloc_type() remains
unchanged because kmalloc_flags() bypasses the kmalloc fastpath.
Do not pass SLAB_ALLOC_NO_RECURSE to kmalloc_flags() in
alloc_slab_obj_exts() and instead use SLAB_ALLOC_NO_OBJ_EXT only when
the objects are allocated from normal kmalloc caches. While this
prevents unbounded recursive allocation of obj_exts, it allows
KMALLOC_NO_OBJ_EXT caches to have sheaves.
Since sheaf allocations specify SLAB_ALLOC_NO_RECURSE that prevents
allocation of both sheaves and obj_exts arrays, the recursion depth
is bounded.
obj_exts arrays for non-kmalloc-normal caches can now have a valid tag.
Do not call mark_obj_codetag_empty() when freeing an obj_exts array to
avoid false warnings. KMALLOC_NO_OBJ_EXT don't need this as they never
allocate those arrays.
Reported-by: Danielle Costantino <dcostantino@meta.com>
Reported-by: Shakeel Butt <shakeel.butt@linux.dev>
Closes: https://lore.kernel.org/linux-mm/20260625230029.703750-1-shakeel.butt@linux.dev [1]
Fixes: 4b87369646 ("mm/slab: add allocation accounting into slab allocation and free paths")
Cc: stable@vger.kernel.org
Link: https://lore.kernel.org/linux-mm/c5c4208d-a6f0-413e-bad9-49be12f12d55@kernel.org [2]
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Link: https://patch.msgid.link/20260713-kmalloc-no-objext-v3-4-47c7bd138de7@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
mem_alloc_profiling_enabled() tells whether memalloc profiling is
currently enabled. However, even when this function returns false,
it can be enabled later.
However, this is not enough. Some optimizations can be applied only when
memalloc profiling is permanently disabled. For example, to skip the
creation of KMALLOC_NO_OBJ_EXT caches at boot time, mem_profiling must
be set to "never", "0" w/ debugging on, or have been shutdown so that
it can no longer be enabled.
Introduce mem_alloc_profiling_permanently_disabled() for this purpose.
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Acked-by: Suren Baghdasaryan <surenb@google.com>
Link: https://patch.msgid.link/20260713-kmalloc-no-objext-v3-3-47c7bd138de7@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Bootstrap caches are created with SLAB_NO_OBJ_EXT to disallow sheaves
and obj_exts.
To allow disabling obj_exts while allowing sheaves, decouple
SLAB_NO_SHEAVES from SLAB_NO_OBJ_EXT. Bootstrap caches now have both
SLAB_NO_SHEAVES and SLAB_NO_OBJ_EXT.
No functional change intended.
Reviewed-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Link: https://patch.msgid.link/20260713-kmalloc-no-objext-v3-2-47c7bd138de7@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
When kmalloc caches are aliased, multiple cache pointers reference
the same kmem_cache. As a result, iterating over kmalloc indices and
bootstrapping sheaves can bootstrap the same cache more than once and
leak memory.
Currently, this could happen when the architecture specifies
minimum alignment for slab caches that is larger than
ARCH_KMALLOC_MINALIGN.
Bootstrap sheaves only when the cache does not have them already.
Add a warning when bootstrap_cache_sheaves() is called for a cache
that already has sheaves enabled.
Fixes: 913ffd3a1b ("slab: handle kmalloc sheaves bootstrap")
Cc: stable@vger.kernel.org
Signed-off-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Suren Baghdasaryan <surenb@google.com>
Link: https://patch.msgid.link/20260713-kmalloc-no-objext-v3-1-47c7bd138de7@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Add a comment to clarify why we don't call kobject_put() when
kobject_init_and_add() fails in sysfs_slab_add().
Per commit 2420baa8e0 ("mm/slab: Allow cache creation to proceed
even if sysfs registration fails"), sysfs failures are treated as
non-fatal and the cache continues to be used. Calling kobject_put()
would trigger slab_kmem_cache_release() which frees the entire
cache structure, so we intentionally skip it.
Suggested-by: Harry Yoo <harry@kernel.org>
Signed-off-by: Hongling Zeng <zenghongling@kylinos.cn>
Acked-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260713070024.153552-1-zenghongling@kylinos.cn
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
We have been moving remote objects to an on-stack array and flushing it
when full. Instead, we can swap them towards the beginning of the
supplied array and bulk-free it just once.
Also add a comment to explain the rationale of freeing remote objects
last, because now it would appear to be simpler to free them first.
Reviewed-by: Pedro Falcato <pfalcato@suse.de>
Reviewed-by: Shengming Hu <hu.shengming@zte.com.cn>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260713-bulk_free_remote-v2-1-24ee24771c2f@kernel.org
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
It has been noted that free_to_pcs_bulk() is difficult to follow, with a
number of goto labels, and this has contributed to two memory leak bugs
in there.
Extract part of the code to __free_to_pcs_batch(), which focuses only on
freeing free-hook-processed local objects to a percpu sheaf, and
returning how many were freed. Zero means a trylock failure or no empty
sheaf available, and thus the caller should fallback to
__kmem_cache_free_bulk().
Make free_to_pcs_bulk() call this in a while loop, removing all goto
labels from the function. __free_to_pcs_batch() retains two rather
straightforward ones.
Reviewed-by: Shengming Hu <hu.shengming@zte.com.cn>
Link: https://patch.msgid.link/20260707-slab-simplify-bulk-pcs-v1-1-4850dbe0d904@kernel.org
Reviewed-by: Hao Li <hao.li@linux.dev>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
The per-cpu slab and per-cpu partial slab mechanisms were removed when
SLUB was fully converted to per-cpu sheaves in Linux 7.0. The cpu_slabs,
slabs_cpu_partial and cpu_partial sysfs attributes were kept as stubs
that always return 0 for backwards compatibility, but their
documentation still described them as if they were functional.
Update the three descriptions to state that the attributes are
deprecated and always read 0, and note that they are retained only for
compatibility. While here, fix a "partialli" typo in the
slabs_cpu_partial description.
Signed-off-by: Seongjun Hong <hsj0512@snu.ac.kr>
Acked-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260701141755.85119-1-hsj0512@snu.ac.kr
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Since the removal of the per-cpu slab in commit 32c894c727 ("slab:
remove struct kmem_cache_cpu"), show_slab_objects() no longer has a
branch handling SO_CPU, so cpu_slabs_show() always produces "0".
Emit "0\n" directly instead, matching the sibling cpu_partial and
slabs_cpu_partial stubs, and remove the now-unused SO_CPU macro and
SL_CPU enum value.
No functional change intended; the cpu_slabs sysfs attribute continues
to read 0.
Signed-off-by: Seongjun Hong <hsj0512@snu.ac.kr>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Reviewed-by: Hao Li <hao.li@linux.dev>
Link: https://patch.msgid.link/20260701140634.71608-1-hsj0512@snu.ac.kr
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
The mempool subsystem historically wrapped its debugging logic inside an
merely defines compile-time defaults for SLUB and caused two flaws:
1. On production kernels where CONFIG_SLUB_DEBUG=y but
CONFIG_SLUB_DEBUG_ON=n, mempool debugging was completely compiled out
at compile time.
2. On kernels with CONFIG_SLUB_DEBUG_ON=y, mempool debugging stayed active
even if a user explicitly disabled slub debugging at boot time.
Clean up this mess by removing the #ifdef and switching to a runtime static
key (mempool_debug_enabled), allowing mempool debugging to be toggled
cleanly via its own boot parameter.
Suggested-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Signed-off-by: Li RongQing <lirongqing@baidu.com>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Harry Yoo <harry@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Hao Li <hao.li@linux.dev>
Cc: Christoph Lameter <cl@gentwo.org>
Cc: David Rientjes <rientjes@google.com>
Cc: Roman Gushchin <roman.gushchin@linux.dev>
Cc: Matthew Wilcox <willy@infradead.org>
Cc: Usama Arif <usama.arif@linux.dev>
Reviewed-by: SeongJae Park <sj@kernel.org>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260604110318.2089-1-lirongqing@baidu.com
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
Currently, alloc_from_pcs() and __slab_alloc_node() both calculate the
NUMA policy independently. Since they are called consecutively in paths
like __kmalloc_nolock_noprof() and slab_alloc_node(), this leads to
redundant code snippets.
Introduce a helper function to resolve the NUMA policy once, eliminating
the duplicated code and reducing execution overhead.
Also remove __slab_alloc_node() function because it is almost empty.
The callers of __slab_alloc_node now call ___slab_alloc() directly.
Additional notes:
Previously, when slab_strict_numa was enabled, alloc_from_pcs() and
__slab_alloc_node() could each resolve the task mempolicy, so
MPOL_INTERLEAVE or MPOL_WEIGHTED_INTERLEAVE could advance the
interleave state twice for a single object allocation attempt.
And each retry will also advance the interleave state.
With this change, the strict NUMA node is resolved once and reused by
both alloc_from_pcs() and ___slab_alloc() in each retry.
This is a behavior change, but it better matches the intent of
selecting one policy node for one allocation attempt.
Signed-off-by: Hao Li <hao.li@linux.dev>
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Link: https://patch.msgid.link/20260624100320.430115-1-hao.li@linux.dev
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
In free_to_pcs_bulk(), when remote_objects[] fills to PCS_BATCH_MAX,
the code jumps to flush_remote to free the batch. If all remote entries
have already been compacted out of p[] via tail swaps while local objects
remain, the flush_remote path returns early since `i < size` no longer
holds. The leftover local objects are then neither cached in the sheaf
nor returned to the slab freelist, causing a memory leak.
For illustration:
size = 64, local objects at p[0..31], remote objects at p[32..63]
After scanning all remotes: i = 32, size = 32
p[0..31] local objects are dropped.
Harry pointed out that, although the logic contains a real leak, it does
not appear to be triggerable with the current in-tree users. To hit this
path, at least PCS_BATCH_MAX objects, currently hardcoded to 32, need to
be collected in remote_objects[]. Looking at current kmem_cache_free_bulk()
users:
* maple_node has sheaf_capacity = 32
* skbuff_head_cache has sheaf_capacity = 28
* panthor and msm drivers have sheaf_capacity = 4
The sheaf capacity is, at least for now, derived purely from the object
size, with the user-requested capacity used as a minimum. Therefore, among
the current users, only maple_node has a sheaf_capacity large enough to
reach PCS_BATCH_MAX.
However, for the bug to trigger in maple_node, all objects in the sheaf
would have to be from remote nodes. In that case, there would be no local
objects left to leak. So this issue was found by code review rather than
from a runtime report, and it does not seem to be triggerable by current
users.
Still, the bug could become reachable with future users, a different sheaf
capacity, or a change to PCS_BATCH_MAX. Fix the logic by freeing a full
remote batch in place during the scan and then continuing to process the
compacted array. This keeps all local objects on the normal fast path,
while the tail path only handles any leftover partial remote batch. The
redundant next_remote_batch jump label is removed as well.
Fixes: 989b09b739 ("slab: skip percpu sheaves for remote object freeing")
Signed-off-by: Shengming Hu <hu.shengming@zte.com.cn>
Link: https://patch.msgid.link/202607062139095043SOsLi6TIf403tcjPf8fm@zte.com.cn
Cc: stable@vger.kernel.org
Reviewed-by: Harry Yoo (Oracle) <harry@kernel.org>
Signed-off-by: Vlastimil Babka (SUSE) <vbabka@kernel.org>
device removal, along with documentation fixes and minor AMD driver
cleanups.
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEoE9b9c3U2JxX98mqbmZLrHqL0iMFAmpBIsgACgkQbmZLrHqL
0iN+tA/+LCXsiysq5XLscCkWL7vwT2rbh/uLBjnBC2YEgfFWOtfOrwnNYlSXEJAo
PE9RRbqCBGXpv0CQ5iah96ff8HdDU6O36+UD9el1jGxUN/UxtD0Q7ibfsJ+gyf4N
ANNDrkTAo6AcDxM/AG9B/bBy5EnNDDH1bURIWM3dohp3CNhRbbSHbyiuf20uFgef
i+iArjk2ePL0tGEzUIEHVsEwGJFVrgAYr/7OZ2dSvpMn8AhI8bj5U3YANwQOnQS6
QlNKZ/t7mgNJ4zGBIhzQcmUlIuLY0GsOhyKhXAcmc7uNumJotyuTcQ7F3Ybky8DY
Pjzt7fc+4wT9fH/B303JfuFY2SXUWHkTXA1zZGfeV8zJrljqfEx+8elJAkRR7ifJ
PVR2w85W88qfW5HfBfiBGstq5aA9aFNQscJQIN//ejjuPbrLPEeIodLte4j8QzRb
uo+8L1n2TiR7NljdAYBrhuUAnoW5F7sx/k0jsqZvOvM48oVzr1pZeNJF3EG5UHfE
9RmOx+VuF4f5p0zMYjvhumCRN3p8fwKNDMxGBTGSD4YQ6+SRTh6HnfQDwv4uqVBn
nC8fBHLEdj4wmUzttYuCOD+ltBN02viyv/bUIQdd+BDjsR5b7q76Rxo3BGkNy1/z
FdXGFs3L26pOY5X+equ8FAfd5V4sfVL2HZO5PYPclP2HsJroHz8=
=V3cD
-----END PGP SIGNATURE-----
Merge tag 'ntb-7.2' of https://github.com/jonmason/ntb
Pull NTB updates from Jon Mason:
"An EPF bug fix to prevent an invalid unmap during device removal,
along with documentation fixes and minor AMD driver cleanups"
* tag 'ntb-7.2' of https://github.com/jonmason/ntb:
ntb: amd: Use named initializer for pci_device_id::driver_data
NTB: fix kernel-doc warnings in ntb.h
NTB: epf: Avoid pci_iounmap() with offset when PEER_SPAD and CONFIG share BAR
ntb_hw_amd: Fix incorrect debug message in link disable path
- Updates to Synaptics RMI4 driver to fix potential OOB accesses in
F30 and F3A keymap handling
- A workaround in Synaptics RMI4 to tolerate buggy firmware on some
touchpads (e.g. ThinkPad T14 Gen 1) that report incomplete register
descriptor structures, preventing probe failures
- A revert of an incorrect register descriptor address calculation in
Synaptics RMI4 driver
- A fix for a regression in HP GSC PS/2 (gscps2) driver where the
receive buffer write index was not advanced, leaving keyboard and
mouse unusable.
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQST2eWILY88ieB2DOtAj56VGEWXnAUCakCsxwAKCRBAj56VGEWX
nCICAQDSgfrAi+4SqTb92EjtdQO+ypluS42mKO75LTitJcS8dAEA1iKmdss8mGww
c4ai4W+UFxori7IqgkoQ7LfaPTOWlA0=
=QP4c
-----END PGP SIGNATURE-----
Merge tag 'input-for-v7.2-rc0-2' of git://git.kernel.org/pub/scm/linux/kernel/git/dtor/input
Pull more input updates from Dmitry Torokhov:
- Updates to Synaptics RMI4 driver to fix potential OOB accesses in F30
and F3A keymap handling
- A workaround in Synaptics RMI4 to tolerate buggy firmware on some
touchpads (e.g. ThinkPad T14 Gen 1) that report incomplete register
descriptor structures, preventing probe failures
- A revert of an incorrect register descriptor address calculation in
Synaptics RMI4 driver
- A fix for a regression in HP GSC PS/2 (gscps2) driver where the
receive buffer write index was not advanced, leaving keyboard and
mouse unusable.
* tag 'input-for-v7.2-rc0-2' of git://git.kernel.org/pub/scm/linux/kernel/git/dtor/input:
Input: gscps2 - advance receive buffer write index
Input: rmi4 - tolerate short register descriptor structure
Revert "Input: rmi4 - fix register descriptor address calculation"
Input: synaptics-rmi4 - bound the F30 keymap to the GPIO/LED count
Input: synaptics-rmi4 - bound the F3A keymap to the GPIO count
Two more fixes that I managed to put into the public branch merged into
next before my first PR but missed to include them in it. The first
change is a relevant change that fixes misconfigurations due to a
variable overflow. The 2nd is only cosmetic but very obviously an
improvement.
-----BEGIN PGP SIGNATURE-----
iQEzBAABCgAdFiEEP4GsaTp6HlmJrf7Tj4D7WH0S/k4FAmpAE2gACgkQj4D7WH0S
/k7tnAgAsjkBQWoav1MF8kyVzew+iIuVvBrtUPimzeMs2gzwC1czi/TPKCvZxGND
XaPESmf51lC9AZuFd/2Wfvb2GydxA/wmsbVB0Q6ZBwLtS6/L4yiv3DpYN2to9yWN
QvVHeVCBdSbMHOQXdG0iFbhMiLyUX5YwCwyZT2cuVUHb4gvNLuDuKbgSXX9Odh1R
/p9C0afNMbdxuj2yRy+S8CM5Rl5v4yfBw6cswKX6w3uA+LnvksWuC8og7tEfFa/F
Vdwy4csWeKHGm7jEP9o6iSlWYy+DMKf5Llop+yFLZV6Db6AXMlEpNOXJsGxEv95y
qMTQFCGhepNWqVuUDDOOeUs60Asjpg==
=0WY9
-----END PGP SIGNATURE-----
Merge tag 'pwm/for-7.2-rc1-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ukleinek/linux
Pull pwm fixes from Uwe Kleine-König:
"Two more fixes that I managed to put into the public branch merged
into next before my first pull request but missed to include them in
it.
The first change is a relevant change that fixes misconfigurations due
to a variable overflow. The second is only cosmetic but very obviously
an improvement"
* tag 'pwm/for-7.2-rc1-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ukleinek/linux:
pwm: rzg2l-gpt: Add missing newlines to dev_err_probe() messages
pwm: rzg2l-gpt: Fix period_ticks type from u32 to u64
Fixes:
- fbcon: fix NULL pointer dereference for a console without vc_data [Ian Bridges]
- fbdev: fix fb_new_modelist to prevent null-ptr-deref in fb_videomode_to_var [Ian Bridges]
Fixes in failure paths:
- pm2fb: unwind write-cache setting on probe failure [Haoxiang Li]
- goldfishfb: fail pan display on base-update timeout [Pengpeng Hou]
- viafb: return error on DMA copy time-out [Pengpeng Hou]
- fbcon: fix out-of-bounds read in error path of fbcon_do_set_font() [Mingyu Wang]
- fbdev: fix modelist use-after-free in store_modes() [Ian Bridges]
Code cleanup:
- vga16fb: clean up platform_device_id table [Uwe Kleine-König]
-----BEGIN PGP SIGNATURE-----
iHUEABYKAB0WIQS86RI+GtKfB8BJu973ErUQojoPXwUCakAKnwAKCRD3ErUQojoP
X27+AQCPvMJUMScjO1dlGKkPnnIFiOWCU6AW1aM2m6T+qWrCTQEAnNi5CCI3JiLi
6NCcyPm7/4R6tSXDm+UcWJ5tyzaNvww=
=EBHb
-----END PGP SIGNATURE-----
Merge tag 'fbdev-for-7.2-rc1-2' of git://git.kernel.org/pub/scm/linux/kernel/git/deller/linux-fbdev
Pull more fbdev updates from Helge Deller:
"Fixes for generic fbdev & fbcon code for the handling of modelists
and preventing a potential NULL ptr dereference in the console code.
Fix missed cleanups in the error path of various fbdev drivers.
And Uwe Kleine-König contributed a cleanup patch to use named
initializers in the vga16fb driver"
* tag 'fbdev-for-7.2-rc1-2' of git://git.kernel.org/pub/scm/linux/kernel/git/deller/linux-fbdev:
fbdev: Fix fb_new_modelist to prevent null-ptr-deref in fb_videomode_to_var
fbcon: fix NULL pointer dereference for a console without vc_data
fbdev: fix use-after-free in store_modes()
fbdev: viafb: return an error when DMA copy times out
fbdev: goldfishfb: fail pan display on base-update timeout
fbdev: fbcon: fix out-of-bounds read in err_out of fbcon_do_set_font()
fbdev: pm2fb: unwind WC setup on probe failure
fbdev: vga16fb: Drop unused assignment of platform_device_id driver data
A collection of small bug fixes accumulated over the last week.
Most are device-specific fixes while there are a few core fixes as
well.
Here are the highlights:
ALSA Core:
- A fix for an uninitialised heap leak in ALSA sequencer core
- A fix for error handling/resource leak in compress-offload API
USB-audio:
- A teardown-ordering fix in USB MIDI 2.0 to prevent use-after-free
- Bounds and length checks for packet data in Native Instruments caiaq
/ Traktor Kontrol input parsers
- Avoidance of expensive kobject path lookups in DualSense controller
matches
- Robustness/memory leak fixes for Qualcomm USB offload driver
- Focusrite Control Protocol (FCP) NULL-pointer dereference fix and a
new device quirk (ISA C8X)
- Device-specific quirks for Yamaha CDS3000 and SC13A
HD-Audio:
- A bunch of quirks and mute/mic-mute LED fixups for various laptops
(Acer, Clevo, Lenovo, HP)
ASoC & SoundWire:
- Avoid failing card registration if the device_link creation fails
- A workaround for SoundWire randconfig build failures by making
helper functions static inline
- Corrected MCLK reference validation for CS530x codecs
- Clean up of untested, problematic guard() macro replacements in
Rockchip SAI driver
- Fix for eDMA maxburst misalignment with channel count in Freescale
ASRC
- Miscellaneous hardware-specific fixes (qcom, rt5650, tlv320aic3x,
tas2781/3)
Others:
- Bounds and length checks for packet data in Apple iSight
-----BEGIN PGP SIGNATURE-----
iQJCBAABCAAsFiEEIXTw5fNLNI7mMiVaLtJE4w1nLE8FAmo+mnEOHHRpd2FpQHN1
c2UuZGUACgkQLtJE4w1nLE92VQ/+ItB+EBTpiba9YQYBrzUzq2R3BiNR/EZjU33G
UMut1zQYQJ53eMmN8yMYc0GMbtk9dCFUAtRGPyQCNEHS6uFw51t3A4wlcXvIu1Sx
kQqtyaDQ2jp98J72ms4WtN42o29MjcFmhBBcTb3Kw12T+OVTYYneccsGPsHqCXsZ
RBjJFpDr0Xo1TfnOy9nt/UNUUIMJEtZ1gGlYBqzQgNoLeYH3+dRKBoX2qVAvhIcL
FJnSGiDgyLpt6uucPAAeIzGHawQXW4ej7XY4S8cLscsB7mY7VEtPFIMx4bN1QYIO
Ioj2P9KLG4/KYOV8oRQ6kzYTwtO7St9Kd/+xpU5Divjxf6TqRGlv/hlQCTBBZPLq
RVUsEiE36UlSuipyruK34KubtVkbqUgUjBiPygFr6cLKb6fc6sjWrK5P8KUtN860
8q1froUK43gwdVcdmLgrMbFCspE+KUp3xzSDh9tcVq6Ffw+otuuC0cJeVG4j+GOf
xntsUqlAX6XSudTvTfa1pqvQmynBqvBy4wW9yrRfvEJ6eJqRlT17Sbs6AzpeLE4k
dpeHlHwHtk5kfEGkYarJ3CEDw1GfHdLfQ6B6lBmCKq6DwnTZbq+lX5C3wB+OXVom
xn5enCuaygVnXs6RF6DP3KlSvLoCJ09BEehkERxVg1uyVnGkioXXwNU8vdjkj2Qd
srITNeo=
=hJ5w
-----END PGP SIGNATURE-----
Merge tag 'sound-fix-7.2-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound
Pull sound fixes from Takashi Iwai:
"A collection of small bug fixes accumulated over the last week.
Most are device-specific fixes while there are a few core fixes as
well.
Here are the highlights:
ALSA Core:
- A fix for an uninitialised heap leak in ALSA sequencer core
- A fix for error handling/resource leak in compress-offload API
USB-audio:
- A teardown-ordering fix in USB MIDI 2.0 to prevent use-after-free
- Bounds and length checks for packet data in Native Instruments
caiaq / Traktor Kontrol input parsers
- Avoidance of expensive kobject path lookups in DualSense controller
matches
- Robustness/memory leak fixes for Qualcomm USB offload driver
- Focusrite Control Protocol (FCP) NULL-pointer dereference fix and a
new device quirk (ISA C8X)
- Device-specific quirks for Yamaha CDS3000 and SC13A
HD-Audio:
- A bunch of quirks and mute/mic-mute LED fixups for various laptops
(Acer, Clevo, Lenovo, HP)
ASoC & SoundWire:
- Avoid failing card registration if the device_link creation fails
- A workaround for SoundWire randconfig build failures by making
helper functions static inline
- Corrected MCLK reference validation for CS530x codecs
- Clean up of untested, problematic guard() macro replacements in
Rockchip SAI driver
- Fix for eDMA maxburst misalignment with channel count in Freescale
ASRC
- Miscellaneous hardware-specific fixes (qcom, rt5650, tlv320aic3x,
tas2781/3)
Others:
- Bounds and length checks for packet data in Apple iSight"
* tag 'sound-fix-7.2-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/tiwai/sound: (46 commits)
ALSA: FCP: Fix NULL pointer dereference in interface lookup
ALSA: hda/realtek: Update Acer Nitro ANV15-41 quirk to enable mute LED
ASoC: fsl_asrc_dma: fix eDMA maxburst misalignment with channel count
ASoC: codecs: pcm512x: only print info once on no sclk
ASoC: tas2781: Update default register address to TAS2563
ALSA: firewire: isight: bound the sample count to the packet payload
ALSA: usb-audio: qcom: Free QMI handle
ALSA: hda: Add Lenovo Legion 7i 16IAX7 17AA3874 quirk
ALSA: usb-audio: avoid kobject path lookup in DualSense match
ALSA: hda/realtek: Add quirk for Acer Nitro ANV15-41
ASoC: soc-core: Don't fail if device_link could not be created
ASoC: rockchip: rockchip_sai: #include <linux/platform_device.h> explicitly
ALSA: seq: Fix uninitialised heap leak in snd_seq_event_dup()
ASoC: rt5575: Use __le32 for SPI burst write address
ASoC: tas2783: Update loaded firmware names to linux-firmware 20260519
ASoC: SDCA: Validate written enum value in ge_put_enum_double()
ASoC: realtek: Add back local call to sdw_show_ping_status()
ASoC: ti: Add back local call to sdw_show_ping_status()
ASoC: max98373: Add back local call to sdw_show_ping_status()
ASoC: es9356: Add back local call to sdw_show_ping_status()
...
Subsystem:
- add rtc_read_next_alarm() to read next expiring timer
Drivers:
- ds1307: handle OSF for ds1337/ds1339/ds3231, add clock provider for ds1307,
fix wday for rx8130
- m41t93: DT support, alarm, clock provider, watchdog support
- mv: add suspend/resume support for wakeup
- pcap: remove driver
- renesas-rtca3: many fixes
-----BEGIN PGP SIGNATURE-----
iQIzBAABCgAdFiEEBqsFVZXh8s/0O5JiY6TcMGxwOjIFAmo+9GoACgkQY6TcMGxw
OjKw5w/9GsA/6bIFf0xusBLhjooYLUCbHoluvgdifG6lLSSRUrkAwbipEhpcWtkp
uyGsjGtv/AxoqycC9OSg+OzkXY3Gw9viKVbSCjrA/PSe4trbcVcPxolYKg+Js20V
tDHbrVgpxsvSetjn1kM/rjVGoL4ywprmatMjkb6xXGoo6NE5IAXBVb+EJXKNWy+c
d/iR+DM0WHTLeNQ8MOxSexOReY4IiDj+Z9dxdZ600UCg54dYFFi06r4om+EX71T+
LrhUVyFxvOUmwMoavRiBiWh9PpLee/fN6z44QPK3nQp1qgvzsLCi90HI/h4ZaF+E
N+vSED2iaahU188bkXTmFNvQHJvipUKkAWDfw/wLJQXKkIjWGhj+RWQmMMeFBCdu
CA3NxiXvup4wPsSW66etz1Z6VJ22UeclNm57bn6rXLJn5t/enTc2c6HH2gsSzj8M
EncNd/yt76Sd9OnVSaM6LsPom+tm/Nd8DKVnORRlPl3p02z7u+GgMPzv9u7DlB5j
MNU4TLHFqL5kSbiSxQ+5bTRiqVGspEQFI9wIpTFibl89hpJYce+aonOeY5ZJDidk
/wyJArMu7S/ZdG0TNBeJ0jFatKoK6nEQe8tjxNRvYLhT2SL1Hcjmo1ab4J9DD5Js
YSoU4iQmjG1gF9Lj1Of+9WFqlRaAwa3mX3TJrxhSmXgl2zXoWKs=
=SCOZ
-----END PGP SIGNATURE-----
Merge tag 'rtc-7.2' of git://git.kernel.org/pub/scm/linux/kernel/git/abelloni/linux
Pull RTC updates from Alexandre Belloni:
"Most of the work and improvements are for features of the m41t93.
The ds1307 also gets support for OSF (Oscillator Stop Flag) for
new variants.
The pcap driver is being removed as the Motorola EZX support was
removed a while ago.
Subsystem:
- add rtc_read_next_alarm() to read next expiring timer
Drivers:
- ds1307: handle OSF for ds1337/ds1339/ds3231, add clock provider for
ds1307, fix wday for rx8130
- m41t93: DT support, alarm, clock provider, watchdog support
- mv: add suspend/resume support for wakeup
- pcap: remove driver
- renesas-rtca3: many fixes"
* tag 'rtc-7.2' of git://git.kernel.org/pub/scm/linux/kernel/git/abelloni/linux: (36 commits)
rtc: ds1307: update reference to removed CONFIG_RTC_DRV_DS1307_HWMON
platform/x86: amd-pmc: Fix S0i3 wakeup with alarmtimer
rtc: s35390a: fix typo in comment
rtc: cmos: unregister HPET IRQ handler on probe failure
rtc: ds1307: Fix off-by-one issue with wday for rx8130
dt-bindings: rtc: ds1307: Add epson,rx8901
rtc: bq32000: add delay between RTC reads
rtc: m41t93: Add watchdog support
rtc: m41t93: Add square wave clock provider support
rtc: m41t93: Add alarm support
rtc: m41t93: migrate to regmap api for register access
rtc: m41t93: add device tree support
dt-bindings: rtc: Add ST m41t93
rtc: ds1307: add support for clock provider in ds1307
rtc: mv: add suspend/resume support for wakeup
rtc: aspeed: add AST2700 compatible
dt-bindings: rtc: add ASPEED AST2700 compatible
rtc: interface: fix typos in rtc_handle_legacy_irq() documentation
rtc: msc313: fix NULL deref in shared IRQ handler at probe
rtc: remove unused pcap driver
...
- Fix a bug where in a specific edge case, file contents en/decryption
could be done with the wrong data unit size.
- Fix the data structure used for keeping track of users that have added
an fscrypt key to be a simple list instead of a 'struct key' keyring.
This fixes issues such as a lockdep report found by syzbot and
possible unintended interactions with the keyctl() system calls.
-----BEGIN PGP SIGNATURE-----
iIoEABYIADIWIQSacvsUNc7UX4ntmEPzXCl4vpKOKwUCaj8bBhQcZWJpZ2dlcnNA
a2VybmVsLm9yZwAKCRDzXCl4vpKOKyhIAP47lR+H783gopiz10Z7dwLQr2EHMfZ1
NcU7Zfq++AzZRQD/Tvv9/daqSh0OJTDNkBskSgVj6z7sRlQVNHY0wrD+QwE=
=LWES
-----END PGP SIGNATURE-----
Merge tag 'fscrypt-for-linus' of git://git.kernel.org/pub/scm/fs/fscrypt/linux
Pull fscrypt fixes from Eric Biggers:
- Fix a bug where in a specific edge case, file contents en/decryption
could be done with the wrong data unit size
- Fix the data structure used for keeping track of users that have
added an fscrypt key to be a simple list instead of a 'struct key'
keyring
This fixes issues such as a lockdep report found by syzbot and
possible unintended interactions with the keyctl() system calls
* tag 'fscrypt-for-linus' of git://git.kernel.org/pub/scm/fs/fscrypt/linux:
fscrypt: Replace mk_users keyring with simple list
fscrypt: Fix key setup in edge case with multiple data unit sizes
Commit 44f9200699 ("Input: gscps2 - use guard notation when
acquiring spinlock") moved the receive loop into gscps2_read_data()
and gscps2_report_data().
While moving the code, it preserved the writes to
buffer[ps2port->append], but omitted the following producer index
update from the original loop:
ps2port->append = (ps2port->append + 1) & BUFFER_SIZE;
As a result, append never advances. Since gscps2_report_data() only
reports bytes while act != append, the receive buffer always appears
empty and no keyboard or mouse data reaches the serio core.
Restore the omitted index update.
Fixes: 44f9200699 ("Input: gscps2 - use guard notation when acquiring spinlock")
Cc: stable@vger.kernel.org # 6.13+
Signed-off-by: Xu Rao <raoxu@uniontech.com>
Link: https://patch.msgid.link/460B5655BA580C60+20260624094739.850306-1-raoxu@uniontech.com
Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>
Some touchpads (e.g. ThinkPad T14 Gen 1) have buggy firmware that reports
a register descriptor structure size that is too small for the number of
registers it claims to have in the presence map. The remaining bytes in
the structure are 0, which with the new strict bounds checking causes the
parser to fail with -EIO, aborting the device probe.
Tolerate such short reads by dropping the remaining (unparseable or
0-size) registers from the list instead of failing the probe,
preventing the driver from trying to use them.
Fixes: 0adb483fbf ("Input: rmi4 - refactor register descriptor parsing")
Reported-by: Barry K. Nathan <barryn@pobox.com>
Tested-by: Barry K. Nathan <barryn@pobox.com>
Cc: stable@vger.kernel.org
Assisted-by: Antigravity:gemini-3.5-flash
Signed-off-by: Dmitry Torokhov <dmitry.torokhov@gmail.com>