Commit Graph

4188 Commits

Author SHA1 Message Date
Linus Torvalds
415f204422 Landlock fix for v7.3-rc5
-----BEGIN PGP SIGNATURE-----
 
 iIYEABYKAC4WIQSVyBthFV4iTW/VU1/l49DojIL20gUCarVI0xAcbWljQGRpZ2lr
 b2QubmV0AAoJEOXj0OiMgvbSEVgA+gNbC9CVCbCo0oufZbVpQwlwtuuEKEVpvxZx
 q7oFYKCsAP9svCujGCXRHOmWhAAwe+wpXNb43l8coFn+pCfA8x3fCw==
 =NOli
 -----END PGP SIGNATURE-----

Merge tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux

Pull Landlock fixes from Mickaël Salaün:
 "This mainly fixes the Landlock tracepoint support merged this cycle so
  that denial and rule events report the intended policy context,
  whether through tracefs or BTF-visible callbacks.

  The size of this all is mainly from propagating the corrected contract
  through event definitions and producers, adding new tests for the
  reported context, and updating the documentation.

  Also improve annotation and fix a GCC 16 build warning"

* tag 'landlock-7.3-rc5' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux:
  landlock: Widen ruleset versions to 64 bits
  landlock: Add counted_by in landlock_domain
  landlock: Fix tracepoint contract documentation
  selftests/landlock: Test network denial context
  selftests/landlock: Test filesystem denial blockers
  landlock: Report the effective signal number
  landlock: Report the actual ptrace tracer
  landlock: Fix network denial trace context
  landlock: Fix rule tracepoint context
  landlock: Fix filesystem denial blocker reporting
  landlock: Fix tracepoint fixed-width type names
  landlock: Work around gcc-16 -Wuninitialized warning
2026-09-24 10:53:59 -07:00
Mickaël Salaün
e7e0a54300
landlock: Widen ruleset versions to 64 bits
Tracepoint consumers use a ruleset ID and version to identify the
successful landlock_add_rule(2) call prefix used to create a domain.
LANDLOCK_MAX_NUM_RULES bounds distinct stored rules, not successful
calls: re-adding already-present rights for an object or port succeeds
without increasing num_rules.  Because every successful call increments
the version, these calls can wrap the 32-bit counter and give different
prefixes the same trace identity.

Widen the counter and its trace fields to 64 bits so the counter cannot
wrap in practice, while preserving the successful-call semantics.
Saturating would alias all subsequent histories, while rejecting a call
at the limit would change otherwise valid syscall behavior solely for
trace metadata.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Fixes: 63747c9477 ("landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints")
Reviewed-by: Günther Noack <gnoack@google.com>
Link: https://patch.msgid.link/20260922132615.1025945-1-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-22 20:52:13 +02:00
Mickaël Salaün
3fa5aa398e
landlock: Fix tracepoint contract documentation
The tracepoint documentation claims that denial and lifecycle events
expose every input needed to reproduce a verdict. Instead document how
denial, ruleset, and domain events identify the denying policy, checked
operation and object, and reason for denial. Direct consumers to generic
tracepoints for additional operational context.

State the reconstruction limits: IDs are boot-local, rule checks have no
request ID, and exported records may be lost or cross-CPU reordered.

Also replace the incorrect BPF_RAW_TRACEPOINT guidance with libbpf
SEC("tp_btf/...") attachment and refer consumers to the event prototypes
for callback argument layouts.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Link: https://patch.msgid.link/20260918185036.608651-10-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 14:13:33 +02:00
Mickaël Salaün
0889db596a
landlock: Report the effective signal number
The signal-scope denial callback identifies its target but not the
effective signal. This loses permission-probe signal zero and makes the
file-owner hook's zero sentinel ambiguous.

Append an int signal argument to the typed-BPF callback. Preserve sig,
including zero, in hook_task_kill(). In hook_file_send_sigiotask(),
translate signum zero to SIGIO at the producer, where its meaning is
known.

Carry the effective signal and target domain ID in a private,
stack-backed context consumed synchronously. This requires no allocation
or task reference in the interrupt-capable file-owner path. Gate this
context and the remaining scope-only domain IDs with CONFIG_TRACEPOINTS.

Keep the tracefs record and audit output unchanged.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Fixes: bb91730f16 ("landlock: Add tracepoints for ptrace and scope denials")
Link: https://patch.msgid.link/20260918185036.608651-7-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:05 +02:00
Mickaël Salaün
7ad69ac633
landlock: Report the actual ptrace tracer
The ptrace denial callback identifies only the tracee. Current is the
tracer during hook_ptrace_access_check(), but it is the tracee during
PTRACE_TRACEME, where the parent is the actual tracer. A consumer
therefore cannot infer both parties from the existing arguments.

Append the actual tracer task to the typed-BPF callback: current for
hook_ptrace_access_check() and parent for hook_ptrace_traceme(). Carry
it with the tracee domain ID in a private ptrace context. Both hooks
keep the selected tasks alive through synchronous dispatch, so no extra
task reference is needed.

Keep the tracefs record unchanged. The new context is available only to
typed BPF, while same_exec continues to describe the tracer that owns
the denying policy.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Fixes: bb91730f16 ("landlock: Add tracepoints for ptrace and scope denials")
Link: https://patch.msgid.link/20260918185036.608651-6-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:03 +02:00
Mickaël Salaün
98b04ab00f
landlock: Fix network denial trace context
Network denial events report source and destination ports reconstructed
from audit data. Their zero values are ambiguous, and neither identifies
the complete endpoint that Landlock checked.

Carry the checked sockaddr and its signed length in a private trace-only
context. For an enabled event, validate the length and copy only the
initialized prefix into zeroed local storage. This prevents a typed BPF
program from reading uninitialized bytes while exposing the socket
family, socket, address, and length.

Replace the source and destination trace-record fields with one signed
port derived from the checked address. A value of -1 means that no port
was checked, zero is a valid port, and positive values use host
endianness. Bind blockers select the bind address; connect and send
blockers select the destination.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Fixes: 01ce260f5c ("landlock: Add landlock_deny_access_fs and landlock_deny_access_net")
Link: https://patch.msgid.link/20260918185036.608651-5-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:02 +02:00
Mickaël Salaün
1a985d3890
landlock: Fix rule tracepoint context
Name each event after the identity it reports. Add-rule events describe
UAPI rule insertion, so rename them after LANDLOCK_RULE_PATH_BENEATH and
LANDLOCK_RULE_NET_PORT. Check-rule events describe matches in internal
rule trees, so rename them after LANDLOCK_KEY_INODE and
LANDLOCK_KEY_NET_PORT. This remains accurate if multiple UAPI rule types
share one lookup and stored rule. Keep denial event names based on
filesystem and network families because they describe final access
decisions.

Use u64 for growable access masks passed by value to add-rule and
check-rule typed BTF callbacks. CO-RE can relocate pointer-reached
fields, but it cannot widen a scalar callback slot declared by a BPF
program. Keep native access_mask_t for internal state and trace records.

For add-rule callbacks, report the normalized per-call contribution
passed to landlock_insert_rule() and expose the complete validated flags
value. Put the ruleset and flags first as a common invocation prefix.
This distinguishes duplicate and effective-zero additions without
recovering arguments from saved syscall registers.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Fixes: 63747c9477 ("landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints")
Fixes: 3f1f106e4c ("landlock: Add tracepoints for rule checking")
Link: https://patch.msgid.link/20260918185036.608651-4-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:02 +02:00
Mickaël Salaün
0de33ca344
landlock: Fix filesystem denial blocker reporting
Filesystem topology denials are rendered with an empty blockers value
because their blocker is identified by the request type instead of an
access mask.

Introduce the private struct landlock_blockers to carry the request type
and final missing access mask to filesystem and network denial
tracepoints. Copy both members into named trace-record fields, then use
the type to print change_topology for topology denials while preserving
symbolic access masks for ordinary denials.

The request type lets typed BPF consumers distinguish topology denials
from access denials. Keeping the native access mask in a pointer-reached
field also lets CO-RE adjust existing programs' load width if
access_mask_t grows.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Fixes: 01ce260f5c ("landlock: Add landlock_deny_access_fs and landlock_deny_access_net")
Link: https://patch.msgid.link/20260918185036.608651-3-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:01 +02:00
Mickaël Salaün
f71ecaece4
landlock: Fix tracepoint fixed-width type names
The new Landlock tracepoints use UAPI-prefixed __u32 and __u64 names for
callback arguments and record fields, including internal IDs that are
not Landlock UAPI values. Typed BPF consumers see callback typedef names
through BTF.

Use the kernel u32 and u64 aliases before release so the tracepoint
contract does not present internal values as Landlock UAPI types. This
changes BTF-visible typedef spelling but not integer widths, calling
conventions, tracefs formats, or record layouts.

Cc: Günther Noack <gnoack@google.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Link: https://patch.msgid.link/20260918185036.608651-2-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-20 11:08:00 +02:00
Linus Torvalds
4aec9ad1c6 dma-mapping fixes for Linux 7.3
A few fixes for the DMA-mapping code:
 - resolved regression in accessing encrypted memory by IOMMU-backed
 devices (Aneesh Kumar K.V),
 - improved failure handling and removed rare bug in swiotlb/highmem
 (Donggeun Yoo).
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQSrngzkoBtlA8uaaJ+Jp1EFxbsSRAUCaquwMwAKCRCJp1EFxbsS
 RMT6AP0elpdaZXNY0KwUBTwU95H604J+donqriepHABIBhIDEQD9GWZqNf/m1gEI
 tR5lHQ3+NGs0Q7Vd2ed1vSe82HQSsgU=
 =QxUC
 -----END PGP SIGNATURE-----

Merge tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux

Pull dma-mapping fixes from Marek Szyprowski:
 "A few fixes for the DMA-mapping code:

   - resolved regression in accessing encrypted memory by IOMMU-backed
     devices (Aneesh Kumar K.V)

   - improved failure handling and removed rare bug in swiotlb/highmem
     (Donggeun Yoo)"

* tag 'dma-mapping-7.3-2026-09-17' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux:
  x86/mm: Don't force unencrypted DMA for IOMMU-backed devices
  dma-mapping: don't trace the DMA address when the allocation fails
  swiotlb: use the adjusted address for the highmem page lookup
  dma-coherent: report a failed reserved memory assignment
2026-09-17 08:03:37 -07:00
Linus Torvalds
22098763a1 tracing fixes for 7.3:
- Don't destroy user event fields when removal fails
 
   User event fields are destroyed before the event is removed from
   visibility. But that can fail leaving the still visible event with no
   fields. Move the destroying of the fields to after the event is
   successfully removed from visibility.
 
 - Initialize function graph state is fork before calling copy_exec_state()
 
   For non-CLONE_VM forks, copy_exec_state() allocates a new task_exec_state.
   If that allocation fails, ftrace_graph_exit_task() will free the tasks
   ret_stack pointer. Since that pointer is still using the parent's
   ret_stack, it mistakenly frees the parent's pointer too.
 
   Call ftrace_graph_init() on the task first which will NULL out the new
   tasks's ret_stack and if the copy fails, it will not free anything.
 
 - Remove FGRAPH_MAX_INDEX
 
   The macro FGRAPH_MAX_INDEX was added but never used. Remove it.
 
 - Save ent_size in function graph printing of nested functions
 
   The function graph tracer needs to look at the next event to see if the
   next event is the return of the current function entry. If it is, it
   prints a single line:
 
     ktime_get();
 
   Otherwise it prints it like a nested function:
 
     tick_nohz_irq_exit() {
       ktime_get();
       kcpustat_irq_exit();
     }
 
   In order to look at the next event, it must save the current event so that
   it has the information to print from it. It saves the event in the
   iterator descriptor called "ent". What it doesn't save is the ent_size of
   the event which is now used to know if the function graph arguments are to
   be printed. The peek doesn't save the size so the size used happens to be
   that of the size of the last event that was seen.
 
   Save the entry event size in the iterator descriptor so that the correct
   size is used.
 
 - Fix several errors with freeing data in the histogram code
 
   The histogram code had a lot of leaked or or incorrect accounting when
   failures happen. Correct them.
 
 - Fix histogram regression of .percent and .graph modifiers
 
   Up until 6.3 histogram values could have "percent" or "graph" modifiers
   that changed how they were printed. But a change that added restricting
   histograms values from being strings, stack traces and other modifiers
   inadvertently prevented them from using the percent and graph modifiers,
   which were legal use cases for values.
 
   Put back the percent and graph modifiers.
 
 - Fix various typos in the comments
 
 - Set the trace_clock before initializing a histogram with clock argument
 
   The histogram API allows the user to specific which trace clock to use via
   a "clock=" string. The histogram is set up first before the clock is
   checked. If the passed in clock is not valid, it exits without fully
   fixing up the histogram leaving it on the list and a use-after-free can
   trigger.
 
   Update the clock argument first and if it fails then exit gracefully
   before the histogram trigger is placed on any lists.
 
 - Restore :mod: trailer after parsing in ftrace_set_clr_event
 
   The function ftrace_set_clr_event() modifies the parse string and needs to
   put it back to what was passed in. It searches for ":mod:" via a strsep()
   but fails to put back the first ':' in the string.
 
   Add back the ':' in the passed in string.
 
 - Take trace_array reference when opening a tracer options file
 
   The options files are dynamically created and some tracers add their own
   options. When a tracer adds their own list of options, the trace_array
   holding them has an array to hold the list of options for each tracer.
   This array increases in size via a krealloc(), and the new entry gets a
   newly allocated array to hold the options of the new tracer being added.
 
   The element in each entry of the tracer's option array holds a pointer
   back to the trace_array, a pointer to the tracer it is associated to, a
   pointer to the flags of the option.
 
   The issue is that these arrays are freed when the trace_array is freed
   when its instance it represents is removed from the instances directory.
   There's a race that an open of one of these options files can happen when
   the instance is being removed.
 
   Add a new helper function to be called by the open function of the options
   file to iterate all existing trace_arrays under a lock and find the one
   that has the given option element in one of it's tracer arrays. If found,
   then update the associated trace_array's reference counter to keep it from
   being freed. If not found, have the open call return -ENODEV.
 
 - Disable interrupts when acquiring the lock in rb_wake_up_waiters()
 
   The function rb_wake_up_waiters() assumes it will be called in interrupt
   context and does not disable irqs when taking cpu_buffer->reader_lock,
   which can be called in hard interrupt context. The issue is in PREEMPT_RT,
   this function is called in thread context leaving this lock open to a
   deadlock.
 
   Take the lock with interrupts disabled.
 
 - Use rcu_assign_pointer() for tmp_ops filter hash
 
   The tmp_ops used in update_ftrace_direct_mod() assigns its filter_hash
   field directly, but that field is annotated as __rcu and sparse complains.
   Assign it with rcu_assign_pointer()
 
 - Fix use-after-free in enable_trigger_private_data_free()
 
   The trace_event_call is accessed through the event_trigger_data's
   trace_event_file pointer to put the trace_event_call on freeing. The issue
   is that the trace_event_file data may have been freed already causing a
   use-after-free. Add a field to the event_trigger_data that points directly
   to the trace_event_call so that it can decrement its reference directly
   without needing to go through the trace_event_file.
 
 - Fix accounting of buffer data remote headers
 
   trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount the
   number of pages is needed for the asked for size as it doesn't take into
   account the meta data on each page. Add a helper function to do the
   calculation properly and use that in these functions.
 
 - Catch nr_page_va overflow in ring_buffer_desc sizing
 
   The number of pages per remote ring buffer is capped by
   ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to
   overflow that field would silently allocate a descriptor smaller than what
   was asked for.
 
 - Do not resize the subbuf order if any per_cpu buffer is disabled
 
   The mmapping of ring buffers disables resizing the subbuffers, but it is
   done per-cpu whereas the subbuf size change is done for all the per_cpu
   buffers under the buffer->mutex. It could change the size of some while
   the mapping is happening on others. Have the resize of the subbuf order
   check all the per_cpu buffers under the lock to see if any of them is
   disabled before starting and causing an inconsistency between buffers that
   are being mapped.
 -----BEGIN PGP SIGNATURE-----
 
 iIoEABYKADIWIQRRSw7ePDh/lE+zeZMp5XQQmuv6qgUCaqbdrBQccm9zdGVkdEBn
 b29kbWlzLm9yZwAKCRAp5XQQmuv6qro9AQDF/j3VW3Uu98lVFI9AB10XYhLDd5nt
 Zpf+3RviNgFpxgEAiE2+4K+4sM2SfaDDh9JMww9MKg1exL+cemE3a+JbBgY=
 =jgYE
 -----END PGP SIGNATURE-----

Merge tag 'trace-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace

Pull tracing fixes from Steven Rostedt:

 - Don't destroy user event fields when removal fails

   User event fields are destroyed before the event is removed from
   visibility. But that can fail leaving the still visible event with no
   fields. Move the destroying of the fields to after the event is
   successfully removed from visibility.

 - Initialize function graph state is fork before calling
   copy_exec_state()

   For non-CLONE_VM forks, copy_exec_state() allocates a new
   task_exec_state. If that allocation fails, ftrace_graph_exit_task()
   will free the tasks ret_stack pointer. Since that pointer is still
   using the parent's ret_stack, it mistakenly frees the parent's
   pointer too.

   Call ftrace_graph_init() on the task first which will NULL out the
   new tasks's ret_stack and if the copy fails, it will not free
   anything.

 - Remove FGRAPH_MAX_INDEX

   The macro FGRAPH_MAX_INDEX was added but never used. Remove it.

 - Save ent_size in function graph printing of nested functions

   The function graph tracer needs to look at the next event to see if
   the next event is the return of the current function entry. If it is,
   it prints a single line:

	ktime_get();

   Otherwise it prints it like a nested function:

	tick_nohz_irq_exit() {
	    ktime_get();
	    kcpustat_irq_exit();
	}

   In order to look at the next event, it must save the current event so
   that it has the information to print from it. It saves the event in
   the iterator descriptor called "ent". What it doesn't save is the
   ent_size of the event which is now used to know if the function graph
   arguments are to be printed. The peek doesn't save the size so the
   size used happens to be that of the size of the last event that was
   seen.

   Save the entry event size in the iterator descriptor so that the
   correct size is used.

 - Fix several errors with freeing data in the histogram code

   The histogram code had a lot of leaked or or incorrect accounting
   when failures happen. Correct them.

 - Fix histogram regression of .percent and .graph modifiers

   Up until 6.3 histogram values could have "percent" or "graph"
   modifiers that changed how they were printed. But a change that added
   restricting histograms values from being strings, stack traces and
   other modifiers inadvertently prevented them from using the percent
   and graph modifiers, which were legal use cases for values.

   Put back the percent and graph modifiers.

 - Fix various typos in the comments

 - Set the trace_clock before initializing a histogram with clock
   argument

   The histogram API allows the user to specific which trace clock to
   use via a "clock=" string. The histogram is set up first before the
   clock is checked. If the passed in clock is not valid, it exits
   without fully fixing up the histogram leaving it on the list and a
   use-after-free can trigger.

   Update the clock argument first and if it fails then exit gracefully
   before the histogram trigger is placed on any lists.

 - Restore :mod: trailer after parsing in ftrace_set_clr_event

   The function ftrace_set_clr_event() modifies the parse string and
   needs to put it back to what was passed in. It searches for ":mod:"
   via a strsep() but fails to put back the first ':' in the string.

   Add back the ':' in the passed in string.

 - Take trace_array reference when opening a tracer options file

   The options files are dynamically created and some tracers add their
   own options. When a tracer adds their own list of options, the
   trace_array holding them has an array to hold the list of options for
   each tracer. This array increases in size via a krealloc(), and the
   new entry gets a newly allocated array to hold the options of the new
   tracer being added.

   The element in each entry of the tracer's option array holds a
   pointer back to the trace_array, a pointer to the tracer it is
   associated to, a pointer to the flags of the option.

   The issue is that these arrays are freed when the trace_array is
   freed when its instance it represents is removed from the instances
   directory. There's a race that an open of one of these options files
   can happen when the instance is being removed.

   Add a new helper function to be called by the open function of the
   options file to iterate all existing trace_arrays under a lock and
   find the one that has the given option element in one of it's tracer
   arrays. If found, then update the associated trace_array's reference
   counter to keep it from being freed. If not found, have the open call
   return -ENODEV.

 - Disable interrupts when acquiring the lock in rb_wake_up_waiters()

   The function rb_wake_up_waiters() assumes it will be called in
   interrupt context and does not disable irqs when taking
   cpu_buffer->reader_lock, which can be called in hard interrupt
   context. The issue is in PREEMPT_RT, this function is called in
   thread context leaving this lock open to a deadlock.

   Take the lock with interrupts disabled.

 - Use rcu_assign_pointer() for tmp_ops filter hash

   The tmp_ops used in update_ftrace_direct_mod() assigns its
   filter_hash field directly, but that field is annotated as __rcu and
   sparse complains. Assign it with rcu_assign_pointer()

 - Fix use-after-free in enable_trigger_private_data_free()

   The trace_event_call is accessed through the event_trigger_data's
   trace_event_file pointer to put the trace_event_call on freeing. The
   issue is that the trace_event_file data may have been freed already
   causing a use-after-free. Add a field to the event_trigger_data that
   points directly to the trace_event_call so that it can decrement its
   reference directly without needing to go through the
   trace_event_file.

 - Fix accounting of buffer data remote headers

   trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount
   the number of pages is needed for the asked for size as it doesn't
   take into account the meta data on each page. Add a helper function
   to do the calculation properly and use that in these functions.

 - Catch nr_page_va overflow in ring_buffer_desc sizing

   The number of pages per remote ring buffer is capped by
   ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to
   overflow that field would silently allocate a descriptor smaller than
   what was asked for.

 - Do not resize the subbuf order if any per_cpu buffer is disabled

   The mmapping of ring buffers disables resizing the subbuffers, but it
   is done per-cpu whereas the subbuf size change is done for all the
   per_cpu buffers under the buffer->mutex. It could change the size of
   some while the mapping is happening on others. Have the resize of the
   subbuf order check all the per_cpu buffers under the lock to see if
   any of them is disabled before starting and causing an inconsistency
   between buffers that are being mapped.

* tag 'trace-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: (25 commits)
  ring-buffer: Check resize_disabled before publishing the new subbuf order
  tracing/remotes: Catch nr_page_va overflow in ring_buffer_desc sizing
  tracing/remotes: Account for ring buffer page header in size calculation
  tracing: Don't dereference trace_event_file in deferred trigger free
  ftrace: Use rcu_assign_pointer() for tmp_ops filter hash
  ring-buffer: Acquire the lock with irqsave in rb_wake_up_waiters()
  tracing: Take trace_array reference when opening a tracer options file
  tracing: Fix ring_buffer_read_page_size() kernel-doc
  tracing: Restore :mod: trailer after parsing in ftrace_set_clr_event()
  tracing: Fix memory corruption from a "STACKTRACE" histogram key
  tracing: Fix memory corruption from the histogram stacktrace modifier
  tracing: Undo the registration when enabling the histogram trigger fails
  tracing: Take the reference before publishing the named histogram trigger
  tracing: Set the trace clock before registering the histogram trigger
  tracing: Fix typo "preceeded" in comment
  tracing: Fix typo "availabe" in comment
  tracing: Let histogram values keep the percent and graph modifiers
  tracing: Keep the entry count when the histogram stats allocation fails
  tracing: Free histogram the field rejected for a bad modifier
  tracing: Free histogram the var ref when its initialization fails
  ...
2026-09-13 12:27:00 -07:00
Hemanth Selam
0999d3e16d tracing: Fix typo "preceeded" in comment
Correct "preceeded" to "Preceded", reported by scripts/checkpatch.pl using
the misspelling list in scripts/spelling.txt.  Only touches comments, no
code changes.

Link: https://patch.msgid.link/20260907065607.36615-1-hemanth.selam@gmail.com
Assisted-by: Cursor:claude-opus-5
Signed-off-by: Hemanth Selam <hemanth.selam@gmail.com>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2026-09-11 13:54:32 -04:00
Linus Torvalds
50d05c7c76 Landlock fix for v7.3-rc3
-----BEGIN PGP SIGNATURE-----
 
 iIYEABYKAC4WIQSVyBthFV4iTW/VU1/l49DojIL20gUCaqF0GxAcbWljQGRpZ2lr
 b2QubmV0AAoJEOXj0OiMgvbSs+YBALj3Ttl+T8cnEmxExfOYnPt6eL+oIsZFo6HU
 zSXUqyiNAQDxtpucp/JgwBNbuk0XA+BfLSVWuw94jdqbPpCrjUW0BA==
 =egGv
 -----END PGP SIGNATURE-----

Merge tag 'landlock-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux

Pull Landlock fixes from Mickaël Salaün:
 "This fixes a use-after-free and a lockdep assert NULL dereferencing,
  and properly truncates too-long strings printed by a Landlock
  tracepoint. Most of the changes are brought by new tests"

* tag 'landlock-7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux:
  landlock: Test trace path output boundaries
  landlock: Bound escaped trace path output
  landlock: Clean up ruleset validation checks
  selftests/landlock: Test abstract socket trace name limits
  landlock: Fix use-after-free of the source's parent directory
2026-09-09 11:00:35 -07:00
Linus Torvalds
5e1287972b vfs-7.3-rc3.fixes
Please consider pulling these changes from the signed vfs-7.3-rc3.fixes tag.
 
 Thanks!
 Christian
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQRAhzRXHqcMeLMyaSiRxhvAZXjcogUCaqFWOQAKCRCRxhvAZXjc
 otLlAP9X02ybdUt9NndBK8LjslDWwB9hOXzPgYsOKYODEqODjQD/aLpbXVEsA1yy
 SLdSDtbtpf+01z4KHorvAakBzk/jrw4=
 =OzBX
 -----END PGP SIGNATURE-----

Merge tag 'vfs-7.3-rc3.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs

Pull vfs fixes from Christian Brauner:

 - netfs:

     - Fix an uninitialized return value in netfs_unbuffered_write()
       when preparing the first subrequest fails

     - For partial unbuffered/DIO writes return the amount transferred
       rather than an error

     - Update i_size with the amount actually written when a partial
       transfer ends in an error

     - Fix a subrequest reference leak when the io_iter ends up empty

     - Handle netfs_alloc_subrequest() failure during unbuffered writes

     - Load all readahead folios into the rolling buffer upfront and
       drop the readahead references once the first subrequest is
       dispatched

     - Mark folios for copy-to-cache while issuing subrequests

     - Fix read progress reporting

 - afs:

     - Add the missing kunmap in the error path of afs_dir_search_bucket()

     - Fix a double kunmap in afs_edit_dir_remove()

     - Don't free an existing server's endpoint state when cleaning up a
       candidate server in afs_lookup_server()

     - Unbind peers removed from a server's address list

 - ufs:

     - Load the cylinder group metadata before creating the root dentry

     - Validate the cylinder group index and rotor positions before
       caching them

     - Treat an unreadable directory block as not empty

 - exec:

     - Close the close-on-exec files before taking exec_update_lock

       Closing a file can block on the filesystem, so a hung filesystem
       blocked everything that takes exec_update_lock and a FUSE server
       inspecting the calling process could deadlock

     - Drop the bprm loader before closing bprm->file in free_bprm()

 - exit: Hold a reference to thread_pid across proc_flush_pid()

 - reboot: Fix a use-after-free on cad_pid

 - nsfs: Keep the namespace tree fields out of the rcu_head used by
   kfree_rcu()

 - nstree: Check listing permission before taking a namespace
   reference in listns()

 - super: Return 0 when a nested thaw drops its hold while other
   freezers remain

 - ext4: Don't set I_METADATA_WRITEBACK during fastcommit replay

 - adfs: Free s_fs_info in ->kill_sb()

 - autofs: Free the inode info allocated in autofs_fill_super() when
   the root inode allocation fails

 - ovl: Return EINVAL instead of EIO on a user namespace mismatch now
   that it's a plain refusal and not an internal error

 - cachefiles: Don't cast the variable-length coherency data to a
   __be64 in the coherency tracepoint

* tag 'vfs-7.3-rc3.fixes' of git://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs: (28 commits)
  nstree: check listing permission before taking a namespace reference
  exec: do_close_on_exec() before taking exec_update_lock
  exit: hold a reference to thread_pid across proc_flush_pid
  fs: autofs: fix memory leak in autofs_fill_super()
  exec: Drop bprm loader before closing bprm->file
  afs: Clear stale peer app data after address list changes
  afs: Fix incorrect free in candidate cleanup in afs_lookup_server()
  afs: Fix double-unmap of directory block
  afs: Fix missing kunmap in afs_dir_search_bucket()
  ovl: return EINVAL instead of EIO in case of mismatched user_ns
  reboot: fix cad_pid use-after-free race
  cachefiles: Fix potential UAF/KASAN warning
  netfs: Fix read progress reporting
  netfs: Mark folios with COPY_TO_CACHE whilst issuing subreqs
  netfs: Fix readahead synchronisation issues by loading all folios upfront
  netfs: break unbuffered write when netfs_alloc_subrequest() fails
  netfs: Fix subreq ref leak
  netfs: Fix i_size update for partial transfer
  netfs: Fix error vs transferred passed to ->ki_complete()
  netfs: Fix unbuffered/DIO write partial transfer error return
  ...
2026-09-09 09:38:03 -07:00
Donggeun Yoo
92c6a8d647 dma-mapping: don't trace the DMA address when the allocation fails
dma_alloc_attrs() passes *dma_handle to trace_dma_alloc() without
checking whether the allocation succeeded. No backend writes it on
failure: dma_direct_alloc(), iommu_dma_alloc() and the dma_map_ops
instances assign it only on the path that returns a buffer. Callers
usually pass an uninitialized automatic variable, so a failed allocation
records whatever the stack held, next to the virt_addr=(null) that marks
the record as an error:

 dma_alloc: dmatrace dir=BIDIRECTIONAL dma_addr=deadbeefdeadbeef
 size=1099511627776 virt_addr=0000000000000000

The device coherent pool path reaches the same call: a non-zero return
from dma_alloc_from_dev_coherent() means the request was handled, not
that it succeeded, so cpu_addr is NULL and dma_handle is untouched once
the pool runs out.

For an allocation event a NULL virt_addr already means the request
failed, so the address field carries nothing. Report 0 for it in the
event class rather than at each call site, which covers dma_alloc_pages()
and dma_alloc_sgt_err() as well.

Fixes: 038eb433dc ("dma-mapping: add tracing for dma-mapping API calls")
Fixes: 68b6dbf1f4 ("dma-mapping: trace more error paths")
Suggested-by: Marek Szyprowski <m.szyprowski@samsung.com>
Signed-off-by: Donggeun Yoo <donggeunyoo.kernel@gmail.com>
Link: https://lore.kernel.org/r/20260907120124.603373-1-donggeunyoo.kernel@gmail.com
Reviewed-by: Sean Anderson <sean.anderson@linux.dev>
Signed-off-by: Marek Szyprowski <m.szyprowski@samsung.com>
2026-09-09 12:02:33 +02:00
Mickaël Salaün
3125751cd1
landlock: Bound escaped trace path output
Filesystem paths may expand fourfold when trace text escapes spaces and
other untrusted bytes.  A sufficiently long representation can exhaust
the shared scratch sequence.  A sibling __print_flags() helper may then
return an unterminated one-past pointer because TP_printk() argument
ordering is unspecified.

Use a fixed budget rather than the scratch space available at call time,
so output does not vary with sibling evaluation order.  Limit an
untrusted string to three quarters of the trace sequence, leaving the
rest for sibling helpers and final event metadata.  Compute and commit
complete escaped output transactionally so an exact fill cannot consume
the terminating NUL or poison the scratch sequence.

For strings that exceed the limit, retain the largest prefix ending at a
complete escape unit, then append a raw UTF-8 ellipsis.  Keep the
helper's existing octal fallback so complete values remain unchanged.
Hex fallback would consume the same four bytes per escaped byte without
increasing the prefix or strengthening the marker.  ESCAPE_NAP renders
every non-ASCII input byte in octal, so legitimate data cannot reproduce
the marker without being escaped.

Cc: Günther Noack <gnoack@google.com>
Link: https://patch.msgid.link/20260907154401.124362-1-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-09-08 11:49:37 +02:00
Linus Torvalds
adf50c47a4 Including fixes from bluetooth.
Previous releases - regressions:
 
   - page_pool: keep frag_offset aligned for odd-sized requests
 
   - sched: fix u32 duplicate handle when node ID pool is exhausted
 
   - udp: create exceptions before socket matching
 
   - igmp: convert struct ip_sf_list to RCU
 
   - ip6_gre: check tunnel info before xmit in ip6gre_tunnel_xmit
 
   - rds: acquire the fastpath locks in rds_conn_shutdown()
 
   - tipc:
     - protect node reset trace dump with node lock
     - fix NULL deref in tipc_named_node_up() on empty publication list
 
   - bluetooth:
       L2CAP: fix out-of-bounds write in l2cap_ecred_connect
       hci_core: fix race condition during device registration
 
   - eth: mlx5e: prevent stale XSK buffer release on refill retries
 
   - eth: bridge: don't truncate the port group walk on teardown
 
 Previous releases - always broken:
 
   - gro: fix nesting of TCP GSO SKBs in skb_gro_receive_list()
 
   - sched: fix skb sizing and action leak on reoffload delete
 
   - tcp: fix use-after-free in do_tcp_getsockopt()
 
   - af_packet: don't cast tpacket_hdr.tp_len to int in tpacket_parse_header().
 
   - sctp: fix soft lockup from unpadded ASCONF-ACK parameter iteration
 
   - iptunnel: fix stale transport header during tunnel decapsulation
 
   - eth: vxlan: fix use-after-free in vxlan_mdb_remote_src_del()
 
   - eth: bonding: fix uninitialized transport header access in alb_determine_nd()
 
 Signed-off-by: Paolo Abeni <pabeni@redhat.com>
 -----BEGIN PGP SIGNATURE-----
 
 iQJGBAABCgAwFiEEg1AjqC77wbdLX2LbKSR5jcyPE6QFAmqZp4kSHHBhYmVuaUBy
 ZWRoYXQuY29tAAoJECkkeY3MjxOkEjkQALMGg903vZ4TfGlzKayFzyhcd5ZC8G4F
 R4M+UgTGRfNuas/1YwjpyOpvOYyFgGZ9xBmYNFdsW0YCzZwu8PxpXgu6WTZ+F4gu
 sDtWoAbN5V6CfY3fdC7IbXTp8t4CX+shQAsVvEp39Y8SJF4AZeMn8N0+Lnu4DlD3
 DAPo/lYSSfvv7RK/5Jvr9FWo7vyoEylfG+LekzGASmWGwhC3h7kWGB4RB4PhJmyq
 vRIj2ZjnzdDxu4N7ZGh+EEu5SBCcLP0e/dIMCDg++HAghDqPJ+7pzbWC1kFtQ0ss
 qOSyws/xMW3D0Rb68tkiikYWRwgvXUsfEL7Jdf2lhC1xDI8ZpxrrnPYYSZS4rsjb
 hjeBwtzRZhv5R0PnNlaZyNpFICIW3XwqP0bYqH/Z/CgwwkKYd+Rp6Tm7hkLpafNx
 Py607x2Ff/L2Aydp8csJEyqFP33QOHAfeW+X/YCo4jTc0zBTMSsOblG/EsPPBdNX
 fvhVkx4NdqAvIdLYm65cdvhe5dZtIhOAhwAMcrGiMIia4vCIsXq3fWbAe6Phtu8T
 KHpQ7Esg/if6blNPpflBuVPsoU+5N6mL7a+zsurqvilJBC5DgYaxwbMlgRSUmJzg
 2rR6IlrOsb0FFITNUv+xgPEGzMGV1rL4wnaMSmVN0tjmFQvcu24XQ17D+ug27+eG
 rD+E2lrX5rNn
 =0sv8
 -----END PGP SIGNATURE-----

Merge tag 'net-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net

Pull networking fixes from Paolo Abeni:
 "Including fixes from bluetooth.

  Previous releases - regressions:

    - page_pool: keep frag_offset aligned for odd-sized requests

    - sched: fix u32 duplicate handle when node ID pool is exhausted

    - udp: create exceptions before socket matching

    - igmp: convert struct ip_sf_list to RCU

    - ip6_gre: check tunnel info before xmit in ip6gre_tunnel_xmit

    - rds: acquire the fastpath locks in rds_conn_shutdown()

    - tipc:
        - protect node reset trace dump with node lock
        - fix NULL deref in tipc_named_node_up() on empty publication
          list

    - bluetooth:
        - L2CAP: fix out-of-bounds write in l2cap_ecred_connect
        - hci_core: fix race condition during device registration

    - eth:
        - mlx5e: prevent stale XSK buffer release on refill retries
        - bridge: don't truncate the port group walk on teardown

  Previous releases - always broken:

    - gro: fix nesting of TCP GSO SKBs in skb_gro_receive_list()

    - sched: fix skb sizing and action leak on reoffload delete

    - tcp: fix use-after-free in do_tcp_getsockopt()

    - af_packet: don't cast tpacket_hdr.tp_len to int in
      tpacket_parse_header()

    - sctp: fix soft lockup from unpadded ASCONF-ACK parameter iteration

    - iptunnel: fix stale transport header during tunnel decapsulation

    - eth:
        - vxlan: fix use-after-free in vxlan_mdb_remote_src_del()
        - bonding: fix uninitialized transport header access in
          alb_determine_nd()"

* tag 'net-7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net: (83 commits)
  net: gro: Fix nesting of TCP GSO SKBs in skb_gro_receive_list()
  net: stmmac: reconfigure RX packet parser table in stmmac_hw_setup() after reset
  net: airoha: enable RX_DONE interrupt for RX queue 31
  net/rds: don't let rds_conn_shutdown() consume a concurrent drop
  net/rds: acquire the fastpath locks in rds_conn_shutdown()
  net/rds: acquire RDS_IN_XMIT in rds_tcp_reset_callbacks()
  net/rds: tcp: don't force RDS_CONN_RESETTING over a concurrent shutdown
  net/rds: clear cp_flags bits individually in rds_conn_path_reset()
  net/rds: use clear_bit_unlock() in release_refill()
  net/rds: use wq_has_sleeper() in release_in_xmit()
  net: usb: qmi_wwan: add Compal EXM-G1x support
  net: macb: exclude software FCS from TX byte statistics
  net: Remove conflicting altnames for dying netns in __dev_change_net_namespace().
  net: bridge: mcast: don't truncate the port group walk on teardown
  bonding: do not clear curr_active_slave prematurely when releasing all slaves
  net: qrtr: Send HELLO message on endpoint register
  octeontx2-af: Fix limiting SRIOV VF count logic
  bonding: alb: fix uninitialized transport header access in alb_determine_nd()
  s390/ctcm: Prevent XID null dereference
  net: psp: do not inherit the Rx association on clone
  ...
2026-09-03 10:18:12 -07:00
David Howells
a67632c8c2
cachefiles: Fix potential UAF/KASAN warning
Currently, trace_cachefiles_coherency() is being passed a pointer to a
__be64 lain over the coherency data in struct cachefiles_xattr so that it
can display the first 8 bytes.  However, the data is of variable length and
could even be 0 bytes.  This could lead to a UAF or KASAN warning.

Fix this by making sure the buffer has room for at least 8 bytes and that
those 8 bytes are pre-cleared.

Further, those bytes are not 8-byte aligned, so fix the tracepoint to
extract the data as four 2-byte words (they are 2-byte aligned) and
reassemble the __be64.  The compiler will convert this into a single 8-byte
load where the CPU supports it.

Fixes: 229105e5cf ("cachefiles: Add auxiliary data trace")
Link: https://sashiko.dev/#/patchset/20260810144746.574036-1-dhowells%40redhat.com
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/20260827134304.2075713-11-dhowells@redhat.com
Acked-by: Paulo Alcantara <pc@manguebit.org>
cc: Paulo Alcantara <pc@manguebit.org>
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-31 09:54:38 +02:00
David Howells
e00827a4d0
netfs: Fix read progress reporting
For really big read RPC ops that span multiple folios, netfslib allows the
filesystem to give progress notifications to wake up the collector thread
to do a collection of folios that have now been fetched, even if the RPC is
still ongoing, thereby allowing the application to make progress.

This works by taking the current rreq->cleaned_to value (which indicates
which folios have been unlocked) and adding the stashed size of the next
folio to it.  cleaned_to, however, is subject to 64-bit tearing on a 32-bit
arch.

Fix this by stashing the next progress notification point as a size_t
(which won't tear) to be added to rreq->start (which won't change), with
the collector thread calculating that from cleaned_to plus the next folio
size.

Further, however, if the folios are small, the collector thread gets
constantly woken up - which has a negative performance impact on the
system.

Fix that too by setting a minimum trigger of 256KiB or the size of the
folio at the front of the queue, whichever is larger.  Note that this has
an issue that different subreqs have different need-to-be-cached
properties; this is solved by a preceding patch that marks the property on
the folios whilst issuing subreqs rather than when collecting them.

Also, make sure rreq->cleaned_to is initialised up front, along with
rreq->collected_to and stream->collected_to.

Fixes: e2d46f2ec3 ("netfs: Change the read result collector to only use one work item")
Link: https://sashiko.dev/#/patchset/20260804100224.2748935-1-dhowells%40redhat.com
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/20260827134304.2075713-10-dhowells@redhat.com
Acked-by: Paulo Alcantara <pc@manguebit.org>
cc: Paulo Alcantara <pc@manguebit.org>
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-31 09:54:38 +02:00
David Howells
533203c418
netfs: Mark folios with COPY_TO_CACHE whilst issuing subreqs
Mark folios with NETFS_FOLIO_COPY_TO_CACHE whilst issuing subreqs rather than
when collecting them.  This means that the collector thread doesn't have to
try and keep track of which subreqs contribute to which folios - and thus
which folios will need to be copied to the cache because at least one byte
wasn't in the cache.  Instead, this is marked on the folios up front and the
collector need only consider the folios.

For PG_private_2-using filesystems, PG_private_2 is set instead of
NETFS_FOLIO_COPY_TO_CACHE, but otherwise it works the same.

The NETFS_RREQ_COPY_TO_CACHE is replaced with NETFS_RREQ_CANCEL_CACHING, which
is now set if caching fails somewhere, thereby causing the collection thread
to cancel the copy-to-cache marks on the remaining folios.

Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/20260827134304.2075713-9-dhowells@redhat.com
Acked-by: Paulo Alcantara <pc@manguebit.org>
cc: Paulo Alcantara (Red Hat) <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: netfs@lists.linux.dev
cc: linux-mm@kvack.org
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-31 09:54:38 +02:00
David Howells
fed0b33e6c
netfs: Fix readahead synchronisation issues by loading all folios upfront
There are some synchronisation issues that derive from the app thread
adding more folios to the rolling buffer whilst the collector thread is
looking at them or trying to clear them, such as determining the setting of
front_folio_order when the next folio hasn't been added yet,

The reason for the rolling buffer approach is that loading the buffer
upfront and then dropping all the refs just acquired is quite a slow
operation, and loading progressively allows some of the cost to be deferred
until after at least some of the I/O is started.

Instead, a better way is to load all the folios into the rolling buffer
upfront - and then drop the refs later, once the I/O is in progress.  (Even
better would be for the refs not to be there at all.)

Fix this by changing the rolling buffer loader to load all the folios
selected by the VM for readahead upfront into the folio queue.  The folio
queue is allocated a batch worth at a time as we don't know how many folios
are involved (the readahead_control struct, alas, has a page count, not a
folio count).

The folio refs acquired from readahead are then dropped in bulk once the
first subrequest is dispatched as it's quite a slow operation.  The
collector waits for NETFS_RREQ_NEED_PUT_RA_REFS to be cleared so that it
doesn't unlock folios before the xarray has been scanned for them.

This simplifies the buffer handling later and isn't noticeably slower as
the xarray doesn't need to be modified and the folios are all already
pre-locked.

Fixes: ee4cdf7ba8 ("netfs: Speed up buffered reading")
Link: https://sashiko.dev/#/patchset/20260824120224.504575-1-dhowells%40redhat.com
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/20260827134304.2075713-8-dhowells@redhat.com
Acked-by: Paulo Alcantara <pc@manguebit.org>
cc: Paulo Alcantara (Red Hat) <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: netfs@lists.linux.dev
cc: linux-mm@kvack.org
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-31 09:54:38 +02:00
Eric Dumazet
7fcc2fe39f net: icmp: avoid invalid transport header access in icmp_send tracepoint
syzbot reported a WARNING triggered by DEBUG_NET_WARN_ON_ONCE():

 WARNING: at skb_transport_header include/linux/skbuff.h:3087 [inline]
 WARNING: at udp_hdr include/linux/udp.h:23 [inline]
 WARNING: at do_trace_event_raw_event_icmp_send include/trace/events/icmp.h:30 [inline]
 WARNING: at trace_event_raw_event_icmp_send+0x48c/0x6ec include/trace/events/icmp.h:11
 Call trace:
  skb_transport_header include/linux/skbuff.h:3087 [inline]
  udp_hdr include/linux/udp.h:23 [inline]
  do_trace_event_raw_event_icmp_send include/trace/events/icmp.h:30 [inline]
  trace_event_raw_event_icmp_send+0x48c/0x6ec include/trace/events/icmp.h:11
  __traceiter_icmp_send include/trace/events/icmp.h:11 [inline]
  __do_trace_icmp_send include/trace/events/icmp.h:11 [inline]
  trace_icmp_send+0x320/0x49c include/trace/events/icmp.h:11
  __icmp_send+0xcfc/0x11d8 net/ipv4/icmp.c:1013
  ipv4_send_dest_unreach net/ipv4/route.c:1280 [inline]
  ipv4_link_failure+0x57c/0x8dc net/ipv4/route.c:1287
  dst_link_failure include/net/dst.h:438 [inline]
  vti_tunnel_xmit+0xe40/0x17a4 net/ipv4/ip_vti.c:307

TP_fast_assign() unconditionally calls udp_hdr(skb) before checking
whether the packet is UDP. Furthermore, __icmp_send() can be invoked
from paths (e.g., link failures, ARP errors, forwarding, AF_PACKET)
where skb->transport_header was never initialized (~0U).

Under CONFIG_DEBUG_NET=y, calling skb_transport_header(skb) triggers
DEBUG_NET_WARN_ON_ONCE(!skb_transport_header_was_set(skb)).

Fix this by:
1. Only parsing transport info when iph->protocol == IPPROTO_UDP.
2. Using skb_header_pointer() at skb_network_offset(skb) + (iph->ihl << 2)
   to safely fetch the UDP header without assuming transport_header is set.

Fixes: db3efdcf70 ("net/ipv4: add tracepoint for icmp_send")
Reported-by: syzbot+6d2762674103618994b0@syzkaller.appspotmail.com
Closes: https://lore.kernel.org/netdev/6a8d5538.91706f20.ef82.0009.GAE@google.com/T/#u
Signed-off-by: Eric Dumazet <edumazet@google.com>
Cc: Peilin He <he.peilin@zte.com.cn>
Cc: xu xin <xu.xin16@zte.com.cn>
Cc: Steven Rostedt <rostedt@goodmis.org>
Reviewed-by: Jiayuan Chen <jiayuan.chen@linux.dev>
Reviewed-by: David Ahern <dsahern@kernel.org>
Link: https://patch.msgid.link/20260825084551.1562967-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-28 14:58:16 -07:00
Linus Torvalds
115bd364ab f2fs-for-7.3-rc1
In this round, key enhancements focus on reducing inode management memory
 overhead, introducing resizable tail sections with unified pinned allocation,
 and boosting I/O throughput via parallel multi-device flushes and asynchronous
 f2fs_write_end_io() execution. We also add dynamic device alias reservations to
 allow on-the-fly space donation from user partitions.
 
 Alongside these features, critical bug fixes resolve folio race conditions,
 lingering dirty flags, dentry and block counter leaks, and potential deadloops
 in f2fs_fsync_node_pages(). Additional stability patches address error-path
 handling across symlink, sync, and rename/unlink operations, prevent pinned file
 fragmentation, and correct segment migration and free section accounting in
 free_segment_range.
 
 Enhancement:
  - reduce memory footprint of ino management
  - support dynamic reserve/release for device aliasing
  - issue multi-device flushes in parallel
  - add a way to run f2fs_write_end_io() asynchronously
  - support resizable tail section and unify pinned allocation
 
 Bug fix:
  - fix to pass folio->index to f2fs_sanity_check_node_footer()
  - fix folio_nr_pages() race after put in large folio invalidate
  - fix to clear dirty flag on folio in error path
  - accurately adjust free_sections during free_segment_range
  - fix to avoid potential deadloop in f2fs_fsync_node_pages()
  - fix the error path in symlink, device alias in rename/unlink,
    f2fs_sync_fs,
  - fix to migrate all curseg types during free_segment_range
  - fix to avoid pinfile fragment on fragment:{block, segment} mode
  - fix valid block count leak on data block allocation failure
  - fix dentry folio leak in find_in_level
  - reject overlapping move range after len expansion
  - fix some bugs related to file pinning, GC functions, i_size.
 
 And, the series includes a number of minor bug fixes.
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE00UqedjCtOrGVvQiQBSofoJIUNIFAmqPybMACgkQQBSofoJI
 UNKp2g/+OP6XZi56hNTqscnKyKrDdVJnOcS/YOe7d1BR+070qmPTpyFjLgjng05K
 exu55rz9vJ3DlpFLsjMEo60DRlDEc5rR4AqymMjqFJH9424ZlxPpdDn6ofCVT0Ck
 D6RTf3y1HFSi4x7//gPQofR9y4MlDrH2Q7NPDriipqbymuNXEjrx/vdr2nq/kHUu
 2lbf7QQs08qYiyDxQcxOFdCdxUTrsEW/tkYZiwgbU2nCJ/eG2R59amgtYJg3SlVt
 xdrf+IaSS7kE5+mGCoBm0WooPpB507kHaoQpZYDj2uueFvEw7nFcSfOCXapOvg8U
 wFkRR0F/rZ4+AW/u4n8Ye7N4a7WjWMTBwfR3WJ2j+arhfn87vJZK6LQQlYwq16l4
 tRcQFcCrKsXHh5HY2OGj8DzTd40zryXujH566YioBCAXU312My1yFjeTJzjobbUW
 TclkfMl689iTr9pqBhIjKT2tTvZFLROYSLk5UBNFNSfA4PtsAAhqlUrE1ck3AuTA
 8kIjLmppfZEBqVuZCF0T1z9TXk0Bg0eM8qHbl/8SdDavmf1pF1BupLqchfdZp6Jr
 4iV1weK4dzmCF6++YDsNlvnBGyTbhiFdN7C5cMlv89PTSTUQZ45Me6JmNwrRDqyU
 MSTyiSSpI+3rAMO0h58PHi+QAOtM0f+S7C24ySkBBerBCbMTK6A=
 =xQ77
 -----END PGP SIGNATURE-----

Merge tag 'f2fs-for-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs

Pull f2fs updates from Jaegeuk Kim:
 "In this round, key enhancements focus on reducing inode management
  memory overhead, introducing resizable tail sections with unified
  pinned allocation, and boosting I/O throughput via parallel
  multi-device flushes and asynchronous f2fs_write_end_io() execution.
  We also add dynamic device alias reservations to allow on-the-fly
  space donation from user partitions.

  Alongside these features, critical bug fixes resolve folio race
  conditions, lingering dirty flags, dentry and block counter leaks, and
  potential deadloops in f2fs_fsync_node_pages(). Additional stability
  patches address error-path handling across symlink, sync, and
  rename/unlink operations, prevent pinned file fragmentation, and
  correct segment migration and free section accounting in
  free_segment_range.

  Enhancements:
   - reduce memory footprint of ino management
   - support dynamic reserve/release for device aliasing
   - issue multi-device flushes in parallel
   - add a way to run f2fs_write_end_io() asynchronously
   - support resizable tail section and unify pinned allocation

  Bug fixes:
   - fix to pass folio->index to f2fs_sanity_check_node_footer()
   - fix folio_nr_pages() race after put in large folio invalidate
   - fix to clear dirty flag on folio in error path
   - accurately adjust free_sections during free_segment_range
   - fix to avoid potential deadloop in f2fs_fsync_node_pages()
   - fix the error path in symlink, device alias in rename/unlink,
     f2fs_sync_fs
   - fix to migrate all curseg types during free_segment_range
   - fix to avoid pinfile fragment on fragment:{block, segment} mode
   - fix valid block count leak on data block allocation failure
   - fix dentry folio leak in find_in_level
   - reject overlapping move range after len expansion
   - fix some bugs related to file pinning, GC functions, i_size

  And, the series includes a number of minor bug fixes"

* tag 'f2fs-for-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/jaegeuk/f2fs: (51 commits)
  f2fs: support resizable tail section and unify pinned allocation
  f2fs: don't leave the hashed inode while it's unlinked
  f2fs: accurately adjust free_sections during free_segment_range
  f2fs: fix to avoid potential deadloop in f2fs_fsync_node_pages()
  f2fs: use adjusted write range after f2fs_write_checks()
  f2fs: fix to propagate error from f2fs_sync_fs()
  f2fs: return symlink writeback errors
  f2fs: fix error handling on device alias check in rename and unlink
  f2fs: fix to reset all pinned status during fggc
  f2fs: use f2fs_{down, up}_(read, write}_trace() for nat_tree_lock
  f2fs: reduce memory footprint of ino management
  f2fs: fix i_size when pinned fallocate partially fails
  f2fs: fix to migrate all curseg types during free_segment_range
  f2fs: avoid setting SBI_NEED_FSCK on transient resize failure
  f2fs: fix to avoid pinfile fragment on fragment:{block, segment} mode
  f2fs: cleanup w/ f2fs_need_rand_{blk, seg, seg_blk}
  f2fs: fix to shrink gc_lock coverage in f2fs_gc_range()
  f2fs: fix to reclaim space in f2fs_allocate_pinning_section()
  f2fs: unify add/remove ino entry API for all ino types
  f2fs: fix to zero post-EOF data when extending file size
  ...
2026-08-28 10:48:48 -07:00
Linus Torvalds
70f5376dbd TTY / Serial driver updates for 7.3-rc1
Here is the "big" set of tty and serial driver updates for 7.3-rc1.  Not
 really all that much happened this development cycle for this subsystem,
 changes in here are:
   - removal of the ipwireless driver as it's no longer used or needed
   - new 8250_mxpcie driver added
   - qcom serial driver updates and additions
   - vt mode validation addition
   - lots of other small serial driver updates and additions
 
 All of these have been in linux-next for weeks with no reported issues.
 
 Signed-off-by: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
 -----BEGIN PGP SIGNATURE-----
 
 iG0EABECAC0WIQT0tgzFv3jCIUoxPcsxR9QN2y37KQUCao2mTw8cZ3JlZ0Brcm9h
 aC5jb20ACgkQMUfUDdst+yk34ACdFfyDYJ0n1JcdskTxdNMBSPRkj7UAoJtMC/2y
 jxCfyfkM18YIuwZD6CvO
 =MlfE
 -----END PGP SIGNATURE-----

Merge tag 'tty-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/tty

Pull TTY / serial driver updates from Greg KH:
 "Here is the "big" set of tty and serial driver updates for 7.3-rc1.

  Not really all that much happened this development cycle for this
  subsystem, changes in here are:

   - removal of the ipwireless driver as it's no longer used or needed

   - new 8250_mxpcie driver added

   - qcom serial driver updates and additions

   - vt mode validation addition

   - lots of other small serial driver updates and additions

  All of these have been in linux-next for weeks with no reported issues"

* tag 'tty-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/gregkh/tty: (97 commits)
  serial: imx: serialize imx_uart_ports[] lifetime
  tty: clear cdev pointer after cdev_add() failure
  tty: skip cdev_del() when no cdev is registered
  serial: core: clear freed pointers on uart_register_driver() failure
  serial: core: do fallible allocations before the console can be registered
  serial: 8250_mxpcie: implement rx_trig_bytes callbacks via MUEx50 RTL
  serial: 8250_mxpcie: introduce per-port private data structure
  serial: 8250: allow UART drivers to override rx_trig_bytes handling
  serial: 8250_mxpcie: add break support for RS485 using MUEx50 features
  serial: 8250: allow low-level drivers to override break control
  serial: 8250_mxpcie: support serial interface mode switching
  serial: 8250_mxpcie: speed up TX using memory-mapped FIFO window
  serial: 8250_mxpcie: speed up RX using memory-mapped FIFO window
  serial: 8250_mxpcie: add custom handle_irq callback
  serial: 8250_mxpcie: offload XON/XOFF flow control to MUEx50 hardware
  serial: 8250_mxpcie: enable automatic RTS/CTS flow control
  serial: 8250_mxpcie: enable enhanced mode and program FIFO trigger levels
  serial: 8250: add Moxa MUEx50 UART port type
  serial: 8250: split Moxa PCIe serial board support out of 8250_pci
  serial: qcom-geni: Use geni_se_set_perf_level() for baud rate perf level
  ...
2026-08-25 10:59:12 -07:00
Linus Torvalds
2f43193b88 dma-mapping updates for Linux 7.3:
- swiotlb: added new configuration option for the default pool size
 (Jagadeesh Pagadala) and reduced overhead for high watermark tracking
 (chenhuguanshen)
 
 - minor code cleanups and improvements (Vova Sharaienko, Honglei Huang
 and Marek Szyprowski)
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQSrngzkoBtlA8uaaJ+Jp1EFxbsSRAUCaoxGQgAKCRCJp1EFxbsS
 RFPEAP0eo9usjFcvh0YKTPh6/mXgqxRuTNQZ7i+2lRGEczKcJQEA7mwkgwpiOaKn
 f++mMVOmsPvl2Y7r/5XqBWywwhyygA0=
 =X8pY
 -----END PGP SIGNATURE-----
mergetag object 04a19b35dc
 type commit
 tag dma-mapping-7.3-2026-08-24-2
 tagger Marek Szyprowski <m.szyprowski@samsung.com> 1787582472 +0200
 
 second dma-mapping update for Linux 7.3:
 
 - important dma-mapping update for confidential-computing, which adds
 proper tracking of the shared DMA state through direct, pool and swiotlb
 paths (Aneesh Kumar K.V)
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQSrngzkoBtlA8uaaJ+Jp1EFxbsSRAUCaoxYzgAKCRCJp1EFxbsS
 RM9bAP4mJuHHzj2DqsKV7QX19uhyzmsHIg+ecjBNRaOdUAgelQD9FsaG/fwrZnRT
 y89H0QUErqLsdmkDqV0zsXfaGzWY4gk=
 =dKjS
 -----END PGP SIGNATURE-----

Merge tags 'dma-mapping-7.3-2026-08-24' and 'dma-mapping-7.3-2026-08-24-2' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux

Pull dma-mapping updates from Marek Szyprowski:

 - swiotlb:
     - new configuration option for the default pool size
       (Jagadeesh Pagadala)
     - reduce overhead for high watermark tracking (chenhuguanshen)

 - minor code cleanups and improvements (Vova Sharaienko, Honglei Huang
   and Marek Szyprowski)

 - add proper tracking of the shared DMA state through direct, pool and
   swiotlb paths (Aneesh Kumar K.V)

   This is important for confidential-computing

* tag 'dma-mapping-7.3-2026-08-24' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux:
  dma/swiotlb: decouple high watermark tracking from CONFIG_DEBUG_FS
  MAINTAINERS: update tree for DMA MAPPING HELPERS
  dma/swiotlb: introduce Kconfig option for compile-time default pool size
  dma-direct: Improve readability of the dma_direct_map_sg() for P2PDMA case
  iommu/dma: simplify dma_iova_destroy() and drop the free_iova helper
  dma-coherent: use KiB in DMA allocation logs
  dma-coherent: fix spacing coding style issue

* tag 'dma-mapping-7.3-2026-08-24-2' of git://git.kernel.org/pub/scm/linux/kernel/git/mszyprowski/linux: (23 commits)
  swiotlb: remove unused SWIOTLB_FORCE flag
  dma: swiotlb: handle set_memory_decrypted() failures
  dma: swiotlb: free dynamic pools from process context
  dma-direct: rename ret to cpu_addr in alloc helpers
  dma-direct: select DMA address encoding from __DMA_ATTR_ALLOC_CC_SHARED
  dma-direct: set decrypted flag for remapped DMA allocations
  dma-direct: make dma_direct_map_phys() honor DMA_ATTR_CC_SHARED
  dma-direct: Move dma_direct_map_phys() to dma/direct.c
  dma-direct: pass attrs to dma_capable() for DMA_ATTR_CC_SHARED checks
  dma-mapping: make dma_pgprot() honor __DMA_ATTR_ALLOC_CC_SHARED
  dma: swiotlb: track pool encryption state and honor DMA_ATTR_CC_SHARED
  dma: swiotlb: pass mapping attributes by reference
  dma-pool: track decrypted atomic pools and select them via attrs
  dma-direct: use __DMA_ATTR_ALLOC_CC_SHARED in alloc/free paths
  dma-mapping: Add internal shared allocation attribute
  coco: arm64: s390: powerpc: Mark secure guests with CC_ATTR_GUEST_MEM_ENCRYPT
  dma-direct: swiotlb: handle swiotlb alloc/free outside __dma_direct_alloc_pages
  s390: Expose protected virtualization through cc_platform_has()
  swiotlb: Preserve allocation virtual address for dynamic pools
  dma: free atomic pool pages by physical address
  ...
2026-08-24 11:35:46 -07:00
Linus Torvalds
918e25291c slab changes for 7.3
-----BEGIN PGP SIGNATURE-----
 
 iQFPBAABCAA5FiEEe7vIQRWZI0iWSE3xu+CwddJFiJoFAmqMSHQbFIAAAAAABAAO
 bWFudTIsMi41KzEuMTIsMiwyAAoJELvgsHXSRYiaC8EH/ihctMRDnqbUxyc3cSIZ
 cuy3ocSSu8UnHzSNarylY8sVYIYflgY6owV7UvaUKmXGYHiHIDdI5NMDzpjola7X
 Ct7mpIuobFHpFhTQMqvYkeEQ3EytS7NPHKs3jd6t2HSOYJdY4lghjRZcxOlo7+GA
 5EJa468TSHswWbT34e3NY7gpoHCXTucR6FBqoVtluLOliQWNMkWbeDQ9hEnMoyNy
 hEdooWDxt8vBuSqVRCZxrgJRK2nauf1P/CHZm4HcNttOQp0iwr8Q7GER7dBpW0jR
 1b1zM9YPH4g6CSll5V6sWSMUNAaDx3Akx3ZyKhd3G2kyP3HkPXKSag/fS8xxCFLN
 hh0=
 =6lTR
 -----END PGP SIGNATURE-----

Merge tag 'slab-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/vbabka/slab

Pull slab updates from Vlastimil Babka:

 - Add kfree_rcu_nolock() that can be used from contexts where spinning
   on a lock might be unsafe, such as a BPF program attached to an
   arbitrary function, or in NMI context. This complements the existing
   kfree_nolock() support (Harry Yoo)

 - Runtime instead of compile-time slabobj_ext sizing.

   Avoid wasting memory when memory allocation profiling is compiled but
   not enabled, with initial partial support to also avoid wasting
   memory for objcg pointers when those are not needed, while profiling
   is enabled (Vlastimil Babka)

 - Various non-urgent fixes, cleanups and optimizations (Hao Li,
   Hongling Zeng, Li RongQing, Li Xiasong, Seongjun Hong, Shengming Hu)

* tag 'slab-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/vbabka/slab: (31 commits)
  mm/slab, kfence, memcg: completely remove obj_ext for kfence objects
  mm/slab: stop allocating objcg pointers when unnecessary
  mm/slab: add cache_ and slab_needs_objcg() helpers
  mm/slab: stop exporting kvfree_rcu_barrier[_on_cache]()
  slub_kunit: extend the test for kfree_rcu_nolock()
  mm/slab: introduce kfree_rcu_nolock()
  mm/slab: introduce struct kvfree_rcu_head for kvfree_rcu batching
  mm/slab: reduce slabobj_ext memory with allocation profiling disabled
  mm/slab: introduce slab_obj_ext_has_codetag()
  mm/slab: allow kfree_rcu_sheaf() on PREEMPT_RT
  mm/slab: extend deferred free mechanism to handle rcu sheaves
  mm/slab: use call_rcu() in unknown context if irqs are enabled
  mm/slab: handle the !allow_spin case in kfree_rcu_sheaf()
  mm/slab: change struct slabobj_ext to a union
  mm/slab: replace slab.stride with obj_exts_in_object
  mm/slab: abstract slabobj_ext.ref access
  mm/slab: abstract slabobj_ext.objcg access
  mm/slab: make slab_obj_ext() determine object index
  mm: move struct slabobj_ext to mm/slab.h
  mm/slab: remove objs_per_slab()
  ...
2026-08-24 10:58:57 -07:00
Linus Torvalds
83684c4e4d RCU updates:
Make expedited grace periods expedite normal RCU callbacks
 
 Miscellaneous fixes:
  * Improve diagnostic output with character task states.
  * Mark accesses to inform KCSAN of concurrency design.
  * Move from kmalloc() to kmalloc_obj().
  * Documentation updates.
  * Improve handling of RCU deferred quiescent states.
  * Clean up unused function arguments and structure fields.
  * Reduce show_rcu_gp_kthreads() stack space.
 
 Tasks RCU updates:
  * Clean up after SRCU re-implementation of Tasks Trace RCU.
  * Mark accesses to inform KCSAN of concurrency design.
  * Add ->lazy_timer status to diagnostic output.
  * Remove an unnecessary memory barrier.
  * Fix a data race, courtesy of KCSAN.
  * Documentation updates.
  * Convert cond_resched_tasks_rcu_qs() from macro to static inline
    function.
 
 SRCU updates:
  * Add Rust helpers for SRCU.
  * Avoid losing queued work at cleanup_srcu_struct() time.
 
 Torture-test updates:
  * Preparation work for immediate RCU priority deboosting.
  * Test RCU readers from real interrupt handlers (as opposed to softirq).
  * Simplify code through use of cpumask_next_wrap().
  * Improve diagnostic output with character task states.
  * Add rcutorture.nwriters parameter to allow lightweight stall testing,
    and rcutorture.stall_only to make doing so easier.
  * Test an RCU Tasks Trace grace period implying an RCU grace period.
  * Make RCU Tasks Trace torturing track reader batches.
  * Fix a data race, courtesy of KCSAN.
  * Plug a shuffle_tmp_mask memory leak on kthread spawn failure.
 -----BEGIN PGP SIGNATURE-----
 
 iQJHBAABCgAxFiEEbK7UrM+RBIrCoViJnr8S83LZ+4wFAmqE5nYTHHBhdWxtY2tA
 a2VybmVsLm9yZwAKCRCevxLzctn7jCoDD/4uM0FYUucaPFp1DcQDSHR/o+UIvqS4
 UBuVNXN3kz0kTM2qWQ4mwsCPDtv2uxmzp+6OEmWpoPtutSujQc1vM9aEMxeEfCDo
 W4PRAJrtXCCfDCZu0xkq+UaXmIF5ajjfFtJIYZxsu6Gv1xR2XtvZqQ58x0MnVXU9
 FfW8XNBhTlXX+2WT9rFxkP4XR6hn1AIY5F9vEIamvu/z3DXwMRHD1wCEJ6BD60qg
 uPIPIIArAC79vidZPK/HBmj0FBqZ0S2NK4uugbkc1xzx1HBfcWA6Y8m+ECkeKbOH
 P4UArtTpwAszvrRAfNNmNe/1bR4fMoGcoLFdvAK9vmc8qpYXKVkZh6XblLUiV/XF
 oo6NKnWeywIQ595RfBzziK8d5coV/ge56P/7Idf+QBUM0XtDTFpwtzmzsYWgdzqi
 Y6s9+t022Eh9013rZ6aMHSNa4Vdffg5P8SjkEWmkqYGIP597kjpRRKYe0y3WGYhy
 wB21LDTi69BFgniytTbH5K0nw1sFbyWOmBpY6ABfDuagGmEDIHzYSw/cI4OW0BMI
 V+ZwpNYY1IPM00GLI76940iLekT6EAV/b06ca0xWum1Am4rR8qwxvCdg4oFCXcGD
 +tomWerTZtK53mkVt+z27iETH8jQD50vdaFYn/WWhQtTeVlYEmzm0qJcyBEp9xWV
 NtLzoF1NZBT8Mw==
 =JeeV
 -----END PGP SIGNATURE-----

Merge tag 'rcu.2026.08.18a' of git://git.kernel.org/pub/scm/linux/kernel/git/rcu/linux

Pull RCU updates from Paul McKenney:
 "Make expedited grace periods expedite normal RCU callbacks

  Miscellaneous fixes:
   - Improve diagnostic output with character task states
   - Mark accesses to inform KCSAN of concurrency design
   - Move from kmalloc() to kmalloc_obj()
   - Documentation updates
   - Improve handling of RCU deferred quiescent states
   - Clean up unused function arguments and structure fields
   - Reduce show_rcu_gp_kthreads() stack space

  Tasks RCU updates:
   - Clean up after SRCU re-implementation of Tasks Trace RCU
   - Mark accesses to inform KCSAN of concurrency design
   - Add ->lazy_timer status to diagnostic output
   - Remove an unnecessary memory barrier
   - Fix a data race, courtesy of KCSAN
   - Documentation updates
   - Convert cond_resched_tasks_rcu_qs() from macro to static inline
     function

  SRCU updates:
   - Add Rust helpers for SRCU
   - Avoid losing queued work at cleanup_srcu_struct() time

  Torture-test updates:
   - Preparation work for immediate RCU priority deboosting
   - Test RCU readers from real interrupt handlers (as opposed to
     softirq)
   - Simplify code through use of cpumask_next_wrap()
   - Improve diagnostic output with character task states
   - Add rcutorture.nwriters parameter to allow lightweight stall
     testing, and rcutorture.stall_only to make doing so easier
   - Test an RCU Tasks Trace grace period implying an RCU grace period
   - Make RCU Tasks Trace torturing track reader batches
   - Fix a data race, courtesy of KCSAN
   - Plug a shuffle_tmp_mask memory leak on kthread spawn failure"

* tag 'rcu.2026.08.18a' of git://git.kernel.org/pub/scm/linux/kernel/git/rcu/linux: (59 commits)
  rcu: Add closing parenthesis in comment in rcu_read_unlock_strict()
  rcutorture: Make {,s}rcu_read_delay() better handle forward-progress testing
  rcutorture: Announce declining to forward-progress test
  torture: Don't leak shuffle_tmp_mask when shuffler kthread fails to start
  rcutorture: Use this_cpu_inc() for rcu_torture_count[] and rcu_torture_batch[]
  rcutorture: Make RCU Tasks Trace track Reader Batches
  rcutorture: Test RCU Tasks Trace GP implying RCU GP
  rcutorture: Add a stall_only module parameter
  rcutorture: Add nwriters module parameter
  rcutorture: Use task_state_to_char() for task-state reporting
  rcutorture: Use cpumask_next_wrap() in rcu_torture_preempt()
  rcutorture: Test RCU readers from hardware interrupt handlers
  rcutorture: Check for immediate deboosting at reader end
  srcu: Queue sdp->work when the delay timer is successfully deleted
  rcu-tasks: Convert cond_resched_tasks_rcu_qs() to static inline
  rcu-tasks: Fix some comments for call_rcu_tasks() and call_rcu_tasks_rude()
  rcu-tasks: Rename tasks_rcu_exit_srcu_stall_timer to tasks_rcu_exit_stall_timer
  rcu: Mark interrupts-enabled accesses to rdp->cpu_no_qs.s
  rcu: Reduce stack usage in show_rcu_gp_kthreads()
  rcu: Mark accesses to ->rcu_urgent_qs and ->rcu_need_heavy_qs
  ...
2026-08-23 18:00:22 -07:00
Chao Yu
46d4246d8d f2fs: use f2fs_{down, up}_(read, write}_trace() for nat_tree_lock
Under heavy workloads or during background GC/fallocate operations,
nat_tree_lock can experience high lock contention between background
readers (e.g. f2fs_get_node_info() in gc_data_segment) and writers
(e.g. flush_nat_entries, set_node_addr, shrinker).

[375067.327986][T13777]  schedule+0x4c/0x114
[375067.327997][T13777]  f2fs_get_node_info+0x438/0x5c4
[375067.328002][T13777]  f2fs_get_inode_page+0x1e0/0x3f0
[375067.328013][T13777]  f2fs_iget+0x88/0x1180
[375067.328024][T13777]  f2fs_lookup+0x168/0x3a8
[375067.328035][T13777]  path_openat+0xa28/0x1b04
[375067.328046][T13777]  do_filp_open+0xac/0x130
[375067.328056][T13777]  do_sys_openat2+0x140/0x21c
[375067.328066][T13777]  __arm64_sys_openat+0x70/0x9c

[375067.330299][T13777]  schedule+0x4c/0x114
[375067.330310][T13777]  schedule_preempt_disabled+0x24/0x40
[375067.330321][T13777]  rwsem_down_write_slowpath+0x3b4/0x9d0
[375067.330332][T13777]  down_write+0x98/0x170
[375067.330343][T13777]  set_node_addr+0x74/0x4b4
[375067.330354][T13777]  f2fs_new_node_page+0xb0/0x280
[375067.330444][T13777]  f2fs_new_inode_page+0x3c/0x64
[375067.330455][T13777]  f2fs_init_inode_metadata+0x4c/0x47c
[375067.330461][T13777]  f2fs_add_regular_entry+0x258/0x5b8
[375067.330471][T13777]  f2fs_add_dentry+0x100/0x158
[375067.330476][T13777]  f2fs_do_add_link+0x84/0x140
[375067.330487][T13777]  f2fs_create+0xec/0x250

[375067.331759][T13777]  schedule+0x4c/0x114
[375067.331770][T13777]  f2fs_down_read+0x9c/0xc4
[375067.331781][T13777]  f2fs_need_inode_block_update+0x20/0x10c
[375067.331792][T13777]  f2fs_do_sync_file+0x478/0x830
[375067.331802][T13777]  f2fs_sync_file+0x2c/0x40

This patch converts nat_tree_lock to use the f2fs_{down,up}_{read,write}_trace
infrastructure.

Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
2026-08-21 23:42:31 +00:00
Linus Torvalds
7199989f3f Landlock update for v7.3-rc1
-----BEGIN PGP SIGNATURE-----
 
 iIYEABYKAC4WIQSVyBthFV4iTW/VU1/l49DojIL20gUCaoa8/RAcbWljQGRpZ2lr
 b2QubmV0AAoJEOXj0OiMgvbSSjMBAIq+/dIjMVUddtyqRlWKMfBQqokU4Dl5JeHB
 Y4Jh61idAQCihXimNu+TOuX1zc80eBR+IzDYaAU8ZGcK079hmGO5Bw==
 =TKoa
 -----END PGP SIGNATURE-----

Merge tag 'landlock-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux

Pull Landlock update from Mickaël Salaün:
 "This improves observability with Landlock tracepoints support, which
  required some refactoring for dedicated domain types and common
  helpers shared with audit code.

  A LANDLOCK_RESTRICT_SELF_NO_NEW_PRIVS flag is also added to improve
  process-wide domain enforcement consistency.

  Whiteout files are now correctly handled and tested, and a few other
  fixes"

* tag 'landlock-7.3-rc1' of git://git.kernel.org/pub/scm/linux/kernel/git/mic/linux: (34 commits)
  landlock: Document tracepoints
  selftests/landlock: Add landlock_enforce_domain trace tests
  selftests/landlock: Add scope and ptrace tracepoint tests
  selftests/landlock: Add network tracepoint tests
  selftests/landlock: Add filesystem tracepoint tests
  selftests/landlock: Add trace event test infrastructure and tests
  landlock: Add tracepoints for ptrace and scope denials
  landlock: Add landlock_deny_access_fs and landlock_deny_access_net
  landlock: Add tracepoints for rule checking
  landlock: Add landlock_enforce_domain tracepoint
  landlock: Add create_domain and free_domain tracepoints
  landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints
  landlock: Add create_ruleset and free_ruleset tracepoints
  landlock: Consolidate access-right and scope names in a shared header
  landlock: Decouple the per-denial logging decision from CONFIG_AUDIT
  landlock: Split denial logging from audit into common framework
  landlock: Split struct landlock_domain from struct landlock_ruleset
  landlock: Move domain query functions to domain.c
  landlock: Prepare ruleset and domain type split
  samples/landlock: Add LANDLOCK_RESTRICT_SELF_NO_NEW_PRIVS to sampler
  ...
2026-08-21 12:28:35 -07:00
Linus Torvalds
ed3b875bea mm.git review status for mm-hotfixes-stable..mm-stable
Everything:
 
 Total patches:       501
 Reviews/patch:       1.66
 Reviewed rate:       70%
 
 Excluding DAMON:
 
 Total patches:       356
 Reviews/patch:       2.26
 Reviewed rate:       90%
 
 Excluding DAMON and selftests:
 
 Total patches:       329
 Reviews/patch:       2.31
 Reviewed rate:       92%
 
 Excluding DAMON, selftests and maple_tree:
 
 Total patches:       328
 Reviews/patch:       2.31
 Reviewed rate:       92%
 
 Summary of patch series in this merge:
 
 - The 2 patch series "mm: drop "sub" prefix from various places" from
   Dev Jain implements some page->folio conversion and a naming cleanup.
 
 - The 2 patch series "mm/kasan: remove redundant initialization for
   kasan_flag_write_only" from Igor Putko provides some KASAN cleanup work.
 
 - The 2 patch series "mm/filemap: reduce unnecessary xarray lookups"
   from Chi Zhiling provides a small speedup in the pagecaache read code.
 
 - The 4 patch series "mm/percpu: Fix possible NOFS/NOIO reclaim
   recursion" from Kaitao Cheng improves a few things in the vmalloc code -
   mainly the avoidance of GFP_KERNEL allocations when the caller asked for
   GFP_NOFS or GFP_NOIO.
 
 - The 3 patch series "mm/kmemleak: avoid soft lockup when scanning task
   stacks" from Breno Leitao avoids a soft lockup watchdog trigger from the
   kmemleak scanning code in extreme situations.
 
 - The 6 patch series "mm/page_owner: misc cleanups" from Ye Liu is a
   collection of unrelated cleanups to the page_owner code.  For some
   reason lots of people have been working on the page_owner code this
   cycle.
 
 - The 4 patch series "mm: convert to walk_page_range_vma() to eliminate
   find_vma()" from Kefeng Wang simplifies and accelerates the page walking
   library function.
 
 - The 3 patch series "mm/migrate: preparatory cleanups for batch copy
   and offload" from Shivank Garg implements cleanups in the migration
   code.
 
 - The 4 patch series "mm/page_owner: add per-fd filter infrastructure
   for print_mode and NUMA filtering" from Zhen Ni provides per-fd
   filtering to page_owner in order to reduce the sometimes vast amount of
   output it can produce.
 
 - The 19 patch series "mm: Refactor bootmem gigantic hugepage
   allocation" from Muchun Song is a "set of fixes and preparatory cleanups
   around bootmem HugeTLB handling, sparse initialization ordering, and
   related vmemmap setup".
 
 - The 4 patch series "mm/zsmalloc: reduce lock contention in zs_free()"
   from Wenchao Hao reduces lock contention in zs_free(), which dominates
   the unmap path under memory pressure on Android (LMK kills) and on x86
   servers running zswap-heavy workloads.  Up to 1.83x improvement in
   microbenchmarking.
 
 - The 2 patch series "move alloc_tag.c file under mm/" from Suren
   Baghdasaryan does that.
 
 - The 6 patch series "samples/damon: handle damon_{start,stop}()
   failures" from SJ Park fixes improper handling of damon_start(),
   damon_stop(), and damon_call() failures across DAMON sample modules to
   prevent potential memory leaks, operation disruptions and use-after-free
   bugs.
 
 - The 11 patch series "mm/damon/sysfs: kobject_del() directories that
   users can create/remove" from SJ Park resolves an issue where delayed
   sysfs directory removal under CONFIG_DEBUG_KOBJECT_RELEASE causes
   creation failures due to duplicate directory names by adding missing
   kobject_del() calls before creating new directories.
 
 - The 3 patch series "mm: cleanup clear_not_present_full_ptes()" from
   David Hildenbrand cleans up the core pte handling code.
 
 - The 3 patch series "selftests/damon: misc fixes for test bugs" from
   Kunwu Chan fixes several bugs in the DAMON selftests.
 
 - The 2 patch series "selftests/damon: fix memcg_path staging handling"
   from Cheng Nie fixes a bug in _damon_sysfs.py for damos_filter
   memcg_path setup, and adds a test case for it in sysfs.py.
 
 - The 2 patch series "selftests/damon: test kdamond refresh_ms" from
   Ruslan Valiyev introduces selftest coverage for DAMON's refresh_ms sysfs
   feature by updating the test control module and verifying that scheme
   stats update automatically without manual intervention.
 
 - The 5 patch series "mm/damon: five misc fixups" from Akinobu Mita
   contains miscellaneous DAMON fixups.
 
 - The 2 patch series "mm/damon/core: detect internal variation above
   max_nr_regions/2" from Jiayuan Chen fixes DAMON's region splitting
   behavior when region counts exceed half the maximum budget by
   dynamically scaling down the split fraction as the limit approaches,
   preventing large regions from staying un-split, and adds corresponding
   KUnit test coverage.
 
 - The 6 patch series "mm: preparatory patches for PMD level swap
   entries" from Usama Arif refactors and cleans up PMD softleaf helpers,
   call sites, and architecture flags to lay the groundwork for a follow-up
   series that introduces PMD page table swap entries.
 
 - The 11 patch series "mm/damon: update, optimize, and clean up doc,
   tests, and code" from SJ Park updates DAMON design and ABI
   documentation, expands unit and selftest coverage, optimizes
   damon_commit_target_regions(), and cleans up recently added sysfs
   interface code for better readability.
 
 - The 2 patch series "mm/vmpressure: reduce CPU, memory and code
   overhead on cgroup v2" from Usama Arif optimizes vmpressure() by
   skipping unnecessary work on cgroup v2 for userspace event notifications
   and refactors v1-only eventfd handling into mm/memcontrol-v1.c to reduce
   memory overhead and code complexity.
 
 - The 10 patch series "selftests/mm: refactor pkey helpers and fix mmap
   error handling" from Hongfu Li refactors pkeys shared tracing and
   assertion helpers into a common file, unifies protection key selftests
   to use consistent diagnostic logging and assertions, and enforces
   standardized MAP_FAILED return checks for mmap() calls across the tests.
 
 - The 18 patch series "mm/damon: optimize out nr_accesses_bp" from SJ
   Park replaces the error-prone, continuously updated nr_accesses_bp field
   in damon_region with an on-demand moving sum function
   (damon_nr_accesses_mvsum()), reducing structure memory overhead and
   avoiding state corruption bugs.
 
 - The 6 patch series "Open HugeTLB allocation routine for more generic
   use" from Ackerley Tng decouples HugeTLB folio allocation from VMA
   dependencies by introducing hugetlb_alloc_folio(), enabling subsystems
   like guest_memfd to allocate HugeTLB folios without standard VMA
   reservations or pseudo-VMAs.
 
 - The 3 patch series "mm/damon: provide pseudo moving sum probe_hits"
   from SJ Park integrates DAMON's probe_hits attribute counter into the
   pseudo moving sum infrastructure, enabling real-time, online monitoring
   without waiting for full aggregation intervals.
 
 - The 18 patch series "mm: Some cleanups for page allocator APIs" from
   Brendan Jackman simplifies and refactors the page allocator entry points
   and flags by unifying allocation paths, adding internal alloc_flags
   arguments, and eliminating redundant __ prefixed alloc_pages variants.
 
 - The 5 patch series "Fix incorrect access of hugetlb pte entries" from
   Dev Jain enforces the consistent use of huge_ptep_get() instead of
   ptep_get() for HugeTLB entries and fixes an unaligned address issue in
   arm64's huge_ptep_get() implementation.
 
 - The 8 patch series "mm/damon: validate all parameters in the core"
   from SJ Park consolidates parameter validation into the DAMON core
   specifically within damon_start() and damon_commit_ctx() to centralize
   error checking, eliminate caller-side redundant checks and to improve
   maintenance efficiency.
 
 - The 3 patch series "tools/mm/page_owner_sort: fix filtering and
   cleanup issues" from Yichong Chen renames is_need() to filter_record()
   for clearer return semantics, fixes per-record allocation memory leaks
   and bounds output copies in search_pattern() to address an existing
   buffer issue.
 
 - The 4 patch series "memcg: bail out reclaim when memcg is dying" from
   Jiayuan Chen mitigates a system-wide stall which occurs when a cgroup is
   removed while one of its memory control files is doing synchronous
   reclaim.
 
 - The 5 patch series "mm/memory-failure: add panic option for
   unrecoverable pages" from Breno Leitao introduces an opt-in
   vm.panic_on_unrecoverable_memory_failure sysctl that immediately panics
   the kernel on unrecoverable memory errors in kernel-owned pages to
   preserve error context and prevent delayed, silent data corruption.
 
 - The 11 patch series "mm/damon: refactor damon_{start,stop,commit}()
   for simple error handling" from SJ Park refactors the DAMON core API
   functions to guarantee that all contexts are fully stopped when
   damon_start(), damon_stop(), or damon_commit() fail, eliminating the
   need for complex and error-prone caller-side cleanup code.
 
 - The 5 patch series "Keep tail page private zero at free and folio
   split" from Zi Yan adds checks to ensure tail_page->private is zero when
   freeing compound or high-order pages and when promoting tail pages
   during large folio splits.  By validating these fields at free and split
   time, it allows the removal of redundant private field clearing inside
   prep_compound_tail().
 
 - The 4 patch series "mm: drop redundant lru_add_drain in anon folio
   reuse paths" from Barry Song eliminates redundant lru_add_drain() calls
   in wp_can_reuse_anon_folio() and do_swap_page() to reduce LRU lock
   contention and system overhead.
 
   By validating folio refcounts against the LRU cache before draining
   and removing unnecessary drains in the swap path, it achieves up to a
   30.5% reduction in drain calls during heavy swap workloads.
 
 - The 3 patch series "mm: clean up folio LRU and swap declarations" from
   Jianyue Wu reorganizes folio LRU and swap code by relocating
   page-cluster state to mm/swap_state.c, renaming mm/swap.c to mm/folio.c,
   and moving MM-internal reclaim declarations into mm/internal.h.
 
 - The 15 patch series "userfaultfd: working set tracking for VM guest
   memory" from Kiryl Shutsemau adds userfaultfd support for tracking the
   working set of VM guest memory, so a VMM can identify hot pages and
   reclaim cold ones to tiered or remote storage.
 
 - The 10 patch series "mm: remove CONFIG_HAVE_BOOTMEM_INFO_NODE (Part
   2)" from David Hildenbrand removes the remaining pieces of
   CONFIG_HAVE_BOOTMEM_INFO_NODE, performing some smaller cleanups around
   freeing of reserved vmemmap pages on the way.
 
 - The 7 patch series "mm/damon: update probe hits for runtime parameter
   commits" from SJ Park ensures that DAMON's probe_hits attribute counter
   is properly updated when monitoring intervals are changed at runtime,
   matching the behavior of nr_accesses.  To achieve this, it refactors and
   renames existing helper functions for shared use, applies the updates to
   probe_hits, and handles edge cases in damon_probe_hits_mvsum() to
   maintain measurement accuracy.
 
 - The 3 patch series "KSM: performance optimizations for rmap_walk_ksm"
   from xu xin resolves a severe KSM reverse-mapping performance bottleneck
   where thousands of split VMAs sharing a single anon_vma cause extended
   lock contention.  By adding an interval-filtering check during the rmap
   walk, it reduces worst-case anon_vma lock hold times from over 500ms
   down to under 2ms, preventing application freezes and latency spikes
   under memory pressure.
 
 - The 3 patch series "mm: split a couple of headers from internal.h"
   from Mike Rapoport splits declarations related to mm_init, memblock,
   vmalloc and sparse into new headers.
 
 - The 2 patch series "KSM: use linear_page_index in collect_procs_ksm()"
   from xu xin applies the interval tree optimization from rmap_walk_ksm()
   to collect_procs_ksm() to avoid iterating over non-matching VMAs during
   KSM memory error handling.  It hoists loop-invariant address
   initialization and restricts the anon_vma_interval_tree_foreach walk to
   a targeted page offset range, reducing redundant checks and improving
   lookup efficiency.
 
 - The 3 patch series "selftests/mm: avoid false failures in hugetlb and
   KSM tests" from Sayali Patil fixes issues in the hugetlb and KSM MM
   selftest categories that can report failures when the prerequisites for
   the tests are not satisfied.
 
 - The 19 patch series "mm/damon: introduce data attributes only
   monitoring" from SJ Park introduces attribute-weighted region management
   in DAMON, allowing users to prioritize specific data attributes (such as
   page sizes or cgroups) over or instead of access monitoring.
 
   By assigning weights to attribute probes, DAMON can completely disable
   access tracking and adjust monitoring regions based on weighted
   probe-hit counters to optimize monitoring quality for attribute-focused
   workloads.
 
 - The 8 patch series "mm/hmm: Add mmap lock-drop support for
   userfaultfd-backed mappings" from Stanislav Kinsburskii extends
   hmm_range_fault() to support userfaultfd-backed regions by allowing the
   mmap lock to be dropped during fault handling via a new
   hmm_range_fault_locked() helper.
 
   By accepting a locked pointer and signaling retry status when lock
   release occurs, it enables page fault resolution in userfaultfd regions
   while preserving backward compatibility for existing callers.
 
 - The 33 patch series "mm: make VMA page offset handling more
   consistent" from Lorenzo Stoakes cleans up and standardizes how
   vma->vm_pgoff is accessed and manipulated across file-backed and
   anonymous mappings in the kernel.
 
   It introduces dedicated helper functions such as vma_start_pgoff(),
   vma_end_pgoff(), vma_set_pgoff() and linear_page_delta() while renaming
   rmap interval tree helpers to better reflect their functionality.
 
   These changes establish a cleaner foundation for future work that will
   unify virtual page offset indexing for all anonymous and CoW'd folios.
 
 - The 3 patch series "mm: handle device-private PMDs in walk callbacks"
   from Usama Arif addresses kernel panics and state corruption caused by
   MM walk callbacks reaching non-present device-private PMD swap entries
   created during HMM migrations.
 
   It ensures that functions which acquire pmd_trans_huge_lock() properly
   recognize device-private PMDs instead of assuming a present THP or a
   standard migration entry.
 
 - The 5 patch series "mm/rmap: Refactor try_to_unmap_one" from Dev Jain
   refactors try_to_unmap_one by modularizing Hugetlb, anonymous-lazyfree,
   and anonymous-swapbacked logic into dedicated functions, laying the
   structural groundwork for batched anonymous large folio unmapping.
 
 - The 4 patch series "Docs/ABI/damon: sysfs ABI document fixes and
   additions" from Song Hu fixes typos and fills in missing entries in the
   DAMON sysfs ABI document.
 
 - The 10 patch series "dax/kmem: atomic whole-device hotplug via sysfs"
   from Gregory Price introduces an atomic sysfs state attribute and
   supporting DAX/MM infrastructure to prevent userland races when
   offlining and removing entire memory regions.
 
   By adding an unplugged state alongside standard online modes, it
   enables whole-device atomic hotplug control while preserving backward
   compatibility.
 
 - The 13 patch series "mm: convert more vm_flags_t users to vma_flags_t"
   from Lorenzo Stoakes continues transitioning the kernel from the
   deprecated vm_flags_t type to vma_flags_t across core memory management
   infrastructure.
 
   It replaces legacy type usage in core functions such as do_mmap(),
   unmapped area allocation, mm->def_vma_flags, and VMA operations like
   mlock, mprotect, and mremap.
 
 - The 2 patch series "Two small patches to clean up mm/mm_slot.h" from
   xu xin refactors mm_slot.h by introducing mm_slot_remove() to unify
   duplicate slot deletion sequences in khugepaged and KSM.  It also adds
   code documentation explaining why mm_slot_lookup and mm_slot_insert must
   remain as preprocessor macros rather than static inline functions.
 
 - The 10 patch series "mm/damon/core: hide core-private struct fields"
   from SJ Park cleans up DAMON core structures by consistently marking
   internal-only fields with private: comment tags to prevent improper
   direct access from outer layers.
 
   It enforces encapsulation across core structures including
   damon_region, damon_target, and damon_ctx and updates DAMON_SYSFS to
   interact through approved access APIs instead of exposing raw struct
   members.
 
 - The 6 patch series "mm/damon: unurgent fixes for infinite loop, NULL
   de-ref and races" from SJ Park addresses potential infinite loops, NULL
   dereferences, and race conditions identified in DAMON.
 
   It fixes an infinite loop triggered by extreme user configurations, a
   NULL pointer dereference within unit tests and minor monitoring
   accuracy degradation caused by subtle runtime races.
 
 - The 2 patch series "mm/page_alloc: fixes for free_pages_nolock() on
   RT/UP" from Brendan Jackman fixes an NMI safety flaw in
   __free_frozen_pages() where freeing pages on non-SMP or PREEMPT_RT
   kernels can bypass can_spin_trylock() checks via non-PCP or isolated
   migration paths.
 
   It also resolves potential kernel crashes and privilege escalation
   risks triggered when BPF tracing runs in NMI context alongside memory
   hotplug or large allocation frees.
 
 - The 4 patch series "mm/page_alloc: couple of followups for recent
   cleanups" from Brendan Jackman cleans up and updates page allocator
   nomenclature, documentation, and debug assertions.
 
   It aligns internal FPI_ flags with the public "nolock" naming
   convention, removes outdated internal implementation details from
   high-level page allocator comments, and eliminates obsolete VM_BUG_ON()
   assertions in allocation paths.
 
 - The 3 patch series "mm/mseal: further cleanups" from Lorenzo Stoakes
   refactors and simplifies the mseal implementation by clarifying API
   boundaries and removing unnecessary code complexity.
 
   It replaces generic do_mseal() usage outside the syscall with a
   dedicated mseal_mmap_page_zero() helper for MMAP_PAGE_ZERO, eliminates
   mm_struct parameters to enforce that sealing applies only to
   current->mm, and streamlines overall logic and comments with no
   functional changes intended.
 
 - The 4 patch series "mm/vmscan: fix swappiness=max and clean up
   per-node proactive reclaim" from Ridong Chen resolves reclaim behavior
   bugs and cleans up function parameters across memory reclaim paths.
 
   It fixes swappiness=max in both standard reclaim and MGLRU so
   unswappable anonymous memory no longer falls back to evicting page
   cache, ensures reclaim_store() returns accurate error codes instead of
   collapsing all failures into -EAGAIN, and removes the obsolete gfp_mask
   parameter from __node_reclaim().
 
 - The 6 patch series "mm: mincore: misc cleanups" from Kefeng Wang
   cleans up and simplifies the mincore code.  Most importantly, it removes
   the historical special behavior that always reports VM_PFNMAP pages as
   non-resident.
 
 - The 2 patch series "mm/huge_memory: drop dead split helper variants"
   from Kiryl Shutsemau implements two trivial cleanups in the folio split
   API.
 
 - The 7 patch series "mm/damon: fix uninitialized DAMOS field and kunit
   exec expectation bugs" from SJ Park resolves minor operational and
   testing bugs in DAMON identified by Sashiko.  It initializes the
   damos->last_applied field to prevent occasional efficiency degradation
   and fixes invalid memory accesses in DAMON KUnit tests during test
   failure handling.
 
 - The 3 patch series "cleanup for stable_page_flags()" from Jinjiang Tu
   cleans up and refactors stable_page_flags() used by /proc/kpageflags
   without altering functionality.
 
   It uses BIT_ULL() to prevent shift-overflow warnings on 64-bit flag
   bits, converts folio-specific flag checks to standard folio_test_*()
   helpers, and removes redundant CONFIG_PAGE_IDLE_FLAG handling.
 
 - The 3 patch series "Batch unmap of uffd-wp file folios" from Dev Jain
   extends batched folio unmapping support to file folios within
   userfaultfd write-protect (uffd-wp) VMAs by adding batching capabilities
   to pte_install_uffd_wp_if_needed().
 
   This removes special-case restrictions on uffd-wp VMAs in
   try_to_unmap_one(), significantly simplifying the function's control
   flow and complexity.
 
 - The 3 patch series "mm/early_ioremap: clarify and clean up
   early_ioremap_reset()" from Sang-Heon Jeon clarifies and cleans up the
   architecture-specific usage of __late_set_fixmap() and
   __late_clear_fixmap() after early_ioremap_reset().
 
   It adds explicit documentation regarding when early_ioremap_reset()
   must be called and removes redundant macro definitions and reset calls
   in the RISC-V and ARM64 architectures.
 
 - The 4 patch series "mm: fix reclaim storms in defrag_mode" from
   Johannes Weiner addresses severe performance regressions, swap storms,
   and spurious OOMs caused by vm.defrag_mode=1 under high memory pressure
   in Meta production.
 
   It updates the page allocator slowpath so non-movable allocation
   requests actively trigger direct reclaim and direct compaction at
   pageblock_order scale, allowing them to claim whole pageblocks rather
   than spinning unproductively.
 
 - The 2 patch series "zram: lockmap tweaks" from Sebastian Siewior
   optimizes and fixes lockdep tracking for zram devices by consolidating
   per-entry lockmaps and isolating lock classes across multiple instances.
 
   It reduces memory overhead by replacing per-entry lockdep_map instances
   with a single map per struct zram, and assigns a dynamic lock_class_key
   to each instance to prevent false deadlock reports when different zram
   devices are backed by distinct filesystems.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTTMBEPP41GrTpTJgfdBJ7gKXxAjgUCaoUJbQAKCRDdBJ7gKXxA
 jqrzAP9WoPU0hiK4qS/kSjhtoZxhjpS5eLSUCy/utKuEvZbfGgEAu1zA+LH+X9Tm
 THK5ex4iUZxiFbXpWfLMxE/Q9PmQYQ8=
 =QTyb
 -----END PGP SIGNATURE-----

Merge tag 'mm-stable-2026-08-18-18-39' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm

Pull MM updates from Andrew Morton:

 - "mm: drop "sub" prefix from various places" (Dev Jain)

   page->folio conversion and a naming cleanup

 - "mm/kasan: remove redundant initialization for kasan_flag_write_only"
   (Igor Putko)

   KASAN cleanup work

 - "mm/filemap: reduce unnecessary xarray lookups" (Chi Zhiling)

   Small speedup in the pagecaache read code

 - "mm/percpu: Fix possible NOFS/NOIO reclaim recursion" (Kaitao Cheng)

   Improve the vmalloc code - mainly the avoidance of GFP_KERNEL
   allocations when the caller asked for GFP_NOFS or GFP_NOIO

 - "mm/kmemleak: avoid soft lockup when scanning task stacks" (Breno
   Leitao)

   Avoid a soft lockup watchdog trigger from the kmemleak scanning code
   in extreme situations

 - "mm/page_owner: misc cleanups" (Ye Liu)

   Cleanups to the page_owner code. For some reason lots of people have
   been working on the page_owner code this cycle.

 - "mm: convert to walk_page_range_vma() to eliminate find_vma()"
   (Kefeng Wang)

   Simplify and accelerate the page walking library function

 - "mm/migrate: preparatory cleanups for batch copy and offload"
   (Shivank Garg)

   Cleanups in the migration code

 - "mm/page_owner: add per-fd filter infrastructure for print_mode and
   NUMA filtering" (Zhen Ni)

   Per-fd filtering to page_owner in order to reduce the sometimes vast
   amount of output it can produce

 - "mm: Refactor bootmem gigantic hugepage allocation" (Muchun Song)

   Fixes and preparatory cleanups around bootmem HugeTLB handling,
   sparse initialization ordering, and related vmemmap setup

 - "mm/zsmalloc: reduce lock contention in zs_free()" (Wenchao Hao)

   Reduce lock contention in zs_free(), which dominates the unmap path
   under memory pressure on Android (LMK kills) and on x86 servers
   running zswap-heavy workloads.

   Up to 1.83x improvement in microbenchmarking.

 - "move alloc_tag.c file under mm/" (Suren Baghdasaryan)

 - "samples/damon: handle damon_{start,stop}() failures" (SJ Park)

   Fix improper handling of damon_start(), damon_stop(), and
   damon_call() failures across DAMON sample modules to prevent
   potential memory leaks, operation disruptions and use-after-free
   bugs

 - "mm/damon/sysfs: kobject_del() directories that users can
   create/remove" (SJ Park)

   Fix delayed sysfs directory removal under DEBUG_KOBJECT_RELEASE
   causeing creation failures due to duplicate directory names by adding
   missing kobject_del() calls before creating new directories

 - "mm: cleanup clear_not_present_full_ptes()" (David Hildenbrand)

   Clean up the core pte handling code

 - "selftests/damon: misc fixes for test bugs" (Kunwu Chan)

   Fix several bugs in the DAMON selftests

 - "selftests/damon: fix memcg_path staging handling" (Cheng Nie)

   Fix a bug in _damon_sysfs.py for damos_filter memcg_path setup, and
   add a test case for it in sysfs.py.

 - "selftests/damon: test kdamond refresh_ms" (Ruslan Valiyev)

   Selftest coverage for DAMON's refresh_ms sysfs feature by updating
   the test control module and verifying that scheme stats update
   automatically without manual intervention

 - "mm/damon: five misc fixups" (Akinobu Mita)

   Miscellaneous DAMON fixups.

 - "mm/damon/core: detect internal variation above max_nr_regions/2"
   (Jiayuan Chen)

   Fix DAMON's region splitting behavior when region counts exceed half
   the maximum budget by dynamically scaling down the split fraction as
   the limit approaches, preventing large regions from staying un-split,
   and add corresponding KUnit test coverage

 - "mm: preparatory patches for PMD level swap entries" (Usama Arif)

   Refactor and clean up PMD softleaf helpers, call sites, and
   architecture flags to lay the groundwork for a follow-up series that
   introduces PMD page table swap entries

 - "mm/damon: update, optimize, and clean up doc, tests, and code" (SJ
   Park)

   Update DAMON design and ABI documentation, expands unit and selftest
   coverage, optimize damon_commit_target_regions(), and clean up
   recently added sysfs interface code for better readability

 - "mm/vmpressure: reduce CPU, memory and code overhead on cgroup v2"
   (Usama Arif)

   Optimize vmpressure() by skipping unnecessary work on cgroup v2 for
   userspace event notifications and refactor v1-only eventfd handling
   into mm/memcontrol-v1.c to reduce memory overhead and code complexity

 - "selftests/mm: refactor pkey helpers and fix mmap error handling"
   (Hongfu Li)

   Refactor pkeys shared tracing and assertion helpers into a common
   file, unify protection key selftests to use consistent diagnostic
   logging and assertions, and enforce standardized MAP_FAILED return
   checks for mmap() calls across the tests

 - "mm/damon: optimize out nr_accesses_bp" (SJ Park)

   Replace the error-prone, continuously updated nr_accesses_bp field in
   damon_region with an on-demand moving sum function, reducing
   structure memory overhead and avoiding state corruption bugs

 - "Open HugeTLB allocation routine for more generic use" (Ackerley Tng)

   Decouple HugeTLB folio allocation from VMA dependencies by
   introducing hugetlb_alloc_folio(), enabling subsystems like
   guest_memfd to allocate HugeTLB folios without standard VMA
   reservations or pseudo-VMAs

 - "mm/damon: provide pseudo moving sum probe_hits" (SJ Park)

   Integrate DAMON's probe_hits attribute counter into the pseudo moving
   sum infrastructure, enabling real-time, online monitoring without
   waiting for full aggregation intervals

 - "mm: Some cleanups for page allocator APIs" (Brendan Jackman)

   Simplify and refactor the page allocator entry points and flags by
   unifying allocation paths, adding internal alloc_flags arguments, and
   eliminating redundant __ prefixed alloc_pages variants.

 - "Fix incorrect access of hugetlb pte entries" (Dev Jain)

   Enforce the consistent use of huge_ptep_get() instead of ptep_get()
   for HugeTLB entries and fixes an unaligned address issue in arm64's
   huge_ptep_get() implementation

 - "mm/damon: validate all parameters in the core" (SJ Park)

   Consolidate parameter validation into the DAMON core specifically
   within damon_start() and damon_commit_ctx() to centralize error
   checking, eliminate caller-side redundant checks and to improve
   maintenance efficiency

 - "tools/mm/page_owner_sort: fix filtering and cleanup issues" (Yichong
   Chen)

   Rename is_need() to filter_record() for clearer return semantics, fix
   per-record allocation memory leaks and bound output copies in
   search_pattern() to address an existing buffer issue

 - "memcg: bail out reclaim when memcg is dying" (Jiayuan Chen)

   Mitigate a system-wide stall which occurs when a cgroup is removed
   while one of its memory control files is doing synchronous reclaim

 - "mm/memory-failure: add panic option for unrecoverable pages" (Breno
   Leitao)

   Introduce an opt-in vm.panic_on_unrecoverable_memory_failure sysctl
   that immediately panics the kernel on unrecoverable memory errors in
   kernel-owned pages to preserve error context and prevent delayed,
   silent data corruption

 - "mm/damon: refactor damon_{start,stop,commit}() for simple error
   handling" (SJ Park)

   Refactor the DAMON core API functions to guarantee that all contexts
   are fully stopped when damon_start(), damon_stop(), or damon_commit()
   fail, eliminating the need for complex and error-prone caller-side
   cleanup code

 - "Keep tail page private zero at free and folio split" (Zi Yan)

   Add checks to ensure tail_page->private is zero when freeing compound
   or high-order pages and when promoting tail pages during large folio
   splits. By validating these fields at free and split time, it allows
   the removal of redundant private field clearing inside
   prep_compound_tail()

 - "mm: drop redundant lru_add_drain in anon folio reuse paths" (Barry
   Song)

   Eliminate redundant lru_add_drain() calls in
   wp_can_reuse_anon_folio() and do_swap_page() to reduce LRU lock
   contention and system overhead

   By validating folio refcounts against the LRU cache before draining
   and removing unnecessary drains in the swap path, it achieves up to a
   30.5% reduction in drain calls during heavy swap workloads

 - "mm: clean up folio LRU and swap declarations" (Jianyue Wu)

   Reorganize folio LRU and swap code by relocating page-cluster state
   to mm/swap_state.c, renaming mm/swap.c to mm/folio.c, and moving
   MM-internal reclaim declarations into mm/internal.h.

 - "userfaultfd: working set tracking for VM guest memory" (Kiryl
   Shutsemau)

   Add userfaultfd support for tracking the working set of VM guest
   memory, so a VMM can identify hot pages and reclaim cold ones to
   tiered or remote storage

 - "mm: remove CONFIG_HAVE_BOOTMEM_INFO_NODE (Part 2)" (David
   Hildenbrand)

   Remove the remaining pieces of CONFIG_HAVE_BOOTMEM_INFO_NODE,
   performing some smaller cleanups around freeing of reserved vmemmap
   pages on the way.

 - "mm/damon: update probe hits for runtime parameter commits" (SJ Park)

   Ensure that DAMON's probe_hits attribute counter is properly updated
   when monitoring intervals are changed at runtime, matching the
   behavior of nr_accesses. To achieve this, it refactors and renames
   existing helper functions for shared use, applies the updates to
   probe_hits, and handles edge cases in damon_probe_hits_mvsum() to
   maintain measurement accuracy.

 - "KSM: performance optimizations for rmap_walk_ksm" (xu xin)

   Resolve a severe KSM reverse-mapping performance bottleneck where
   thousands of split VMAs sharing a single anon_vma cause extended lock
   contention.

   By adding an interval-filtering check during the rmap walk, it
   reduces worst-case anon_vma lock hold times from over 500ms down to
   under 2ms, preventing application freezes and latency spikes under
   memory pressure.

 - "mm: split a couple of headers from internal.h" (Mike Rapoport)

   Split declarations related to mm_init, memblock, vmalloc and sparse
   into new headers

 - "KSM: use linear_page_index in collect_procs_ksm()" (xu xin)

   Apply the interval tree optimization from rmap_walk_ksm() to
   collect_procs_ksm() to avoid iterating over non-matching VMAs during
   KSM memory error handling.

   It hoists loop-invariant address initialization and restricts the
   anon_vma_interval_tree_foreach walk to a targeted page offset range,
   reducing redundant checks and improving lookup efficiency.

 - "selftests/mm: avoid false failures in hugetlb and KSM tests" (Sayali
   Patil)

   Fix issues in the hugetlb and KSM MM selftest categories that can
   report failures when the prerequisites for the tests are not
   satisfied

 - "mm/damon: introduce data attributes only monitoring" (SJ Park)

   Introduce attribute-weighted region management in DAMON, allowing
   users to prioritize specific data attributes (such as page sizes or
   cgroups) over or instead of access monitoring.

   By assigning weights to attribute probes, DAMON can completely
   disable access tracking and adjust monitoring regions based on
   weighted probe-hit counters to optimize monitoring quality for
   attribute-focused workloads.

 - "mm/hmm: Add mmap lock-drop support for userfaultfd-backed mappings"
   (Stanislav Kinsburskii)

   Extend hmm_range_fault() to support userfaultfd-backed regions by
   allowing the mmap lock to be dropped during fault handling via a new
   hmm_range_fault_locked() helper.

   By accepting a locked pointer and signaling retry status when lock
   release occurs, it enables page fault resolution in userfaultfd
   regions while preserving backward compatibility for existing callers.

 - "mm: make VMA page offset handling more consistent" (Lorenzo Stoakes)

   Clean up and standardize how vma->vm_pgoff is accessed and
   manipulated across file-backed and anonymous mappings in the kernel

   It introduces dedicated helper functions such as vma_start_pgoff(),
   vma_end_pgoff(), vma_set_pgoff() and linear_page_delta() while
   renaming rmap interval tree helpers to better reflect their
   functionality.

   These changes establish a cleaner foundation for future work that
   will unify virtual page offset indexing for all anonymous and CoW'd
   folios.

 - "mm: handle device-private PMDs in walk callbacks" (Usama Arif)

   Address kernel panics and state corruption caused by MM walk
   callbacks reaching non-present device-private PMD swap entries
   created during HMM migrations

   It ensures that functions which acquire pmd_trans_huge_lock()
   properly recognize device-private PMDs instead of assuming a present
   THP or a standard migration entry.

 - "mm/rmap: Refactor try_to_unmap_one" (Dev Jain)

   Refactor try_to_unmap_one by modularizing Hugetlb,
   anonymous-lazyfree, and anonymous-swapbacked logic into dedicated
   functions, laying the structural groundwork for batched anonymous
   large folio unmapping.

 - "Docs/ABI/damon: sysfs ABI document fixes and additions" (Song Hu)

   Fix typos and fills in missing entries in the DAMON sysfs ABI
   document

 - "dax/kmem: atomic whole-device hotplug via sysfs" (Gregory Price)

   Introduce an atomic sysfs state attribute and supporting DAX/MM
   infrastructure to prevent userland races when offlining and removing
   entire memory regions

   By adding an unplugged state alongside standard online modes, it
   enables whole-device atomic hotplug control while preserving backward
   compatibility.

 - "mm: convert more vm_flags_t users to vma_flags_t" (Lorenzo Stoakes)

   Continue transitioning the kernel from the deprecated vm_flags_t type
   to vma_flags_t across core memory management infrastructure.

   It replaces legacy type usage in core functions such as do_mmap(),
   unmapped area allocation, mm->def_vma_flags, and VMA operations like
   mlock, mprotect, and mremap.

 - "Two small patches to clean up mm/mm_slot.h" (xu xin)

   Refactor mm_slot.h by introducing mm_slot_remove() to unify duplicate
   slot deletion sequences in khugepaged and KSM. It also adds code
   documentation explaining why mm_slot_lookup and mm_slot_insert must
   remain as preprocessor macros rather than static inline functions.

 - "mm/damon/core: hide core-private struct fields" (SJ Park)

   Clean up DAMON core structures by consistently marking internal-only
   fields with private: comment tags to prevent improper direct access
   from outer layers.

   It enforces encapsulation across core structures including
   damon_region, damon_target, and damon_ctx and updates DAMON_SYSFS to
   interact through approved access APIs instead of exposing raw struct
   members.

 - "mm/damon: unurgent fixes for infinite loop, NULL de-ref and races"
   (SJ Park)

   Address potential infinite loops, NULL dereferences, and race
   conditions identified in DAMON

   It fixes an infinite loop triggered by extreme user configurations, a
   NULL pointer dereference within unit tests and minor monitoring
   accuracy degradation caused by subtle runtime races.

 - "mm/page_alloc: fixes for free_pages_nolock() on RT/UP" (Brendan
   Jackman)

   Fix an NMI safety flaw in __free_frozen_pages() where freeing pages
   on non-SMP or PREEMPT_RT kernels can bypass can_spin_trylock() checks
   via non-PCP or isolated migration paths.

   It also resolves potential kernel crashes and privilege escalation
   risks triggered when BPF tracing runs in NMI context alongside memory
   hotplug or large allocation frees.

 - "mm/page_alloc: couple of followups for recent cleanups" (Brendan
   Jackman)

   Clean up and update page allocator nomenclature, documentation, and
   debug assertions.

   It aligns internal FPI_ flags with the public "nolock" naming
   convention, removes outdated internal implementation details from
   high-level page allocator comments, and eliminates obsolete
   VM_BUG_ON() assertions in allocation paths.

 - "mm/mseal: further cleanups" (Lorenzo Stoakes)

   Refactor and simplify the mseal implementation by clarifying API
   boundaries and removing unnecessary code complexity.

   It replaces generic do_mseal() usage outside the syscall with a
   dedicated mseal_mmap_page_zero() helper for MMAP_PAGE_ZERO,
   eliminates mm_struct parameters to enforce that sealing applies only
   to current->mm, and streamlines overall logic and comments with no
   functional changes intended.

 - "mm/vmscan: fix swappiness=max and clean up per-node proactive
   reclaim" (Ridong Chen)

   Resolve reclaim behavior bugs and clean up function parameters across
   memory reclaim paths

   It fixes swappiness=max in both standard reclaim and MGLRU so
   unswappable anonymous memory no longer falls back to evicting page
   cache, ensures reclaim_store() returns accurate error codes instead
   of collapsing all failures into -EAGAIN, and removes the obsolete
   gfp_mask parameter from __node_reclaim().

 - "mm: mincore: misc cleanups" (Kefeng Wang)

   Clean up and simplifies the mincore code. Most importantly, it
   removes the historical special behavior that always reports VM_PFNMAP
   pages as non-resident.

 - "mm/huge_memory: drop dead split helper variants" (Kiryl Shutsemau)

   Two trivial cleanups in the folio split API

 - "mm/damon: fix uninitialized DAMOS field and kunit exec expectation
   bugs" (SJ Park)

   Resolve minor operational and testing bugs in DAMON identified by
   Sashiko. It initializes the damos->last_applied field to prevent
   occasional efficiency degradation and fixes invalid memory accesses
   in DAMON KUnit tests during test failure handling.

 - "cleanup for stable_page_flags()" (Jinjiang Tu)

   Clean up and refactor stable_page_flags() used by /proc/kpageflags
   without altering functionality.

   It uses BIT_ULL() to prevent shift-overflow warnings on 64-bit flag
   bits, converts folio-specific flag checks to standard folio_test_*()
   helpers, and removes redundant CONFIG_PAGE_IDLE_FLAG handling.

 - "Batch unmap of uffd-wp file folios" (Dev Jain)

   Extend batched folio unmapping support to file folios within
   userfaultfd write-protect (uffd-wp) VMAs by adding batching
   capabilities to pte_install_uffd_wp_if_needed().

   This removes special-case restrictions on uffd-wp VMAs in
   try_to_unmap_one(), significantly simplifying the function's control
   flow and complexity.

 - "mm/early_ioremap: clarify and clean up early_ioremap_reset()"
   (Sang-Heon Jeon)

   Clarify and clean up the architecture-specific usage of
   __late_set_fixmap() and __late_clear_fixmap() after
   early_ioremap_reset()

   It adds explicit documentation regarding when early_ioremap_reset()
   must be called and removes redundant macro definitions and reset
   calls in the RISC-V and ARM64 architectures.

 - "mm: fix reclaim storms in defrag_mode" (Johannes Weiner)

   Address severe performance regressions, swap storms, and spurious
   OOMs caused by vm.defrag_mode=1 under high memory pressure in Meta
   production

   It updates the page allocator slowpath so non-movable allocation
   requests actively trigger direct reclaim and direct compaction at
   pageblock_order scale, allowing them to claim whole pageblocks rather
   than spinning unproductively.

 - "zram: lockmap tweaks" (Sebastian Siewior)

   Optimize and fix lockdep tracking for zram devices by consolidating
   per-entry lockmaps and isolate lock classes across multiple instances

   This reduces memory overhead by replacing per-entry lockdep_map
   instances with a single map per struct zram, and assigns a dynamic
   lock_class_key to each instance to prevent false deadlock reports
   when different zram devices are backed by distinct filesystems.

* tag 'mm-stable-2026-08-18-18-39' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (501 commits)
  selftests/mm: thuge-gen: fix test_shmget() for PAGE_SIZE check
  selftests/mm: unpoison pages in memory-failure teardown
  mm/shmem: downgrade final i_blocks check in shmem_evict_inode() to pr_warn()
  mm/khugepaged: replace mutex_lock/mutex_unlock usage with guard macro
  mm/zsmalloc: fix release order of locks in zs_page_migrate()
  Documentation: zram: remove sections numbering
  ksm: stop iterating VMAs when ksm_test_exit returns true
  mm: fold userfaultfd_rwp() to false without CONFIG_ARCH_HAS_PTE_PROTNONE
  mm/migrate: report RCU-tasks quiescent states in migrate_pages_batch()
  zram: use a custom key for each zram object
  zram: move lockmap to be per-zram instead per table
  selftests/mm: fix gup_longterm EINVAL error message
  mm: page_alloc: fix non-movable reclaim storm in defrag_mode
  mm: page_alloc: move capture_control to the page allocator
  mm: compaction: support non-movable compaction for pageblock requests
  mm: page_alloc: __GFP_FS lockdep annotation for direct compaction
  hugetlb: evaluate subpool free state while locked
  mm/damon: remove trailing semicolons after function definitions
  mm/damon/ops-common: prevent migration fallback to non-target nodes
  mm/damon: update outdated comment about DAMOS filter handling
  ...
2026-08-20 18:17:08 -07:00
Linus Torvalds
50c44fea13 for-7.3-tag
-----BEGIN PGP SIGNATURE-----
 
 iQJPBAABCgA5FiEE8rQSAMVO+zA4DBdWxWXV+ddtWDsFAmqEFbYbFIAAAAAABAAO
 bWFudTIsMi41KzEuMTIsMiwyAAoJEMVl1fnXbVg7/R0P/jNxH2bHnc3yNLTuSuIA
 hj6QNSeTF6Q8bY5HGLbxIURr1npcODXGAFSZK5/vONZWgjSVy8j65gmciOPrD6Aq
 YF9zjFY4JuRfltx2E1aOKLyeozq/LPPs/3VBcfTHv2vWiYEabkpvZrNu9Fu7iX7c
 9TXSHzuBGp+Ao9faj9i0V/6en7hRxA6NW00FoIxvyANglJqcV35NV68zTuUeY62C
 8pFoDpDMV2ePQZW9l4EUsgqsEL6/PF362HmVyWUREYtnfpqJuzTADMV5NH29jXN1
 ZHw/P8La7gkHnzaJqDK1Nhs1L2vYLAdtIi5XdCMRCzHfuT2s7Wpo2VuXlpN77bP+
 nR7ZMxbg6zAIvxAc246bYEUQ6LNxVTZ4+7gSDJXUbWdpM1zYoHYLPAvelsCvrpq+
 2RJgChcEPNPHpQ4TXHLLERaJ6IvYFX1zfH+tU4yOGrapsF9wvVWpliR5dyJOGbj3
 RDOmZdV1XHZzstb5muBgigC2Tqh/z4h2Inl3ShucxDpZUUAl7wXNAuogfUV88ioh
 /zHqcSJVFYiSFi7m4Ml//TQgbtezUwFPHWdVoUybqHxAiBzC1HZ2/HBAyn93Cd8y
 PbNYoF7D3bU7UQ7g+aJj9lbOm4ttamg00R3D3JGfYneufUOJWGAnditQv0xakORG
 RXbq24AW0Im6O9zNU+E/iaCT
 =rFrX
 -----END PGP SIGNATURE-----

Merge tag 'for-7.3-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux

Pull btrfs updates from David Sterba:
 "This is the summer edition of btrfs changes, smaller than usual. Yet,
  there are performance improvements in various areas or for specific
  workloads and some notable changes like removing space cache v1 code
  or mount option reduction.

  User visible changes:

   - free space v1 disabled by default; the v2 (free space tree) is mkfs
     default since 5.15, filesystems with v1 still work but could be
     slightly slower due to lack of block group caching

   - mount option 'rescue=usebackuproot' requires read-only mount, it's
     too risky to allow writable mount

   - remove standalone mount option 'usebackuproot', deprecated in 5.9

   - preserve constraints of NODATASUM and NODATACOW when chattr and
     mount options may change the attributes

   - remove arbitrary limitation of 4KiB for page size when allowing
     block sizes smaller than page

   - print messages when pinned block groups affect swap activation

  Performance improvements:

   - use iomap bounce buffer for direct io instead of a fall back to
     buffered io; past correctness vs speed trade-offs dropped
     performance to ~50% of theoretical maximum, now it's ~95%,
     effectively doubled

   - replace xarray with local LRU list for tracking inhibited
     extent buffers, restored performance to pre-inhibition state
     (relatively ~3x)

   - remove unnecessary 1 jiffy delay in "non-SSD" mode with multiple
     logging tasks, decrease latency, throughput increased ~5x on sample
     workload

   - skip hole detection during full fsync for files without holes
     and lots of extents, reduce run time ~5x on sample workload
     (microsecond ranges)

   - reduce locking around extent readahead so it does not slow down
     other tasks using an overlapping range

   - enhance extent buffer allocation modes to allow NOWAIT semantics
     in some cases

  Notable fixes:

   - write-protect folios during writeback, prevent concurrent mmap
     and compress/checksumming/etc undesired interactions

   - in zoned mode, handle transient overcommit full instead of going
     read-only

   - fix possible deadlock between defragmentation and delayed
     allocation reservations

   - handle remaining iputs at umount time

   - fix lockdep warning between device scan locking and log mutex

   - add workaround for degenerate RAID56 device count modes (2 and 3)
     not supported by the parity calculation library

   - restore check that subvolume is not read-only when changing ACLs

   - retry reading verity data colliding with up-to-date status changes

  Core:

   - simplify raid56 stripe handling by using contiguous virtual
     allocations

   - in zoned mode, fix various metadata write issues in writeback or
     unmount

   - space reservation fixes

   - remove unused data structure members

   - more auto-freeing conversions

   - error pointer values are printed using %pe format"

* tag 'for-7.3-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux: (72 commits)
  btrfs: skip hole detection during full fsync for files without holes
  btrfs: add extra ASSERT()s to make sure the folio size is correct
  btrfs: use GFP_NOWAIT for tree block readahead
  btrfs: enable unlocked NOFAIL retry for eb allocations
  btrfs: add struct btrfs_eb_prealloc
  btrfs: factor init_extent_buffer from __alloc_extent_buffer
  btrfs: qgroup: fix a wrong length calculation in qgroup_free_reserved_data()
  btrfs: add validation for extent states
  btrfs: use aligned range for locking in reflink
  btrfs: use aligned range for locking in extent_fiemap()
  btrfs: zoned: don't clobber the extent buffer when zeroing it out
  btrfs: zoned: drop stranded dirty metadata buffers at unmount
  btrfs: zoned: drop stranded dirty metadata on transaction abort
  btrfs: zoned: flush active metadata block group at btree_writepages() start
  btrfs: convert reflink.c to use btrfs_inode as parameters
  btrfs: use simple booleans for log_commit field in struct btrfs_root
  btrfs: check for exit condition after waking in wait_log_commit()
  btrfs: move condition for log commit wait into wait_log_commit()
  btrfs: remove log batch counter use for fsync
  btrfs: stop sleeping for one jiffy in non-ssd mounts during log commit
  ...
2026-08-20 13:00:40 -07:00
Linus Torvalds
11260c335e sched_ext: Changes for v7.3
This depends on the arena argument support in the BPF tree and should be
 pulled after the scheduler core and BPF pulls. The patches based on bpf-next
 were kept on a separate branch which was merged into for-7.3 just now. The
 same merged result was in linux-next for several days.
 
 Most of this cycle completes the enqueue-path support for hierarchical
 sub-scheduling, which makes sub-scheduler support feature complete: a root
 BPF scheduler can now hand a cgroup subtree over to a nested sub-scheduler
 together with revocable CPU grants, and the sub-scheduler owns all
 scheduling decisions for its tasks on those CPUs.
 
 Development volume was high and a number of changes plugging holes in the
 new support landed late in the cycle. Also included are core scheduling
 fixes that were completed too late for the v7.2 release and are routed
 through this pull request.
 
 - Sub-scheduler CPU delegation:
 
   - Parent schedulers now grant and revoke per-CPU capabilities (enqueueing,
     preemption, CPU frequency control) on their children, enforced on every
     path a scheduler can reach a CPU through. Previously only dispatching
     could be delegated; this lets sub-schedulers fully schedule their CPUs.
 
   - Rescue execution: a task whose scheduler doesn't have access to the CPUs
     the task needs to run on starved until the watchdog ejected the whole
     scheduler. The kernel now runs such tasks directly on a small bandwidth
     budget, turning a scheduler-killing failure into bounded degradation.
 
   - Cgroup integration: tasks migrating across a sub-scheduler boundary
     weren't re-homed to the new owner, causing wrong-scheduler scheduling
     and a use-after-free. Sub-schedulers now take over their cgroup subtree
     and receive its cgroup callbacks.
 
   - Arena objects now cross the kernel/BPF boundary as typed pointer
     arguments, translated transparently by the BPF tree's new arena argument
     support, replacing untyped arguments with manual translation.
 
   - scx_qmap now demonstrates full hierarchical sub-scheduling.
 
 - Robustness improvements: the abort path is now NMI-safe, fixing deadlocks
   when errors are raised from NMI context and making hardlockup recovery
   direct. Reenqueue loops that could monopolize a CPU ahead of the watchdog
   now eject the offending scheduler, and stalls are blamed on the scheduler
   actually responsible.
 
 - Hardening: BPF-writable arena memory is validated before kernel use, and
   task slice and vtime writes got explicit synchronization rules, closing
   corruption vectors open to buggy or malicious schedulers.
 
 - Core scheduling: sched_ext dispatching can drop the rq lock inside the
   core-wide pick, which let interleaving selections corrupt each other's
   state and hard-hang the machine. The selection now restarts when the lock
   was released. The task ordering callback was also invoked with its
   arguments swapped, and the default ordering is updated to work across
   sub-scheduler boundaries. The fixes are marked for stable.
 
 - Other fixes headed for stable: a task init leak on fork failure during
   enable, tooling compat macros that silently failed to detect newer
   kernels, and a crash on reenqueueing against a destroyed dispatch queue.
 
 - Tooling: scx_pair moves off deprecated callbacks, and the deprecated
   scx_bpf_cpu_rq() kfunc is removed.
 -----BEGIN PGP SIGNATURE-----
 
 iIQEABYKACwWIQTfIjM1kS57o3GsC/uxYfJx3gVYGQUCaoOA7w4cdGpAa2VybmVs
 Lm9yZwAKCRCxYfJx3gVYGQWqAP9Sy8GwS7dRdGze/eHwYlDBt5U9ayd2ntR0Z+H1
 1Hd23AEA5kYPaEN68OgCXh/XqmFljkvEgEisXtMtw8XsZA+pHQA=
 =OwIm
 -----END PGP SIGNATURE-----

Merge tag 'sched_ext-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext

Pull sched_ext updates from Tejun Heo:
 "Most of this cycle completes the enqueue-path support for hierarchical
  sub-scheduling, which makes sub-scheduler support feature complete: a
  root BPF scheduler can now hand a cgroup subtree over to a nested
  sub-scheduler together with revocable CPU grants, and the
  sub-scheduler owns all scheduling decisions for its tasks on those
  CPUs.

  Development volume was high and a number of changes plugging holes in
  the new support landed late in the cycle. Also included are core
  scheduling fixes that were completed too late for the v7.2 release and
  are routed through this pull request.

  Sub-scheduler CPU delegation:

   - Parent schedulers now grant and revoke per-CPU capabilities
     (enqueueing, preemption, CPU frequency control) on their children,
     enforced on every path a scheduler can reach a CPU through.
     Previously only dispatching could be delegated; this lets
     sub-schedulers fully schedule their CPUs.

   - Rescue execution: a task whose scheduler doesn't have access to the
     CPUs the task needs to run on starved until the watchdog ejected
     the whole scheduler. The kernel now runs such tasks directly on a
     small bandwidth budget, turning a scheduler-killing failure into
     bounded degradation.

   - Cgroup integration: tasks migrating across a sub-scheduler boundary
     weren't re-homed to the new owner, causing wrong-scheduler
     scheduling and a use-after-free. Sub-schedulers now take over their
     cgroup subtree and receive its cgroup callbacks.

   - Arena objects now cross the kernel/BPF boundary as typed pointer
     arguments, translated transparently by the BPF tree's new arena
     argument support, replacing untyped arguments with manual
     translation.

   - scx_qmap now demonstrates full hierarchical sub-scheduling.

  Other fixes and updates:

   - Robustness improvements: the abort path is now NMI-safe, fixing
     deadlocks when errors are raised from NMI context and making
     hardlockup recovery direct. Reenqueue loops that could monopolize a
     CPU ahead of the watchdog now eject the offending scheduler, and
     stalls are blamed on the scheduler actually responsible.

   - Hardening: BPF-writable arena memory is validated before kernel
     use, and task slice and vtime writes got explicit synchronization
     rules, closing corruption vectors open to buggy or malicious
     schedulers.

   - Core scheduling: sched_ext dispatching can drop the rq lock inside
     the core-wide pick, which let interleaving selections corrupt each
     other's state and hard-hang the machine. The selection now restarts
     when the lock was released. The task ordering callback was also
     invoked with its arguments swapped, and the default ordering is
     updated to work across sub-scheduler boundaries. The fixes are
     marked for stable.

   - Other fixes headed for stable: a task init leak on fork failure
     during enable, tooling compat macros that silently failed to detect
     newer kernels, and a crash on reenqueueing against a destroyed
     dispatch queue.

   - Tooling: scx_pair moves off deprecated callbacks, and the
     deprecated scx_bpf_cpu_rq() kfunc is removed"

* tag 'sched_ext-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext: (144 commits)
  sched_ext: Drop the dead SCX_DEQ_CORE_SCHED_EXEC test in dequeue_task_scx()
  sched_ext: Make core-sched task ordering hierarchy-aware
  sched_ext: Use runnable_at for the default core-sched task ordering
  sched_ext: Fix inverted ops.core_sched_before() invocation
  sched_ext: Move the config-off sub-cap kfunc stubs into sub.c
  sched_ext: Rename balance-era identifiers to dispatch terms
  sched_ext: Drop the stale keep_prev fixup in dispatch_pick()
  sched_ext: Keep kick_sync waiting on the rq's own CPU
  sched_ext: Make SCHED_CLASS_EXT select GENERIC_ALLOCATOR
  sched_ext/scx_flatcg: Fix cvtime true-up on slice expiry
  sched_ext: Don't BUG_ON a destroyed DSQ in process_deferred_reenq_users
  sched_ext: Fix scx_bpf_dsq_move_to_local___v2 compat detection
  sched_ext: Make scx_bpf_events() read the calling scheduler's counters
  sched_ext: Drop unlocked scx_rq_clock_invalidate() from scx_root_disable()
  selftests/sched_ext: Fix flaky ddsp failure tests on busy systems
  selftests/sched_ext: Make numa idle validation race-free
  sched_ext: Fix scx_bpf_dsq_reenq___compat kfunc extern prototype
  sched_ext/scx_flatcg: expire cached hweights on weight changes
  sched_ext: Fix exit_task leak on fork failure during enable
  sched_ext: fix stale references in doc comments
  ...
2026-08-20 11:01:37 -07:00
Linus Torvalds
91ec203513 Networking changes for 7.3.
Core & protocols
 ----------------
 
  - A few steps lowering rtnl_lock dependence:
    - per-netns netdev unregistration for select SW drivers
      (e.g. veth, ipvlan, tunnels)
    - rtnl_lock-less FIB rule changes (RTM_NEWRULE and RTM_DELRULE)
    - prepare software drivers and TC qdiscs for rtnl_lock-less GET
 
  - Support BIG TCP (>64kB TSO) in UDP tunnels (vxlan, geneve).
 
  - Support buffers larger than PAGE_SIZE in devmem zero-copy API.
 
  - Improve MPTCP handling of extreme memory pressure handling,
    when out-of-order queue had to be pruned.
 
  - Report the per-group user count via RTM_GETMULTICAST.
 
  - Expose the route deletion reason in RTM_DELROUTE.
 
  - Add a SO_RIGHTS_NOTRUNC option to UNIX sockets to enable more useful
    handling of LSM denials when receiving SCM_RIGHTS messages: instead
    of truncating the message at the first blocked fd, keep every fd slot
    and store the LSM errno in the blocked slot.
 
  - IPv6 Segment Routing - support looking up the post-encap SID
    (address) in a different/specified routing table.
 
  - Support PRP RedBox (interlink) creation.
 
  - Support per-nexthop UDP dst port in VXLAN.
 
  - Continue converting getsockopt callbacks in a number of protocols
    to iov_iter.
 
 Ethernet
 --------
 
  - Marge initial CXL support for AMD/Solarflare NICs (shared branch
    with the CXL tree).
 
  - New drivers:
    - ADIN1140 10BASE-T1S MACPHY
    - Initial skeleton of Intel iXD and ZTE Dinghai drivers.
 
  - High-speed NICs:
    - AMD/Pensando:
      - support firmware flashing
    - Cisco (enic):
      - SR-IOV V2 admin channel and MBOX protocol
    - Huawei (hns3):
      - support for ethtool pfc_prevention_tout
    - nVidia/Mellanox:
      - support sharing bandwidth control across interfaces of
        the same device
    - Marvell (octeontx2-pf):
      - link RQ page pools to netdev for Netlink stats
    - Google vNIC:
      - XDP metadata support for DQ RDA
    - Microsoft vNIC:
      - support forcing full-page RX buffers
 
  - Other NICs:
    - Synopsys IP:
      - eic7700: support for eth1
    - Microchip (lan743x):
      - support for RMII interface
    - Wangxun:
      - support for ethtool -G and -C for VFs
      - add Tx timeout and PCIe error handling
    - Intel (igb/igc):
      - RSS key get/set support
      - support for forcing link speed without auto-negotiation
 
  - Switches:
    - NXP (dpaa2):
      - support bonding/LAG offload
    - Mediatek:
      - mt7530: EN7528 support
      - initial support for MT7628
    - Micrel (ksz8/9):
      - refactoring work to move towards library model
      - PTP support for KSZ8463
    - nVidia/Mellanox:
      - support rtnl-lock-less ethtool callbacks
    - Realtek:
      - rtl8366rb: use generic RTL83xx code
      - support SGMII and HSGMII for RTL8367S
 
  - PHYs:
    - Airoha:
      - EcoNet EN7528 PHY support
    - DAPU Telecom
      - DAPU Telecom DAP8211R(I) Gigabit PHY support
    - Realtek:
      - support RTL8261C_CG
      - support RTL8261D
 
 Wireless
 --------
 
  - nl80211: per-link statistics support for multi-link operation
 
  - mac80211: AQL/airtime-fairness support for multicast
 
  - Merge Peripheral Authentication Service (PAS) / TEE support
    for ath12k (shared branch with the firmware/qcom tree).
 
  - New drivers:
    - mm81x for Morse Micro Long-Range S1G devices
    - nxpwifi for NXP devices (mostly forked off from mwifiex)
 
  - Driver changes:
    - Broadcom (brcmfmac):
      - DPP support, some Cypress part update
    - MediaTek (mt76):
      - mt7928 support
      - mt7925 NAN support
      - mt7996 AP powersave improvements
    - Qualcomm (ath12k):
      - much kernel infrastructure integration work
      - AHB platform MultiPD support
    - Realtek (rt89):
      - LED support
      - RTL8922DE support
      - dual-BT coex for RTL8922D
    - Intel:
      - new FW version support
 
 Bluetooth
 ---------
 
  - HCI: add support for Shorter Connection Interval (SCI) feature.
 
  - af_bluetooth: add minimal context analysis annotations.
 
  - Driver changes:
    - Intel:
      - add Bluetooth SAR revision 2 support
      - add vendor_reset PCI sysfs for PLDR
    - Mediatek:
      - add USB IDs for MT7902 and MT7922 devices
    - Realtek:
      - add USB IDs for 8761CU and 8852BE devices
    - NXP:
      - add M.2 Bluetooth device support using pwrseq
 
 Misc
 ----
 
  - DPLL support for manual/numerical oscillator control (NCO)
    (implement in zl3073x).
 
  - MCTP support for MCTP over USB v1.1 (DMTF DSP0283).
 
  - Power-over-Ethernet: support Realtek PSE controllers.
 
  - Remove the IBM EHEA driver.
 
  - Remove tulip/xircom_cb driver.
 
 Signed-off-by: Jakub Kicinski <kuba@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEE6jPA+I1ugmIBA4hXMUZtbf5SIrsFAmqEwJ4ACgkQMUZtbf5S
 Irsegw//fmHJae525nxg3DHoXhrUz8EDDOVoLH6oyWyLQnh5bmbReAY/+oWA4m54
 3KKKO0b2rtgRvmY/7rnjAt3bjecYgCSjvZT7I+NosB0QbbBYc14PtHfYig9HffYm
 uCXfNJOk+aJ2QK4ncEvU2SjgE89Ya7cC+yARFBAwYx4zi/Qx24RB+ziOyvkQ8ksX
 atvMOZrnhwqvYUFOwnOLNHTpvdxB/ZsNwWY6iXcx6EYp9xrtPusbh3FlushWkwxH
 8cI/dNla44TcIKXAzRn0znRdgiEVmCMyHvOv7LKaOfy8P3I+knmuIf/mScYQqOEF
 T143HdXhVSBZFRtLtFKXIja/KsvCjX9lCeMn/2ak0brQDUREcacXxYbuZKDsNAAK
 zXt/+5qAcm/mO8W1gKR9Ulfli5bhFN4HKXgXMLjo5ucPtzfPxFN7HGxTiC3Cxv1v
 lSXexKaj74pNBVFmADrb5jWbq7oG+GzIdjzx3ycvm2q39Fr4nJ2SzrSPPNwc/ItQ
 IHv3tGLQKXlr8dl0+p2mDkRInmHXrawVNsB1UgN8E/jtcwT2QMwyWOV6s5G3uEDl
 a+0U/XsrPvDYBTUCRs/KaOJQGB90QkzLe9DATt159mf+rPzAX2/oCDo8xIEe+kWV
 aivP+YutFfMH/CSC9PMuvdLE2KmoPY4mibAeE4/4AYLKtJnc/yU=
 =zDto
 -----END PGP SIGNATURE-----

Merge tag 'net-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next

Pull networking updates from Jakub Kicinski:
 "One of the 'small improvements all over the place' releases for us.

  It's hard to draw any direct comparisons because summer vacations
  disrupted our patch processing (and presumably - generation) quite a
  bit.

  Quick and dirty count suggests we (Paolo and I) merged a very similar
  number of net (632) and net-next (648) patches. This is not telling
  the full story either because 1/3 to 1/2 of the net-next patches also
  *seem* like AI-driven low priority fixes, cleanups and clarifications.

  We are completely overwhelmed, of course. The glimmer of hope is that
  we secured sufficient LLM budget and access (thank you Meta!) to run
  reviews with multiple frontier models on each patch. This eliminates
  some hallucinations. That said, in terms of review, the LLMs can only
  do so much.

  The sad truth is that our APIs (especially for rare events like PCIe
  errors, timeouts etc) have always been racy, and now LLMs don't let us
  ignore that. I expect our direction for the next release will be to
  tweak the reviews a little bit more, but start shifting focus to
  letting the LLMs take care of the busy work - managing patchwork,
  automating common process complaints, editing commit messages, and
  maybe applying patches which already got "reviewed-by" tags from
  people we trust...

  Core & protocols:

   - A few steps lowering rtnl_lock dependence:
      - per-netns netdev unregistration for select SW drivers (e.g.
        veth, ipvlan, tunnels)
      - rtnl_lock-less FIB rule changes (RTM_NEWRULE and RTM_DELRULE)
      - prepare software drivers and TC qdiscs for rtnl_lock-less GET

   - Support BIG TCP (>64kB TSO) in UDP tunnels (vxlan, geneve)

   - Support buffers larger than PAGE_SIZE in devmem zero-copy API

   - Improve MPTCP handling of extreme memory pressure handling, when
     out-of-order queue had to be pruned

   - Report the per-group user count via RTM_GETMULTICAST

   - Expose the route deletion reason in RTM_DELROUTE

   - Add a SO_RIGHTS_NOTRUNC option to UNIX sockets to enable more
     useful handling of LSM denials when receiving SCM_RIGHTS messages:
     instead of truncating the message at the first blocked fd, keep
     every fd slot and store the LSM errno in the blocked slot

   - IPv6 Segment Routing - support looking up the post-encap SID
     (address) in a different/specified routing table

   - Support PRP RedBox (interlink) creation

   - Support per-nexthop UDP dst port in VXLAN

   - Continue converting getsockopt callbacks in a number of protocols
     to iov_iter

  Ethernet:

   - Merge initial CXL support for AMD/Solarflare NICs (shared branch
     with the CXL tree)

   - New drivers:
      - ADIN1140 10BASE-T1S MACPHY
      - Initial skeleton of Intel iXD and ZTE Dinghai drivers

   - High-speed NICs:
      - AMD/Pensando:
         - support firmware flashing
      - Cisco (enic):
         - SR-IOV V2 admin channel and MBOX protocol
      - Huawei (hns3):
         - support for ethtool pfc_prevention_tout
      - nVidia/Mellanox:
         - support sharing bandwidth control across interfaces
           of the same device
      - Marvell (octeontx2-pf):
         - link RQ page pools to netdev for Netlink stats
      - Google vNIC:
         - XDP metadata support for DQ RDA
      - Microsoft vNIC:
         - support forcing full-page RX buffers

   - Other NICs:
      - Synopsys IP:
         - eic7700: support for eth1
      - Microchip (lan743x):
         - support for RMII interface
      - Wangxun:
         - support for ethtool -G and -C for VFs
         - add Tx timeout and PCIe error handling
      - Intel (igb/igc):
         - RSS key get/set support
         - support for forcing link speed without auto-negotiation

   - Switches:
      - NXP (dpaa2):
         - support bonding/LAG offload
      - Mediatek:
         - mt7530: EN7528 support
         - initial support for MT7628
      - Micrel (ksz8/9):
         - refactoring work to move towards library model
         - PTP support for KSZ8463
      - nVidia/Mellanox:
         - support rtnl-lock-less ethtool callbacks
      - Realtek:
         - rtl8366rb: use generic RTL83xx code
         - support SGMII and HSGMII for RTL8367S

   - PHYs:
      - Airoha:
         - EcoNet EN7528 PHY support
      - DAPU Telecom
         - DAPU Telecom DAP8211R(I) Gigabit PHY support
      - Realtek:
         - support RTL8261C_CG
         - support RTL8261D

  Wireless:

   - nl80211: per-link statistics support for multi-link operation

   - mac80211: AQL/airtime-fairness support for multicast

   - Merge Peripheral Authentication Service (PAS) / TEE support for
     ath12k (shared branch with the firmware/qcom tree)

   - New drivers:
      - mm81x for Morse Micro Long-Range S1G devices
      - nxpwifi for NXP devices (mostly forked off from mwifiex)

   - Driver changes:
      - Broadcom (brcmfmac):
         - DPP support, some Cypress part update
      - MediaTek (mt76):
         - mt7928 support
         - mt7925 NAN support
         - mt7996 AP powersave improvements
      - Qualcomm (ath12k):
         - much kernel infrastructure integration work
         - AHB platform MultiPD support
      - Realtek (rt89):
         - LED support
         - RTL8922DE support
         - dual-BT coex for RTL8922D
      - Intel:
         - new FW version support

  Bluetooth:

   - HCI: add support for Shorter Connection Interval (SCI) feature

   - af_bluetooth: add minimal context analysis annotations

   - Driver changes:
      - Intel:
         - add Bluetooth SAR revision 2 support
         - add vendor_reset PCI sysfs for PLDR
      - Mediatek:
         - add USB IDs for MT7902 and MT7922 devices
      - Realtek:
         - add USB IDs for 8761CU and 8852BE devices
      - NXP:
         - add M.2 Bluetooth device support using pwrseq

  Misc:

   - DPLL support for manual/numerical oscillator control (NCO)
     (implement in zl3073x)

   - MCTP support for MCTP over USB v1.1 (DMTF DSP0283)

   - Power-over-Ethernet: support Realtek PSE controllers

   - Remove the IBM EHEA driver

   - Remove tulip/xircom_cb driver"

* tag 'net-next-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net-next: (1433 commits)
  net/mlx5e: do not HW-GRO coalesce small frames
  net: openvswitch: fix nf_connlabels leak in ovs_ct_init
  net: add missing ref_tracker_dir_exit() to alloc_netdev_mqs()
  net: openvswitch: fix flow mask use-after-free on flow deletion
  sctp: stop processing a packet once its association is deleted
  dpll: zl3073x: add PTP clock support
  dpll: zl3073x: add channel ToD, phase step and TIE operations
  dpll: zl3073x: scale poll interval proportionally to timeout
  ptp: vmclock: prevent read-only mappings from becoming writable
  ipv4: reject undersized MTUs in ip_do_fragment()
  bonding: initialize err for empty target lists
  net: dsa: initial support for MT7628 embedded switch
  net: dsa: initial MT7628 tagging driver
  net: phy: mediatek: add phy driver for MT7628 built-in Fast Ethernet PHYs
  dt-bindings: net: dsa: add MT7628 ESW
  net: pse-pd: realtek-pse-mcu: add UART transport
  net: pse-pd: realtek-pse-mcu: add I2C transport
  net: pse-pd: add Realtek PSE MCU core
  dt-bindings: net: pse-pd: add bindings for Realtek PSE MCU
  vsock: use sock_error() to consume sk_err after a failed connect
  ...
2026-08-20 08:16:04 -07:00
Linus Torvalds
307b9ddbbc spi: Updates for v7.3
Along with a lot of driver specific work we've got a couple of core
 features here.  The bigger one is that we've now got support for
 instantiating devices from sysfs similarly to how it's already done for
 I2C, this is used with development boards with non-enumerable expansion
 headers since SPI devices need to be manually specified.  We also have
 support for the DQS signal on higher end flash devices.
 
  - Support for intantiating devices from sysfs, useful for development
    boards with non-enumerable plugin modules, from Vishwaroop A.
  - Support for DQS in spi-mem, an additional signal used by flash
    devices to avoid clock skew from Miquel Raynal.
  - Support for more advanced SPI modes on DesignWare controllers from
    Sudip Mukherjee.
  - Changes from Jisheng Zhang to update to modern methods of specifying
    the PM callbacks.
  - Fixes for DMA mapping error handling, plus KUnit tests for this, from
    Honghui Jiang.
  - Substantial cleanup and performance work in the nxp-spi driver.
  - Support for Microchip LAN969x, Nuvoton MA35D1 QSPI, Qualcomm SA8255p
    and SA8797P, and StarFive JHB100 SFC.
 
 There is a trivial add/add conflict with the KUnit tree in their
 all_tests.config.
 -----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCgAdFiEEreZoqmdXGLWf4p/qJNaLcl1Uh9AFAmqDOgcACgkQJNaLcl1U
 h9BKxQf/QznwKffXtoEL4ZFVBckmuYXRzbkHLoX1v7upZ3QFEPG8K6rRySaXufDI
 BX3/W14CfjXANIwWCr7llQmXLBoyOq8faz8U2h/aLnBpbI4WzZELfw49Lh9JQVk5
 dxVwWQ674UFuyRC4B8QaPqMFdbn8a1CORa5Rqg/dyenxwxgsVCyl2qz+PTLVdCGH
 IWO1WW6ZISDJ5YovzMcHM4KBgut1gO6FDtz0DEj2GQC+JpCl8oHyF7E/RQ2E+Nv+
 LWb6nCLfp4z3/68zQ2UrcXa1crdGsdEIFJFwV8FUYdmv24TSD+/pClczK57h422v
 Ros+Krb7fOTE+YCXqFcZgg8IGupZsw==
 =5wbJ
 -----END PGP SIGNATURE-----

Merge tag 'spi-v7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/broonie/spi

Pull spi updates from Mark Brown:
 "Along with a lot of driver specific work we've got a couple of core
  features here. The bigger one is that we've now got support for
  instantiating devices from sysfs similarly to how it's already done
  for I2C, this is used with development boards with non-enumerable
  expansion headers since SPI devices need to be manually specified. We
  also have support for the DQS signal on higher end flash devices.

   - Support for instantiating devices from sysfs, useful for
     development boards with non-enumerable plugin modules, from
     Vishwaroop A.

   - Support for DQS in spi-mem, an additional signal used by flash
     devices to avoid clock skew from Miquel Raynal.

   - Support for more advanced SPI modes on DesignWare controllers from
     Sudip Mukherjee.

   - Changes from Jisheng Zhang to update to modern methods of
     specifying the PM callbacks.

   - Fixes for DMA mapping error handling, plus KUnit tests for this,
     from Honghui Jiang.

   - Substantial cleanup and performance work in the nxp-spi driver.

   - Support for Microchip LAN969x, Nuvoton MA35D1 QSPI, Qualcomm
     SA8255p and SA8797P, and StarFive JHB100 SFC"

* tag 'spi-v7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/broonie/spi: (132 commits)
  spi: Add KUnit coverage for DMA mapping error paths
  spi: Clear current DMA devices when unmapping a message
  spi: Move __spi_unmap_msg() before __spi_map_msg()
  spi: Fix DMA mapping ownership on partial map failure
  spi: dt-bindings: sun6i: Add compatibles for A733's SPI controllers
  spi: ma35d1-qspi: Use the existing update helper
  spi: ma35d1-qspi: Add DTR support
  spi: ma35d1-qspi: Allow several command bytes
  spi: ma35d1-qspi: Move speed setting to bus configuration
  spi: ma35d1-qspi: Remove redundant reset operation
  spi: dw: Remove shadowed dws in dw_spi_setup()
  spi: img-spfi: don't disable runtime PM on DMA deferred probe
  spi: mtk-nor: Propagate errors from IRQ request
  spi: mtk-nor: Propagate errors from optional IRQ lookup
  spi: spi-qpic-snand: Handle Macronix quad read opcode 0x6b
  spi: spi-qpic-snand: add quad mode support
  spi: spi-qpic-snand: move command mapping helper
  spi: hisi-sfc-v3xx: Propagate errors from optional IRQ lookup
  spi: meson-spifc: use devm_pm_runtime_set_active_enabled
  spi: sprd-adi: Fix probe succeeding without registering the controller
  ...
2026-08-19 09:47:41 -07:00
Linus Torvalds
abea5c3493 i2c for v7.3
Core and helpers:
 - support bus recovery with single-ended GPIOs
 - acpi: clean up resource handling
 - acpi: force ELAN1300 to 100 kHz
 - algo-bit: allow consumers to skip the optional bus test
 
 Drivers:
 - use generic bus frequency definitions in nomadik,
   octeon-core, microchip-corei2c, k1, davinci and pnx
 - i2c-gpio: support multiple buses sharing the same SCL line
 - qup: propagate clock enable failures
 - spacemit: configure SCL timing and clean up clock handling
 - amd-asf: guard against oversized firmware length
 
 qcom-geni:
 - add tracepoints for bus setup, interrupts and errors
 - use dedicated completion events for abort and reset
 - distinguish address and data NACK handling
 - cancel transfers before falling back to abort
 - simplify runtime PM and resource management
 - refactor resource and serial engine initialization
 
 DT bindings:
 - convert Altera bindings to DT schema
 - convert Axxia bindings to DT schema
 
 New support:
 - R-Car Gen5 and R-Car X5H
 - Axiado AX3005
 - Qualcomm Nord SA8797P
 - Qualcomm SA8255p
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQScDfrjQa34uOld1VLaeAVmJtMtbgUCaoA//QAKCRDaeAVmJtMt
 bvzYAQDPhRd2O6BruFmCkxKKvlrxZ2AaHQX/ZHLNqLAq/oX4NwD8CSecgva9faf/
 bC+kTGoTMBhje3FEJFufHKe2YtmJdgg=
 =u6az
 -----END PGP SIGNATURE-----

Merge tag 'i2c-7.3-part1' of git://git.kernel.org/pub/scm/linux/kernel/git/andi.shyti/linux

Pull i2c updates from Andi Shyti:
 "The main changes are support for shared SCL lines in i2c-gpio, a
  larger qcom-geni update covering tracing and transfer recovery and
  support for R-Car Gen5.

  The rest is mostly smaller driver, core and DT binding updates.

  Core and helpers:
   - support bus recovery with single-ended GPIOs
   - acpi: clean up resource handling
   - acpi: force ELAN1300 to 100 kHz
   - algo-bit: allow consumers to skip the optional bus test

  Drivers:
   - use generic bus frequency definitions in nomadik, octeon-core,
     microchip-corei2c, k1, davinci and pnx
   - i2c-gpio: support multiple buses sharing the same SCL line
   - qup: propagate clock enable failures
   - spacemit: configure SCL timing and clean up clock handling
   - amd-asf: guard against oversized firmware length

  qcom-geni:
   - add tracepoints for bus setup, interrupts and errors
   - use dedicated completion events for abort and reset
   - distinguish address and data NACK handling
   - cancel transfers before falling back to abort
   - simplify runtime PM and resource management
   - refactor resource and serial engine initialization

  DT bindings:
   - convert Altera bindings to DT schema
   - convert Axxia bindings to DT schema

  New support:
   - R-Car Gen5 and R-Car X5H
   - Axiado AX3005
   - Qualcomm Nord SA8797P
   - Qualcomm SA8255p"

* tag 'i2c-7.3-part1' of git://git.kernel.org/pub/scm/linux/kernel/git/andi.shyti/linux: (33 commits)
  i2c: core: support recovery for single-ended GPIOs
  i2c: rcar: add R-Car Gen5 support
  dt-bindings: i2c: rcar-i2c: Document R-Car X5H support
  i2c: i2c-gpio: Enhance driver for buses with shared SCL
  i2c: algo: bit: Allow to skip bit test
  i2c: qcom-geni: Add trace events for Qualcomm GENI I2C driver
  i2c: qcom-geni: trace: Add trace events for Qualcomm GENI I2C
  i2c: qup: Propagate clock enable failures
  i2c: qcom-geni: distinguish address-phase and data-phase NACK
  i2c: qcom-geni: use dedicated completions for abort and reset events
  i2c: qcom-geni: use cancel command before abort on transfer timeout
  dt-bindings: i2c: cdns: add Axiado AX3005 I2C variant
  i2c: qcom-geni: Use devm_pm_runtime_enable() for PM management
  dt-bindings: i2c: qcom,sa8255p-geni-i2c: Add compatible for Nord SA8797P
  i2c: nomadik: Use generic definitions for bus frequencies
  i2c: octeon-core: Use generic definitions for bus frequencies
  i2c: microchip-corei2c: Use generic definitions for bus frequencies
  i2c: k1: Use generic definitions for bus frequencies
  i2c: davinci: Use generic definitions for bus frequencies
  i2c: pnx: Use generic definitions for bus frequencies
  ...
2026-08-19 09:23:13 -07:00
Linus Torvalds
dfa35434d7 Locking updates for v7.3:
Futexes:
 
  - Use runtime constants for futex_hash computation
    (K Prateek Nayak, Peter Zijlstra)
 
  - Optimise the size check get_futex_key() (Sebastian Andrzej Siewior)
 
  - Avoid private hash use-after-free on final put (Felix Hoffmann)
 
  - Tell kmemleak we're not leaking __futex_queues (Peter Zijlstra)
 
 Rust integration updates:
 
  - Implement refcounted interrupt disable and SpinLockIrq for Rust
    (Boqun Feng, Heiko Carstens, Joel Fernandes, Lyude Paul)
 
  - Rust sync: add helpers for mb, dma_mb and friends;
    add generic memory barriers and use LKMM atomics
    instead of Rust atomics in the revocable code (Gary Guo)
 
  - Add abstraction and integrate synchronize_rcu() (Philipp Stanner)
 
 Lock debugging:
 
  - Add qspinlock contended_release tracepoint
    (Dmitry Ilvokhin, Peter Zijlstra)
 
  - Enable the printing of held locks of remote running tasks and print
    task CPU (Ingo Molnar)
 
  - percpu-rwsem: Annotate intentional data race in readers_active_check()
    (Sun Shaojie)
 
 Misc fixes and updates by Boqun Feng, Peter Zijlstra, Fangrui Song,
 Naveen Kumar Chaudhary and Thomas Huth.
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqC2KMRHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1gNwg//awvTQONfhPanAyTgl7CLDSlMSHdqmlyh
 Ue0/Q8Ef1Cy4jwXY2FE2A0b1VcM6cGpDPoryVdg/wMdUXRNwinzAEXmxIkRy9kve
 4LybrZwDShgLxJ7pJ6KKhgjgDiat8EdYmOwCBEE3LnP7AYhkAb8BFetA3YZJvzPa
 KfA2BRYCgvBTid6yOAuXWm55Ev92AczOBamBzTxCadcaDGtNGXtQO6LfnqiQDOav
 X5tVoANBeaQtSs1+LxE41WdNOiRoBuy0IFFvXtZRal6PZYuGGmZ5tbQvscD099em
 haVwQyzDHQrqzglv71M0KRTXvYzdGveMRg/Au1SQnuLO3V6Vd5rMQ1g7I2M9Ln0f
 Pg+tlRvQ77mLoqcgrtl0W/u0fRR4eDkiJ1pmG+98oniPwau23RdbFhC0vKFz3ikF
 WHMgk3/9TcULylgF1Tj6QLmNrBY3Vx8LBdsFjhflEw7bG4cW42D91npmXIiEDE6K
 tJc9CcaVdyE75o59z2Dtjj+qQVBlNPlfKQFXFL7p3jU/gFw2SzYuqon66X3kGmr0
 mKJ9UNJdkLdiCjxS/QiMcDeYhwJksJqxFBkH50z3Kzmo84JsSpUFkoa6GM4aSiGn
 HEwgC0Q7oOXVNIKUBYk5QaRW0HSk55hbsX2TWkvpeBYkE1zXshVZCCmgpaTSJgb5
 oFmiwrGfUjo=
 =slqC
 -----END PGP SIGNATURE-----

Merge tag 'locking-core-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull locking updates from Ingo Molnar:
 "Futexes:

   - Use runtime constants for futex_hash computation (K Prateek Nayak,
     Peter Zijlstra)

   - Optimise the size check get_futex_key() (Sebastian Andrzej Siewior)

   - Avoid private hash use-after-free on final put (Felix Hoffmann)

   - Tell kmemleak we're not leaking __futex_queues (Peter Zijlstra)

  Rust integration updates:

   - Implement refcounted interrupt disable and SpinLockIrq for Rust
     (Boqun Feng, Heiko Carstens, Joel Fernandes, Lyude Paul)

   - Rust sync: add helpers for mb, dma_mb and friends; add generic
     memory barriers and use LKMM atomics instead of Rust atomics in the
     revocable code (Gary Guo)

   - Add abstraction and integrate synchronize_rcu() (Philipp Stanner)

  Lock debugging:

   - Add qspinlock contended_release tracepoint (Dmitry Ilvokhin, Peter
     Zijlstra)

   - Enable the printing of held locks of remote running tasks and print
     task CPU (Ingo Molnar)

   - percpu-rwsem: Annotate intentional data race in readers_active_check()
     (Sun Shaojie)

  Misc fixes and updates by Boqun Feng, Peter Zijlstra, Fangrui Song,
  Naveen Kumar Chaudhary and Thomas Huth"

* tag 'locking-core-2026-08-17' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip: (44 commits)
  rust: sync: Introduce SpinLockIrq::lock_with() and friends
  rust: sync: Add SpinLockIrq
  rust: sync: Use super::* in spinlock.rs
  rust: helper: Add spin_{un,}lock_irq_{enable,disable}() helpers
  rust: Introduce interrupt module
  s390/preempt: Enable HAS_SEPARATE_PREEMPT_RESCHED_BITS
  arm64: sched/preempt: Enable HAS_SEPARATE_PREEMPT_RESCHED_BITS
  preempt: Introduce HAS_SEPARATE_PREEMPT_RESCHED_BITS
  sched: Avoid signed comparison of preempt_count() in __cant_migrate()
  sched: Remove the unused preempt_offset parameter of __cant_sleep()
  locking: Switch to _irq_{disable,enable}() variants in cleanup guards
  irq: Add KUnit test for refcounted interrupt enable/disable
  irq,spin_lock: Add counted interrupt disabling/enabling
  openrisc: Include <linux/cpumask.h> in smp.h
  preempt: Introduce __preempt_count_{sub,add}_return()
  preempt: Introduce HARDIRQ_DISABLE_BITS
  preempt: Track NMI nesting to separate per-CPU counter
  futex: Tell kmemleak we're not leaking __futex_queues
  x86/paravirt: Trace contended_release on unlock
  tracing/lock: Use TRACE_EVENT_FN() for contended_release
  ...
2026-08-18 13:07:17 -07:00
Geliang Tang
91424c4513 mptcp: remove unused data_ack from struct mptcp_ext
The data_ack and data_ack32 fields in struct mptcp_ext are no longer used
anywhere. Remove them from the structure and update mptcp_dump_mpext()
trace helper accordingly. Drop the data_ack field from the trace entry
and the corresponding output in TP_printk().

Signed-off-by: Geliang Tang <tanggeliang@kylinos.cn>
Reviewed-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Signed-off-by: Matthieu Baerts (NGI0) <matttbe@kernel.org>
Link: https://patch.msgid.link/20260812-net-next-mptcp-misc-feat-7-3-v1-2-1905a818f6cb@kernel.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-17 17:25:49 -07:00
Mickaël Salaün
bb91730f16
landlock: Add tracepoints for ptrace and scope denials
Scope and ptrace denials follow a different code path (a domain
hierarchy check) than access-right denials, so they need dedicated
tracepoints with type-specific TP_PROTO arguments.  Complete the denial
coverage with:
- landlock_deny_ptrace: ptrace access denied by a domain hierarchy
  mismatch.
- landlock_deny_scope_signal: signal delivery denied by
  LANDLOCK_SCOPE_SIGNAL.
- landlock_deny_scope_abstract_unix_socket: abstract unix socket access
  denied by LANDLOCK_SCOPE_ABSTRACT_UNIX_SOCKET.

TP_PROTO passes the raw kernel object (struct task_struct or struct
sock) for eBPF BTF access; the comm and sun_path string fields use
__print_untrusted_str() because they hold untrusted input.  Unlike the
deny_access events, these omit the blockers field: each maps to exactly
one denial type named by the event, so the bitmask would always be zero.
Like the deny_access events they carry same_exec and logged.

Audit logs the task-targeted denials with generic field names (opid,
ocomm), but a strongly typed trace event can use role-prefixed names
(tracee_pid/tracee_comm, target_pid/target_comm) that match the mainline
task-name convention (sched_process_fork's parent_comm/child_comm) and
say whose name each field holds; a bare comm= would collide across
events.  The abstract-unix-socket event reports peer_pid instead, a
tracepoint-only field with no audit counterpart.

A scope or ptrace verdict compares the subject domain against the other
party's domain, so each event also reports that other party's Landlock
domain (tracee_domain=, target_domain=, or peer_domain=); the subject
domain= alone does not let a consumer redo domain_is_scoped() or
domain_ptrace().  It is reported as a scalar ID rather than a domain
pointer: a domain object is immutable, but the other task can replace
its credential and free the domain that credential referenced, so a
stored foreign pointer could dangle before the event is consumed.  The
scalar ID also honors the tracepoint no-nullable-pointer rule, since the
other party is frequently unsandboxed.  Passing the foreign domain
hierarchy object so an eBPF consumer could walk the other party's
ancestry live would lengthen the RCU section on the shared denial path
and needs a deferred refcount put, so it is left as a future
enhancement.  The relational domain-ID field (tracee_domain,
target_domain, or peer_domain) is trace-only and is not added to audit
records, so audit's denial format is unchanged by this series.

Cc: Günther Noack <gnoack@google.com>
Cc: Justin Suess <utilityemal77@gmail.com>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Tingmao Wang <m@maowtm.org>
Link: https://patch.msgid.link/20260811094338.288094-14-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-08-17 10:17:16 +02:00
Mickaël Salaün
01ce260f5c
landlock: Add landlock_deny_access_fs and landlock_deny_access_net
Add per-type tracepoints emitted from landlock_log_denial() when an
access is denied: landlock_deny_access_fs for filesystem denials and
landlock_deny_access_net for network denials.  They use the "deny_"
prefix (rather than "check_") to mark that they fire only on a denial,
and they complement the check_rule events by making the
denial-by-absence case explicit (when no rule matches, no check_rule
event fires).

Unlike the audit records, these events fire regardless of the audit
configuration and the domain's log flags: the user's "disable logging"
intent applies to audit records, not to kernel tracing.  The logged
field records whether the domain's log policy would submit the denial to
audit; it is the decision computed once by landlock_log_denial() and
passed to both the audit and the tracing emitter, so a stateless ftrace
filter can select the audit-visible denials with logged==1.

TP_PROTO passes the denying hierarchy node, not the task's current
domain, so domain_id reports the specific node that blocked the access,
matching audit record semantics. (check_rule instead passes the current
domain, which it needs to size its per-layer array.) same_exec is also
passed explicitly because it is computed from the credential bitmask and
is not derivable from the hierarchy pointer alone.  The denial field is
named blockers to match the audit record field.

The filesystem path comes from the request's audit data.  Its type
selects which union member holds the object, exactly as
dump_common_audit_data() selects it (a path, a file's path, an ioctl
op's path, or a bare dentry); reading the wrong member would dereference
garbage, so every reachable type has an explicit case and an unexpected
one is flagged with WARN_ONCE() instead of misread.  Path-backed types
resolve via d_absolute_path() (as landlock_add_rule_fs does) and the
bare-dentry case via dentry_path_raw().

The inode number is read defensively.  A filesystem denial can carry a
negative dentry (no backing inode), for example a denied creation, so
the event mirrors the guard in dump_common_audit_data() and reports
inode 0 rather than dereferencing a NULL inode.  The sibling fs
tracepoints do not need the guard: a dentry that matches a rule during
an access check, or one opened to add a rule, always has a backing
inode.  Landlock tracepoints are reachable by unprivileged sandboxees,
so a denial on a negative dentry with the event enabled must not fault
the kernel.

Cc: Günther Noack <gnoack@google.com>
Cc: Justin Suess <utilityemal77@gmail.com>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Tingmao Wang <m@maowtm.org>
Link: https://patch.msgid.link/20260811094338.288094-13-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-08-17 10:17:16 +02:00
Mickaël Salaün
3f1f106e4c
landlock: Add tracepoints for rule checking
Merge landlock_find_rule() into landlock_unmask_layers() so rule
pointers stay inside the domain implementation while unmask checking
gets the matched rule it needs for the check_rule tracepoint.
landlock_unmask_layers() now takes a landlock_id and the domain instead
of a rule pointer.  A rename or link evaluates the same dentry against
both renamed parents, so this path now looks the rule up once per
parent; collapsing that back to a single lookup is left to a follow-up.

Emit, via the per-type wrappers unmask_layers_fs() and
unmask_layers_net(), the rights each matching rule grants at every
domain layer.  The events carry this as a dynamic per-layer array (up to
LANDLOCK_MAX_NUM_LAYERS entries) reserved from the trace ring buffer,
not the caller's stack, and rendered symbolically per layer.  A
WARN_ON_ONCE() in __trace_landlock_fill_layers() flags a rule whose
layer levels fall outside the domain range or are unsorted, a
cannot-happen case; the zero-filled slots keep the rendered output and
the array bounds safe regardless.

Setting allowed_parent2 to true for non-dom-check requests when
get_inode_id() returns false preserves the pre-refactoring behavior: a
negative dentry (no backing inode) has no matching rule, so the access
is allowed at this path component.  Before the refactoring,
landlock_unmask_layers() with a NULL rule produced this result as a side
effect; now the caller must set it explicitly.

Name the trace-only check_rule fields so each printk label equals its
ring-buffer field name and works directly as an ftrace filter: the
request field is labelled access_request= and the per-layer array is
named grants.  Values audit also logs keep audit's label (domain=,
ruleset=) so a single filter works across trace and audit.

Cc: Günther Noack <gnoack@google.com>
Cc: Justin Suess <utilityemal77@gmail.com>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Tingmao Wang <m@maowtm.org>
Link: https://patch.msgid.link/20260811094338.288094-12-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-08-17 10:17:15 +02:00
Mickaël Salaün
132d84b16b
landlock: Add landlock_enforce_domain tracepoint
The landlock_create_domain event records that a domain was created,
once, before thread-sync.  It cannot tell which threads end up enforcing
it: a successful landlock_restrict_self(2) with
LANDLOCK_RESTRICT_SELF_TSYNC applies the domain to the caller and every
eligible sibling.  Creation (the operation) and enforcement (the
per-thread outcome) are distinct.

Add landlock_enforce_domain(domain, complete, process_wide), emitted
once per thread the domain is applied to, strictly after that thread's
commit_creds(), so it fires only for a thread that is enforcing the
domain, never speculatively; an aborted operation emits none.  The
lifecycle now reads create -> enforce* -> free.

The two booleans name properties, not the implementation:
- complete: marks the single event that concludes the operation.  It
  names the outcome, the set is now enforced, not which thread
  finishes, which the contract leaves unspecified.
- process_wide: means every eligible thread of the process is
  covered.  It is set race-free by either establishing path,
  thread-sync or a single-threaded process, so
  complete && process_wide is the whole-process-enforced guarantee.

The requesting thread and source ruleset are not repeated here: they are
on create_domain (joined via domain->hierarchy->id) and on the immutable
domain->hierarchy->details.  Source ruleset means the ruleset_id and
ruleset_version recorded on create_domain, not the ruleset object, which
the caller may close before enforcement.

Cc: Günther Noack <gnoack@google.com>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Tingmao Wang <m@maowtm.org>
Link: https://patch.msgid.link/20260811094338.288094-11-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-08-17 10:17:15 +02:00
Mickaël Salaün
67567f03a4
landlock: Add create_domain and free_domain tracepoints
Add a landlock_create_domain tracepoint emitted from
landlock_restrict_self() after the new domain is created, so a consumer
can correlate the source ruleset with the resulting domain.  The
flags-only path (ruleset_fd == -1) creates no domain and emits no event.

Move the ruleset lock acquisition from landlock_merge_ruleset() to the
caller so the lock is held across both the merge and the tracepoint
emission, giving an eBPF program a consistent ruleset snapshot.  Release
it before the thread-sync: holding ruleset->lock across
landlock_restrict_sibling_threads() would deadlock a sibling blocked on
the same lock.  The event therefore fires before the (rare) thread-sync
failure path; when that path aborts the just-created domain, the
matching free_domain event fires so the create/free pair stays balanced.

Add a landlock_free_domain tracepoint that fires when a domain's
hierarchy node is freed.  The hierarchy node is the lifecycle boundary
because it represents the domain's identity and outlives the domain's
access masks, which may still be active in descendant domains.

A domain freed without ever being committed to a credential was never
visible to user space, so free_domain is suppressed for it.  This is
tracked by a new landlock_log_status value, LANDLOCK_LOG_UNCOMMITTED,
which is also the zero value so a hierarchy whose initialization failed
defaults to not observable.  A hierarchy is born UNCOMMITTED and is
promoted to LANDLOCK_LOG_PENDING (or LANDLOCK_LOG_DISABLED when logging
is off) right after its create_domain event fires; a thread-sync failure
does not reset it, so an aborted domain that already emitted
create_domain still emits the matching free_domain.  Promoting right
after the event, rather than at commit_creds() time, avoids a race: on a
successful thread-sync the sibling threads commit the new domain in
lockstep before landlock_restrict_self() returns, so the shared domain
may already have moved to LANDLOCK_LOG_RECORDED through a plain store,
and a late promotion would race that store and could unbalance the
domain allocation and deallocation audit records.

Cc: Günther Noack <gnoack@google.com>
Cc: Justin Suess <utilityemal77@gmail.com>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Tingmao Wang <m@maowtm.org>
Link: https://patch.msgid.link/20260811094338.288094-10-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-08-17 10:17:14 +02:00
Mickaël Salaün
63747c9477
landlock: Add landlock_add_rule_fs and landlock_add_rule_net tracepoints
Add tracepoints for Landlock rule addition, landlock_add_rule_fs for
filesystem rules and landlock_add_rule_net for network rules, so trace
consumers can correlate filesystem objects and network ports with their
rulesets.  Both are emitted under the ruleset lock (asserted in
TP_fast_assign) so an eBPF program reads the ruleset, including the rule
just inserted, in a consistent snapshot.

Add a version field to struct landlock_ruleset, gated on
CONFIG_TRACEPOINTS like the id field and incremented under the ruleset
lock on each successful landlock_add_rule(2), including when it only
extends an existing rule's access rights.  It fills the existing 4-byte
hole after usage, so the struct does not grow.  Pairing the ruleset ID
with the version lets a later restrict_self event record the exact
ruleset revision merged into a domain.

Resolve the filesystem rule's absolute path with d_absolute_path()
rather than the d_path() audit uses: d_absolute_path() produces
namespace-independent paths that do not depend on the tracer's chroot
state, making trace output deterministic regardless of mount namespace
configuration.  Distinguish the error cases as "<too_long>"
(-ENAMETOOLONG) and "<unreachable>" (anonymous files or detached
mounts).

Also add __trace_print_untrusted_str(), a static inline helper in the
header guarded by CREATE_TRACE_POINTS: it escapes separators, quotes,
backslashes, and non-printable bytes via string_escape_mem() so an
untrusted string (the path here, process names in later denial events)
cannot inject field separators or control characters into the ftrace
text output.

Cc: Christian Brauner <brauner@kernel.org>
Cc: Günther Noack <gnoack@google.com>
Cc: Justin Suess <utilityemal77@gmail.com>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Tingmao Wang <m@maowtm.org>
Link: https://patch.msgid.link/20260811094338.288094-9-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-08-17 10:17:14 +02:00
Mickaël Salaün
b4540a72be
landlock: Add create_ruleset and free_ruleset tracepoints
Add the first Landlock tracepoints, for ruleset lifecycle:
landlock_create_ruleset fires from the landlock_create_ruleset() syscall
handler, and landlock_free_ruleset fires in free_ruleset() before the
ruleset is freed.

These tracepoints, and the ones added by the following commits, share a
common design.  Rather than one polymorphic event distinguished by a
status field (as audit uses a shared record type with a "status="
field), each lifecycle transition and denial type gets its own event
with a type-safe TP_PROTO, giving precise ftrace filtering by event name
and type-safe eBPF access.  TP_PROTO passes the object pointer and the
fields are read from it in TP_fast_assign, so an eBPF program reads the
full object state (rules, access masks, hierarchy) via BTF from a single
pointer rather than from the flattened TP_STRUCT__entry fields.  The
whole cost is paid only when a tracer is attached; the static branch is
not taken otherwise.  Trace fields carry the bare access-right and scope
names (read_file), reusing the audit name tables; audit prepends the
category (fs.read_file), which the trace event name already conveys.
The trace header's DOC comment documents the consistency and locking
guarantees these events share.

create_ruleset needs no lock because the ruleset is not yet shared (its
file descriptor is not yet installed).  The deallocation events use the
"free_" prefix, not "drop_", because they fire when the object is
actually freed.

Add trace.c, built for CONFIG_TRACEPOINTS, which defines
CREATE_TRACE_POINTS, and extend CONFIG_SECURITY_LANDLOCK_LOG to also be
selected by CONFIG_TRACEPOINTS so the common log framework is available
to a tracepoints-only build.

Add an id field to struct landlock_ruleset, gated on CONFIG_TRACEPOINTS
and assigned from landlock_get_id_range() at creation.  Only the
tracepoints consume it (audit identifies domains, not rulesets), so it
does not exist in an audit-only build.  The Landlock ID is a stable u64
that names the ruleset across the trace stream and uses the same scheme
as audit, so a ruleset can be correlated between trace and audit
records.

Cc: Günther Noack <gnoack@google.com>
Cc: Justin Suess <utilityemal77@gmail.com>
Cc: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Steven Rostedt <rostedt@goodmis.org>
Cc: Tingmao Wang <m@maowtm.org>
Link: https://patch.msgid.link/20260811094338.288094-8-mic@digikod.net
Signed-off-by: Mickaël Salaün <mic@digikod.net>
2026-08-17 10:17:13 +02:00
Jakub Kicinski
3da8c3c8b8 Merge git://git.kernel.org/pub/scm/linux/kernel/git/netdev/net
Cross-merge networking fixes after downstream PR (net-7.2-rc8).

No conflicts.

Adjacent changes:

drivers/net/ethernet/wangxun/ngbe/ngbe_main.c
  5f3a13e0bb ("net: ngbe: fix NULL pointer dereference in non-MSI-X interrupt enabling")
  d661abdc30 ("net: ngbe: correct misleading interrupt comment")

drivers/net/ipvlan/ipvlan_main.c
  e16e960d55 ("ipvlan: inherit needed_headroom and needed_tailroom from phy_dev")
  00a40d8092 ("ipvlan: Support per-netns netdev unregistration.")

Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-13 11:00:14 -07:00
Gao Xiang
949eb95d5b
cachefiles,netfs: sunset ondemand mode
It was an effort to enhance fscache as a kernel cache for lazy
pulling (at least according to previous Incremental FS discussion [1])
and EROFS over fscache was the in-tree user of this mode.

fscache has since evolved to be netfslib-oriented, serving network
filesystem inodes via the netfs library, but EROFS never acts as a
network filesystem and we need to cache golden filesystem images rather
than individual EROFS inodes.

Since EROFS over fscache is now removed, clean up netfs/fscache/
cachefiles upstream too.

[1] https://lore.kernel.org/r/CAOQ4uxi4dzxArY24YO=+kBCK2gGoq3Ptb8WkzCqSogPgU_R3dQ@mail.gmail.com

[dh] Fixed up comments on:
https://sashiko.dev/#/patchset/20260716103030.3065561-1-dhowells%40redhat.com
https://sashiko.dev/#/patchset/20260722130218.78958-1-dhowells%40redhat.com

Signed-off-by: Gao Xiang <xiang@kernel.org>
Signed-off-by: David Howells <dhowells@redhat.com>
Link: https://patch.msgid.link/1046393.1786544127@warthog.procyon.org.uk
cc: Paulo Alcantara <pc@manguebit.org>
cc: netfs@lists.linux.dev
cc: linux-erofs@lists.ozlabs.org
cc: bpf@vger.kernel.org
cc: linux-fsdevel@vger.kernel.org
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
2026-08-12 16:24:26 +02:00
Tejun Heo
872a8f6b08 Merge branch 'master' of git://git.kernel.org/pub/scm/linux/kernel/git/bpf/bpf-next into for-7.3-arena-args
Pull bpf-next d114bb9893 ("Merge branch
'add-arena-argument-support-to-kfuncs-and-struct_ops'") to make the __arena
and __arena__nullable kfunc and struct_ops argument suffixes available. The
suffixed arguments will be used to convert sched_ext kfuncs and struct_ops
callbacks that currently pass arena pointers as scalars and rebase them by
hand.
2026-08-10 12:38:03 -10:00
Filipe Manana
12d2f44bdf btrfs: use simple booleans for log_commit field in struct btrfs_root
We are using atomic types for the log_commit array of struct btrfs_root
but all we need is simple booleans. The log_commit array elements are
always protected by the root's log_mutex, both for writes and reads, so
we can use a simple boolean. The use of atomics if from the very early
days of the log tree code where the access to the fields was not protected
by any lock.

So switch to simple booleans, which results in cheaper code and slightly
reduces the object size too.

Reviewed-by: Boris Burkov <boris@bur.io>
Reviewed-by: Jeff Layton <jlayton@kernel.org>
Signed-off-by: Filipe Manana <fdmanana@suse.com>
Signed-off-by: David Sterba <dsterba@suse.com>
2026-08-07 19:17:18 +02:00
Dmitry Ilvokhin
b359800c69 tracing/lock: Use TRACE_EVENT_FN() for contended_release
queued_spin_unlock() gates its contended_release trace call behind a
static branch, so a NOP sits on the unlock path even while the
tracepoint is disabled. Removing that requires replacing the unlock
implementation only while contended_release is enabled, which needs a
callback when the tracepoint is toggled.

Convert contended_release to TRACE_EVENT_FN() and add weak no-op
arch_contended_release_trace_reg()/arch_contended_release_trace_unreg()
hooks.

The default hooks are empty, so this is a no-op until an architecture
overrides them.

No functional change intended.

Signed-off-by: Dmitry Ilvokhin <d@ilvokhin.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Acked-by: Juergen Gross <jgross@suse.com>
Link: https://patch.msgid.link/1c2fcccfb584c075c02890c484f22c76a1948bf1.1785778551.git.d@ilvokhin.com
2026-08-07 17:58:10 +02:00
Linus Torvalds
6c68fa601b for-7.2-rc6-fixup-worker-tag
-----BEGIN PGP SIGNATURE-----
 
 iQJPBAABCgA5FiEE8rQSAMVO+zA4DBdWxWXV+ddtWDsFAmp03m0bFIAAAAAABAAO
 bWFudTIsMi41KzEuMTIsMiwyAAoJEMVl1fnXbVg7fJgP/1IK0RIWKQ0aa7Y8iDis
 xsJ+Z6qu1vlY/61WGm5s4+Q2orkNBsNl2pNI0EWOt1uKMKmvrfY5V2/fPGolaxFK
 Kj2Okw77uRCesWDhB+lCp4ldgazvge6OzqoueF5oPYmP22Upf1D3rOlxBPEpuemD
 k+nFxFQaHP82T5GC55pZ5xS6mE15tIBNv7DjTmQcTmFzGAYjocx+NaoNKcu8v3EX
 QzfqJa+AWaK2a84sodXwAtCzInD13q0BnofFYwwZ9gGJwJ1m8dC/NqBexA2fa/zF
 4Oa5S5BxwTYpghE8iFqM0LMXsaMU9g1z7jXeNT+IQDTNPk9Vom+R4+BeXfANc6rO
 W2g0zmduzisI14I9XoxPgMTJemCvxa2SLELpWc1wyvG1lHuLEgCHgqMlR/S/Y7TY
 b+yCdDiDer7NtDV4nuxtsI9ZbF+xSxXI9MOaPMnU9IRk/2eMh0mKorxd1PtTc6rY
 y2y/pxm3R2w4TZ2L5sB9qIvgWngH+fNJVsprkQF1ZRrlKOHIphXQ9a/hIM3AqRlF
 b7aWhs/2MJGmZTzVKFSPzKOnxcVuioG08R2SWwXtJrUbQiSQQp2IOriMockJFEtV
 pyHOcga7+D73eop/oc6GMTGRaeG2A/zBWdzBj7DFqti36lGcaVHewU0ho+P74du9
 dnmWDKC+acbp+vdqWPITcCqZ
 =Ier0
 -----END PGP SIGNATURE-----

Merge tag 'for-7.2-rc6-fixup-worker-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux

Pull Btrfs Fixes 2: Electric Boogaloo from David Sterba:
 "This brings back the fixup worker infrastructure.

  It's a mechanism to detect pages/folios that are marked dirty without
  filesystem knowledge and require COW fixup. The consequence of not
  doing so is silent data loss.

  The first patch covers the scenarios in detail, also reflecting folio
  API port and subpage block size support added in recent years. The
  original fixup worker was only for pages.

  The patch is relatively big, half of the code is debugging and support
  code, the rest is the core design around the detection and fix.

  The second patch handles an unlikely case when there's work left
  during unmount"

* tag 'for-7.2-rc6-fixup-worker-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux:
  btrfs: flush the fixup workers during close_ctree
  btrfs: trigger cow fixup via dirty_folio()
2026-08-06 13:29:15 -07:00