Commit Graph

1482885 Commits

Author SHA1 Message Date
Dave Airlie
94f69bfa18 amd-drm-fixes-7.3-2026-09-17:
amdgpu:
 - SMU 14.x fix
 - DC IRQ fix
 - Runtime PM fix for P2P
 - RAS fix
 - PCIe reporting fix
 - DCN 6 fix
 - Device removal fix
 - DC MALL fix
 
 amdkfd:
 - GC 12.x fixes
 - Boundary checks
 - Mapping clear fix
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQQgO5Idg2tXNTSZAr293/aFa7yZ2AUCaqxIpQAKCRC93/aFa7yZ
 2HcZAP9ZA5cERbp6QfT0a1tT3kDoMP02BKev5/XUNWEJjdgOYAEAsCj7UFE4oKjb
 0996gK/lJvqq9lOgFTJnwl/z2jwytA0=
 =Ye34
 -----END PGP SIGNATURE-----

Merge tag 'amd-drm-fixes-7.3-2026-09-17' of https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes

amd-drm-fixes-7.3-2026-09-17:

amdgpu:
- SMU 14.x fix
- DC IRQ fix
- Runtime PM fix for P2P
- RAS fix
- PCIe reporting fix
- DCN 6 fix
- Device removal fix
- DC MALL fix

amdkfd:
- GC 12.x fixes
- Boundary checks
- Mapping clear fix

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260917201213.3880863-1-alexander.deucher@amd.com
2026-09-18 11:24:43 +10:00
Dave Airlie
cd011719ba Merge tag 'drm-intel-fixes-2026-09-17' of https://gitlab.freedesktop.org/drm/i915/kernel into drm-fixes
drm/i915 fixes for 7.3-rc4:
- Revert a commit touching registers that don't necessarily exist
- Check for negative numbers before passing to BIT()

Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Jani Nikula <jani.nikula@intel.com>
Link: https://patch.msgid.link/3f86e0ede95fb3d52053934ac43c5271812428f5@intel.com
2026-09-18 11:08:19 +10:00
Dave Airlie
c24f824f0b Couple shrinker related fixes plus a series of patches fixing several
xe_mmio_gem issues around fault handler and destroy path.
 -----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCgAdFiEEbSBwaO7dZQkcLOKj+mJfZA7rE8oFAmqr550ACgkQ+mJfZA7r
 E8p3UggAtIEI+BFbK7zeELK3zEtO85WQnIvUP/PprLu9Ga7Vvl+quXxZrWCHYcIE
 VwcK/5y2BGVKBmR1Vwm+PNBmtlbrNvIPGTn9u0oXB1bWVKqPT5o5MjY7lxdWhK95
 MkLFo7rsQM791DAEUeo6IzbdHd6K9k2At8Yzz5+1zt97zxGmXSNpGRxV6Ee8/crU
 e/B1cQSfY8I4sjhAormTLZ13M6vRh1Yoy5P21UdCqIU/Eq/PjpkW2bcmM6+8K+qO
 2a4y2CHIQnd+2bR4nEnuHYIa6DVuXEZIHh5v/8amenuZqBgVj7OcxUKiiZ4TnhMY
 bGe1UFyy0OWyEJ/X4ddGfCggpiFi6w==
 =0FwX
 -----END PGP SIGNATURE-----

Merge tag 'drm-xe-fixes-2026-09-17' of https://gitlab.freedesktop.org/drm/xe/kernel into drm-fixes

Couple shrinker related fixes plus a series of patches fixing several
xe_mmio_gem issues around fault handler and destroy path.

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Rodrigo Vivi <rodrigo.vivi@intel.com>
Link: https://patch.msgid.link/aqvoKaPuLLPBGayA@intel.com
2026-09-18 11:08:01 +10:00
Francis Marlou Pacaro
2ac2fe765e drm/amd/display: fix MALL hysteresis timer underflow at high refresh rates
dcn30_apply_idle_power_optimizations() derives the MALL frame cache
hysteresis timer with

	tmr_delay = (uint32_t)(div_u64(..., denom) - 64LL);

div_u64() returns a u64, so when the quotient is smaller than 64 the
subtraction wraps instead of going negative and tmr_delay ends up huge.
The loop that follows tries to squeeze it into the 6 bit register field
by doubling denom, but that only makes the quotient smaller, so tmr_delay
can never converge.  tmr_scale is bumped past 3 and the function gives up
with

	/* Delay exceeds range of hysteresis timer */
	ASSERT(false);

even though the requested delay is too *short* to encode, not too long.

With mall_additional_timer_percent left at its default of 0, the quotient
drops below 64 once the refresh rate used for the calculation goes above
~243 Hz.  Every DCN 3.0 display above that loses MALL static screen
entirely and splats a WARN once per boot.  Reproduced on Navi 23
(RX 6600) driving 1920x1080, resetting /sys/kernel/debug/clear_warn_once
between modes:

	refresh   MALL       ASSERT
	144 Hz    enabled    no
	240 Hz    enabled    no
	280 Hz    skipped    yes
	360 Hz    skipped    yes

Commit 3bb68cec4d ("drm/amd/display: Add Overflow check to skip MALL")
already covered the other end of the range, where a large stutter period
makes the delay too long to encode.  Cover the short end by clamping to
0, which selects the shortest hysteresis the register can express,
65.28us * 64 = ~4.18ms.  That is marginally longer than what the formula
asks for at these refresh rates, and erring long is the safe direction:
it only delays MALL entry, it can never enter early.

The numerator does not change between iterations, only denom does, so
compute it once and keep both call sites inside 100 columns.

The genuinely out of range case at very low refresh rates still reaches
the ASSERT, which is where it belongs.

Fixes: 52f2e83e2f ("drm/amdgpu/display: add MALL support (v2)")
Signed-off-by: Francis Marlou Pacaro <pacaro.francis.marlou.n@gmail.com>
Reviewed-by: Leo Li <sunpeng.li@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 387550e53e1405f1f960b62b22f8783db17c8e1d)
2026-09-17 11:59:24 -04:00
Chengjun Yao
5155002b03 drm/amdgpu: fix rmmio iounmap skipped on device removal
amdgpu_pci_remove() calls drm_dev_unplug() before fini_sw(), so
drm_dev_enter() is already false there and the iounmap() guarded by it
is skipped. This .remove path runs on both hot-unplug and plain rmmod,
so the register BAR ioremap mapping leaks one instance per unload.

Unmap rmmio unconditionally (guard only on non-NULL) and drop the now
unused idx.

Fixes: 62d5f9f711 ("drm/amdgpu: Unmap MMIO mappings when device is not unplugged")
Signed-off-by: Chengjun Yao <Chengjun.Yao@amd.com>
Reviewed-by: Asad Kamal <asad.kamal@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit dd6f86a97260e5207d3329ad03aa89fdad61b1e6)
Cc: stable@vger.kernel.org
2026-09-17 11:58:48 -04:00
Mario Limonciello
7f9caa70ae drm/amdgpu: Skip KFD mapping clear before initialization
amdgpu_amdkfd_clear_kfd_mapping() assumes that a non-NULL kfd_dev
has a fully populated node array. This is not true when KFD device
initialization fails after probe.

For example, kgd2kfd_device_init() sets num_nodes before checking
PCIe atomics support. On Polaris systems without the required atomics,
it returns before allocating nodes[0], but the kfd_dev remains attached
to the amdgpu device. A later GPU reset then dereferences nodes[0]->id.

Require the authoritative KFD initialization flag before walking the
node array, matching the existing KFD reset and teardown paths.

Fixes: 70cadefcc6 ("drm/amdgpu: unmap all user mappings of framebuffer and doorbell before mode1 reset")
Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5833
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 4ac1835823c47903fbb278bbf474773c46f59edc)
Cc: stable@vger.kernel.org
2026-09-17 11:58:04 -04:00
Srinivasan Shanmugam
0d2f4cfa56 drm/amd/display: Fix NULL dereference in dcn50/dcn60 init_hw
dc->clk_mgr is checked for NULL earlier in dcn50_init_hw() and
dcn60_init_hw(), but dcn50_initialize_min_clocks() and
dcn401_initialize_min_clocks() are called without any guard,
causing Smatch to report potential NULL dereferences.

Guard both call sites with the same pattern used throughout
both functions:

  if (dc->clk_mgr && dc->clk_mgr->funcs)

Also fix dcn50_initialize_min_clocks() which calls
get_dispclk_from_dentist without checking the function pointer,
unlike the dcn401 equivalent which guards that call.

Fix kernel-doc in dcn60_hwseq.c by adding missing parameter descriptions
for @probe in dcn60_update_probe_status() and @type in
is_probe_measurement_type_for_hubbub().

Fixes: 7f7d7ea1fa ("drm/amd/display: Add new sources for DCN6")
Reported-by: Dan Carpenter <error27@gmail.com>
Cc: Aurabindo Pillai <aurabindo.pillai@amd.com>
Cc: Ivan Lipski <ivan.lipski@amd.com>
Cc: Dan Wheeler <daniel.wheeler@amd.com>
Cc: Roman Li <roman.li@amd.com>
Cc: Alex Hung <alex.hung@amd.com>
Cc: Tom Chung <chiahsuan.chung@amd.com>
Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 325c9a827cdd748e126eafeadaffc556204773d2)
2026-09-17 11:57:29 -04:00
David Francis
8ee521b8b1 drm/amdkfd: Avoid integer underflow in EOP ring size calculation.
The low 6 bits of cp_hqd_eop_control store the base-2 logarithm
of the EOP ring size. This was calculated as

order_base_2(q->eop_ring_buffer_size / 4) - 1

But order_base_2 can in theory return 0, so this could underflow
(although in practice the ring buffer size cannot be less than 4096).

Change this to

order_base_2(q->eop_ring_buffer_size / 8)

using properties of logarithms.

Also add to the above comment to make the mathematics more clear.

Reviewed-by: Kent Russell <kent.russell@amd.com>
Signed-off-by: David Francis <David.Francis@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit f0f43fcf8b2b3a924cad9444340921c96ed5f634)
Cc: stable@vger.kernel.org
2026-09-17 11:56:45 -04:00
David Francis
c883d0a132 drm/amdkfd: Avoid integer underflow with ffs in EOP ring size calc
The low 6 bits of cp_hqd_eop_control store the base-2 logarithm
of the EOP ring size. This was calculated as

ffs(q->eop_ring_buffer_size / sizeof(unsigned int)) - 1 - 1

But ffs can in theory return 1 or 0, so this could underflow
(although in practice the ring buffer size cannot be less than 4096).

Change this to

ffs(q->eop_ring_buffer_size / sizeof(unsigned int) / 4)

using properties of logarithms.

Reviewed-by: Kent Russell <kent.russell@amd.com>
Signed-off-by: David Francis <David.Francis@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 4f18c56630383c14bfc6b2d65f88f2f895d2121a)
Cc: stable@vger.kernel.org
2026-09-17 11:56:02 -04:00
Mario Limonciello
04de4007d3 drm/amdgpu: Fix GPU PCIe link capability reporting
Commit eb53125a7a ("drm/amd: Add dedicated helper for
amdgpu_device_find_parent()") made amdgpu_device_gpu_bandwidth() query
the first device outside the dGPU. That is the host side of the
physical link, not the GPU side.

As a result, the ASIC and platform capability masks can both be based
on the host port. drm_amdgpu_info_device then exposes the host
capabilities to userspace, such as Gen5 x16 for a Gen4 x8 GPU.

Cache both ends of the physical link during device initialization.
Use link_dev for the GPU capability and link_partner for the platform
capability and _PR3 detection.

Reported-by: "Marek Olšák" <maraeo@gmail.com>
Closes: https://lore.kernel.org/amd-gfx/CAAxE2A4VhsAzzO1QjBjUg+NgnbD04ZzMyN6xsUJxjKJHH6hxiw@mail.gmail.com/
Suggested-by: Lijo Lazar <lijo.lazar@amd.com>
Fixes: eb53125a7a ("drm/amd: Add dedicated helper for amdgpu_device_find_parent()")
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 7ea6a47224e2c6e89a3a682d7fbaace4817a55aa)
Cc: stable@vger.kernel.org
2026-09-17 11:55:14 -04:00
Dmitriy Chumachenko
723d4dc628 drm/amdgpu: check ras and obj before dereference
nbio_v7_9_handle_ras_controller_intr_no_bifring() dereferences ras and obj
without checking either for NULL. Both amdgpu_ras_get_context() and
amdgpu_ras_find_obj() can return NULL, e.g. during the window between
adev->nbio.ras being set (early in amdgpu_ras_init(), by design, to
enable the fatal-error interrupt as soon as possible) and the PCIE_BIF
ras object actually being created in RAS late_init. Any interrupt in that
window crashes in hard-IRQ context.

This is analogous to commit d190b459b2 ("drm/amdgpu: the warning
dereferencing obj for nbio_v7_4"), which fixed the same issue in the
nbio_v7_4 handler.

Found by Linux Verification Center (linuxtesting.org) with SVACE.

Fixes: 7692e1ee24 ("drm/amdgpu: add RAS fatal error handler for NBIO v7.9")
Reviewed-by: Tao Zhou <tao.zhou1@amd.com>
Signed-off-by: Dmitriy Chumachenko <Dmitry.Chumachenko@cyberprotect.ru>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit c7071767a50a32ed727cf800ac84372429e3b4b3)
2026-09-17 11:54:31 -04:00
Mike Lothian
636139603b drm/amdgpu: hold a runtime PM reference for P2P dma-buf attachments
amdgpu_dma_buf_map() adds VRAM to the allowed domains for a peer2peer
attachment.  GTT is only a fallback placement when VRAM is preferred, so
ttm_bo_validate() migrates the buffer from GTT into VRAM.  While the
exporting device is runtime suspended its SDMA rings are down and the
move fails:

  amdgpu: Move buffer fallback to memcpy unavailable

An importer on a second GPU reaches this holding no runtime PM
reference on the exporter, e.g. a compositor on the APU submitting a
frame that references a buffer exported by an idle dGPU:

  amdgpu_cs_ioctl -> amdgpu_cs_parser_bos -> amdgpu_cs_bo_validate
    -> ttm_bo_validate -> amdgpu_bo_move -> dma_buf_map_attachment
      -> amdgpu_dma_buf_map -> ttm_bo_validate -> amdgpu_bo_move

Pinning a dma-buf into VRAM has the same requirement, which
commit 030631e97b ("drm/amdgpu: revert "take runtime pm reference
when we attach a buffer" v2") called out as the one case that would
need the reference back.

Take it in attach and drop it in detach.  pm_runtime_get_if_active()
never resumes the device, so it cannot deadlock against the reservation
taken during resume, which is why the old pm_runtime_get_sync() had to
go.  If the device is not active, clear peer2peer instead: the buffer
then stays in GTT, which remains accessible while the GPU is powered
down.  If runtime PM is disabled, take a plain reference so the put in
detach stays balanced.

Fixes: 030631e97b ("drm/amdgpu: revert "take runtime pm reference when we attach a buffer" v2")
Suggested-by: Christian König <christian.koenig@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Mike Lothian <mike@fireburn.co.uk>
Assisted-by: Claude:Opus-5 [Claude Code]
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 062ff15e30a48d14fb7d7558eba84f8dc97197f0)
Cc: stable@vger.kernel.org
2026-09-17 11:54:05 -04:00
Vladimir Marioukhine
5f28bb1c2c drm/amdkfd: implement restore_mqd callbacks for GFX12/12.1
kfd_mqd_manager_v12.c (GFX 12.0) and kfd_mqd_manager_v12_1.c (GFX 12.1)
do not implement restore_mqd callbacks, leaving the function pointers
NULL and causing CRIU restore to return -EOPNOTSUPP on GFX12.

Implement restore_mqd for both compute and SDMA queues in
kfd_mqd_manager_v12.c and kfd_mqd_manager_v12_1.c, modeled after the
GFX 11 implementation with the following improvements:
- update cp_mqd_base_addr_lo/hi to the newly allocated MQD address,
  fixing a pre-existing gap shared with v11 where the in-MQD copy
  still pointed at the old checkpoint-time address after restore
- memset the full allocation before memcpy for compute queues to avoid
  stale data in the GTT sub-allocator tail; SDMA MQDs use sizeof(*m)
  since they are packed at mqd_size stride in a shared BO

checkpoint_mqd registration is deferred to a follow-up patch that also
implements get_checkpoint_info, so that checkpoint and restore are
enabled together as a complete and testable unit.

Note: GFX12.1 restore handles XCC0 only. Multi-XCC CRIU restore is
currently unreachable due to a separate validation issue in
kfd_criu_restore_queue(). A pr_warn_once() is emitted if a multi-XCC
device is encountered.

Signed-off-by: Vladimir Marioukhine <Vladimir.Marioukhine@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit b1f9601237d050f5df478464cf51bf1fff29a256)
Cc: stable@vger.kernel.org
2026-09-17 11:53:24 -04:00
Leo Li
63e19ef3dd drm/amd/display: Atomize IRQ register read/modify/write ops
[Why]

The OTG_GLOBAL_SYNC_STATUS register controls various HW IRQ sources for
the output timing generator (OTG). VUPDATE_NO_LOCK is one of them.

To enable the IRQ, driver sets the VUPDATE_NO_LOCK_EN bit in the
GLOBAL_SYNC_STATUS register.

To ack the IRQ after it fires, the driver sets the VUPDATE_NO_LOCK_CLEAR
bit in the same GLOBAL_SYNC_STATUS register.

The bit sets are done through read/modify/write operations, which are
not atomic. Thus, the following race is possible:

    Thread A:                       IRQ handler:
                                    *HW IRQ fires*
    # IRQ disable
    val = read(GLOBAL_SYNC_STATUS)
    unset(val, VUPDATE_NO_LOCK_EN)
    write(val, GLOBAL_SYNC_STATUS)
                                    # ACK reads VUPDATE_NO_LOCK_EN unset
                                    val1 = read(GLOBAL_SYNC_STATUS)
                                    set(val1, VUPDATE_NO_LOCK_CLEAR)
    # IRQ enable
    val = read(GLOBAL_SYNC_STATUS)
    set(val, VUPDATE_NO_LOCK_EN)
    write(val, GLOBAL_SYNC_STATUS)
                                    # BAD! clears VUPDATE_NO_LOCK_EN
                                    write(val1, GLOBAL_SYNC_STATUS)

Regarding the tagged Fixes: change, it appears the change made this race
more likely to occur. Since VUPDATE_NO_LOCK is now the sole IRQ source
for vblank handling, a single race on high refresh panels can lead to a
time out.

[How]

The GLOBAL_SYNC_STATUS register is only one example, other IRQ control
registers also share the same scheme. On top of GLOBAL_SYNC_STATUS,
let's clean up those as well.

To keep things simple, Let's atomize the IRQ rmw ops via a single
driver-wide spinlock. Due to the small scope of this lock, it is
unlikely to cause noticeable overhead on top of all the existing locking
within the IRQ set/handle paths.

Since DM is responsible for locking, wrap dc_interrupt_set/ack with the
spinlock in the new amdgpu_dm_irq_set/ack functions. Migrate/drop all
references in DM to dc_interrupt_set/ack to use amdgpu_dm_irq_set/ack
instead.

Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5616
Fixes: c87e6635d2 ("drm/amd/display: consolidate DCN vblank/flip handling onto vupdate_no_lock")
Reviewed-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Leo Li <sunpeng.li@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 70de0a0216583a53c946155f8c8adedfdca6b4e7)
Cc: stable@vger.kernel.org
2026-09-17 11:52:06 -04:00
Kevin Wang
9413959fa9 drm/amd/pm: report energy accumulator for smu 14.0.3
add energy accumulator on pmfw 0x00685000 and above version.

Signed-off-by: Kevin Wang <kevin.wang@amd.com>
Reviewed-by: Kenneth Feng <kenneth.feng@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 4aa733ab15b303e2a40e2985ac21a0e01f24cc4a)
2026-09-17 11:51:15 -04:00
Dave Airlie
b15e54761a Merge tag 'drm-msm-fixes-2026-09-16' of https://gitlab.freedesktop.org/drm/msm into drm-fixes
Fixes for v7.3-rc4:

DT:
- Corrected indentation

Core:
- Marked fbdev as system memory

GPU:
- Fixed autosuspend cleanup on teardown
- a750: fix timestamps
- Increase GMU fw init timeout
- Misc fixes/cleanups

DPU:
- Fixed clock rounding, unbreaking newest platforms
- Cleared pending flush state

DP:
- Skip PUSH_IDLE when link was never enabled
- Fixed bandwidth checks

HDMI:
- Fixed runtime PM cleanup on probe failure

Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Rob Clark <rob.clark@oss.qualcomm.com>
Link: https://patch.msgid.link/CACSVV021rZsjmiPEyR_LR7L=k7=DRV0QdsA50BF_Gd0s2DXaCw@mail.gmail.com
2026-09-17 19:05:21 +10:00
Raag Jadav
f0e9f963a3
drm/xe/i2c: Disable IRQ on unbind
Currently, struct xe_i2c is freed before SGUnit IRQ is disabled in unbind
path, leaving a potential UAF in case I2C IRQ is hit during this small
window. Explicitly disable I2C IRQ in xe_i2c_remove() and fix this.

Fixes: 0bb78ce099 ("drm/xe/i2c: Wire up reset/postinstall for I2C IRQ")
Signed-off-by: Raag Jadav <raag.jadav@intel.com>
Reviewed-by: Heikki Krogerus <heikki.krogerus@linux.intel.com>
Link: https://patch.msgid.link/20260911121547.2407261-1-raag.jadav@intel.com
Signed-off-by: Matt Roper <matthew.d.roper@intel.com>
(cherry picked from commit 8ba5c8b8ab3fd362267c11df2cd5a90ee46f6e24)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-16 08:46:16 -04:00
Nemesa Garg
a26204be58 Revert "drm/i915/display: Clear SEL_FETCH_PLANE_CTL on plane disable"
This reverts commit 7f1172a2ac.

This commit replaced the crtc_state->enable_psr2_sel_fetch guard in
icl_plane_disable_sel_fetch_arm() and i9xx_cursor_disable_sel_fetch_arm()
with HAS_PSR2_SEL_FETCH(). This is a display version check and
says nothing about the pipe, so every plane and cursor disable on a
display 12+ platform started writing SEL_FETCH_PLANE_CTL() /
SEL_FETCH_CUR_CTL(), including on pipes that do not implement them.
It shows up as an unclaimed register access on pipes driving HDMI where
selective fetch was never enabled.

The stale selective fetch enable bit that commit addressed is handled
in the next patch.

Fixes: 7f1172a2ac ("drm/i915/display: Clear SEL_FETCH_PLANE_CTL on plane disable")
Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16876
Signed-off-by: Nemesa Garg <nemesa.garg@intel.com>
Reviewed-by: Jouni Högander <jouni.hogander@intel.com>
Signed-off-by: Suraj Kandpal <suraj.kandpal@intel.com>
Link: https://patch.msgid.link/20260909110332.3528029-2-nemesa.garg@intel.com
(cherry picked from commit d393529394167e0f5f706657eebe84d8529ce4fc)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
2026-09-16 12:55:22 +03:00
Luca Coelho
ee415ce8cb drm/i915/display: check configuration index before shifting
The calc_allowed_config_filter() function passes the return value of
iter_pos_to_idx() directly to BIT(), but the helper can return -1 for
an invalid iterator.

The iterator already rejects negative indices before doing a
configuration, so this should not matter in normal flows.  In any
case, for robustness, check the index explicitly and warn if it is
negative, avoiding an undefined shift.

Fixes: 39e30bdf2f ("drm/i915/dp_link_caps: Add link configuration iterator")
Reviewed-by: Imre Deak <imre.deak@intel.com>
Link: https://patch.msgid.link/20260908100659.113555-1-luciano.coelho@intel.com
Signed-off-by: Luca Coelho <luciano.coelho@intel.com>
(cherry picked from commit fe05cb9b9fb0ecc10409c4c6133257214b6cd8c8)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
2026-09-15 11:14:13 +03:00
Krzysztof Kozlowski
a15fac810c dt-bindings: display/msm: Use consistent indentation in the example
Correct indentation in the examples to consistent 2- or 4-spaces
indentation to fix dt-check-style warnings ("example 0
[indent-consistent] indent mismatch ...").  Preferred is 4-spaces, but
re-indenting entire example just for that is too much churn.

Signed-off-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/753054/
Link: https://lore.kernel.org/r/20260913123331.100293-4-krzysztof.kozlowski@oss.qualcomm.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-15 02:09:28 +03:00
Guangshuo Li
f4fae975db drm/msm/hdmi_phy: fix runtime PM cleanup on probe failure
msm_hdmi_phy_probe() enables runtime PM before enabling the PHY
resources and initializing the PLL, but failures from either operation
return without calling the matching pm_runtime_disable().

The remove path disables runtime PM, but it is not called when probe
fails. As a result, runtime PM remains enabled after an unsuccessful
probe.

Route failures after pm_runtime_enable() through a common error path
and disable runtime PM before returning.

This issue was found by manual code inspection.

Fixes: 15b4a45238 ("drm/msm/hdmi: Create a separate HDMI PHY driver")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Reviewed-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/753043/
Link: https://lore.kernel.org/r/20260913085814.1509352-1-lgs201920130244@gmail.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-15 02:09:28 +03:00
Dmitry Baryshkov
2028280686 drm/msm/dsi: round the byte clock rate after reparenting to the PHY PLL
DSI 6G v2.9 hosts (SM8650, SM8750, Kaanapali, etc.) reparent the byte and
pixel RCGs to the DSI PHY PLL at runtime from
dsi_link_clk_set_rate_6g_v2_9(), after the PHY has been enabled. However
dsi_calc_clk_rate_6g() runs earlier, in order to compute the bit clock
request for the PHY. At that point the byte RCG still has its reset
parent (XO), so clk_round_rate() returns a bogus rate, which then ends up
in the PHY bit clock request and the PLL gets programmed to a wrong
frequency, breaking the panel.

Move the rounding to dsi_link_clk_set_rate_6g(), which is called after
the RCGs have been reparented to the PLL. Storing the rounded rate at
this point still makes later link_clk_set_rate() calls no-ops in the
CCF. Derive the byte interface clock rate from the rounded byte clock
rate, otherwise it would keep requesting the idealized rate and
retrigger the PLL on every transfer.

Reported-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Reported-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>
Fixes: 6cd33b6f41 ("drm/msm/dsi: round 6G byte clock rate to the PLL-achievable value")
Assisted-by: LLM
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Reviewed-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com>
Tested-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com> # SM6115P J606F
Tested-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Reviewed-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/750496/
Link: https://lore.kernel.org/r/20260903-fix-eliza-dsi-v1-1-3474a6c9f2e0@oss.qualcomm.com
2026-09-15 02:09:28 +03:00
Shuicheng Lin
985862be16
drm/xe/shrinker: Take a runtime PM ref before shrinking non-system memory
__xe_shrinker_walk() walks the SYSTEM and TT LRUs without a runtime PM
reference.  Shrinking a bo outside system memory invalidates its GPU
mappings, which needs the device resumed, so while it is runtime
suspended the page table zap trips an assert and the TLB invalidation
returns -ENODEV:

  WARNING: drivers/gpu/drm/xe/xe_bo.c:770 at xe_bo_move_notify+0x1fc/0x450 [xe]
   xe_bo_shrink+0x20f/0x2b0 [xe]
   __xe_shrinker_walk+0x174/0x410 [xe]
   xe_shrinker_scan+0x10c/0x1e0 [xe]
   do_shrink_slab+0x176/0x7e0
   drop_caches_sysctl_handler+0x9c/0xf0

Take a reference before walking a memory type other than XE_PL_SYSTEM
and stop there if it cannot be acquired.  Reuse the shrinker's existing
acquire path, which resumes the device directly where reclaim allows
that and otherwise queues the PM worker for a later scan.  Stop the walk
once the scan target is met, so a satisfied scan does not wake the
device.  System memory is still reclaimed while the device is suspended.

Gate this on xe_device_is_l2_flush_optimized(), the same condition under
which xe_bo_trigger_rebind() issues the invalidation for a non-fault-mode
vm, so reclaim is unaffected elsewhere.  The System CCS copy already has
its own reference in xe_bo_shrink().

Only a non-fault-mode vm can reach this, since a fault-mode vm requires
LR mode and that holds a runtime PM reference for the vm's lifetime.

Reproduced with igt@xe_madvise@dontneed-before-exec while the GPU is
runtime suspended.

v2: simplify needs_rpm check. (Matt)
    retarget Fixes tag since the issue occurs with the non-fault-mode
    path added by 4e7ebff69a.
v3: handle this in xe_shrinker.c instead of xe_bo.c (Thomas)
v4: stop the walk once the scan target is met. (Sashiko)
v5: rebase on the freed page accounting fix. (Sashiko)
v6: reuse the shrinker acquire path so runtime pm can be resumed
    directly instead of always queueing a worker. (Thomas)
v7: replace xe_pm_runtime_put() with xe_shrinker_runtime_pm_put(). (Thomas)

Fixes: 4e7ebff69a ("drm/xe/xe3p_lpg: flush shrinker bo cachelines manually")
Assisted-by: Claude:claude-opus-5
Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Link: https://patch.msgid.link/20260909162102.1097006-3-shuicheng.lin@intel.com
Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
(cherry picked from commit 628f92b28bf4c371c10207daf6fc4caee0c0db2e)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Shuicheng Lin
3c90e42a01
drm/xe/shrinker: Return the freed page count through a parameter
__xe_shrinker_walk() and xe_shrinker_walk() return either the number of
pages freed or a negative error, so the two cannot be reported at once.
On error the pages already freed are dropped, and since xe_shrinker_scan()
only accumulates non-negative returns while *scanned is updated by
pointer, the shrinker tells mm that it scanned without freeing.

Accumulate the count into a caller-provided counter and return only the
status, so an error no longer discards what the walk had freed.

Fixes: 00c8efc318 ("drm/xe: Add a shrinker for xe bos")
Assisted-by: Claude:claude-opus-5
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Link: https://patch.msgid.link/20260909162102.1097006-2-shuicheng.lin@intel.com
Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
(cherry picked from commit d7aac1a0235a6ce41e30cec385e2db8c33dad12d)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Ilia Levi
d0c0952878
drm/xe/mmio_gem: fix destroy flow
xe_mmio_gem_destroy() currently frees the GEM object directly, bypassing
reference counting.  Since existing VMAs hold a reference and the fault
handler accesses the object through vma->vm_private_data, this is
use-after-free.  Additionally, nothing prevents the fault handler from
installing PTEs to the real MMIO after destroy.

Fix this with proper synchronization and refcounting. Also, do not set
vm_pgoff to zero. Many DRM drivers do this because helpers like
dma_mmap_pages() interpret vm_pgoff as an intra-buffer page offset;
leaving the DRM fake offset there would break these helpers.
Those drivers can get away with zeroing it because they map eagerly -
all PTEs are established before mmap returns, so vm_pgoff is never
consulted again. Our driver does not use such helpers and the newly
introduced call to drm_vma_node_unmap() relies on vm_pgoff being untouched.

v2: (Matt Auld)
- use dma_resv lock to serialize fault handler with destroy
- SIGBUS on access after destroy

Fixes: 1ffcf8b8ae ("drm/xe: Support for mmap-ing mmio regions")
Assisted-by: GitHub-Copilot:claude-opus-4.6
Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-16-matthew.auld@intel.com
(cherry picked from commit fb2ee38bab8025ad6a7a9cbb4635c5a178e4a7bc)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Ilia Levi
de40d31275
drm/xe/mmio_gem: cache the dummy page per object
Currently, when the fault handler provides a dummy page, it
allocates a new one on every invocation and ties its lifetime to
the drm_device via drmm_add_action_or_reset(). Concurrent faults
after hot-unplug therefore accumulate pages that persist until
device teardown.

Cache a single dummy page in the xe_mmio_gem object and use dma_resv
lock to protect its allocation. Free it with the object.

v2: use dma_resv lock to protect the allocation (Matt Auld)

Assisted-by: GitHub-Copilot:claude-opus-4.6
Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-15-matthew.auld@intel.com
(cherry picked from commit 8bf6213f9831e46313af4722a1ee6db1b7596378)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Shuicheng Lin
37fcbd7b2f
drm/xe/mmio_gem: Revoke drm_vma_node on xe_mmio_gem destroy
xe_mmio_gem_create() calls drm_vma_node_allow() but nothing ever calls
drm_vma_node_revoke(). The drm_vma_offset_file rb-tree entry allocated
by drm_vma_node_allow() is not freed by drm_gem_object_release(), so
it is leaked on every create/destroy cycle.

Add a struct drm_file * parameter to xe_mmio_gem_destroy() and call
drm_vma_node_revoke() from there, mirroring the drm_vma_node_allow()
call in xe_mmio_gem_create().

Fixes: 1ffcf8b8ae ("drm/xe: Support for mmap-ing mmio regions")
Suggested-by: Ilia Levi <ilia.levi@intel.com>
Assisted-by: Claude:claude-opus-4.6
Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
Reviewed-by: Ilia Levi <ilia.levi@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-14-matthew.auld@intel.com
(cherry picked from commit 32f0cb250598456d812fb7ca57a040282858323d)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Ilia Levi
0c50663403
drm/xe/mmio_gem: simplify fault handler loop
Make the iteration over the addresses in the VMA more explicit.
No functional change, as the VMA matches the GEM object exactly.

Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-13-matthew.auld@intel.com
(cherry picked from commit 6666ca9192f3bdf839aad33b9e1c9ebb7a222a29)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Ilia Levi
819f189265
drm/xe/mmio_gem: use write-back mapping for dummy page
Currently vmf_insert_pfn() maps the dummy page as UC, inheriting the
VMA's page protection which was set for the real MMIO region. This
conflicts with the direct map's WB mapping of the same page, creating a
cache type alias which is architecturally undefined on some platforms.

Use vmf_insert_pfn_prot() with a WB pgprot instead. Also simplify to
fault in the requested page instead of the whole VMA.

Fixes: 1ffcf8b8ae ("drm/xe: Support for mmap-ing mmio regions")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260525125801.975038-6-ilia.levi%40intel.com
Assisted-by: GitHub-Copilot:claude-opus-4.6
Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-12-matthew.auld@intel.com
(cherry picked from commit 1e8e28e35df0e77ae1b22fc091c1f422f62fa5e9)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 08:49:00 -04:00
Ilia Levi
247a82da6f
drm/xe/mmio_gem: forbid VMA split
The fault handler assumes it always operates on a VMA spanning the entire
GEM object. This does not hold when the VMA has been split, e.g. by a
partial munmap or mprotect. In that case the handler may map wrong
physical pages or cause SIGBUS.

Handle this by forbidding VMA split, as partial unmaps are not deemed
useful for MMIO GEMs.

Suggested-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Fixes: 1ffcf8b8ae ("drm/xe: Support for mmap-ing mmio regions")
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-11-matthew.auld@intel.com
(cherry picked from commit f3391a0b12d7bf826a0b21600d2f294f3dce4c14)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 08:48:53 -04:00
Saim Shujah
a5b5cc9099 drm/msm/dpu: clear pending peripheral flush state
dpu_hw_ctl_clear_pending_flush() resets the cached per-block state after a
flush transaction, but misses pending_periph_flush_mask.

The peripheral flush updater accumulates interface bits in this mask. A
later transaction which sets the top-level peripheral flush bit can write
stale interface bits to CTL_PERIPH_FLUSH together with the current state.

Peripheral flush support was added after the helper started clearing every
individual pending flush mask. Clear the peripheral mask together with the
other cached child masks.

Fixes: 64f7b81f03 ("drm/msm/dpu: add support of new peripheral flush mechanism")
Cc: stable@vger.kernel.org
Signed-off-by: Saim Shujah <saimzst@gmail.com>
Patchwork: https://patchwork.freedesktop.org/patch/748968/
Link: https://lore.kernel.org/r/20260828065440.140410-1-saimzst@gmail.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:46:13 +03:00
Linus Torvalds
fd73f4a665 Linux 7.3-rc3 2026-09-13 14:38:02 -07:00
William Bright
58995b11df drm/msm/dp: fix link bandwidth check when wide bus is enabled
msm_dp_display_mode_valid() halves the pixel clock when either YUV420 or
wide bus is in use, then uses that halved value both for the controller
pixel clock limit and for the DP link bandwidth check.

Only YUV420 halves the data crossing the link. Wide bus widens the
internal DPU to DP interface to two pixels per clock, halving the
controller clock. Every pixel is still transmitted, so the link
bandwidth requirement remains.

As a result, modes needing up to twice the available link bandwidth pass
validation. On the IMDT QCS8550 SBC (rev5 with CYPD6125), where DP runs
over USB-C alt mode where only two lanes are available, 3840x2160@60 was
accepted despite needing 9.6 Gbps against the 8.64 Gbps the link can
carry.

Use a separate link pixel clock that is only halved for YUV420 for the
bandwidth calculation, leaving the wide bus halving to apply solely to
the controller pixel clock limit. With this, 4k@60 is correctly rejected
and 4k@30 selected instead.

Fixes: df9cf852ca ("drm/msm/dp: account for widebus and yuv420 during mode validation")
Assisted-by: Claude:claude-opus-5
Signed-off-by: William Bright <william.bright@imd-tec.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/746145/
Link: https://lore.kernel.org/r/20260812-msm-dp-link-bw-v1-1-b0e3ce1190be@imd-tec.com
[DB: dropped useless comment]
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:24:21 +03:00
Guangshuo Li
6fbbf1e152 drm/msm/adreno: fix autosuspend cleanup during teardown
adreno_gpu_init() calls pm_runtime_use_autosuspend(), but
adreno_gpu_cleanup() does not call the matching
pm_runtime_dont_use_autosuspend() during teardown.

If the autosuspend delay is set to a negative value while autosuspend
is enabled, the runtime PM core increments usage_count to prevent
runtime suspend. Without calling pm_runtime_dont_use_autosuspend()
during teardown, this reference is not dropped and usage_count remains
unbalanced.

The documentation for pm_runtime_use_autosuspend() also notes that it
is important to undo it with pm_runtime_dont_use_autosuspend() at
driver exit time, unless runtime PM was initially enabled with
devm_pm_runtime_enable().

Add the missing pm_runtime_dont_use_autosuspend() call to
adreno_gpu_cleanup().

This issue was found by manual code inspection.

Fixes: eeb754746b ("drm/msm/gpu: use pm-runtime")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/745110/
Link: https://lore.kernel.org/r/20260808131624.2854412-1-lgs201920130244@gmail.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:17:13 +03:00
Jesse Casco
e249a6e2a1 drm/msm/dp: skip PUSH_IDLE when the link was never enabled
msm_dp_display_atomic_enable() returns early when link training fails,
leaving ->power_on false and the main link down.
msm_dp_display_atomic_disable() nevertheless writes DP_STATE_CTRL_PUSH_IDLE
and waits for an idle-pattern completion that cannot arrive, so every failed
enable is followed by "PUSH_IDLE pattern timedout".

Every other step of the teardown is already gated on that flag:
msm_dp_display_disable(), called from .atomic_post_disable(), returns early
on !power_on. The PUSH_IDLE write is the only one that is not, so the
controller's runtime-PM reference is then dropped without the link having
been taken down.

On glymur (Snapdragon X2 Elite) the consequence is not a warning. The SoC
does not survive it: TrustZone force-stops the SOCCP and ADSP remote
processors and the machine resets silently about 50 ms later, with no oops
and no panic. On an ASUS Zenbook A16 (UX3607OA), whose eDP panel does not
currently train, this reproduces without any compositor or GPU involvement:

  # eDP enable has already failed with "Failed link training (rc=-104)"
  echo 1 > /sys/class/graphics/fb0/blank

  [535.645455] === marker ===
  [535.694833] qcom_q6v5_pas d00000.remoteproc: fatal error received: \
                 sys_m_smsm.c:512:TZ force stop
  [535.694875] remoteproc remoteproc0: crash detected in soccp: type fatal error
  [535.728857] qcom_q6v5_pas 6800000.remoteproc: fatal error received: \
                 sys_m_smsm.c:783:err fatal notification received from TZ
  <SoC reset>

Gate the PUSH_IDLE write on ->power_on so the disable path is consistent
with the rest of the teardown. With this applied the same sequence is
harmless and the machine stays up; without it, it resets every time.

The unconditional write dates back to the original DP driver
(c943b4948b ("drm/msm/dp: add displayPort driver support")), but the
surrounding code has been restructured several times since, so no Fixes:
tag is offered.

Note that the eDP link-training failure that exposes this on the A16 is a
separate problem in the glymur eDP PHY and is reported separately; this
change is about not damaging the machine when training fails, for whatever
reason.

Tested on ASUS Zenbook A16 (UX3607OA), Snapdragon X2 Elite Extreme, on
linux-next next-20260803 and next-20260807. The machine has since been
running next-20260807 with this patch as its daily driver.

Assisted-by: Anthropic:Claude-Opus-5
Signed-off-by: Jesse Casco <jesse.casco@gmail.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/745167/
Link: https://lore.kernel.org/r/20260808171325.133041-1-jesse.casco@gmail.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:16:20 +03:00
Dmitry Baryshkov
ea9dadeac7 drm/msm: mark the fbdev framebuffer as system memory
msm_fbdev_driver_fbdev_probe() points screen_buffer at a kernel virtual
mapping of the GEM object and uses the deferred sysmem fb ops, but never
sets FBINFO_VIRTFB. The framebuffer core then assumes the memory is not
in the virtual address space and warns on the first console draw:

  fb0: sys_fillrect: framebuffer is not in virtual address space.

The drm_fbdev_dma, drm_fbdev_shmem and drm_fbdev_ttm helpers all set the
flag for system memory. Do the same here.

Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Assisted-by: Claude:claude-opus-5
Patchwork: https://patchwork.freedesktop.org/patch/743293/
Link: https://lore.kernel.org/r/20260730-drm-msm-fbinfo-virt-v1-1-a27099a6dc58@oss.qualcomm.com
Acked-by: Rob Clark <robin.clark@oss.qualcomm.com> # on IRC
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:11:24 +03:00
Linus Torvalds
22098763a1 tracing fixes for 7.3:
- Don't destroy user event fields when removal fails
 
   User event fields are destroyed before the event is removed from
   visibility. But that can fail leaving the still visible event with no
   fields. Move the destroying of the fields to after the event is
   successfully removed from visibility.
 
 - Initialize function graph state is fork before calling copy_exec_state()
 
   For non-CLONE_VM forks, copy_exec_state() allocates a new task_exec_state.
   If that allocation fails, ftrace_graph_exit_task() will free the tasks
   ret_stack pointer. Since that pointer is still using the parent's
   ret_stack, it mistakenly frees the parent's pointer too.
 
   Call ftrace_graph_init() on the task first which will NULL out the new
   tasks's ret_stack and if the copy fails, it will not free anything.
 
 - Remove FGRAPH_MAX_INDEX
 
   The macro FGRAPH_MAX_INDEX was added but never used. Remove it.
 
 - Save ent_size in function graph printing of nested functions
 
   The function graph tracer needs to look at the next event to see if the
   next event is the return of the current function entry. If it is, it
   prints a single line:
 
     ktime_get();
 
   Otherwise it prints it like a nested function:
 
     tick_nohz_irq_exit() {
       ktime_get();
       kcpustat_irq_exit();
     }
 
   In order to look at the next event, it must save the current event so that
   it has the information to print from it. It saves the event in the
   iterator descriptor called "ent". What it doesn't save is the ent_size of
   the event which is now used to know if the function graph arguments are to
   be printed. The peek doesn't save the size so the size used happens to be
   that of the size of the last event that was seen.
 
   Save the entry event size in the iterator descriptor so that the correct
   size is used.
 
 - Fix several errors with freeing data in the histogram code
 
   The histogram code had a lot of leaked or or incorrect accounting when
   failures happen. Correct them.
 
 - Fix histogram regression of .percent and .graph modifiers
 
   Up until 6.3 histogram values could have "percent" or "graph" modifiers
   that changed how they were printed. But a change that added restricting
   histograms values from being strings, stack traces and other modifiers
   inadvertently prevented them from using the percent and graph modifiers,
   which were legal use cases for values.
 
   Put back the percent and graph modifiers.
 
 - Fix various typos in the comments
 
 - Set the trace_clock before initializing a histogram with clock argument
 
   The histogram API allows the user to specific which trace clock to use via
   a "clock=" string. The histogram is set up first before the clock is
   checked. If the passed in clock is not valid, it exits without fully
   fixing up the histogram leaving it on the list and a use-after-free can
   trigger.
 
   Update the clock argument first and if it fails then exit gracefully
   before the histogram trigger is placed on any lists.
 
 - Restore :mod: trailer after parsing in ftrace_set_clr_event
 
   The function ftrace_set_clr_event() modifies the parse string and needs to
   put it back to what was passed in. It searches for ":mod:" via a strsep()
   but fails to put back the first ':' in the string.
 
   Add back the ':' in the passed in string.
 
 - Take trace_array reference when opening a tracer options file
 
   The options files are dynamically created and some tracers add their own
   options. When a tracer adds their own list of options, the trace_array
   holding them has an array to hold the list of options for each tracer.
   This array increases in size via a krealloc(), and the new entry gets a
   newly allocated array to hold the options of the new tracer being added.
 
   The element in each entry of the tracer's option array holds a pointer
   back to the trace_array, a pointer to the tracer it is associated to, a
   pointer to the flags of the option.
 
   The issue is that these arrays are freed when the trace_array is freed
   when its instance it represents is removed from the instances directory.
   There's a race that an open of one of these options files can happen when
   the instance is being removed.
 
   Add a new helper function to be called by the open function of the options
   file to iterate all existing trace_arrays under a lock and find the one
   that has the given option element in one of it's tracer arrays. If found,
   then update the associated trace_array's reference counter to keep it from
   being freed. If not found, have the open call return -ENODEV.
 
 - Disable interrupts when acquiring the lock in rb_wake_up_waiters()
 
   The function rb_wake_up_waiters() assumes it will be called in interrupt
   context and does not disable irqs when taking cpu_buffer->reader_lock,
   which can be called in hard interrupt context. The issue is in PREEMPT_RT,
   this function is called in thread context leaving this lock open to a
   deadlock.
 
   Take the lock with interrupts disabled.
 
 - Use rcu_assign_pointer() for tmp_ops filter hash
 
   The tmp_ops used in update_ftrace_direct_mod() assigns its filter_hash
   field directly, but that field is annotated as __rcu and sparse complains.
   Assign it with rcu_assign_pointer()
 
 - Fix use-after-free in enable_trigger_private_data_free()
 
   The trace_event_call is accessed through the event_trigger_data's
   trace_event_file pointer to put the trace_event_call on freeing. The issue
   is that the trace_event_file data may have been freed already causing a
   use-after-free. Add a field to the event_trigger_data that points directly
   to the trace_event_call so that it can decrement its reference directly
   without needing to go through the trace_event_file.
 
 - Fix accounting of buffer data remote headers
 
   trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount the
   number of pages is needed for the asked for size as it doesn't take into
   account the meta data on each page. Add a helper function to do the
   calculation properly and use that in these functions.
 
 - Catch nr_page_va overflow in ring_buffer_desc sizing
 
   The number of pages per remote ring buffer is capped by
   ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to
   overflow that field would silently allocate a descriptor smaller than what
   was asked for.
 
 - Do not resize the subbuf order if any per_cpu buffer is disabled
 
   The mmapping of ring buffers disables resizing the subbuffers, but it is
   done per-cpu whereas the subbuf size change is done for all the per_cpu
   buffers under the buffer->mutex. It could change the size of some while
   the mapping is happening on others. Have the resize of the subbuf order
   check all the per_cpu buffers under the lock to see if any of them is
   disabled before starting and causing an inconsistency between buffers that
   are being mapped.
 -----BEGIN PGP SIGNATURE-----
 
 iIoEABYKADIWIQRRSw7ePDh/lE+zeZMp5XQQmuv6qgUCaqbdrBQccm9zdGVkdEBn
 b29kbWlzLm9yZwAKCRAp5XQQmuv6qro9AQDF/j3VW3Uu98lVFI9AB10XYhLDd5nt
 Zpf+3RviNgFpxgEAiE2+4K+4sM2SfaDDh9JMww9MKg1exL+cemE3a+JbBgY=
 =jgYE
 -----END PGP SIGNATURE-----

Merge tag 'trace-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace

Pull tracing fixes from Steven Rostedt:

 - Don't destroy user event fields when removal fails

   User event fields are destroyed before the event is removed from
   visibility. But that can fail leaving the still visible event with no
   fields. Move the destroying of the fields to after the event is
   successfully removed from visibility.

 - Initialize function graph state is fork before calling
   copy_exec_state()

   For non-CLONE_VM forks, copy_exec_state() allocates a new
   task_exec_state. If that allocation fails, ftrace_graph_exit_task()
   will free the tasks ret_stack pointer. Since that pointer is still
   using the parent's ret_stack, it mistakenly frees the parent's
   pointer too.

   Call ftrace_graph_init() on the task first which will NULL out the
   new tasks's ret_stack and if the copy fails, it will not free
   anything.

 - Remove FGRAPH_MAX_INDEX

   The macro FGRAPH_MAX_INDEX was added but never used. Remove it.

 - Save ent_size in function graph printing of nested functions

   The function graph tracer needs to look at the next event to see if
   the next event is the return of the current function entry. If it is,
   it prints a single line:

	ktime_get();

   Otherwise it prints it like a nested function:

	tick_nohz_irq_exit() {
	    ktime_get();
	    kcpustat_irq_exit();
	}

   In order to look at the next event, it must save the current event so
   that it has the information to print from it. It saves the event in
   the iterator descriptor called "ent". What it doesn't save is the
   ent_size of the event which is now used to know if the function graph
   arguments are to be printed. The peek doesn't save the size so the
   size used happens to be that of the size of the last event that was
   seen.

   Save the entry event size in the iterator descriptor so that the
   correct size is used.

 - Fix several errors with freeing data in the histogram code

   The histogram code had a lot of leaked or or incorrect accounting
   when failures happen. Correct them.

 - Fix histogram regression of .percent and .graph modifiers

   Up until 6.3 histogram values could have "percent" or "graph"
   modifiers that changed how they were printed. But a change that added
   restricting histograms values from being strings, stack traces and
   other modifiers inadvertently prevented them from using the percent
   and graph modifiers, which were legal use cases for values.

   Put back the percent and graph modifiers.

 - Fix various typos in the comments

 - Set the trace_clock before initializing a histogram with clock
   argument

   The histogram API allows the user to specific which trace clock to
   use via a "clock=" string. The histogram is set up first before the
   clock is checked. If the passed in clock is not valid, it exits
   without fully fixing up the histogram leaving it on the list and a
   use-after-free can trigger.

   Update the clock argument first and if it fails then exit gracefully
   before the histogram trigger is placed on any lists.

 - Restore :mod: trailer after parsing in ftrace_set_clr_event

   The function ftrace_set_clr_event() modifies the parse string and
   needs to put it back to what was passed in. It searches for ":mod:"
   via a strsep() but fails to put back the first ':' in the string.

   Add back the ':' in the passed in string.

 - Take trace_array reference when opening a tracer options file

   The options files are dynamically created and some tracers add their
   own options. When a tracer adds their own list of options, the
   trace_array holding them has an array to hold the list of options for
   each tracer. This array increases in size via a krealloc(), and the
   new entry gets a newly allocated array to hold the options of the new
   tracer being added.

   The element in each entry of the tracer's option array holds a
   pointer back to the trace_array, a pointer to the tracer it is
   associated to, a pointer to the flags of the option.

   The issue is that these arrays are freed when the trace_array is
   freed when its instance it represents is removed from the instances
   directory. There's a race that an open of one of these options files
   can happen when the instance is being removed.

   Add a new helper function to be called by the open function of the
   options file to iterate all existing trace_arrays under a lock and
   find the one that has the given option element in one of it's tracer
   arrays. If found, then update the associated trace_array's reference
   counter to keep it from being freed. If not found, have the open call
   return -ENODEV.

 - Disable interrupts when acquiring the lock in rb_wake_up_waiters()

   The function rb_wake_up_waiters() assumes it will be called in
   interrupt context and does not disable irqs when taking
   cpu_buffer->reader_lock, which can be called in hard interrupt
   context. The issue is in PREEMPT_RT, this function is called in
   thread context leaving this lock open to a deadlock.

   Take the lock with interrupts disabled.

 - Use rcu_assign_pointer() for tmp_ops filter hash

   The tmp_ops used in update_ftrace_direct_mod() assigns its
   filter_hash field directly, but that field is annotated as __rcu and
   sparse complains. Assign it with rcu_assign_pointer()

 - Fix use-after-free in enable_trigger_private_data_free()

   The trace_event_call is accessed through the event_trigger_data's
   trace_event_file pointer to put the trace_event_call on freeing. The
   issue is that the trace_event_file data may have been freed already
   causing a use-after-free. Add a field to the event_trigger_data that
   points directly to the trace_event_call so that it can decrement its
   reference directly without needing to go through the
   trace_event_file.

 - Fix accounting of buffer data remote headers

   trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount
   the number of pages is needed for the asked for size as it doesn't
   take into account the meta data on each page. Add a helper function
   to do the calculation properly and use that in these functions.

 - Catch nr_page_va overflow in ring_buffer_desc sizing

   The number of pages per remote ring buffer is capped by
   ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to
   overflow that field would silently allocate a descriptor smaller than
   what was asked for.

 - Do not resize the subbuf order if any per_cpu buffer is disabled

   The mmapping of ring buffers disables resizing the subbuffers, but it
   is done per-cpu whereas the subbuf size change is done for all the
   per_cpu buffers under the buffer->mutex. It could change the size of
   some while the mapping is happening on others. Have the resize of the
   subbuf order check all the per_cpu buffers under the lock to see if
   any of them is disabled before starting and causing an inconsistency
   between buffers that are being mapped.

* tag 'trace-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace: (25 commits)
  ring-buffer: Check resize_disabled before publishing the new subbuf order
  tracing/remotes: Catch nr_page_va overflow in ring_buffer_desc sizing
  tracing/remotes: Account for ring buffer page header in size calculation
  tracing: Don't dereference trace_event_file in deferred trigger free
  ftrace: Use rcu_assign_pointer() for tmp_ops filter hash
  ring-buffer: Acquire the lock with irqsave in rb_wake_up_waiters()
  tracing: Take trace_array reference when opening a tracer options file
  tracing: Fix ring_buffer_read_page_size() kernel-doc
  tracing: Restore :mod: trailer after parsing in ftrace_set_clr_event()
  tracing: Fix memory corruption from a "STACKTRACE" histogram key
  tracing: Fix memory corruption from the histogram stacktrace modifier
  tracing: Undo the registration when enabling the histogram trigger fails
  tracing: Take the reference before publishing the named histogram trigger
  tracing: Set the trace clock before registering the histogram trigger
  tracing: Fix typo "preceeded" in comment
  tracing: Fix typo "availabe" in comment
  tracing: Let histogram values keep the percent and graph modifiers
  tracing: Keep the entry count when the histogram stats allocation fails
  tracing: Free histogram the field rejected for a bad modifier
  tracing: Free histogram the var ref when its initialization fails
  ...
2026-09-13 12:27:00 -07:00
Linus Torvalds
d681d7ef61 Merge misc regression fixes that seem to have fallen through the cracks
Thorsten continues to track regressions, and reporting on known issues
with fixes that don't seem to make any progress.

I'm going to do an rc3 release later today - let's not keep these known
issues pending for yet another rc for no obvious reason.

Reported-by: Thorsten Leemhuis <regressions@leemhuis.info>
Link: https://lore.kernel.org/all/46403cf8-9a81-4596-87eb-dde58ae4c5db@leemhuis.info/

* regressions:
  media: ipu-bridge: do not use the CVS device lookup for IVSC
  wifi: mt76: mt792x: fix NULL dereference in ACPI SAR init during probe
  wifi: mt76: mt7921: skip unknown CLC firmware records
2026-09-13 10:18:23 -07:00
Sergey Zagursky
856c562c94 media: ipu-bridge: do not use the CVS device lookup for IVSC
Since commit c6b1b34b50 ("media: pci: intel: Add CVS support for IPU
bridge driver") the internal camera no longer works on laptops where the
sensor sits behind an IVSC, for example a Dell XPS 16 9640 (IPU6,
INTC10CF, ov02c10):

  intel-ipu6 0000:00:05.0: Found supported sensor OVTI02C1:00
  intel-ipu6 0000:00:05.0: Connected 1 cameras
  ivsc_csi intel_vsc-92335fcf-3203-4472-af93-7b4453ac29da: mei-csi probed
      without device fwnode!

No sensor subdevice is registered, the media graph has no sensor entity
and userspace finds no camera at all.

ipu_bridge_get_ivsc_csi_dev() first looks for the platform device named
"intel_vsc" and returns its mei-csi child. That device is created by
mei_vsc, which on this machine only appears once the LJCA USB bridge and
its SPI controller have probed, about a second after the IPU6 probe that
runs the bridge:

  07:59:29.297  platform INTC10CF:00 created (ACPI scan)
  07:59:41      intel-ipu6 probe -> ipu_bridge_init()
  07:59:42.391  platform intel_vsc created (mei_vsc)

The commit above added two fallbacks for CVS which match on the ACPI
companion alone. They are reached for every entry of ivsc_acpi_ids[],
IVSC IDs included. The IVSC ACPI device has two physical nodes:

  INTC10CF:00/physical_node  -> platform/INTC10CF:00  (no driver bound)
  INTC10CF:00/physical_node1 -> platform/intel_vsc    (mei_vsc)

so bus_find_device_by_acpi_dev(&platform_bus_type, adev) returns the bare
platform device. ipu_bridge_instantiate_ivsc() then attaches the IVSC
software node to that device instead of to the mei-csi client, the bridge
reports success, and the probe is never retried. mei_csi later probes
without a fwnode, the CSI-2 link is never described, and the sensor ACPI
device, which has an honoured _DEP on the IVSC device, is never
enumerated.

Before those fallbacks existed the lookup returned NULL here, the bridge
failed with -ENODEV and the probe was retried once the IVSC device had
shown up.

Skip those fallbacks for IVSC devices, keying on the IVSC IDs rather than
the CVS ones: new CVS IDs keep being added, whereas the IVSC list is
complete. CVS binds a driver to the ACPI device itself, so matching on the
companion stays unambiguous there.

Fixes: c6b1b34b50 ("media: pci: intel: Add CVS support for IPU bridge driver")
Link: https://lore.kernel.org/linux-media/20260901194526.6369-1-gvozdoder@gmail.com/
Cc: stable@vger.kernel.org
Assisted-by: Claude Code:claude-opus-5
Signed-off-by: Sergey Zagursky <gvozdoder@gmail.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2026-09-13 10:15:16 -07:00
Devin Wittmayer
7825de3f75 wifi: mt76: mt792x: fix NULL dereference in ACPI SAR init during probe
Some laptops carry a MediaTek power table in their firmware, and the
driver reads it to set a transmit limit for each frequency range.  It
only fills in the ranges themselves when it registers the device.

The startup step that does this existed already, but it never programmed
anything.  Two recent commits made it run a regulatory update instead,
which sets the limits on the way through, long before registration.

As a result, on a machine that has the table the driver reads through an
empty pointer and the interface never appears:

  BUG: kernel NULL pointer dereference, address: 0000000000000004
  RIP: 0010:mt792x_init_acpi_sar_power
  Call Trace:
   mt7921_set_tx_sar_pwr
   mt7921_mcu_regd_update
   mt7921_regd_update
   mt7921_run_firmware
   mt7921e_mcu_init
   mt7921_init_work

Skip it when the ranges are missing. They are applied again once the
device is up, which is where they came from before.

Reported-by: Klara Modin <klarasmodin@gmail.com>
Closes: https://lore.kernel.org/linux-wireless/aoyxqHYvSuaBeubf@soda.int.kasm.eu/
Fixes: 9b80bd9cab ("wifi: mt76: mt7921: add regulatory wiphy self manager support")
Fixes: e9f3f1cc13 ("wifi: mt76: mt7925: add regulatory wiphy self manager support")
Signed-off-by: Devin Wittmayer <lucid_duck@justthetip.ca>
Tested-by: David Gow <david@davidgow.net>
Tested-by: Klara Modin <klarasmodin@gmail.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2026-09-13 10:15:16 -07:00
Laxman Acharya Padhya
1a296bfd3e wifi: mt76: mt7921: skip unknown CLC firmware records
Treat an out-of-range CLC index as newer firmware rather than a
malformed image. linux-firmware 20260810 ships MT7922 records with
idx 3, and rejecting them made mt7921e fail to probe.

Keep the record-length checks, and report those as errors so a
truncated table is visible instead of a silent retry loop.

Fixes: 9417c5818a ("wifi: mt76: mt7921: validate CLC firmware records")
Reported-by: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com>
Signed-off-by: Laxman Acharya Padhya <acharyalaxman8848@gmail.com>
Reviewed-by: Junjie Cao <junjie.cao@intel.com>
Tested-by: Mikhail Gavrilov <mikhail.v.gavrilov@gmail.com>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2026-09-13 10:15:16 -07:00
David Carlier
d860c67c05 ring-buffer: Check resize_disabled before publishing the new subbuf order
ring_buffer_subbuf_order_set() stores the new order and only then walks
the CPUs, returning -EBUSY if any of them has resizing disabled. A user
mapped buffer has resizing disabled, and __rb_map_vma() reads
buffer->subbuf_order without buffer->mutex, so an mmap of an already
mapped CPU racing the failing order change sizes the mapping with the
new order and inserts pages past the sub-buffer into the VMA.

Check the CPUs before storing the new order.

Cc: stable@vger.kernel.org
Fixes: 117c39200d ("ring-buffer: Introducing ring-buffer mapping functions")
Link: https://patch.msgid.link/20260912103938.1127021-1-devnexen@gmail.com
Signed-off-by: David Carlier <devnexen@gmail.com>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2026-09-13 13:06:43 -04:00
Vincent Donnefort
d059d8bf2c tracing/remotes: Catch nr_page_va overflow in ring_buffer_desc sizing
The number of pages per remote ring buffer is capped by
ring_buffer_desc::nr_page_va (32 bits). A buffer_size large enough to
overflow that field would silently allocate a descriptor smaller than
what was asked for.

Return SIZE_MAX from trace_buffer_desc_size() on nr_page_va overflow.

Link: https://patch.msgid.link/20260911193937.602202-3-vdonnefort@google.com
Fixes: 2e67fabd8b ("ring-buffer: Introduce ring-buffer remotes")
Signed-off-by: Vincent Donnefort <vdonnefort@google.com>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2026-09-13 13:06:43 -04:00
Vincent Donnefort
442ffa742d tracing/remotes: Account for ring buffer page header in size calculation
trace_buffer_desc_size() and trace_remote_alloc_buffer() undercount the
required pages because every ring buffer page contains a header
(BUF_PAGE_HDR_SIZE). Account for that header to ensure allocated remote
ring buffers aren't smaller than requested by the user.

The newly introduced helper __calc_nr_pages_ring_buffer_desc() can
return a value that overflows the descriptor nr_pages field (32 bits).

Link: https://patch.msgid.link/20260911193937.602202-2-vdonnefort@google.com
Fixes: 2e67fabd8b ("ring-buffer: Introduce ring-buffer remotes")
Signed-off-by: Vincent Donnefort <vdonnefort@google.com>
Signed-off-by: Steven Rostedt <rostedt@goodmis.org>
2026-09-13 13:06:29 -04:00
Linus Torvalds
180534c09b Rust fixes for v7.3 (2nd)
Toolchain and infrastructure:
 
  - Work around a 'bindgen' 0.73.2 bug that emits an 'allow' attribute
    for 'unnecessary_transmutes', which is unknown in older compilers.
 
  - Clean 'clippy::as_underscore' lints in generated code by the new
    'bindgen' 0.73.0+ releases.
 
  - Clean new 'clippy::needless_range_loop' lint for the upcoming Rust
    1.100.0 (expected 2026-11-12).
 
 'kernel' crate:
 
  - 'num' module: fix soundness issue in 'Bounded' by sealing the
    'Integer' trait.
 
 'pin-init' crate:
 
  - Fix unreachable warning for the upcoming Rust 1.100.0 (expected
    2026-11-12) due to 'Infallible' becoming an alias of '!'.
 
 Samples:
 
  - Add missing newlines in 'pr_*!'s macro calls.
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEPjU5OPd5QIZ9jqqOGXyLc2htIW0FAmqmuA8ACgkQGXyLc2ht
 IW0ZVRAAk75N61v8xzY5dsQjA0O0ivCxDBqrPnFYYOq9jWWwKR4XF8zfX7dxPzFG
 48NHlQ9s3XEOSfmoVdaab9DMz8l2gCLMcCUOqmEGZtf1ORlFqCn7m0OMXfsidgx9
 YIWYSAySpjaQ27bg8+uvbBlBmD2KaE6zBlrAKbvdC9dJBOMfLjEnT3wtzkRkROzo
 WJMyx+OjIk0kmFNMUPBV/J+VWyxP5IAl8C5xK/hl3L+tf0VeQWkn82f7zzoGfwRV
 xLuIybzlxF2QK6D8OSf+SpxIqgl1fCDxh2rzWyNBJKbdGn1fMTTY7Ci6rM2DK853
 PjmQWtlkrYIOnO7k2qdCebOOv8wOBKE1hNpK+23mkEUbsZjWPNgSHVuf6X098NuH
 GEk5okH6+1e2w80dSRfUjKPY2omYhNoq4/4KEC+0IcV3xV+9FLq1uo9K/eOEr0Cf
 z430H31YnollXCWUx56QJZ7p3r0dITwhKHPE9pfKB53yWZelTEboRuD1zRKYEO4E
 0f+bDuLOAJeCzSX66YteZ+DiWphNB4OX49TGKRJpo9gWKn6+29TKIbJjAuk/oAu9
 FVkE5//WAu8El5++1W0YE4AM5eVvhrarH4vo6q44o6bd3acNHHJDzwUR8fu3FSYA
 skFb9vcRAq6vUPlzotlQMbuCP1PfBuFIhJdfqDcsVzGJO8oQPuk=
 =/75K
 -----END PGP SIGNATURE-----

Merge tag 'rust-fixes-7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ojeda/linux

Pull Rust fixes from Miguel Ojeda:
 "Toolchain and infrastructure:

   - Work around a 'bindgen' 0.73.2 bug that emits an 'allow' attribute
     for 'unnecessary_transmutes', which is unknown in older compilers

   - Clean 'clippy::as_underscore' lints in generated code by the new
     'bindgen' 0.73.0+ releases

   - Clean new 'clippy::needless_range_loop' lint for the upcoming Rust
     1.100.0 (expected 2026-11-12)

  'kernel' crate:

   - 'num' module: fix soundness issue in 'Bounded' by sealing the
     'Integer' trait

  'pin-init' crate:

   - Fix unreachable warning for the upcoming Rust 1.100.0 (expected
     2026-11-12) due to 'Infallible' becoming an alias of '!'

  Samples:

   - Add missing newlines in 'pr_*!'s macro calls"

* tag 'rust-fixes-7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ojeda/linux:
  rust: allow `unknown_lints` in generated bindings for Rust < 1.88
  rust: allow `clippy::as_underscore` in the generated bindings
  rust: num: seal Integer
  drm/panic: clean new `clippy::needless_range_loop` lint for Rust 1.100.0
  rust: samples: add missing newlines in rust_print_main
  rust: pin-init: use irrefutable pattern for `stack_pin_init`
2026-09-13 09:28:28 -07:00
Linus Torvalds
6a0b3fb48d Bootconfig fixes for v7.3-rc3
- bootconfig: Fix integer overflow and truncation vulnerabilities in size checks
   . tools/bootconfig: Fix integer overflow and truncation in size checks.
     Fix size check bypasses caused by integer overflow and truncation
     when parsing initrd or standalone bootconfig files, preventing
     buffer overflow and out-of-bounds writes in the userspace tool.
   . bootconfig: Fix integer overflow in initrd size check.
     Fix pointer arithmetic wrap-around in get_boot_config_from_initrd()
     when handling crafted huge size values, preventing fatal kernel
     page faults during early boot.
 -----BEGIN PGP SIGNATURE-----
 
 iQFPBAABCgA5FiEEh7BulGwFlgAOi5DV2/sHvwUrPxsFAmqml0UbHG1hc2FtaS5o
 aXJhbWF0c3VAZ21haWwuY29tAAoJENv7B78FKz8bXecH/jH1wLtkeeDumrR+5OGn
 hbLnTDryprnhXBP7gKmYfcVRJzF9HZ1Ro12R8ea4N/NJieUi+EDQQ/yn6TdIpV3z
 AU4In+zKT/q2hF3R1rmYuYxEMo9Po+dxgoB3BxKdwh9aDz8kPxQGP2/0Q/vjVMvZ
 5YosoEGYtNW6NpovVK+nMkYY0TwGXtft3tdGvbdMFToGf73EgeDA7POgdCXYgiP2
 D6equcjf7mRyBxzApCXzEEBynmHI6JTbZ6w0HGWN2bU9iwk/a/jiSBGGFrWopar6
 YhQdySy7mmyHZaQUOiz6M6kYvmF0PaLsiA0NMmZ6B3uKPv+wc7OShSoiknGRgBJc
 Vc4=
 =TCCT
 -----END PGP SIGNATURE-----

Merge tag 'bootconfig-fixes-v7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace

Pull bootconfig fixes from Masami Hiramatsu:
 "Fix integer overflow and truncation in size checks.

   - Fix size check bypasses caused by integer overflow and truncation
     when parsing initrd or standalone bootconfig files, preventing
     buffer overflow and out-of-bounds writes in the userspace tool.

   - Fix pointer arithmetic wrap-around in get_boot_config_from_initrd()
     when handling crafted huge size values, preventing fatal kernel
     page faults during early boot"

* tag 'bootconfig-fixes-v7.3-rc3' of git://git.kernel.org/pub/scm/linux/kernel/git/trace/linux-trace:
  bootconfig: Fix integer overflow in initrd size check
  tools/bootconfig: Fix integer overflow and truncation in size checks
2026-09-13 09:16:36 -07:00
Linus Torvalds
c874ace034 Misc timer fixes:
- Fix clockevents replacement race when a broadcast
    device is replaced which may trigger a BUG() crash
    (朱恺乾 - Zhu Kaiqian)
 
  - Fix potential timerqueue ordering bug when rearming
    a queued timer with nonzero slack (Andrea Parri)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqmXaERHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1jM9RAAnEShNuh27uYj3oxVgaLo+Dhsk1AYzNnr
 WVS/Cz8HfFXMEVnOsT3CibcB6p5wydTHms/8248GZWWUuM1HtL/7zcUHWGFgs72S
 WRTcw/ouzvdAKQfxlH2j96uApMWwnWnv9XRfpFel9bgIIK1POL9g0JJmcuFK5uNf
 5aYvzkdLv3SKHR0BnrIF4a6jqq1Shf2sZPDXHmJDE/k/He28zvFjiOlIEChxfCfh
 qUEzY036hSU0RbAONSbn88bj7dc10/Xuck/iW3WVW8cOqtVxw79biAoumUe9jqwE
 hX3B6rvDEPvaEOPDm2PgUlrapFukjfImu7K9rDljbFMX1jF6eb7ZQMk4ftJXL+rM
 M0RPCdrS2ZrVOKt3VFIYRH7ZzFNwtE+RHPZSD6lpVgia6xpgi6yY++AzTeCn+VpK
 3AmkxMg3xHOLkISyCRUlmtTn3Cis6O7+9+9dEad24dh5mkQM7Tr6nzprYeg3fgpR
 z714UKjOUvBNBtxCjdZl5/c/i8mb0IaH4DmT+/V6mIXWoHchbqgw0Or7G7XMm5XM
 M1J+4RJrGhhg1eUTb254PWi/OixuXZ8XgcB1wwAiJFMTJY9YqBExKRAdZ+Al1DJP
 FgHDEPyvElpeh5XFFqf8Ft9xXOTn1CSEc6G+dCO0MswcikHnw4UnmLGn9z3js2AS
 SRlK75Dm49s=
 =SyXt
 -----END PGP SIGNATURE-----

Merge tag 'timers-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull timer fixes from Ingo Molnar:

 - Fix clockevents replacement race when a broadcast
   device is replaced which may trigger a BUG() crash
   (朱恺乾 - Zhu Kaiqian)

 - Fix potential timerqueue ordering bug when rearming
   a queued timer with nonzero slack (Andrea Parri)

* tag 'timers-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  hrtimer: Use hard expiry when updating timers on the same base
  tick/broadcast: Plug clockevents replacement race
2026-09-13 09:10:38 -07:00
Linus Torvalds
b2a8a7669e Miscellaneous scheduler fixes:
- Fix EEVDF se->max_slice value on enqueueing (Vincent Guittot)
 
  - Fix EEVDF augmented rb-trees re-balancing with
    multiple fields (Vincent Guittot)
 
  - In proxy scheduling, account cgroup CPU time to the execution
    context, not the scheduling context (Hui Su)
 
  - Likewise, call wq_worker_tick() for the execution context,
    not the scheduling context (Hui Su)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqmXEMRHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1hchg/6A1gRkn7T5+K957U8wpB9vjtV9cKVdWpI
 XFDGm60ylFUmU388Xb8mmrbDgmej6RpX6C4ccppygM3196w513tB+Zr8w6jbSszk
 ddgWwfwi58FFBJZTH7JDqeJ64wvrl8KId44yM6k2JdXATxh2DGF0w+YdsA+M5HVJ
 EJbjACYhePdK27wvQDtj1poDfAyiabqEnv7w62dhEU9I+ikmcPAyrhmqU0yFDNUR
 sNozsDQnEJrHtllGHpr3FVxYRqob6lOtG+86VSiZ8F6i2kA3p/451mpMyyCOMUrF
 kZlBIryLG0gylXIensqLox+z2ZIE4nUL0OX3o7mC+MLNERdWvsdgHi9AcZlIoFpJ
 wMPBLENnnGbilmwhXjk0pL655rlVVUGwaTV4T9Pk5D5qew6B9LGe8uqiLo8U/qih
 1o5Lf3ZnUi0o8XHMfNkwQ3Y0m1S7CbgJYKItE+ec+2QifmKGD5dOKo5WWDKszywF
 Zb9ScP2fMKideS/JEX1/+jvLcpDmV+HE3mC58Muek96fbJG1IQ4bL5tu+fqO85rM
 68dJymtMrkoeegmq4jqERt758sZnyv2QbDmr1Kc1G9vSKNAeahYldPU2gQ6qFBnO
 VCMvI2BUmfiia9A6NBKKgpjijc4zWydyuaQ2zdvEBddYnLWKZ5Ge26lLrjraNtdB
 GtHalAP1gVs=
 =zJHp
 -----END PGP SIGNATURE-----

Merge tag 'sched-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull scheduler fixes from Ingo Molnar:

 - Fix EEVDF se->max_slice value on enqueueing (Vincent Guittot)

 - Fix EEVDF augmented rb-trees re-balancing with multiple
   fields (Vincent Guittot)

 - In proxy scheduling, account cgroup CPU time to the execution
   context, not the scheduling context (Hui Su)

 - Likewise, call wq_worker_tick() for the execution context,
   not the scheduling context (Hui Su)

* tag 'sched-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  sched/core: Call wq_worker_tick() for the execution context
  sched: Account cgroup CPU time to the execution context
  sched/eevdf: Fix rb augmented with multi fields
  sched/eevdf: Fix augmented max_slice
2026-09-13 09:03:22 -07:00
Linus Torvalds
85855f85de Miscellaneous perf events fixes:
- Fix sched_cb_list corruption on PMU callbacks that
    invoke list_del() during perf_event_overflow()
    calls (Thomas Richter)
 
  - Fix PEBS pt_regs->flags snapshot data that was
    regressed with the introduction of adaptive
    PEBS v4 support (Dapeng Mi)
 
  - Fix possible drain_pebs() re-entry bug
    when intel_pmu_drain_pebs_buffer() is called from
    process context (Dapeng Mi)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqmWoURHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1hzhA//TBu2jx/c2r5tafuW5mpPJH+3eGcMWUh9
 KfVLt/FMzgt+sOtzM/eIQ2Sk4VAachTkSLjh6MQR1m9g4jXTVr5rH8lJlFXqW/74
 v5XhaEryXTSOn6zpxXpKrVGYFq6RfdTijtmrAK+7ac6HChgRrMa0Eb5yVvPauLXE
 +Y0RugHjF75c4iXapb68osWF+7EoVKGqLPZjdQ12D6wga7+1DRTZWV37hOXHtu1M
 GbajBMTKF5Q4QCffsyY9PDks86dKLDrv8z7XxGNYz4pdcnwD4bMdYfu8FxohowaO
 EXx9b0dxK1RobcsMo+wHZojTD0i4ySSIbN7lJU6dHopMFJ3nKgCDYuPvzov9T/Dm
 8eFrzlQbr0m9+NqndFomjv4so1WGF1Q2DHqEOJQgSeHNECpUAx7mzIzv/AJiJm1P
 AC3S8PJT5+AFQZtySqV6nI8UyzyMgoDo3EYdp3oKHq/B6SGDlAUHr3k8e5bpI8ss
 JbIvyo1RH6DrB+FstHHve7zn5ueYmjxPQRS8NdIRrBKAeMryC7DYEDU7ZUwkAJS9
 jJZE7wM0zhLbF5EPVEA3+rhrh1wZ0mUFsWJwWbn7QwIpSm331J8AEKdpKFa3k5op
 FH9PusGL/SKeVwo1nReExLJWghdlxhbHAC8QhLMHTKOrhoqoUzpp5hRQ9rgx5z7N
 hksCg2xZC+U=
 =mKBo
 -----END PGP SIGNATURE-----

Merge tag 'perf-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull perf events fixes from Ingo Molnar

 - Fix sched_cb_list corruption on PMU callbacks that
   invoke list_del() during perf_event_overflow()
   calls (Thomas Richter)

 - Fix PEBS pt_regs->flags snapshot data that
   regressed with the introduction of adaptive
   PEBS v4 support (Dapeng Mi)

 - Fix possible drain_pebs() re-entry bug when
   intel_pmu_drain_pebs_buffer() is called from
   process context (Dapeng Mi)

* tag 'perf-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  perf/x86/intel: Prevent drain_pebs() reentry
  perf/x86/intel: Correct pt_regs->flags update for PEBS path
  perf/core: Allow list_del during perf_event_overflow()
2026-09-13 08:44:54 -07:00
Linus Torvalds
feb66eea6b Fix misc objtool bugs:
- Fix potential klp-build allocation leak in
    cleanup functionality handling kzalloc() failure
    (Yafang Shao)
 
  - Fix KLP checksum false positives triggering with
    GCC, caused by quirks in string literal symbol
    generation (Josh Poimboeuf)
 
 Signed-off-by: Ingo Molnar <mingo@kernel.org>
 -----BEGIN PGP SIGNATURE-----
 
 iQJFBAABCgAvFiEEBpT5eoXrXCwVQwEKEnMQ0APhK1gFAmqmWJIRHG1pbmdvQGtl
 cm5lbC5vcmcACgkQEnMQ0APhK1hGRA//fpBKxCoMv13E2ZLyzwWgz8nGkApXCmyg
 +cmmWM8uhNHfH9dA5d7Ipp6ziQcsob0cJ9QM48VM+PdJ1b46Dh41WPb7z9IA+kjG
 smV9wnH4dnfXqtFEUbGpzVc9GVv5tP5ZATqZe05rwlbgk8jQpbsr2EhoyAHShg7J
 Nqw0CmqFhnP3lKGjhU31UkwusFtI0F/m/tTlwT6n/EumpAPgcdiLo7d4I7Mx9d1g
 R0xwNy5OJGUci9bxYU97T6p5aRc4Kkq3XwNHyZcpJNoVjsXphYxSc2Rf/V4QPCTJ
 p8weOOBevYk/fScbq7v1LbflUTUvyjh25CQDwz0VUrSxXHrsAKCAiS60D3Fhate+
 lLtNnRDxCYDlNW50+sB0ch8WhhHpEqKBpnAdkdI4SUIQbAruvSO2s3Ygrs/JIirw
 CoXTzN5yFvSxGw/7DrTDhtrqwGRlCKYvAridbd12tEXGlkOaDB1KZnsazlhqdv1j
 1DjnP7SJu+mzDoBJXwz2Fw3TRWzwqx5w9leAFPjm/AaWWqFo/HvspJEKhh4RfhS9
 Ulb7wO4ftV6R0hdUS6Q48vSZdGr2Mu3U3cLku051kpdXNCZ4z+2liz9vhZgTx9JP
 8X6wMn7kUmg7jc3PR3p0i4Hgar6VYlddJ0BKrSzz5lghi2sA6HDhaB6I9gl7iafj
 /fJd8EjglqQ=
 =g7d3
 -----END PGP SIGNATURE-----

Merge tag 'objtool-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip

Pull objtool fixes from Ingo Molnar:

 - Fix potential klp-build allocation leak in cleanup
   functionality handling kzalloc() failure (Yafang Shao)

 - Fix KLP checksum false positives triggering with GCC, caused
   by quirks in string literal symbol generation (Josh Poimboeuf)

* tag 'objtool-urgent-2026-09-13' of git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip:
  objtool/klp: Fix checksums for constant pool references
  klp-build: Fix wrong index in funcs cleanup error path
2026-09-13 08:37:11 -07:00