Commit Graph

129133 Commits

Author SHA1 Message Date
Dave Airlie
71f370e9ee Two ttm fixes for ttm_tt_swapout(), one page-alignment and one overflow
fix for dma-buf, a drm_pending_vblank_event leak fix for drm,
 suspend/resume fixes for nouveau, one out-of-bounds access fix for gud,
 a use-after-free fix for vc4, a fence signaling fix, a race condition
 fix for sched, planes formats fixes for verisilicon, and add the blend
 mode property for loongson
 -----BEGIN PGP SIGNATURE-----
 
 iJUEABMJAB0WIQTkHFbLp4ejekA/qfgnX84Zoj2+dgUCaqvVAAAKCRAnX84Zoj2+
 dtThAX9JC7SWXw+n4o9EUsxNqvlpviwe8AS7aJR++rq/sZM0an7zl0aCJu/E8p+/
 JRkO+X8BfjE+2CX8RWKH63ed95vA+9DF8JfryBTSJiSz4forppFfwH7uovT3nqq2
 eOon6nJQzg==
 =Utup
 -----END PGP SIGNATURE-----

Merge tag 'drm-misc-fixes-2026-09-17' of https://gitlab.freedesktop.org/drm/misc/kernel into drm-fixes

Two ttm fixes for ttm_tt_swapout(), one page-alignment and one overflow
fix for dma-buf, a drm_pending_vblank_event leak fix for drm,
suspend/resume fixes for nouveau, one out-of-bounds access fix for gud,
a use-after-free fix for vc4, a fence signaling fix, a race condition
fix for sched, planes formats fixes for verisilicon, and add the blend
mode property for loongson

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Maxime Ripard <self@mripard.dev>
Link: https://patch.msgid.link/aqvVENQ4ksJEIcdb@houat
2026-09-19 06:49:36 +10:00
Dave Airlie
94f69bfa18 amd-drm-fixes-7.3-2026-09-17:
amdgpu:
 - SMU 14.x fix
 - DC IRQ fix
 - Runtime PM fix for P2P
 - RAS fix
 - PCIe reporting fix
 - DCN 6 fix
 - Device removal fix
 - DC MALL fix
 
 amdkfd:
 - GC 12.x fixes
 - Boundary checks
 - Mapping clear fix
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQQgO5Idg2tXNTSZAr293/aFa7yZ2AUCaqxIpQAKCRC93/aFa7yZ
 2HcZAP9ZA5cERbp6QfT0a1tT3kDoMP02BKev5/XUNWEJjdgOYAEAsCj7UFE4oKjb
 0996gK/lJvqq9lOgFTJnwl/z2jwytA0=
 =Ye34
 -----END PGP SIGNATURE-----

Merge tag 'amd-drm-fixes-7.3-2026-09-17' of https://gitlab.freedesktop.org/drm/amdgpu/kernel into drm-fixes

amd-drm-fixes-7.3-2026-09-17:

amdgpu:
- SMU 14.x fix
- DC IRQ fix
- Runtime PM fix for P2P
- RAS fix
- PCIe reporting fix
- DCN 6 fix
- Device removal fix
- DC MALL fix

amdkfd:
- GC 12.x fixes
- Boundary checks
- Mapping clear fix

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260917201213.3880863-1-alexander.deucher@amd.com
2026-09-18 11:24:43 +10:00
Dave Airlie
cd011719ba Merge tag 'drm-intel-fixes-2026-09-17' of https://gitlab.freedesktop.org/drm/i915/kernel into drm-fixes
drm/i915 fixes for 7.3-rc4:
- Revert a commit touching registers that don't necessarily exist
- Check for negative numbers before passing to BIT()

Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Jani Nikula <jani.nikula@intel.com>
Link: https://patch.msgid.link/3f86e0ede95fb3d52053934ac43c5271812428f5@intel.com
2026-09-18 11:08:19 +10:00
Dave Airlie
c24f824f0b Couple shrinker related fixes plus a series of patches fixing several
xe_mmio_gem issues around fault handler and destroy path.
 -----BEGIN PGP SIGNATURE-----
 
 iQEzBAABCgAdFiEEbSBwaO7dZQkcLOKj+mJfZA7rE8oFAmqr550ACgkQ+mJfZA7r
 E8p3UggAtIEI+BFbK7zeELK3zEtO85WQnIvUP/PprLu9Ga7Vvl+quXxZrWCHYcIE
 VwcK/5y2BGVKBmR1Vwm+PNBmtlbrNvIPGTn9u0oXB1bWVKqPT5o5MjY7lxdWhK95
 MkLFo7rsQM791DAEUeo6IzbdHd6K9k2At8Yzz5+1zt97zxGmXSNpGRxV6Ee8/crU
 e/B1cQSfY8I4sjhAormTLZ13M6vRh1Yoy5P21UdCqIU/Eq/PjpkW2bcmM6+8K+qO
 2a4y2CHIQnd+2bR4nEnuHYIa6DVuXEZIHh5v/8amenuZqBgVj7OcxUKiiZ4TnhMY
 bGe1UFyy0OWyEJ/X4ddGfCggpiFi6w==
 =0FwX
 -----END PGP SIGNATURE-----

Merge tag 'drm-xe-fixes-2026-09-17' of https://gitlab.freedesktop.org/drm/xe/kernel into drm-fixes

Couple shrinker related fixes plus a series of patches fixing several
xe_mmio_gem issues around fault handler and destroy path.

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Rodrigo Vivi <rodrigo.vivi@intel.com>
Link: https://patch.msgid.link/aqvoKaPuLLPBGayA@intel.com
2026-09-18 11:08:01 +10:00
Francis Marlou Pacaro
2ac2fe765e drm/amd/display: fix MALL hysteresis timer underflow at high refresh rates
dcn30_apply_idle_power_optimizations() derives the MALL frame cache
hysteresis timer with

	tmr_delay = (uint32_t)(div_u64(..., denom) - 64LL);

div_u64() returns a u64, so when the quotient is smaller than 64 the
subtraction wraps instead of going negative and tmr_delay ends up huge.
The loop that follows tries to squeeze it into the 6 bit register field
by doubling denom, but that only makes the quotient smaller, so tmr_delay
can never converge.  tmr_scale is bumped past 3 and the function gives up
with

	/* Delay exceeds range of hysteresis timer */
	ASSERT(false);

even though the requested delay is too *short* to encode, not too long.

With mall_additional_timer_percent left at its default of 0, the quotient
drops below 64 once the refresh rate used for the calculation goes above
~243 Hz.  Every DCN 3.0 display above that loses MALL static screen
entirely and splats a WARN once per boot.  Reproduced on Navi 23
(RX 6600) driving 1920x1080, resetting /sys/kernel/debug/clear_warn_once
between modes:

	refresh   MALL       ASSERT
	144 Hz    enabled    no
	240 Hz    enabled    no
	280 Hz    skipped    yes
	360 Hz    skipped    yes

Commit 3bb68cec4d ("drm/amd/display: Add Overflow check to skip MALL")
already covered the other end of the range, where a large stutter period
makes the delay too long to encode.  Cover the short end by clamping to
0, which selects the shortest hysteresis the register can express,
65.28us * 64 = ~4.18ms.  That is marginally longer than what the formula
asks for at these refresh rates, and erring long is the safe direction:
it only delays MALL entry, it can never enter early.

The numerator does not change between iterations, only denom does, so
compute it once and keep both call sites inside 100 columns.

The genuinely out of range case at very low refresh rates still reaches
the ASSERT, which is where it belongs.

Fixes: 52f2e83e2f ("drm/amdgpu/display: add MALL support (v2)")
Signed-off-by: Francis Marlou Pacaro <pacaro.francis.marlou.n@gmail.com>
Reviewed-by: Leo Li <sunpeng.li@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 387550e53e1405f1f960b62b22f8783db17c8e1d)
2026-09-17 11:59:24 -04:00
Chengjun Yao
5155002b03 drm/amdgpu: fix rmmio iounmap skipped on device removal
amdgpu_pci_remove() calls drm_dev_unplug() before fini_sw(), so
drm_dev_enter() is already false there and the iounmap() guarded by it
is skipped. This .remove path runs on both hot-unplug and plain rmmod,
so the register BAR ioremap mapping leaks one instance per unload.

Unmap rmmio unconditionally (guard only on non-NULL) and drop the now
unused idx.

Fixes: 62d5f9f711 ("drm/amdgpu: Unmap MMIO mappings when device is not unplugged")
Signed-off-by: Chengjun Yao <Chengjun.Yao@amd.com>
Reviewed-by: Asad Kamal <asad.kamal@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit dd6f86a97260e5207d3329ad03aa89fdad61b1e6)
Cc: stable@vger.kernel.org
2026-09-17 11:58:48 -04:00
Mario Limonciello
7f9caa70ae drm/amdgpu: Skip KFD mapping clear before initialization
amdgpu_amdkfd_clear_kfd_mapping() assumes that a non-NULL kfd_dev
has a fully populated node array. This is not true when KFD device
initialization fails after probe.

For example, kgd2kfd_device_init() sets num_nodes before checking
PCIe atomics support. On Polaris systems without the required atomics,
it returns before allocating nodes[0], but the kfd_dev remains attached
to the amdgpu device. A later GPU reset then dereferences nodes[0]->id.

Require the authoritative KFD initialization flag before walking the
node array, matching the existing KFD reset and teardown paths.

Fixes: 70cadefcc6 ("drm/amdgpu: unmap all user mappings of framebuffer and doorbell before mode1 reset")
Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5833
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 4ac1835823c47903fbb278bbf474773c46f59edc)
Cc: stable@vger.kernel.org
2026-09-17 11:58:04 -04:00
Srinivasan Shanmugam
0d2f4cfa56 drm/amd/display: Fix NULL dereference in dcn50/dcn60 init_hw
dc->clk_mgr is checked for NULL earlier in dcn50_init_hw() and
dcn60_init_hw(), but dcn50_initialize_min_clocks() and
dcn401_initialize_min_clocks() are called without any guard,
causing Smatch to report potential NULL dereferences.

Guard both call sites with the same pattern used throughout
both functions:

  if (dc->clk_mgr && dc->clk_mgr->funcs)

Also fix dcn50_initialize_min_clocks() which calls
get_dispclk_from_dentist without checking the function pointer,
unlike the dcn401 equivalent which guards that call.

Fix kernel-doc in dcn60_hwseq.c by adding missing parameter descriptions
for @probe in dcn60_update_probe_status() and @type in
is_probe_measurement_type_for_hubbub().

Fixes: 7f7d7ea1fa ("drm/amd/display: Add new sources for DCN6")
Reported-by: Dan Carpenter <error27@gmail.com>
Cc: Aurabindo Pillai <aurabindo.pillai@amd.com>
Cc: Ivan Lipski <ivan.lipski@amd.com>
Cc: Dan Wheeler <daniel.wheeler@amd.com>
Cc: Roman Li <roman.li@amd.com>
Cc: Alex Hung <alex.hung@amd.com>
Cc: Tom Chung <chiahsuan.chung@amd.com>
Signed-off-by: Srinivasan Shanmugam <srinivasan.shanmugam@amd.com>
Reviewed-by: Alex Hung <alex.hung@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 325c9a827cdd748e126eafeadaffc556204773d2)
2026-09-17 11:57:29 -04:00
David Francis
8ee521b8b1 drm/amdkfd: Avoid integer underflow in EOP ring size calculation.
The low 6 bits of cp_hqd_eop_control store the base-2 logarithm
of the EOP ring size. This was calculated as

order_base_2(q->eop_ring_buffer_size / 4) - 1

But order_base_2 can in theory return 0, so this could underflow
(although in practice the ring buffer size cannot be less than 4096).

Change this to

order_base_2(q->eop_ring_buffer_size / 8)

using properties of logarithms.

Also add to the above comment to make the mathematics more clear.

Reviewed-by: Kent Russell <kent.russell@amd.com>
Signed-off-by: David Francis <David.Francis@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit f0f43fcf8b2b3a924cad9444340921c96ed5f634)
Cc: stable@vger.kernel.org
2026-09-17 11:56:45 -04:00
David Francis
c883d0a132 drm/amdkfd: Avoid integer underflow with ffs in EOP ring size calc
The low 6 bits of cp_hqd_eop_control store the base-2 logarithm
of the EOP ring size. This was calculated as

ffs(q->eop_ring_buffer_size / sizeof(unsigned int)) - 1 - 1

But ffs can in theory return 1 or 0, so this could underflow
(although in practice the ring buffer size cannot be less than 4096).

Change this to

ffs(q->eop_ring_buffer_size / sizeof(unsigned int) / 4)

using properties of logarithms.

Reviewed-by: Kent Russell <kent.russell@amd.com>
Signed-off-by: David Francis <David.Francis@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 4f18c56630383c14bfc6b2d65f88f2f895d2121a)
Cc: stable@vger.kernel.org
2026-09-17 11:56:02 -04:00
Mario Limonciello
04de4007d3 drm/amdgpu: Fix GPU PCIe link capability reporting
Commit eb53125a7a ("drm/amd: Add dedicated helper for
amdgpu_device_find_parent()") made amdgpu_device_gpu_bandwidth() query
the first device outside the dGPU. That is the host side of the
physical link, not the GPU side.

As a result, the ASIC and platform capability masks can both be based
on the host port. drm_amdgpu_info_device then exposes the host
capabilities to userspace, such as Gen5 x16 for a Gen4 x8 GPU.

Cache both ends of the physical link during device initialization.
Use link_dev for the GPU capability and link_partner for the platform
capability and _PR3 detection.

Reported-by: "Marek Olšák" <maraeo@gmail.com>
Closes: https://lore.kernel.org/amd-gfx/CAAxE2A4VhsAzzO1QjBjUg+NgnbD04ZzMyN6xsUJxjKJHH6hxiw@mail.gmail.com/
Suggested-by: Lijo Lazar <lijo.lazar@amd.com>
Fixes: eb53125a7a ("drm/amd: Add dedicated helper for amdgpu_device_find_parent()")
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 7ea6a47224e2c6e89a3a682d7fbaace4817a55aa)
Cc: stable@vger.kernel.org
2026-09-17 11:55:14 -04:00
Dmitriy Chumachenko
723d4dc628 drm/amdgpu: check ras and obj before dereference
nbio_v7_9_handle_ras_controller_intr_no_bifring() dereferences ras and obj
without checking either for NULL. Both amdgpu_ras_get_context() and
amdgpu_ras_find_obj() can return NULL, e.g. during the window between
adev->nbio.ras being set (early in amdgpu_ras_init(), by design, to
enable the fatal-error interrupt as soon as possible) and the PCIE_BIF
ras object actually being created in RAS late_init. Any interrupt in that
window crashes in hard-IRQ context.

This is analogous to commit d190b459b2 ("drm/amdgpu: the warning
dereferencing obj for nbio_v7_4"), which fixed the same issue in the
nbio_v7_4 handler.

Found by Linux Verification Center (linuxtesting.org) with SVACE.

Fixes: 7692e1ee24 ("drm/amdgpu: add RAS fatal error handler for NBIO v7.9")
Reviewed-by: Tao Zhou <tao.zhou1@amd.com>
Signed-off-by: Dmitriy Chumachenko <Dmitry.Chumachenko@cyberprotect.ru>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit c7071767a50a32ed727cf800ac84372429e3b4b3)
2026-09-17 11:54:31 -04:00
Mike Lothian
636139603b drm/amdgpu: hold a runtime PM reference for P2P dma-buf attachments
amdgpu_dma_buf_map() adds VRAM to the allowed domains for a peer2peer
attachment.  GTT is only a fallback placement when VRAM is preferred, so
ttm_bo_validate() migrates the buffer from GTT into VRAM.  While the
exporting device is runtime suspended its SDMA rings are down and the
move fails:

  amdgpu: Move buffer fallback to memcpy unavailable

An importer on a second GPU reaches this holding no runtime PM
reference on the exporter, e.g. a compositor on the APU submitting a
frame that references a buffer exported by an idle dGPU:

  amdgpu_cs_ioctl -> amdgpu_cs_parser_bos -> amdgpu_cs_bo_validate
    -> ttm_bo_validate -> amdgpu_bo_move -> dma_buf_map_attachment
      -> amdgpu_dma_buf_map -> ttm_bo_validate -> amdgpu_bo_move

Pinning a dma-buf into VRAM has the same requirement, which
commit 030631e97b ("drm/amdgpu: revert "take runtime pm reference
when we attach a buffer" v2") called out as the one case that would
need the reference back.

Take it in attach and drop it in detach.  pm_runtime_get_if_active()
never resumes the device, so it cannot deadlock against the reservation
taken during resume, which is why the old pm_runtime_get_sync() had to
go.  If the device is not active, clear peer2peer instead: the buffer
then stays in GTT, which remains accessible while the GPU is powered
down.  If runtime PM is disabled, take a plain reference so the put in
detach stays balanced.

Fixes: 030631e97b ("drm/amdgpu: revert "take runtime pm reference when we attach a buffer" v2")
Suggested-by: Christian König <christian.koenig@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Mike Lothian <mike@fireburn.co.uk>
Assisted-by: Claude:Opus-5 [Claude Code]
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 062ff15e30a48d14fb7d7558eba84f8dc97197f0)
Cc: stable@vger.kernel.org
2026-09-17 11:54:05 -04:00
Vladimir Marioukhine
5f28bb1c2c drm/amdkfd: implement restore_mqd callbacks for GFX12/12.1
kfd_mqd_manager_v12.c (GFX 12.0) and kfd_mqd_manager_v12_1.c (GFX 12.1)
do not implement restore_mqd callbacks, leaving the function pointers
NULL and causing CRIU restore to return -EOPNOTSUPP on GFX12.

Implement restore_mqd for both compute and SDMA queues in
kfd_mqd_manager_v12.c and kfd_mqd_manager_v12_1.c, modeled after the
GFX 11 implementation with the following improvements:
- update cp_mqd_base_addr_lo/hi to the newly allocated MQD address,
  fixing a pre-existing gap shared with v11 where the in-MQD copy
  still pointed at the old checkpoint-time address after restore
- memset the full allocation before memcpy for compute queues to avoid
  stale data in the GTT sub-allocator tail; SDMA MQDs use sizeof(*m)
  since they are packed at mqd_size stride in a shared BO

checkpoint_mqd registration is deferred to a follow-up patch that also
implements get_checkpoint_info, so that checkpoint and restore are
enabled together as a complete and testable unit.

Note: GFX12.1 restore handles XCC0 only. Multi-XCC CRIU restore is
currently unreachable due to a separate validation issue in
kfd_criu_restore_queue(). A pr_warn_once() is emitted if a multi-XCC
device is encountered.

Signed-off-by: Vladimir Marioukhine <Vladimir.Marioukhine@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit b1f9601237d050f5df478464cf51bf1fff29a256)
Cc: stable@vger.kernel.org
2026-09-17 11:53:24 -04:00
Leo Li
63e19ef3dd drm/amd/display: Atomize IRQ register read/modify/write ops
[Why]

The OTG_GLOBAL_SYNC_STATUS register controls various HW IRQ sources for
the output timing generator (OTG). VUPDATE_NO_LOCK is one of them.

To enable the IRQ, driver sets the VUPDATE_NO_LOCK_EN bit in the
GLOBAL_SYNC_STATUS register.

To ack the IRQ after it fires, the driver sets the VUPDATE_NO_LOCK_CLEAR
bit in the same GLOBAL_SYNC_STATUS register.

The bit sets are done through read/modify/write operations, which are
not atomic. Thus, the following race is possible:

    Thread A:                       IRQ handler:
                                    *HW IRQ fires*
    # IRQ disable
    val = read(GLOBAL_SYNC_STATUS)
    unset(val, VUPDATE_NO_LOCK_EN)
    write(val, GLOBAL_SYNC_STATUS)
                                    # ACK reads VUPDATE_NO_LOCK_EN unset
                                    val1 = read(GLOBAL_SYNC_STATUS)
                                    set(val1, VUPDATE_NO_LOCK_CLEAR)
    # IRQ enable
    val = read(GLOBAL_SYNC_STATUS)
    set(val, VUPDATE_NO_LOCK_EN)
    write(val, GLOBAL_SYNC_STATUS)
                                    # BAD! clears VUPDATE_NO_LOCK_EN
                                    write(val1, GLOBAL_SYNC_STATUS)

Regarding the tagged Fixes: change, it appears the change made this race
more likely to occur. Since VUPDATE_NO_LOCK is now the sole IRQ source
for vblank handling, a single race on high refresh panels can lead to a
time out.

[How]

The GLOBAL_SYNC_STATUS register is only one example, other IRQ control
registers also share the same scheme. On top of GLOBAL_SYNC_STATUS,
let's clean up those as well.

To keep things simple, Let's atomize the IRQ rmw ops via a single
driver-wide spinlock. Due to the small scope of this lock, it is
unlikely to cause noticeable overhead on top of all the existing locking
within the IRQ set/handle paths.

Since DM is responsible for locking, wrap dc_interrupt_set/ack with the
spinlock in the new amdgpu_dm_irq_set/ack functions. Migrate/drop all
references in DM to dc_interrupt_set/ack to use amdgpu_dm_irq_set/ack
instead.

Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5616
Fixes: c87e6635d2 ("drm/amd/display: consolidate DCN vblank/flip handling onto vupdate_no_lock")
Reviewed-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Leo Li <sunpeng.li@amd.com>
Signed-off-by: Chenyu Chen <chen-yu.chen@amd.com>
Tested-by: Daniel Wheeler <daniel.wheeler@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 70de0a0216583a53c946155f8c8adedfdca6b4e7)
Cc: stable@vger.kernel.org
2026-09-17 11:52:06 -04:00
Kevin Wang
9413959fa9 drm/amd/pm: report energy accumulator for smu 14.0.3
add energy accumulator on pmfw 0x00685000 and above version.

Signed-off-by: Kevin Wang <kevin.wang@amd.com>
Reviewed-by: Kenneth Feng <kenneth.feng@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 4aa733ab15b303e2a40e2985ac21a0e01f24cc4a)
2026-09-17 11:51:15 -04:00
Dave Airlie
b15e54761a Merge tag 'drm-msm-fixes-2026-09-16' of https://gitlab.freedesktop.org/drm/msm into drm-fixes
Fixes for v7.3-rc4:

DT:
- Corrected indentation

Core:
- Marked fbdev as system memory

GPU:
- Fixed autosuspend cleanup on teardown
- a750: fix timestamps
- Increase GMU fw init timeout
- Misc fixes/cleanups

DPU:
- Fixed clock rounding, unbreaking newest platforms
- Cleared pending flush state

DP:
- Skip PUSH_IDLE when link was never enabled
- Fixed bandwidth checks

HDMI:
- Fixed runtime PM cleanup on probe failure

Signed-off-by: Dave Airlie <airlied@redhat.com>
From: Rob Clark <rob.clark@oss.qualcomm.com>
Link: https://patch.msgid.link/CACSVV021rZsjmiPEyR_LR7L=k7=DRV0QdsA50BF_Gd0s2DXaCw@mail.gmail.com
2026-09-17 19:05:21 +10:00
Huacai Chen
5535d5e61a drm/loongson: Create blend mode property for cursor plane
After commit 860e748bdd ("drm: ensure blend mode supported if pixel
format with alpha exposed") we get warnings at boot:

loongson 0000:00:06.1: [drm] [PLANE:41:ls-cursor-plane-0] pixel format with alpha exposed but blend mode not setup. Please fix.
loongson 0000:00:06.1: [drm] [PLANE:46:ls-cursor-plane-1] pixel format with alpha exposed but blend mode not setup. Please fix.

The reason is the cursor plane supports color formats with alpha but the
driver doesn't create blend mode property, which triggers the warning in
validate_blend_mode_for_alpha_formats().

The loongson DC HW doesn't support DRM_MODE_BLEND_PREMULTI, so create
blend mode property with DRM_MODE_BLEND_COVERAGE for cursor planes since
it is the only one implemented in the driver.

Signed-off-by: Huacai Chen <chenhuacai@loongson.cn>
Reviewed-by: Jianmin Lv <lvjianmin@loongson.cn>
Signed-off-by: Icenowy Zheng <zhengxingda@iscas.ac.cn>
Link: https://patch.msgid.link/20260903084212.3621540-1-chenhuacai@loongson.cn
2026-09-17 15:12:26 +08:00
Raag Jadav
f0e9f963a3
drm/xe/i2c: Disable IRQ on unbind
Currently, struct xe_i2c is freed before SGUnit IRQ is disabled in unbind
path, leaving a potential UAF in case I2C IRQ is hit during this small
window. Explicitly disable I2C IRQ in xe_i2c_remove() and fix this.

Fixes: 0bb78ce099 ("drm/xe/i2c: Wire up reset/postinstall for I2C IRQ")
Signed-off-by: Raag Jadav <raag.jadav@intel.com>
Reviewed-by: Heikki Krogerus <heikki.krogerus@linux.intel.com>
Link: https://patch.msgid.link/20260911121547.2407261-1-raag.jadav@intel.com
Signed-off-by: Matt Roper <matthew.d.roper@intel.com>
(cherry picked from commit 8ba5c8b8ab3fd362267c11df2cd5a90ee46f6e24)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-16 08:46:16 -04:00
Nemesa Garg
a26204be58 Revert "drm/i915/display: Clear SEL_FETCH_PLANE_CTL on plane disable"
This reverts commit 7f1172a2ac.

This commit replaced the crtc_state->enable_psr2_sel_fetch guard in
icl_plane_disable_sel_fetch_arm() and i9xx_cursor_disable_sel_fetch_arm()
with HAS_PSR2_SEL_FETCH(). This is a display version check and
says nothing about the pipe, so every plane and cursor disable on a
display 12+ platform started writing SEL_FETCH_PLANE_CTL() /
SEL_FETCH_CUR_CTL(), including on pipes that do not implement them.
It shows up as an unclaimed register access on pipes driving HDMI where
selective fetch was never enabled.

The stale selective fetch enable bit that commit addressed is handled
in the next patch.

Fixes: 7f1172a2ac ("drm/i915/display: Clear SEL_FETCH_PLANE_CTL on plane disable")
Closes: https://gitlab.freedesktop.org/drm/i915/kernel/-/work_items/16876
Signed-off-by: Nemesa Garg <nemesa.garg@intel.com>
Reviewed-by: Jouni Högander <jouni.hogander@intel.com>
Signed-off-by: Suraj Kandpal <suraj.kandpal@intel.com>
Link: https://patch.msgid.link/20260909110332.3528029-2-nemesa.garg@intel.com
(cherry picked from commit d393529394167e0f5f706657eebe84d8529ce4fc)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
2026-09-16 12:55:22 +03:00
Icenowy Zheng
6c62dfd282 drm/verisilicon: remove ARGB formats from primary plane
As the blending of the primary plane is currently explicitly disabled
(and it's not possible on DC8000), remove the ARGB formats from the
primary plane format tables.

Signed-off-by: Icenowy Zheng <zhengxingda@iscas.ac.cn>
Reviewed-by: Thomas Zimmermann <tzimmermann@suse.de>
Link: https://patch.msgid.link/20260910095000.3505878-2-zhengxingda@iscas.ac.cn
2026-09-16 17:41:25 +08:00
Icenowy Zheng
e5d43d7e92 drm/verisilicon: add primary modifier for format tables
Currently the format tables are only used for the primary plane.

Add primary modifiers to names related to the tables.

Signed-off-by: Icenowy Zheng <zhengxingda@iscas.ac.cn>
Reviewed-by: Thomas Zimmermann <tzimmermann@suse.de>
Link: https://patch.msgid.link/20260910095000.3505878-1-zhengxingda@iscas.ac.cn
2026-09-16 17:41:22 +08:00
Icenowy Zheng
5a83606d9f drm/verisilicon: set blend mode for the cursor plane
Blend mode properties are now required to expose pixel formats w/ alpha.

Experiments show that the fixed blending mode for the cursor seems to be
COVERAGE:

- With a cursor plane filled with R=G=0, B=0xff, A=0x40, the cursor is
  visible on a pure-white background, which means the background is
  multiplied.
- With a cursor plane filled with R=G=B=0xff, A=0x40, the cursor isn't
  pure white and non-white patterns can be see through, which means the
  cursor is multiplied.

Add a fixed COVERAGE blend mode property for the cursor plane.

Signed-off-by: Icenowy Zheng <zhengxingda@iscas.ac.cn>
Reviewed-by: Thomas Zimmermann <tzimmermann@suse.de>
Link: https://patch.msgid.link/20260910094904.3502741-1-zhengxingda@iscas.ac.cn
2026-09-16 17:39:55 +08:00
Maxime Ripard
666f12ae9f
Merge fdo/drm/drm-fixes into drm-misc-fixes
Backmerging to get drm-misc-fixes up to v7.3-rc3.

Signed-off-by: Maxime Ripard <mripard@kernel.org>
2026-09-16 09:43:25 +02:00
Tvrtko Ursulin
2ab510e631 drm/sched: Fix virtual runtime race
Prevent pushing a new job to an entity seeing it being the first in the
queue, and hence entering the drm_sched_rq_add_entity() path, if the pop
side in drm_sched_entity_pop_job() has just de-queued the job but not yet
updated the saved virtual time.

Restoring the unsaved virtual time, which is at this point not a delta but
still an absolute value, pushes the said entity to the rear of the run queue
for a potentially very long time.

We close this race by pulling the locked sections out to encompass both
the queue push/pop and corresponding rbtree management.

This is aligned with the future direction to replace the current lockless
job queue with one of the fully locked standard list primitives.

Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Fixes: 2fa4d8e2c1 ("drm/sched: Add fair scheduling policy")
Suggested-by: Luke.Wildhardt@proton.me # via Claude Opus
Tested-by: Luke.Wildhardt@proton.me
Cc: Christian König <christian.koenig@amd.com>
Cc: Danilo Krummrich <dakr@kernel.org>
Cc: Philipp Stanner <phasta@kernel.org>
Cc: Pierre-Eric Pelloux-Prayer <pierre-eric.pelloux-prayer@amd.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Vitaly Prosyak <vitaly.prosyak@amd.com>
Cc: stable@vger.kernel.org # v7.2+
[phasta: commit title]
Signed-off-by: Philipp Stanner <phasta@kernel.org>
Link: https://patch.msgid.link/20260915150557.62847-1-tvrtko.ursulin@igalia.com
2026-09-16 09:17:51 +02:00
Luca Coelho
ee415ce8cb drm/i915/display: check configuration index before shifting
The calc_allowed_config_filter() function passes the return value of
iter_pos_to_idx() directly to BIT(), but the helper can return -1 for
an invalid iterator.

The iterator already rejects negative indices before doing a
configuration, so this should not matter in normal flows.  In any
case, for robustness, check the index explicitly and warn if it is
negative, avoiding an undefined shift.

Fixes: 39e30bdf2f ("drm/i915/dp_link_caps: Add link configuration iterator")
Reviewed-by: Imre Deak <imre.deak@intel.com>
Link: https://patch.msgid.link/20260908100659.113555-1-luciano.coelho@intel.com
Signed-off-by: Luca Coelho <luciano.coelho@intel.com>
(cherry picked from commit fe05cb9b9fb0ecc10409c4c6133257214b6cd8c8)
Signed-off-by: Jani Nikula <jani.nikula@intel.com>
2026-09-15 11:14:13 +03:00
Guangshuo Li
f4fae975db drm/msm/hdmi_phy: fix runtime PM cleanup on probe failure
msm_hdmi_phy_probe() enables runtime PM before enabling the PHY
resources and initializing the PLL, but failures from either operation
return without calling the matching pm_runtime_disable().

The remove path disables runtime PM, but it is not called when probe
fails. As a result, runtime PM remains enabled after an unsuccessful
probe.

Route failures after pm_runtime_enable() through a common error path
and disable runtime PM before returning.

This issue was found by manual code inspection.

Fixes: 15b4a45238 ("drm/msm/hdmi: Create a separate HDMI PHY driver")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Reviewed-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/753043/
Link: https://lore.kernel.org/r/20260913085814.1509352-1-lgs201920130244@gmail.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-15 02:09:28 +03:00
Dmitry Baryshkov
2028280686 drm/msm/dsi: round the byte clock rate after reparenting to the PHY PLL
DSI 6G v2.9 hosts (SM8650, SM8750, Kaanapali, etc.) reparent the byte and
pixel RCGs to the DSI PHY PLL at runtime from
dsi_link_clk_set_rate_6g_v2_9(), after the PHY has been enabled. However
dsi_calc_clk_rate_6g() runs earlier, in order to compute the bit clock
request for the PHY. At that point the byte RCG still has its reset
parent (XO), so clk_round_rate() returns a bogus rate, which then ends up
in the PHY bit clock request and the PLL gets programmed to a wrong
frequency, breaking the panel.

Move the rounding to dsi_link_clk_set_rate_6g(), which is called after
the RCGs have been reparented to the PLL. Storing the rounded rate at
this point still makes later link_clk_set_rate() calls no-ops in the
CCF. Derive the byte interface clock rate from the rounded byte clock
rate, otherwise it would keep requesting the idealized rate and
retrigger the PLL on every transfer.

Reported-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Reported-by: Krzysztof Kozlowski <krzysztof.kozlowski@oss.qualcomm.com>
Fixes: 6cd33b6f41 ("drm/msm/dsi: round 6G byte clock rate to the PLL-achievable value")
Assisted-by: LLM
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Reviewed-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com>
Tested-by: Konrad Dybcio <konrad.dybcio@oss.qualcomm.com> # SM6115P J606F
Tested-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Reviewed-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/750496/
Link: https://lore.kernel.org/r/20260903-fix-eliza-dsi-v1-1-3474a6c9f2e0@oss.qualcomm.com
2026-09-15 02:09:28 +03:00
Karl Mehltretter
073a30d75f
drm/vc4: Use managed KMS polling to fix UAF on unbind
vc4_kms_load() calls drm_kms_helper_poll_init() but the driver provides
no matching drm_kms_helper_poll_fini(). The output poll work stays
scheduled after unbind and runs on the freed drm_device:

  # modprobe vc4; rmmod vc4; sleep 10
  BUG: KASAN: slab-use-after-free in delayed_work_timer_fn
  BUG: KASAN: slab-use-after-free in drm_client_dev_hotplug [drm]
  Workqueue: events output_poll_execute [drm_kms_helper]
  Allocated by task 171: __devm_drm_dev_alloc
  Freed by task 262 (rmmod): drm_dev_put / component_del

Use drmm_kms_helper_poll_init() so polling is finalized with the device,
as other drivers do.

Fixes: c8b75bca92 ("drm/vc4: Add KMS support for Raspberry Pi.")
Assisted-by: Claude:claude-fable-5
Signed-off-by: Karl Mehltretter <kmehltretter@gmail.com>
Link: https://patch.msgid.link/20260822143110.68594-1-kmehltretter@gmail.com
Reviewed-by: Maíra Canal <mcanal@igalia.com>
Signed-off-by: Maíra Canal <mcanal@igalia.com>
2026-09-14 12:11:42 -03:00
Shuicheng Lin
985862be16
drm/xe/shrinker: Take a runtime PM ref before shrinking non-system memory
__xe_shrinker_walk() walks the SYSTEM and TT LRUs without a runtime PM
reference.  Shrinking a bo outside system memory invalidates its GPU
mappings, which needs the device resumed, so while it is runtime
suspended the page table zap trips an assert and the TLB invalidation
returns -ENODEV:

  WARNING: drivers/gpu/drm/xe/xe_bo.c:770 at xe_bo_move_notify+0x1fc/0x450 [xe]
   xe_bo_shrink+0x20f/0x2b0 [xe]
   __xe_shrinker_walk+0x174/0x410 [xe]
   xe_shrinker_scan+0x10c/0x1e0 [xe]
   do_shrink_slab+0x176/0x7e0
   drop_caches_sysctl_handler+0x9c/0xf0

Take a reference before walking a memory type other than XE_PL_SYSTEM
and stop there if it cannot be acquired.  Reuse the shrinker's existing
acquire path, which resumes the device directly where reclaim allows
that and otherwise queues the PM worker for a later scan.  Stop the walk
once the scan target is met, so a satisfied scan does not wake the
device.  System memory is still reclaimed while the device is suspended.

Gate this on xe_device_is_l2_flush_optimized(), the same condition under
which xe_bo_trigger_rebind() issues the invalidation for a non-fault-mode
vm, so reclaim is unaffected elsewhere.  The System CCS copy already has
its own reference in xe_bo_shrink().

Only a non-fault-mode vm can reach this, since a fault-mode vm requires
LR mode and that holds a runtime PM reference for the vm's lifetime.

Reproduced with igt@xe_madvise@dontneed-before-exec while the GPU is
runtime suspended.

v2: simplify needs_rpm check. (Matt)
    retarget Fixes tag since the issue occurs with the non-fault-mode
    path added by 4e7ebff69a.
v3: handle this in xe_shrinker.c instead of xe_bo.c (Thomas)
v4: stop the walk once the scan target is met. (Sashiko)
v5: rebase on the freed page accounting fix. (Sashiko)
v6: reuse the shrinker acquire path so runtime pm can be resumed
    directly instead of always queueing a worker. (Thomas)
v7: replace xe_pm_runtime_put() with xe_shrinker_runtime_pm_put(). (Thomas)

Fixes: 4e7ebff69a ("drm/xe/xe3p_lpg: flush shrinker bo cachelines manually")
Assisted-by: Claude:claude-opus-5
Cc: Tejas Upadhyay <tejas.upadhyay@intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Link: https://patch.msgid.link/20260909162102.1097006-3-shuicheng.lin@intel.com
Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
(cherry picked from commit 628f92b28bf4c371c10207daf6fc4caee0c0db2e)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Shuicheng Lin
3c90e42a01
drm/xe/shrinker: Return the freed page count through a parameter
__xe_shrinker_walk() and xe_shrinker_walk() return either the number of
pages freed or a negative error, so the two cannot be reported at once.
On error the pages already freed are dropped, and since xe_shrinker_scan()
only accumulates non-negative returns while *scanned is updated by
pointer, the shrinker tells mm that it scanned without freeing.

Accumulate the count into a caller-provided counter and return only the
status, so an error no longer discards what the walk had freed.

Fixes: 00c8efc318 ("drm/xe: Add a shrinker for xe bos")
Assisted-by: Claude:claude-opus-5
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Link: https://patch.msgid.link/20260909162102.1097006-2-shuicheng.lin@intel.com
Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
(cherry picked from commit d7aac1a0235a6ce41e30cec385e2db8c33dad12d)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Ilia Levi
d0c0952878
drm/xe/mmio_gem: fix destroy flow
xe_mmio_gem_destroy() currently frees the GEM object directly, bypassing
reference counting.  Since existing VMAs hold a reference and the fault
handler accesses the object through vma->vm_private_data, this is
use-after-free.  Additionally, nothing prevents the fault handler from
installing PTEs to the real MMIO after destroy.

Fix this with proper synchronization and refcounting. Also, do not set
vm_pgoff to zero. Many DRM drivers do this because helpers like
dma_mmap_pages() interpret vm_pgoff as an intra-buffer page offset;
leaving the DRM fake offset there would break these helpers.
Those drivers can get away with zeroing it because they map eagerly -
all PTEs are established before mmap returns, so vm_pgoff is never
consulted again. Our driver does not use such helpers and the newly
introduced call to drm_vma_node_unmap() relies on vm_pgoff being untouched.

v2: (Matt Auld)
- use dma_resv lock to serialize fault handler with destroy
- SIGBUS on access after destroy

Fixes: 1ffcf8b8ae ("drm/xe: Support for mmap-ing mmio regions")
Assisted-by: GitHub-Copilot:claude-opus-4.6
Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-16-matthew.auld@intel.com
(cherry picked from commit fb2ee38bab8025ad6a7a9cbb4635c5a178e4a7bc)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Ilia Levi
de40d31275
drm/xe/mmio_gem: cache the dummy page per object
Currently, when the fault handler provides a dummy page, it
allocates a new one on every invocation and ties its lifetime to
the drm_device via drmm_add_action_or_reset(). Concurrent faults
after hot-unplug therefore accumulate pages that persist until
device teardown.

Cache a single dummy page in the xe_mmio_gem object and use dma_resv
lock to protect its allocation. Free it with the object.

v2: use dma_resv lock to protect the allocation (Matt Auld)

Assisted-by: GitHub-Copilot:claude-opus-4.6
Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-15-matthew.auld@intel.com
(cherry picked from commit 8bf6213f9831e46313af4722a1ee6db1b7596378)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Shuicheng Lin
37fcbd7b2f
drm/xe/mmio_gem: Revoke drm_vma_node on xe_mmio_gem destroy
xe_mmio_gem_create() calls drm_vma_node_allow() but nothing ever calls
drm_vma_node_revoke(). The drm_vma_offset_file rb-tree entry allocated
by drm_vma_node_allow() is not freed by drm_gem_object_release(), so
it is leaked on every create/destroy cycle.

Add a struct drm_file * parameter to xe_mmio_gem_destroy() and call
drm_vma_node_revoke() from there, mirroring the drm_vma_node_allow()
call in xe_mmio_gem_create().

Fixes: 1ffcf8b8ae ("drm/xe: Support for mmap-ing mmio regions")
Suggested-by: Ilia Levi <ilia.levi@intel.com>
Assisted-by: Claude:claude-opus-4.6
Signed-off-by: Shuicheng Lin <shuicheng.lin@intel.com>
Reviewed-by: Ilia Levi <ilia.levi@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-14-matthew.auld@intel.com
(cherry picked from commit 32f0cb250598456d812fb7ca57a040282858323d)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Ilia Levi
0c50663403
drm/xe/mmio_gem: simplify fault handler loop
Make the iteration over the addresses in the VMA more explicit.
No functional change, as the VMA matches the GEM object exactly.

Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-13-matthew.auld@intel.com
(cherry picked from commit 6666ca9192f3bdf839aad33b9e1c9ebb7a222a29)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 09:09:01 -04:00
Ilia Levi
819f189265
drm/xe/mmio_gem: use write-back mapping for dummy page
Currently vmf_insert_pfn() maps the dummy page as UC, inheriting the
VMA's page protection which was set for the real MMIO region. This
conflicts with the direct map's WB mapping of the same page, creating a
cache type alias which is architecturally undefined on some platforms.

Use vmf_insert_pfn_prot() with a WB pgprot instead. Also simplify to
fault in the requested page instead of the whole VMA.

Fixes: 1ffcf8b8ae ("drm/xe: Support for mmap-ing mmio regions")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Closes: https://sashiko.dev/#/patchset/20260525125801.975038-6-ilia.levi%40intel.com
Assisted-by: GitHub-Copilot:claude-opus-4.6
Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-12-matthew.auld@intel.com
(cherry picked from commit 1e8e28e35df0e77ae1b22fc091c1f422f62fa5e9)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 08:49:00 -04:00
Ilia Levi
247a82da6f
drm/xe/mmio_gem: forbid VMA split
The fault handler assumes it always operates on a VMA spanning the entire
GEM object. This does not hold when the VMA has been split, e.g. by a
partial munmap or mprotect. In that case the handler may map wrong
physical pages or cause SIGBUS.

Handle this by forbidding VMA split, as partial unmaps are not deemed
useful for MMIO GEMs.

Suggested-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Ilia Levi <ilia.levi@intel.com>
Fixes: 1ffcf8b8ae ("drm/xe: Support for mmap-ing mmio regions")
Reviewed-by: Matthew Auld <matthew.auld@intel.com>
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Link: https://patch.msgid.link/20260908165046.1393557-11-matthew.auld@intel.com
(cherry picked from commit f3391a0b12d7bf826a0b21600d2f294f3dce4c14)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
2026-09-14 08:48:53 -04:00
Saim Shujah
a5b5cc9099 drm/msm/dpu: clear pending peripheral flush state
dpu_hw_ctl_clear_pending_flush() resets the cached per-block state after a
flush transaction, but misses pending_periph_flush_mask.

The peripheral flush updater accumulates interface bits in this mask. A
later transaction which sets the top-level peripheral flush bit can write
stale interface bits to CTL_PERIPH_FLUSH together with the current state.

Peripheral flush support was added after the helper started clearing every
individual pending flush mask. Clear the peripheral mask together with the
other cached child masks.

Fixes: 64f7b81f03 ("drm/msm/dpu: add support of new peripheral flush mechanism")
Cc: stable@vger.kernel.org
Signed-off-by: Saim Shujah <saimzst@gmail.com>
Patchwork: https://patchwork.freedesktop.org/patch/748968/
Link: https://lore.kernel.org/r/20260828065440.140410-1-saimzst@gmail.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:46:13 +03:00
William Bright
58995b11df drm/msm/dp: fix link bandwidth check when wide bus is enabled
msm_dp_display_mode_valid() halves the pixel clock when either YUV420 or
wide bus is in use, then uses that halved value both for the controller
pixel clock limit and for the DP link bandwidth check.

Only YUV420 halves the data crossing the link. Wide bus widens the
internal DPU to DP interface to two pixels per clock, halving the
controller clock. Every pixel is still transmitted, so the link
bandwidth requirement remains.

As a result, modes needing up to twice the available link bandwidth pass
validation. On the IMDT QCS8550 SBC (rev5 with CYPD6125), where DP runs
over USB-C alt mode where only two lanes are available, 3840x2160@60 was
accepted despite needing 9.6 Gbps against the 8.64 Gbps the link can
carry.

Use a separate link pixel clock that is only halved for YUV420 for the
bandwidth calculation, leaving the wide bus halving to apply solely to
the controller pixel clock limit. With this, 4k@60 is correctly rejected
and 4k@30 selected instead.

Fixes: df9cf852ca ("drm/msm/dp: account for widebus and yuv420 during mode validation")
Assisted-by: Claude:claude-opus-5
Signed-off-by: William Bright <william.bright@imd-tec.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/746145/
Link: https://lore.kernel.org/r/20260812-msm-dp-link-bw-v1-1-b0e3ce1190be@imd-tec.com
[DB: dropped useless comment]
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:24:21 +03:00
Guangshuo Li
6fbbf1e152 drm/msm/adreno: fix autosuspend cleanup during teardown
adreno_gpu_init() calls pm_runtime_use_autosuspend(), but
adreno_gpu_cleanup() does not call the matching
pm_runtime_dont_use_autosuspend() during teardown.

If the autosuspend delay is set to a negative value while autosuspend
is enabled, the runtime PM core increments usage_count to prevent
runtime suspend. Without calling pm_runtime_dont_use_autosuspend()
during teardown, this reference is not dropped and usage_count remains
unbalanced.

The documentation for pm_runtime_use_autosuspend() also notes that it
is important to undo it with pm_runtime_dont_use_autosuspend() at
driver exit time, unless runtime PM was initially enabled with
devm_pm_runtime_enable().

Add the missing pm_runtime_dont_use_autosuspend() call to
adreno_gpu_cleanup().

This issue was found by manual code inspection.

Fixes: eeb754746b ("drm/msm/gpu: use pm-runtime")
Cc: stable@vger.kernel.org
Signed-off-by: Guangshuo Li <lgs201920130244@gmail.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/745110/
Link: https://lore.kernel.org/r/20260808131624.2854412-1-lgs201920130244@gmail.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:17:13 +03:00
Jesse Casco
e249a6e2a1 drm/msm/dp: skip PUSH_IDLE when the link was never enabled
msm_dp_display_atomic_enable() returns early when link training fails,
leaving ->power_on false and the main link down.
msm_dp_display_atomic_disable() nevertheless writes DP_STATE_CTRL_PUSH_IDLE
and waits for an idle-pattern completion that cannot arrive, so every failed
enable is followed by "PUSH_IDLE pattern timedout".

Every other step of the teardown is already gated on that flag:
msm_dp_display_disable(), called from .atomic_post_disable(), returns early
on !power_on. The PUSH_IDLE write is the only one that is not, so the
controller's runtime-PM reference is then dropped without the link having
been taken down.

On glymur (Snapdragon X2 Elite) the consequence is not a warning. The SoC
does not survive it: TrustZone force-stops the SOCCP and ADSP remote
processors and the machine resets silently about 50 ms later, with no oops
and no panic. On an ASUS Zenbook A16 (UX3607OA), whose eDP panel does not
currently train, this reproduces without any compositor or GPU involvement:

  # eDP enable has already failed with "Failed link training (rc=-104)"
  echo 1 > /sys/class/graphics/fb0/blank

  [535.645455] === marker ===
  [535.694833] qcom_q6v5_pas d00000.remoteproc: fatal error received: \
                 sys_m_smsm.c:512:TZ force stop
  [535.694875] remoteproc remoteproc0: crash detected in soccp: type fatal error
  [535.728857] qcom_q6v5_pas 6800000.remoteproc: fatal error received: \
                 sys_m_smsm.c:783:err fatal notification received from TZ
  <SoC reset>

Gate the PUSH_IDLE write on ->power_on so the disable path is consistent
with the rest of the teardown. With this applied the same sequence is
harmless and the machine stays up; without it, it resets every time.

The unconditional write dates back to the original DP driver
(c943b4948b ("drm/msm/dp: add displayPort driver support")), but the
surrounding code has been restructured several times since, so no Fixes:
tag is offered.

Note that the eDP link-training failure that exposes this on the A16 is a
separate problem in the glymur eDP PHY and is reported separately; this
change is about not damaging the machine when training fails, for whatever
reason.

Tested on ASUS Zenbook A16 (UX3607OA), Snapdragon X2 Elite Extreme, on
linux-next next-20260803 and next-20260807. The machine has since been
running next-20260807 with this patch as its daily driver.

Assisted-by: Anthropic:Claude-Opus-5
Signed-off-by: Jesse Casco <jesse.casco@gmail.com>
Reviewed-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/745167/
Link: https://lore.kernel.org/r/20260808171325.133041-1-jesse.casco@gmail.com
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:16:20 +03:00
Dmitry Baryshkov
ea9dadeac7 drm/msm: mark the fbdev framebuffer as system memory
msm_fbdev_driver_fbdev_probe() points screen_buffer at a kernel virtual
mapping of the GEM object and uses the deferred sysmem fb ops, but never
sets FBINFO_VIRTFB. The framebuffer core then assumes the memory is not
in the virtual address space and warns on the first console draw:

  fb0: sys_fillrect: framebuffer is not in virtual address space.

The drm_fbdev_dma, drm_fbdev_shmem and drm_fbdev_ttm helpers all set the
flag for system memory. Do the same here.

Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
Assisted-by: Claude:claude-opus-5
Patchwork: https://patchwork.freedesktop.org/patch/743293/
Link: https://lore.kernel.org/r/20260730-drm-msm-fbinfo-virt-v1-1-a27099a6dc58@oss.qualcomm.com
Acked-by: Rob Clark <robin.clark@oss.qualcomm.com> # on IRC
Signed-off-by: Dmitry Baryshkov <dmitry.baryshkov@oss.qualcomm.com>
2026-09-14 00:11:24 +03:00
Linus Torvalds
180534c09b Rust fixes for v7.3 (2nd)
Toolchain and infrastructure:
 
  - Work around a 'bindgen' 0.73.2 bug that emits an 'allow' attribute
    for 'unnecessary_transmutes', which is unknown in older compilers.
 
  - Clean 'clippy::as_underscore' lints in generated code by the new
    'bindgen' 0.73.0+ releases.
 
  - Clean new 'clippy::needless_range_loop' lint for the upcoming Rust
    1.100.0 (expected 2026-11-12).
 
 'kernel' crate:
 
  - 'num' module: fix soundness issue in 'Bounded' by sealing the
    'Integer' trait.
 
 'pin-init' crate:
 
  - Fix unreachable warning for the upcoming Rust 1.100.0 (expected
    2026-11-12) due to 'Infallible' becoming an alias of '!'.
 
 Samples:
 
  - Add missing newlines in 'pr_*!'s macro calls.
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEPjU5OPd5QIZ9jqqOGXyLc2htIW0FAmqmuA8ACgkQGXyLc2ht
 IW0ZVRAAk75N61v8xzY5dsQjA0O0ivCxDBqrPnFYYOq9jWWwKR4XF8zfX7dxPzFG
 48NHlQ9s3XEOSfmoVdaab9DMz8l2gCLMcCUOqmEGZtf1ORlFqCn7m0OMXfsidgx9
 YIWYSAySpjaQ27bg8+uvbBlBmD2KaE6zBlrAKbvdC9dJBOMfLjEnT3wtzkRkROzo
 WJMyx+OjIk0kmFNMUPBV/J+VWyxP5IAl8C5xK/hl3L+tf0VeQWkn82f7zzoGfwRV
 xLuIybzlxF2QK6D8OSf+SpxIqgl1fCDxh2rzWyNBJKbdGn1fMTTY7Ci6rM2DK853
 PjmQWtlkrYIOnO7k2qdCebOOv8wOBKE1hNpK+23mkEUbsZjWPNgSHVuf6X098NuH
 GEk5okH6+1e2w80dSRfUjKPY2omYhNoq4/4KEC+0IcV3xV+9FLq1uo9K/eOEr0Cf
 z430H31YnollXCWUx56QJZ7p3r0dITwhKHPE9pfKB53yWZelTEboRuD1zRKYEO4E
 0f+bDuLOAJeCzSX66YteZ+DiWphNB4OX49TGKRJpo9gWKn6+29TKIbJjAuk/oAu9
 FVkE5//WAu8El5++1W0YE4AM5eVvhrarH4vo6q44o6bd3acNHHJDzwUR8fu3FSYA
 skFb9vcRAq6vUPlzotlQMbuCP1PfBuFIhJdfqDcsVzGJO8oQPuk=
 =/75K
 -----END PGP SIGNATURE-----

Merge tag 'rust-fixes-7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ojeda/linux

Pull Rust fixes from Miguel Ojeda:
 "Toolchain and infrastructure:

   - Work around a 'bindgen' 0.73.2 bug that emits an 'allow' attribute
     for 'unnecessary_transmutes', which is unknown in older compilers

   - Clean 'clippy::as_underscore' lints in generated code by the new
     'bindgen' 0.73.0+ releases

   - Clean new 'clippy::needless_range_loop' lint for the upcoming Rust
     1.100.0 (expected 2026-11-12)

  'kernel' crate:

   - 'num' module: fix soundness issue in 'Bounded' by sealing the
     'Integer' trait

  'pin-init' crate:

   - Fix unreachable warning for the upcoming Rust 1.100.0 (expected
     2026-11-12) due to 'Infallible' becoming an alias of '!'

  Samples:

   - Add missing newlines in 'pr_*!'s macro calls"

* tag 'rust-fixes-7.3-2' of git://git.kernel.org/pub/scm/linux/kernel/git/ojeda/linux:
  rust: allow `unknown_lints` in generated bindings for Rust < 1.88
  rust: allow `clippy::as_underscore` in the generated bindings
  rust: num: seal Integer
  drm/panic: clean new `clippy::needless_range_loop` lint for Rust 1.100.0
  rust: samples: add missing newlines in rust_print_main
  rust: pin-init: use irrefutable pattern for `stack_pin_init`
2026-09-13 09:28:28 -07:00
Rob Clark
ba970587a0 drm/msm/a6xx+: Increase GMU FW init timeout
We were using 10ms, kgsl uses 100ms.  In practice it is usually takes
less than 10ms, but very occasionally goes a bit above 10ms, leading to
a "GMU firmware inialization timed out", from which point things go
south.

There are probably some things that could be done to speed up init, like
increasing GMU freq.  But to be safe, increase the timeout to match
kgsl.

Signed-off-by: Rob Clark <robin.clark@oss.qualcomm.com>
Reviewed-by: Akhil P Oommen <akhilpo@oss.qualcomm.com>
Patchwork: https://patchwork.freedesktop.org/patch/753014/
Message-ID: <20260912150915.28700-1-robin.clark@oss.qualcomm.com>
2026-09-12 10:01:54 -07:00
Sajal Gupta
59ced288fc
drm/gud: fix out-of-bounds write in gud_plane_atomic_check()
The plane property loop uses req->properties[num_properties + i] as write
index while simultaneously incrementing `num_properties` inside the loop.
At iteration i, num_properties has also incremented by i, so the write
is done at `initial_num_properties + 2*i`, skipping every other index and
advancing by 2 per iteration.

With just 2 connector and 32 plane properties the last write happens at
index 64, one slot past the end of the 64-slot (indices 0–63)
allocation. A USB device can trigger OOB by advertising the maximum
number of properties.

Fix by dropping the redundant `+ i`; num_properties is already the correct
running index, as gud_connector_fill_properties() fills the preceding
slots.

Fixes: 40e1a70b4a ("drm: Add GUD USB Display driver")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://sashiko.dev/#/patchset/20260821071812.16500-1-sajal2005gupta%40gmail.com?part=1
Signed-off-by: Sajal Gupta <sajal2005gupta@gmail.com>
Cc: <stable@vger.kernel.org>
Acked-by: Ruben Wauters <rubenru09@aol.com>
Signed-off-by: Ruben Wauters <rubenru09@aol.com>
Link: https://patch.msgid.link/20260902123254.36987-1-sajal2005gupta@gmail.com
2026-09-12 13:56:44 +01:00
Sophie D
effce1cb87
drm/gud: Ignore damage clips in full update mode
When running in full update mode, previously small updates (such as
moving the mouse across the screen) would cause many full frames to be
generated. This would bog down the bus and lower the effective framerate
significantly - I was seeing a drop from 60 FPS to 2 FPS.

Set ignore_damage_clips in full update mode so the damage iterator
yields a single full-plane rectangle instead of one per clip.

Fixes: 73cfd166e0 ("drm/gud: Replace simple display pipe with DRM atomic helpers")
Cc: <stable@vger.kernel.org> # 6.18.x
Signed-off-by: Sophie D <patches@scd31.com>
Reviewed-by: Thomas Zimmermann <tzimmermann@suse.de>
Acked-by: Ruben Wauters <rubenru09@aol.com>
Signed-off-by: Ruben Wauters <rubenru09@aol.com>
Link: https://patch.msgid.link/20260910014910.8564-1-patches@scd31.com
2026-09-12 13:52:10 +01:00
Neil Armstrong
a2b6583798 drm/msm/a6xx: Use CX AO Counter register for timestamp on a750 GPUs
The a750 uses the GMU CX AO Counters instead of the GMU_ALWAYS_ON_COUNTER
register on A6xx and other A7xx GPUs, use it when running a A750 GPU.

The GMU_ALWAYS_ON_COUNTER at offset 0x1f888 doesn't seem to exist
on the SM8650 A750 GMU and returns 0, but the CX AO counter at offset
0x1f880 returns some proper timestamp data.

Signed-off-by: Neil Armstrong <neil.armstrong@linaro.org>
Patchwork: https://patchwork.freedesktop.org/patch/752235/
Message-ID: <20260909-topic-sm8650-gmu-a750-timestamp-reg-v2-2-091d74958951@linaro.org>
Signed-off-by: Rob Clark <robin.clark@oss.qualcomm.com>
2026-09-11 11:36:41 -07:00
Neil Armstrong
ed9ad3d418 drm/msm/a6xx: Add CX AO Counter registers used for a750 GPUs
The a750 uses the CX AO Counters instead of the GMU_ALWAYS_ON_COUNTER
register on A6xx and other A7xx GPUs.

Signed-off-by: Neil Armstrong <neil.armstrong@linaro.org>
Patchwork: https://patchwork.freedesktop.org/patch/752234/
Message-ID: <20260909-topic-sm8650-gmu-a750-timestamp-reg-v2-1-091d74958951@linaro.org>
Signed-off-by: Rob Clark <robin.clark@oss.qualcomm.com>
2026-09-11 11:36:41 -07:00
George Emmanuel Thomas
3163cd7253 drm/msm: remove stale perf counter XML TODO
The adreno_perfcntrs macro already takes the XML file stem as its second
argument, allowing each perf counter JSON file to select the appropriate
register XML. The a2xx and a5xx entries already pass their respective
XML file stems.

Remove the stale TODO comment.

Signed-off-by: George Emmanuel Thomas <georgeemmanuelthomas@gmail.com>
Patchwork: https://patchwork.freedesktop.org/patch/746745/
Message-ID: <20260815164335.158958-1-georgeemmanuelthomas@gmail.com>
Signed-off-by: Rob Clark <robin.clark@oss.qualcomm.com>
2026-09-11 11:33:27 -07:00
Lyude Paul
bbb9293c9b drm/nouveau/gsp: Increase delay for magic sleep in r535_gsp_fini()
As it turns out, Turing isn't the only architecture that needs this. On
this Dell Precision 7780 with an AD103 GPU, along with pretty much every
other laptop I tested, runtime PM is still somewhat unreliable. At first
glance it seems as if it's fixed, but lowering the autosuspend delay to
500ms and then doing a stress test of suspend/resume cycles on the GPU ends
up causing everything to start timing out.

After quite a lot of digging, I eventually landed back on this magic
timeout in r535_gsp_fini(). As it turns out, increasing the timeout ends up
fixing the runtime PM issues as far as I can tell, even during intense
stress testing.

Unfortunately after spending quite a bit of time trying to dig through
OpenRM to figure out what this magic sleep is actually doing, I've also
come up short with any reasonable explanation. In lieu of that, I'm going
to include the observations I did make while trying to figure this out in
hopes someone eventually does figure this out:

* The magic sleep has to occur after fbsr is initialized. Performing it at
  any time before that doesn't appear to work.
* In situations where runtime PM starts getting flaky, some rather
  interesting visual effects end up happening on occasion before the GPU
  fully falls over. In particular, squares that look like the result of an
  incomplete blitting operation to a tiled buffer end up showing up on
  applications like vkcube. Interestingly enough, they remain in precisely
  the same place between runtime PM cycles until the GPU falls over - even
  when restarting vkcube multiple times, and even when vkcube is actively
  updating the screen. Even more interestingly, they're not limited to a
  specific framebuffer - you can see the squares changing as the cube
  rotates around.
  We cannot however, say that this is likely to be a incomplete fbsr
  operation. The magic sleep happens before fbsr is actually saved (which
  happens on the GSP unload), so it's something else.
* During a short bit of testing with a desktop that I have, the magic sleep
  seemed to make no difference to whether or not suspend/resume works. It
  seems to generally work almost always. So we can assume this is likely
  exclusive to runtime PM, not S3.

As well, here's a list of the things I tried before settling on the magic
sleep:

* Hooking up NV2080_CTRL_CMD_INTERNAL_GCX_ENTRY_PREREQUISITE and then
  blocking runtime PM until OpenRM signals that GC6/GCOFF is ready appears
  to make no difference.
* Hooking up some (maybe not all, unsure about that part) bits of comptag
  saving including:
  * Fetching static memsys information from GSP
  * Adding the size of the comptag storage to the fbsr data
  * Adding a GA103+ workaround for disabling raw compression mode during
    fbsr (it doesn't seem like it applies for any systems I tried it on
    anyhow)
  * Setting bPreserveVideoMemoryAllocations=1 in GspSystemInfo

So, until we can figure this out properly - just sleep for longer.

Signed-off-by: Lyude Paul <lyude@redhat.com>
Fixes: 53dac06238 ("drm/nouveau/gsp: add support for 570.144")
Cc: <stable@vger.kernel.org> # v6.16+
Reviewed-by: Dave Airlie <airlied@redhat.com>
Link: https://patch.msgid.link/20260814194542.781955-5-lyude@redhat.com
(cherry picked from commit 09b47186a4164f3aaa3591313f80794443117342)
Signed-off-by: Lyude Paul <lyude@redhat.com>
2026-09-11 13:26:18 -04:00