Commit Graph

18270 Commits

Author SHA1 Message Date
Wentao Liang
2b86ab1bd6 drm/amdgpu: Fix runtime PM leak in amdgpu_debugfs_test_ib_show()
amdgpu_debugfs_test_ib_show() resumes the device with
pm_runtime_get_sync() before taking the reset domain semaphore with
down_write_killable().  If the write lock acquisition is interrupted,
the function returns without calling pm_runtime_put_autosuspend(),
leaking the runtime PM reference acquired for the device and keeping
the GPU awake.

Drop the runtime PM reference on the interrupted down_write_killable()
error path before returning.

Fixes: 6049db43d6 ("drm/amdgpu: change reset lock from mutex to rw_semaphore")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit ec30a576c2d4c0364549e6c04218f50704ef56c8)
Cc: stable@vger.kernel.org
2026-09-23 15:53:05 -04:00
Wentao Liang
b4f7b4459b drm/amdgpu: Fix last_update fence leak in amdgpu_vm_init()
amdgpu_vm_init() initializes vm->last_update, vm->last_unlocked and
vm->last_tlb_flush with references to the stub fence taken via
dma_fence_get_stub().  The error label at the end of the function
releases the last_unlocked and last_tlb_flush references with
dma_fence_put(), but the reference stored in vm->last_update is never
dropped, so whenever the page table root creation, the reservation of
the root BO or the PASID registration fails, the stub fence reference
leaks.

Drop the vm->last_update reference together with the other stub fence
references on the error path.

Fixes: 187916e6ed ("drm/amdgpu: install stub fence into potential unused fence pointers")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit e7979c84fc05a176bdf855ee664871b1648404c9)
Cc: stable@vger.kernel.org
2026-09-23 15:53:05 -04:00
Wentao Liang
a997baa611 drm/amdgpu: Fix acpi device leak in amdgpu_acpi_enumerate_xcc()
amdgpu_acpi_enumerate_xcc() looks up each XCC ACPI device with
acpi_dev_get_first_match_dev(), which takes a reference to the device.
The reference is dropped with acpi_dev_put() after the XCC info is
initialized, but if the kzalloc_obj() allocation of the XCC info fails
the function returns -ENOMEM without releasing the reference, leaking
the last reference to the ACPI device.

Drop the ACPI device reference on the allocation failure path before
returning.

Fixes: 4d5275ab0b ("drm/amdgpu: Add parsing of acpi xcc objects")
Reviewed-by: Lijo Lazar <lijo.lazar@amd.com>
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 9211ef48b31ec66999cf55e04d0cbc60cd855fd5)
Cc: stable@vger.kernel.org
2026-09-23 15:53:05 -04:00
Wentao Liang
aea841bc62 drm/amdgpu: Fix vmid_wait fence leak in amdgpu_ring_init()
amdgpu_ring_init() initializes ring->vmid_wait with a reference to the
stub fence taken via dma_fence_get_stub().  When a later step of the
initialization fails, e.g. amdgpu_fence_driver_init_ring(), a writeback
slot allocation or the ring buffer allocation, the function returns an
error without releasing the stub fence reference and the reference is
leaked if the ring is torn down without amdgpu_ring_fini().

Move the stub fence assignment to the end of the initialization, right
before the ring is registered with the GPU scheduler, where no further
failure is possible.  The stub fence is only consumed by command
submission handling in amdgpu_ids.c once the ring is up and running, so
nothing reads it during the error-prone part of the initialization.

Fixes: 48e9fbd1a2 ("drm/amdgpu: initialize the vmid_wait with the stub fence")
Signed-off-by: Wentao Liang <vulab@iscas.ac.cn>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit f2b96986851203e9c50ca0d13aaa3581ca3e8ebd)
Cc: stable@vger.kernel.org
2026-09-23 15:53:05 -04:00
Sunil Khatri
6b13ddbf5b drm/amdgpu/vcn4.0.3: fix video_timeout unit mismatch in jpeg reset wait
vcn_v4_0_3_reset_jpeg_pre_helper() passes adev->video_timeout directly
to amdgpu_fence_wait_polling(), whose timeout parameter is documented
and implemented in usecs (busy-wait loop decrementing by udelay(2)).

adev->video_timeout is set in jiffies by
amdgpu_device_get_job_timeout_settings(), via msecs_to_jiffies().
Passing it unconverted means the intended ~2s wait for outstanding
JPEG fences to complete before the JPEG queue is torn down actually
lasts only a couple of microseconds (HZ jiffies interpreted as usecs),
so pending jobs are almost never given a real chance to finish before
the reset path forces completion in the following helper.

Convert the jiffies value to usecs with jiffies_to_usecs() before
passing it to amdgpu_fence_wait_polling().

Fixes: d25c67fd9d ("drm/amdgpu/vcn4.0.3: rework reset handling")
Cc: Jesse.Zhang <Jesse.Zhang@amd.com>
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 5feabbd673c10ebee22b880e4d812f08974d2ef7)
Cc: stable@vger.kernel.org
2026-09-23 15:53:05 -04:00
Sunil Khatri
f952ed353a drm/amdgpu/vcn5.0.1: fix video_timeout unit mismatch in jpeg reset wait
vcn_v5_0_1_reset_jpeg_pre_helper() passes adev->video_timeout directly
to amdgpu_fence_wait_polling(), whose timeout parameter is documented
and implemented in usecs (busy-wait loop decrementing by udelay(2)).

adev->video_timeout is set in jiffies by
amdgpu_device_get_job_timeout_settings(), via msecs_to_jiffies().
Passing it unconverted means the intended ~2s wait for outstanding
JPEG fences to complete before the JPEG queue is torn down actually
lasts only a couple of microseconds (HZ jiffies interpreted as usecs),
so pending jobs are almost never given a real chance to finish before
the reset path forces completion in the following helper.

Convert the jiffies value to usecs with jiffies_to_usecs() before
passing it to amdgpu_fence_wait_polling().

Fixes: fab47d2db5 ("drm/amdgpu/vcn5.0.1: rework reset handling")
Cc: Jesse.Zhang <Jesse.Zhang@amd.com>
Assisted-by: Claude:claude-sonnet-5
Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit b8334fec8b90ebffcaa01001a23edca9f29a05e9)
Cc: stable@vger.kernel.org
2026-09-23 15:53:05 -04:00
Sunil Khatri
cd195f1616 drm/amdgpu/userq: fix double jiffies conversion in hang detect timeout
Function amdgpu_userq_start_hang_detect_work() calls msecs_to_jiffies()
on adev->gfx_timeout/compute_timeout/sdma_timeout before arming
hang_detect_work. These timeout values already hold jiffies values from
amdgpu_device_get_job_timeout_settings() at device init.

This silently shrinks the real hang-detect deadline to (2 * HZ) ms
instead of the intended timeout. e.g. 500ms instead of the 2000ms
default on a CONFIG_HZ=250 kernel, only coincidentally correct at
HZ=1000. The shortened window is easily exceeded by ordinary
fence-completion latency, causing hang_detect_work to fire and
trigger a per-queue or full GPU reset for queues that are not
actually hung.

Pass the jiffies value directly to queue_delayed_work() instead of
converting it a second time.

Fixes: fc3336be9c ("drm/amd/amdgpu: Add independent hang detect work for user queue fence")
Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 13d44ca033cb74756c2aef0ade54a75cdf2f6271)
Cc: stable@vger.kernel.org
2026-09-23 15:53:05 -04:00
Prike Liang
3022bdfe3e drm/amdgpu: move userq fence wait out of signalling section
The eviction fence suspend worker waits for every pending userq fence
from inside a dma_fence_begin_signalling() critical section. Waiting on
another DMA fence while responsible for signalling one violates the
cross-driver fence contract and is reported by lockdep as a
dma_fence_map dependency.

Move the wait before dma_fence_begin_signalling(). Keep userq_mutex held
so queue lifetime remains stable while inspecting last_fence.

Fixes: fc61df1516 ("drm/amdgpu: annotate eviction fence signaling path")
Signed-off-by: Prike Liang <Prike.Liang@amd.com>
Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 3bd4fbc5ed89621340b5cd249869092691a9c81f)
Cc: stable@vger.kernel.org
2026-09-23 15:53:05 -04:00
Chengjun Yao
5155002b03 drm/amdgpu: fix rmmio iounmap skipped on device removal
amdgpu_pci_remove() calls drm_dev_unplug() before fini_sw(), so
drm_dev_enter() is already false there and the iounmap() guarded by it
is skipped. This .remove path runs on both hot-unplug and plain rmmod,
so the register BAR ioremap mapping leaks one instance per unload.

Unmap rmmio unconditionally (guard only on non-NULL) and drop the now
unused idx.

Fixes: 62d5f9f711 ("drm/amdgpu: Unmap MMIO mappings when device is not unplugged")
Signed-off-by: Chengjun Yao <Chengjun.Yao@amd.com>
Reviewed-by: Asad Kamal <asad.kamal@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit dd6f86a97260e5207d3329ad03aa89fdad61b1e6)
Cc: stable@vger.kernel.org
2026-09-17 11:58:48 -04:00
Mario Limonciello
7f9caa70ae drm/amdgpu: Skip KFD mapping clear before initialization
amdgpu_amdkfd_clear_kfd_mapping() assumes that a non-NULL kfd_dev
has a fully populated node array. This is not true when KFD device
initialization fails after probe.

For example, kgd2kfd_device_init() sets num_nodes before checking
PCIe atomics support. On Polaris systems without the required atomics,
it returns before allocating nodes[0], but the kfd_dev remains attached
to the amdgpu device. A later GPU reset then dereferences nodes[0]->id.

Require the authoritative KFD initialization flag before walking the
node array, matching the existing KFD reset and teardown paths.

Fixes: 70cadefcc6 ("drm/amdgpu: unmap all user mappings of framebuffer and doorbell before mode1 reset")
Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5833
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 4ac1835823c47903fbb278bbf474773c46f59edc)
Cc: stable@vger.kernel.org
2026-09-17 11:58:04 -04:00
Mario Limonciello
04de4007d3 drm/amdgpu: Fix GPU PCIe link capability reporting
Commit eb53125a7a ("drm/amd: Add dedicated helper for
amdgpu_device_find_parent()") made amdgpu_device_gpu_bandwidth() query
the first device outside the dGPU. That is the host side of the
physical link, not the GPU side.

As a result, the ASIC and platform capability masks can both be based
on the host port. drm_amdgpu_info_device then exposes the host
capabilities to userspace, such as Gen5 x16 for a Gen4 x8 GPU.

Cache both ends of the physical link during device initialization.
Use link_dev for the GPU capability and link_partner for the platform
capability and _PR3 detection.

Reported-by: "Marek Olšák" <maraeo@gmail.com>
Closes: https://lore.kernel.org/amd-gfx/CAAxE2A4VhsAzzO1QjBjUg+NgnbD04ZzMyN6xsUJxjKJHH6hxiw@mail.gmail.com/
Suggested-by: Lijo Lazar <lijo.lazar@amd.com>
Fixes: eb53125a7a ("drm/amd: Add dedicated helper for amdgpu_device_find_parent()")
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 7ea6a47224e2c6e89a3a682d7fbaace4817a55aa)
Cc: stable@vger.kernel.org
2026-09-17 11:55:14 -04:00
Dmitriy Chumachenko
723d4dc628 drm/amdgpu: check ras and obj before dereference
nbio_v7_9_handle_ras_controller_intr_no_bifring() dereferences ras and obj
without checking either for NULL. Both amdgpu_ras_get_context() and
amdgpu_ras_find_obj() can return NULL, e.g. during the window between
adev->nbio.ras being set (early in amdgpu_ras_init(), by design, to
enable the fatal-error interrupt as soon as possible) and the PCIE_BIF
ras object actually being created in RAS late_init. Any interrupt in that
window crashes in hard-IRQ context.

This is analogous to commit d190b459b2 ("drm/amdgpu: the warning
dereferencing obj for nbio_v7_4"), which fixed the same issue in the
nbio_v7_4 handler.

Found by Linux Verification Center (linuxtesting.org) with SVACE.

Fixes: 7692e1ee24 ("drm/amdgpu: add RAS fatal error handler for NBIO v7.9")
Reviewed-by: Tao Zhou <tao.zhou1@amd.com>
Signed-off-by: Dmitriy Chumachenko <Dmitry.Chumachenko@cyberprotect.ru>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit c7071767a50a32ed727cf800ac84372429e3b4b3)
2026-09-17 11:54:31 -04:00
Mike Lothian
636139603b drm/amdgpu: hold a runtime PM reference for P2P dma-buf attachments
amdgpu_dma_buf_map() adds VRAM to the allowed domains for a peer2peer
attachment.  GTT is only a fallback placement when VRAM is preferred, so
ttm_bo_validate() migrates the buffer from GTT into VRAM.  While the
exporting device is runtime suspended its SDMA rings are down and the
move fails:

  amdgpu: Move buffer fallback to memcpy unavailable

An importer on a second GPU reaches this holding no runtime PM
reference on the exporter, e.g. a compositor on the APU submitting a
frame that references a buffer exported by an idle dGPU:

  amdgpu_cs_ioctl -> amdgpu_cs_parser_bos -> amdgpu_cs_bo_validate
    -> ttm_bo_validate -> amdgpu_bo_move -> dma_buf_map_attachment
      -> amdgpu_dma_buf_map -> ttm_bo_validate -> amdgpu_bo_move

Pinning a dma-buf into VRAM has the same requirement, which
commit 030631e97b ("drm/amdgpu: revert "take runtime pm reference
when we attach a buffer" v2") called out as the one case that would
need the reference back.

Take it in attach and drop it in detach.  pm_runtime_get_if_active()
never resumes the device, so it cannot deadlock against the reservation
taken during resume, which is why the old pm_runtime_get_sync() had to
go.  If the device is not active, clear peer2peer instead: the buffer
then stays in GTT, which remains accessible while the GPU is powered
down.  If runtime PM is disabled, take a plain reference so the put in
detach stays balanced.

Fixes: 030631e97b ("drm/amdgpu: revert "take runtime pm reference when we attach a buffer" v2")
Suggested-by: Christian König <christian.koenig@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Mike Lothian <mike@fireburn.co.uk>
Assisted-by: Claude:Opus-5 [Claude Code]
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 062ff15e30a48d14fb7d7558eba84f8dc97197f0)
Cc: stable@vger.kernel.org
2026-09-17 11:54:05 -04:00
Thadeu Lima de Souza Cascardo
829157e762 Revert "drm/amdgpu: debugfs: avoid extra EOLs in amdgpu_gem_info"
This reverts commit c119d05a36.

It removes the newline even when there are no fences attached to a
struct dma_resv, leading to multiple BOs being output on the same line,
making the debug file less readable, not more as the commit intended.

Signed-off-by: Thadeu Lima de Souza Cascardo <cascardo@igalia.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit a2aafaeb2be13ed3c893e6a44a3a5d26b251ae6a)
Cc: stable@vger.kernel.org
2026-09-10 12:54:59 -04:00
Arunpravin Paneer Selvam
87ceb8cba7 drm/amdgpu: skip the VMID 0 flush for VRAM
Clear-on-release only runs on VRAM, which amdgpu_ttm_map_buffer() reaches
via its direct MC address without programming a GART window, yet the wipe
still forces a VMID 0 flush. On GFX11 (e.g. Navi33) that spurious SDMA
flush can wedge the engine; only flush when a GART window is actually used.

v2: Let amdgpu_ttm_map_buffer() return whether the VMID 0 flush is needed,
    and drive the clear and copy paths from that. (Christian)
v3: Make the vm_needs_flush output parameter mandatory instead of
    allowing NULL. (Christian)

Fixes: a68c7eaa7a ("drm/amdgpu: Enable clear page functionality")
Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5413
Cc: Christian König <christian.koenig@amd.com>
Signed-off-by: Arunpravin Paneer Selvam <Arunpravin.PaneerSelvam@amd.com>
Reviewed-by: Christian König <christian.koenig@amd.com>
Reviewed-by: Timur Kristóf <timur.kristof@gmail.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit a306e406e570b74318ff7d80e5b07b540ca1d3a9)
Cc: stable@vger.kernel.org
2026-09-10 12:52:51 -04:00
Linus Torvalds
1fc5a74b10 kmalloc_obj conversions for v7.3-rc2
- Run scripts/coccinelle/api/kmalloc_objs.cocci for v7.3
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQRSPkdeREjth1dHnSE2KwveOeQkuwUCapuwVwAKCRA2KwveOeQk
 u5EeAP9TS7K4iVlw3KlZHuLIK2q+CQfALPepcu+ME2lO5dta4gEAxCTi0ZXmU7OT
 XbmWUd+DTkKNYCBW8E6Lvn72Er13uQs=
 =ZtN4
 -----END PGP SIGNATURE-----

Merge tag 'kmalloc_obj-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux

Pull kmalloc_obj conversions from Kees Cook:
 "Another run of the Coccinelle script for converting kmalloc()
  family of allocations to kmalloc_obj() via the existing rules
  in scripts/coccinelle/api/kmalloc_objs.cocci"

* tag 'kmalloc_obj-v7.3-rc2' of git://git.kernel.org/pub/scm/linux/kernel/git/kees/linux:
  treewide: refresh kmalloc_obj() conversions
  drm/amd/display: Fix harmless type mismatch in allocation
2026-09-05 20:45:18 -07:00
Kees Cook
3a2c4d55e3 treewide: refresh kmalloc_obj() conversions
This is another run of the Coccinelle script for converting kmalloc()
family of allocations to kmalloc_obj() via the existing rules in
scripts/coccinelle/api/kmalloc_objs.cocci

This catches both the set of kmalloc() uses added since the first
kmalloc_obj() conversions in v7.0 and adds a large group missed in the
first pass due to Coccinelle not interacting well with the cleanup.h
scoped_...() family of macros[1]. I worked around this with spatch's
"--macro-file" argument to a file with all the scoped_...() macros mapped
to Coccinelle's YACFE_ITERATOR[2] as that was the closest viable control
flow indicator I could find.

Build tested allmodconfig on x86, arm64, arm, loongarch, mips, powerpc,
riscv, and s390 with no new warnings.

Link: https://lore.kernel.org/lkml/202609021314.8A9C0B8@keescook/ [1]
Link: https://github.com/coccinelle/coccinelle/blob/master/standard.h [2]
Signed-off-by: Kees Cook <kees+treewide@kernel.org>
2026-09-04 21:37:00 -07:00
Sunil Khatri
8b4a4193f3 drm/amdgpu/userq: dont overwrite the error of subsequent map call
If a queue fails to map that we need to return the error code back
to the caller and not overwrite with a success specifically.

Accumulate the failure and return that.

Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 42a0197d10039e9518c0324c43331eb22b44d5f8)
2026-09-02 17:01:16 -04:00
Kanala Ramalingeswara Reddy
a26301203a drm/amdgpu: Skip accessing psp rum time db for APUs
Psp runtime DB is for dGPUs only.

Signed-off-by: Kanala Ramalingeswara Reddy <Kanala.RamalingeswaraReddy@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit dce8195027f146467c9378efb2bb1b0859cb735e)
Cc: stable@vger.kernel.org
2026-09-02 17:00:42 -04:00
Sunil Khatri
49a74a2388 drm/amdgpu: update the fw version for gfx12 userqueues
Update to the latest stable fw versions where userqueues
is working as it is expected with major fixes.

Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 69fa36e3ac92f2544ee7a1b719ec212b8247a2da)
Cc: stable@vger.kernel.org
2026-09-02 17:00:16 -04:00
Sunil Khatri
c748dd03df drm/amdgpu: update the fw version for gfx11 userqueues
Update to the latest stable fw versions where userqueues
is working as it is expected with major fixes.

Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit d50201b891604ab97f305d4a20d888ba93305b48)
Cc: stable@vger.kernel.org
2026-09-02 16:59:48 -04:00
Sunil Khatri
3b5c4f4a47 drm/amdgpu: fix byte/dword unit mismatch in coredump IB dump
In amdgpu_devcoredump_print_ibs(), the NO_CPU_ACCESS VRAM path passed
cursor.start/4 and cursor.size/4 to amdgpu_device_mm_access(), but that
function's pos/size parameters are byte offsets/lengths (confirmed by
amdgpu_ttm_vram_mm_access() and leading to wrong size calculation.

Similarly with that change the off index needs to be calculated
based on dword since that is a u32 type.

Fixes: 7b15fc2d1f ("drm/amdgpu: dump job ibs in the devcoredump")
Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Vitaly Prosyak <vitaly.prosyak@amd.com>
Acked-by: Christian König <christian.koenig@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 1bd613b0ed98a23575b18674c94b8b3392614681)
Cc: stable@vger.kernel.org
2026-09-02 16:59:25 -04:00
Sunil Khatri
90ce19bd11 drm/amdgpu: fix Idle BOs list in VM debugfs status info
amdgpu_debugfs_vm_bo_status_info() prints the "Idle BOs" section by
iterating lists->needs_update, the same list already printed just
above under "Moved BOs". struct amdgpu_vm_bo_status has a dedicated
idle list, populated whenever a BO's state machine settles, but it
was never read here, so genuinely idle BOs never show up in the
debugfs output and the "Idle BOs" section duplicates "Moved BOs"
instead.

Iterate lists->idle for the "Idle BOs" section.

Fixes: 4cdbba5a16 ("drm/amdgpu: restructure VM state machine v4")
Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 451bfc778a8c364841837def00ba15936f72762b)
Cc: stable@vger.kernel.org
2026-09-02 16:34:08 -04:00
Sunil Khatri
d6e16df7df drm/amdgpu: use AMDGPU_GPU_PAGE_SHIFT instead of PAGE_SHIFT
For different address types the variable PAGE_SHIFT might
not work well and it's better to use the GPU specific one

Signed-off-by: Sunil Khatri <sunil.khatri@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 3494b77d10375e0f9ab784e9b20763339844b55b)
Cc: stable@vger.kernel.org
2026-09-02 16:33:37 -04:00
Amber Lin
b428f83c7c drm/amdgpu: Update queue reset support version
Update queue reset required MES version for MES 12.1 to 0x7b since we
change the implementation from detect-and-reset method to
per-queue-reset method.

Signed-off-by: Amber Lin <amber.lin@amd.com>
Reviewed-by: Michael Chen <michael.chen@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 2160a5cbf0b7917adce4b55421306b614b4a2c8f)
2026-09-02 16:32:47 -04:00
Alex Deucher
7346a046c6 drm/amdgpu/gfx8: only apply compute quantums to KCQs
Don't apply to KIQ.  Seems to cause problems on KIQ
on some ARM platforms.

Closes: https://gitlab.freedesktop.org/drm/amd/-/work_items/5658
Fixes: 91cf34bc5a ("drm/amdgpu/gfx8: align mqd settings with KFD")
Reviewed-by: Jesse Zhang <jesse.zhang@amd.com>
Reviewed-by: Kent Russell <kent.russell@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit 6aae7bab029cdccae9a7157facfe36bfc35fc940)
Cc: stable@vger.kernel.org
2026-09-02 16:31:18 -04:00
Mario Limonciello
bd1f08246b drm/amdgpu: restrict BAR0 fallback read to SR-IOV VFs only
The BAR0 fallback read path was introduced as a workaround for SR-IOV VFs
where the VRAM aperture is not available during early init. Restrict this
workaround to only SR-IOV VFs where it's needed.

Reported-by: gloveless@jqluv.com
Fixes: cba4928cdf ("drm/amdgpu: reduce early full GPU access during SR-IOV init")
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260826185102.2269511-1-mario.limonciello@amd.com
Signed-off-by: Mario Limonciello <mario.limonciello@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
(cherry picked from commit d8a0affd207c813bd063fa2c27786f449eaf92b8)
2026-09-02 16:29:24 -04:00
Linus Torvalds
a99d741df7 drm next/fixes for 7.3-rc1
core:
 - use drm_warn instead of warn
 
 msm:
 - Bindings:
   - Added Shikra support
   - Document a840, a704, a722
 - Core:
   - Use drm_client buffers for fbdev emulation
   - teardown fixes
   - ARM32 DMA fixup
   - Remove objects from evict list when re-validated
   - Bunch of corner case and error path fixes
 - DPU:
   - Dropped dev_pm_opp_set_rate(0) preventing burnout
   - Fixed SSPP offsets of Kaanapali
 - DP:
   - Dropped dev_pm_opp_set_rate(0) preventing burnout
   - Cleaned up core code in preparation for MST support
   - Fixed prepare() to let Pipewire continue in case of the unplugged cable
 - GPU:
   - Add support for a704
   - Add support for a722
 - HDMI:
   - Simplifed register access
 
 amdgpu:
 - eGPU fixes
 - Runtime PM fix
 - UserQ fixes
 - Backlight fix
 - Discovery sysfs fix
 - Reset handling fixes
 - Buffer func handling fix for xgmi
 - VCN boundary check fix
 - DC lut handling fixes
 - MES fixes
 - UVD fix
 - VCE 3 fix
 - Enforce isolation fix
 - HPD fix for VGA/LVDS
 - DML fix
 - DCN 6 fixes
 - DC gpu reset fix
 
 amdkfd:
 - Fix return value
 - CU occupancy for GFX 11
 - CU occupancy for GFX 12/12.1
 - Queue bounds checking fix
 - SVM fixes
 - CRIU bounds checking fix
 
 radeon:
 - iMac display fix
 
 xe:
 - error message cleanups
 - i2c global register definitions as dependency for xe/i2c fixes
 - Media workardound
 - Add CCS to gt_idle debugfs print
 - Page fault related fix
 - i2c related fixes
 - System Controller mailbox bit fix
 -----BEGIN PGP SIGNATURE-----
 
 iQIzBAABCgAdFiEEEKbZHaGwW9KfbeusDHTzWXnEhr4FAmqR+2wACgkQDHTzWXnE
 hr5VPRAAhJpCnUEOUnhiBRfQB3lqHuKV4N9XQGVoaAZHYLxNkZD6I99ScpIznEr9
 sHG0ViBqz2PHyM2XoJzPZDzm8Us0moVUMMdV56IH7h11l8E7Az2Sd/Ji+AQvZEBT
 /qh9Py0/fjibfDm0ueMROFRhuD8RA2sJqzkGMZUBvErh+zmEQvkIDkIT6A5RpHD8
 B3XEGN+UCxPBzCnKNixyNDgY2i4ipFAe0MDj9+Mh0b9BM9BoV8+Eb7sjkBz+ROMH
 tl57Mjjd37FaYM9MgtEt9m7eOBf266V9Xb9tbIfqShwO1aZa0Vfzuih/3Ck53Nwc
 VkwepP7NHZZKFHIxFgcRCVjzyA3HpZJmLot8YYOyZDJOG6KKp2X0ZUFBog5Uiv8o
 7exd95FLUtIsnUtmLxlermOYJwrokgcVoigcalhLCB2+KZ5/BznutzcCSvrd8vI2
 LzEGnLMgzuEKeOmiWZjagID51TYwHfPogULUlwWT4XB62dKih7/uselCObj4m/Lc
 kRSbiXWz1pxisfiFRsK2GECwsbEb/8qOYgSOQBgppuPbSWpX+2P4IApomRBesQ8E
 Db+i6hptauOwT0/1ZBo8gkOfb0XIWsjg9iV6pVuUIS80hOlr7Zj1MR7mMrWgL0nM
 Chs4/Nq8jWU1ESbmKPBINIXCGOGuev4lJO/pBjfySLpTUhWOQcw=
 =wbEZ
 -----END PGP SIGNATURE-----

Merge tag 'drm-next-2026-08-29' of https://gitlab.freedesktop.org/drm/kernel

Pull more drm updates from Dave Airlie:
 "As mentioned last week, an msm pull request fell down the side of the
  couch or whatever the email equivalent of that is. This has the msm
  next stuff + the usual fixes for amd/intel.

  core:
   - use drm_warn instead of warn

  msm:
   - Bindings:
      - Added Shikra support
      - Document a840, a704, a722
   - Core:
      - Use drm_client buffers for fbdev emulation
      - teardown fixes
      - ARM32 DMA fixup
      - Remove objects from evict list when re-validated
      - Bunch of corner case and error path fixes
   - DPU:
      - Dropped dev_pm_opp_set_rate(0) preventing burnout
      - Fixed SSPP offsets of Kaanapali
   - DP:
      - Dropped dev_pm_opp_set_rate(0) preventing burnout
      - Cleaned up core code in preparation for MST support
      - Fixed prepare() to let Pipewire continue in case of the unplugged cable
   - GPU:
      - Add support for a704
      - Add support for a722
   - HDMI:
      - Simplifed register access

  amdgpu:
   - eGPU fixes
   - Runtime PM fix
   - UserQ fixes
   - Backlight fix
   - Discovery sysfs fix
   - Reset handling fixes
   - Buffer func handling fix for xgmi
   - VCN boundary check fix
   - DC lut handling fixes
   - MES fixes
   - UVD fix
   - VCE 3 fix
   - Enforce isolation fix
   - HPD fix for VGA/LVDS
   - DML fix
   - DCN 6 fixes
   - DC gpu reset fix

  amdkfd:
   - Fix return value
   - CU occupancy for GFX 11
   - CU occupancy for GFX 12/12.1
   - Queue bounds checking fix
   - SVM fixes
   - CRIU bounds checking fix

  radeon:
   - iMac display fix

  xe:
   - error message cleanups
   - i2c global register definitions as dependency for xe/i2c fixes
   - Media workardound
   - Add CCS to gt_idle debugfs print
   - Page fault related fix
   - i2c related fixes
   - System Controller mailbox bit fix"

* tag 'drm-next-2026-08-29' of https://gitlab.freedesktop.org/drm/kernel: (121 commits)
  drm/xe/sysctrl: Read mailbox phase bit from hardware
  drm/xe/i2c: Keep the i2c controller always enabled
  drm/xe/i2c: Fix the interrupt handling
  i2c: designware: Global register definitions
  drm/xe: Reject page faults from non-fault-mode scratch VMs
  drm/xe/xe_gt_idle: Add CCS to the powergating info print
  drm/xe: Do not apply WA 14025883347 to media 3503
  drm/amd/display: fix dc_lock leak on GPU reset error paths
  drm/amd/display: Fix redundant GPUVMEnable checks in dcn6 flip schedule
  drm/amd/display: Fix wrong bytes-per-pixel value for dml2_422_packed_10
  drm/amdkfd: guard against NULL restore_mqd in CRIU queue restore
  drm/amdgpu/userq: fix lock missing for userq fence error set
  drm/amdkfd: Fix the case that vm range is hole at svm_migrate_copy_to_vram
  drm/amdkfd: Fix error path at svm_migrate_copy_to_ram
  drm/amd/display: Log details when failing to register HPD IRQ
  drm/amd/display: Fix HPD consideration for VGA/LVDS connectors on DCE
  drm/amdgpu: clamp the isolation index for rings outside a partition
  drm/amdkfd: Reject zero-sized AQL queue allocations after size halving
  drm/amdgpu: Fix VCE 3 ring align_mask
  drm/kfd: Add CU occupancy support to GFX12.1
  ...
2026-08-28 16:37:55 -07:00
Linus Torvalds
18fbf5151d mm.git review status for linus..mm-stable
Everything:
 
 Total patches:       171
 Reviews/patch:       1.83
 Reviewed rate:       82%
 
 Excluding selftests:
 
 Total patches:       149
 Reviews/patch:       1.77
 Reviewed rate:       80%
 
 Excluding selftests and maple_tree:
 
 Total patches:       129
 Reviews/patch:       1.99
 Reviewed rate:       89%
 
 Summary of patch series in this merge:
 
 - "mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff"
   (Lorenzo Stoakes):
 
   Index MAP_PRIVATE file-backed folios by their anonymous page offset to
   resolve confusion around reverse mapping for zeroed and CoW'd
   file-backed memory.
 
   Use this new VMA anonymous page offset tracking to eliminate index
   conflicts and lay the foundation for scalable CoW performance
   improvements.
 
 - "promote mapped executable folios after first usage for MGLRU" (Baolin
   Wang):
 
   Make MGLRU's protection of mapped executable file folios more
   reliable.  Follow the classical LRU's logic, promoting mapped executable
   file folios after their first usage to give executable code a better
   chance to stay in memory and improve workload performance.
 
 - "mm: vmscan: fix node reclaim ignoring swappiness parameter" (Ridong Chen):
 
   Fix per-node proactive reclaim interface's ignoring the swappiness
   parameter when CONFIG_MEMCG is disabled by consolidating sc_swappiness()
   into a single function that checks proactive_swappiness regardless of
   kernel configuration.
 
 - "mm/vmscan: reduce lru_lock contention via vmstat-derived scan-balance
   cost" (Usama Arif):
 
   Reduce lru_lock contention in the reclaim path by deriving
   scan-balance costs from vmstat counters rather than lock-acquired
   producer updates.
 
   Read and decay these cost signals on the reclaim side under a
   dedicated per-lruvec lock, reducing total LRU lock wait time by over 60%
   without impacting scan throughput.
 
 - "zram: fix zram issues reported by sashiko" (Sergey Senozhatsky):
 
   Fix two low-risk zram bugs which Sashiko spotted in drive-by review.
 
 - "Honor XA_FLAGS_ACCOUNT in xas_split_alloc() and charge to folio's
   memcg" (Zi Yan):
 
   Fix xas_split_alloc() by enabling target folio memcg charging during
   splits and adding the missing __GFP_ACCOUNT flag for proper XArray node
   memory accounting.
 
 - "selftests/mm: use pattern matching in .gitignore" (Pratyush Mallick):
 
   Replace hardcoded binary names in selftests/mm/.gitignore with a
   generic pattern-matching rule to automatically ignore generated test
   files and avoid manual updates when adding new tests.
 
 - "mm/page_ext: remove pgdat_page_ext_init()" (Sang-Heon Jeon):
 
   Make the incompatibility between FLATMEM and NUMA explicit in
   mm/Kconfig and remove the unused pgdat_page_ext_init() function.
 
 - "zram: fix zstd error paths and add parameter validation" (Haoqin Huang):
 
   Clean up zram compression backends by removing redundant error
   cleanup, adding parameter and dictionary validation, auto-prefixing
   algorithm error logs, and resetting parameters prior to
   reinitialization.
 
 - "zram: fix stale scan bounds after reinitialization" (Longlong Xia):
 
   Prevent out-of-bounds slot accesses during concurrent zram resets by
   moving table scan bound calculations under dev_lock in writeback_store()
   and read_block_state().
 
 - "add anon mTHP collapse test cases" (Baolin Wang):
 
   Extend selftests helper functions to support arbitrary page orders and
   add new test cases and options for mTHP collapse in khugepaged.
 
 - "selftests/mm: Handle unsupported and transient test conditions"
   (Muhammad Usama Anjum):
 
   Update MM selftests to report a SKIP status instead of a failure when
   required kernel or filesystem features are unsupported, while adding
   retry logic for transient page migration errors.
 
 - "mm/zswap: Fixes and improves the zswap shrink" (Hao Jia):
 
   Fix the missing zswap global shrinker when CONFIG_MEMCG is disabled
   and extend shrink_memcg() to support batch writeback for improved
   writeback efficiency.
 
 - "alloc_tag: introduce IOCTL-based filtering for MAP" (Suren Baghdasaryan):
 
   Introduce an IOCTL-based binary interface for memory allocation
   profiling that enables kernel-side filtering before per-CPU counter
   aggregation.
 
   This eliminates the text-parsing overhead of /proc/allocinfo and
   provides up to a 20x speedup by transferring only filtered allocation
   data to userspace.
 
 - "better block swap batching and a different take on swap_ops v5"
   (Christoph Hellwig):
 
   Refactor block swap I/O to use swap_iocb for batching instead of
   single-bio requests and rebase the swap_ops interface, achieving faster
   swap throughput during kernel builds.
 
 - "mm: kmemleak: reduce transient false positives by confirming leaks"
   (Catalin Marinas):
 
   Reduce false-positive kmemleak reports by combining two kmemleak
   enhancements that add a second confirmation scan and a configurable
   minimum unreferenced scan count module parameter.
 
 - "mm: kmemleak: default min_unref_scans to 2 for verbose kernels"
   (Breno Leitao):
 
   Auto-scanning kernels can generate false-positive memory leak reports
   on single scans, so this patch defaults min_unref_scans to 2 when
   CONFIG_DEBUG_KMEMLEAK_VERBOSE is enabled to require a second confirming
   scan.
 
 - "swap_ops updates" (Christoph Hellwig):
 
   Batching I/O for synchronous swap devices causes performance
   regressions and filesystem-based swap suffers from double-indirection
   overhead.  This series resolves both issues by reintroducing per-folio
   writes for synchronous swap and allowing filesystems to directly export
   their own swap_ops.
 
 - "mm/khugepaged: several cleanups" (Nico Pache):
 
   khugepaged accumulated redundant state-checking patterns and outdated
   comments following mTHP integration.  Introduce dedicated helpers for
   PTE validation and event counting while refreshing the internal
   documentation.
 
 - "maple_tree: lock checking and clean ups" (Liam Howlett):
 
   Syzbot reports incorrectly blame memory management exit paths for
   locking bugs, maple tree erase operations risk allocation failures
   without gfp flags and internal documentation lacks clarity.
 
   Improve lock error detection, update docs, fix race and allocation
   edge cases and optimize erase allocations using a fallback to GFP_KERNEL
   | GFP_NOFAIL.
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQTTMBEPP41GrTpTJgfdBJ7gKXxAjgUCao9nJQAKCRDdBJ7gKXxA
 jk/9AQDlfevYJuSJmzAI8bt8ISG+/TfXMtIZC/MdbHqtQVYWPQD8Cvm3DUZsdGB/
 Gloq/HBFuMPgE8p2pwUIthdgnTPNvAc=
 =c+Nb
 -----END PGP SIGNATURE-----

Merge tag 'mm-stable-2026-08-26-15-22' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm

Pull more MM updates from Andrew Morton:

 - "mm/rmap: index MAP_PRIVATE file-backed folios by anonymous pgoff"
   (Lorenzo Stoakes)

   Index MAP_PRIVATE file-backed folios by their anonymous page offset
   to resolve confusion around reverse mapping for zeroed and CoW'd
   file-backed memory.

   Use this new VMA anonymous page offset tracking to eliminate index
   conflicts and lay the foundation for scalable CoW performance
   improvements.

 - "promote mapped executable folios after first usage for MGLRU"
   (Baolin Wang)

   Make MGLRU's protection of mapped executable file folios more
   reliable. Follow the classical LRU's logic, promoting mapped
   executable file folios after their first usage to give executable
   code a better chance to stay in memory and improve workload
   performance.

 - "mm: vmscan: fix node reclaim ignoring swappiness parameter" (Ridong
   Chen)

   Fix per-node proactive reclaim interface's ignoring the swappiness
   parameter when CONFIG_MEMCG is disabled by consolidating
   sc_swappiness() into a single function that checks
   proactive_swappiness regardless of kernel configuration.

 - "mm/vmscan: reduce lru_lock contention via vmstat-derived
   scan-balance cost" (Usama Arif)

   Reduce lru_lock contention in the reclaim path by deriving
   scan-balance costs from vmstat counters rather than lock-acquired
   producer updates.

   Read and decay these cost signals on the reclaim side under a
   dedicated per-lruvec lock, reducing total LRU lock wait time by over
   60% without impacting scan throughput.

 - "zram: fix zram issues reported by sashiko" (Sergey Senozhatsky)

   Fix two low-risk zram bugs which Sashiko spotted in drive-by review.

 - "Honor XA_FLAGS_ACCOUNT in xas_split_alloc() and charge to folio's
   memcg" (Zi Yan)

   Fix xas_split_alloc() by enabling target folio memcg charging during
   splits and adding the missing __GFP_ACCOUNT flag for proper XArray
   node memory accounting.

 - "selftests/mm: use pattern matching in .gitignore" (Pratyush Mallick)

   Replace hardcoded binary names in selftests/mm/.gitignore with a
   generic pattern-matching rule to automatically ignore generated test
   files and avoid manual updates when adding new tests.

 - "mm/page_ext: remove pgdat_page_ext_init()" (Sang-Heon Jeon)

   Make the incompatibility between FLATMEM and NUMA explicit in
   mm/Kconfig and remove the unused pgdat_page_ext_init() function.

 - "zram: fix zstd error paths and add parameter validation" (Haoqin
   Huang)

   Clean up zram compression backends by removing redundant error
   cleanup, adding parameter and dictionary validation, auto-prefixing
   algorithm error logs, and resetting parameters prior to
   reinitialization.

 - "zram: fix stale scan bounds after reinitialization" (Longlong Xia)

   Prevent out-of-bounds slot accesses during concurrent zram resets by
   moving table scan bound calculations under dev_lock in
   writeback_store() and read_block_state().

 - "add anon mTHP collapse test cases" (Baolin Wang)

   Extend selftests helper functions to support arbitrary page orders
   and add new test cases and options for mTHP collapse in khugepaged.

 - "selftests/mm: Handle unsupported and transient test conditions"
   (Muhammad Usama Anjum)

   Update MM selftests to report a SKIP status instead of a failure when
   required kernel or filesystem features are unsupported, while adding
   retry logic for transient page migration errors.

 - "mm/zswap: Fixes and improves the zswap shrink" (Hao Jia)

   Fix the missing zswap global shrinker when CONFIG_MEMCG is disabled
   and extend shrink_memcg() to support batch writeback for improved
   writeback efficiency.

 - "alloc_tag: introduce IOCTL-based filtering for MAP" (Suren
   Baghdasaryan)

   Introduce an IOCTL-based binary interface for memory allocation
   profiling that enables kernel-side filtering before per-CPU counter
   aggregation.

   This eliminates the text-parsing overhead of /proc/allocinfo and
   provides up to a 20x speedup by transferring only filtered allocation
   data to userspace.

 - "better block swap batching and a different take on swap_ops v5"
   (Christoph Hellwig)

   Refactor block swap I/O to use swap_iocb for batching instead of
   single-bio requests and rebase the swap_ops interface, achieving
   faster swap throughput during kernel builds.

 - "mm: kmemleak: reduce transient false positives by confirming leaks"
   (Catalin Marinas)

   Reduce false-positive kmemleak reports by combining two kmemleak
   enhancements that add a second confirmation scan and a configurable
   minimum unreferenced scan count module parameter.

 - "mm: kmemleak: default min_unref_scans to 2 for verbose kernels"
   (Breno Leitao)

   Auto-scanning kernels can generate false-positive memory leak reports
   on single scans, so this patch defaults min_unref_scans to 2 when
   CONFIG_DEBUG_KMEMLEAK_VERBOSE is enabled to require a second
   confirming scan.

 - "swap_ops updates" (Christoph Hellwig)

   Batching I/O for synchronous swap devices causes performance
   regressions and filesystem-based swap suffers from double-indirection
   overhead. This series resolves both issues by reintroducing per-folio
   writes for synchronous swap and allowing filesystems to directly
   export their own swap_ops.

 - "mm/khugepaged: several cleanups" (Nico Pache)

   khugepaged accumulated redundant state-checking patterns and outdated
   comments following mTHP integration. Introduce dedicated helpers for
   PTE validation and event counting while refreshing the internal
   documentation.

 - "maple_tree: lock checking and clean ups" (Liam Howlett)

   Syzbot reports incorrectly blame memory management exit paths for
   locking bugs, maple tree erase operations risk allocation failures
   without gfp flags and internal documentation lacks clarity.

   Improve lock error detection, update docs, fix race and allocation
   edge cases and optimize erase allocations using a fallback to
   GFP_KERNEL | GFP_NOFAIL.

* tag 'mm-stable-2026-08-26-15-22' of git://git.kernel.org/pub/scm/linux/kernel/git/akpm/mm: (172 commits)
  selftests/proc: make proc-maps-race work with READ_IMPLIES_EXEC
  memcg: move LRU size accounting on reparenting instead of copying it
  mm/vmscan: fix comment logic in balance_pgdat
  maple_tree: add helper mas_make_walkable()
  maple_tree: avoid extra gap calculation
  maple_tree: fix argument name in header
  maple_tree: change two GFP flags in tests
  maple_tree: document erase and allocations better
  maple_tree: avoid mas_erase() and mtree_erase() failures
  maple_tree: document that erase may use GFP_KERNEL for allocations
  maple_tree: catch race in mas_alloc_cyclic()
  maple_tree: add bulk parent set helper
  maple_tree: micro optimisation of mas_wr_store_type()
  maple_tree: optimise mas_wr_node_store() when not in rcu mode
  maple_tree: use prefetched value in mas_wr_store_type()
  maple_tree: clarify comments on mas_nomem()
  maple_tree: drop MAPLE_ALLOC_SLOTS
  maple_tree: drop dead code from mas_extend_spanning_null()
  maple_tree: documentation fix
  maple_tree: add write lock checking with lockdep sequence numbers
  ...
2026-08-27 09:17:06 -07:00
Prike Liang
a04ea08ddb drm/amdgpu/userq: fix lock missing for userq fence error set
amdgpu_userq_fence_driver() and amdgpu_userq_fence_driver_destroy()
don't acquire the dma_fence spinlock, so locking the dma_fence lock
before test the signaled state and set error state to avoid missing
lock assert error.

Signed-off-by: Prike Liang <Prike.Liang@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:22:14 -04:00
Xiang Liu
b309005666 drm/amdgpu: clamp the isolation index for rings outside a partition
adev->isolation[] has one slot per partition, but a ring that is not
assigned to one keeps AMDGPU_XCP_NO_PARTITION, which is ~0, so indexing
the array with it is out of bounds. SDMA submissions hit this on both
the isolation enforcement and the VM flush path and trip UBSAN.

Fall back to the first slot the way the cleaner shader path already
does, and stop taking the address before the ring type check that makes
it relevant.

Cc: stable@vger.kernel.org
Signed-off-by: Xiang Liu <xiang.liu@amd.com>
Reviewed-by: Hawking Zhang <Hawking.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:19:30 -04:00
Sunday Clement
40ba09e111 drm/amdkfd: Reject zero-sized AQL queue allocations after size halving
KFD_IOC_ALLOC_MEMORY_OF_GPU with flag
KFD_IOC_ALLOC_MEM_FLAGS_AQL_QUEUE_MEM and size=1 triggers the AQL
wraparound workaround (size >>= 1), reducing size to 0. The resulting
zero passes through PAGE_ALIGN(0) = 0 without validation, bypassing the
per-process VRAM quota check in reserve_mem_limit()
(vram_used + 0 > vram_available is always false).

The fix adds post-halving zero-size validation in the primary
allocation path (amdgpu_amdkfd_gpuvm.c). The check happens after size
halving but before reserve_mem_limit(), and uses err_alignment_size
error path to properly clean up the allocated kgd_mem structure and
mutex.

Cc: stable@vger.kernel.org
Signed-off-by: Sunday Clement <Sunday.Clement@amd.com>
Reviewed-by: Alex Deucher <Alexander.Deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:18:44 -04:00
David Rosca
2ee9836545 drm/amdgpu: Fix VCE 3 ring align_mask
The largest frame is 20 dwords, so 0xf mask is too small.
This was always wrong, but we were lucky with the VCE_CMD_END
commands inserted after fence and vm_flush.

Fixes: 8897ea8c76 ("drm/amdgpu: Implement insert_end for VCE 3")
Cc: stable@vger.kernel.org
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: David Rosca <david.rosca@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:17:35 -04:00
David Belanger
fb62f7f031 drm/kfd: Add CU occupancy support to GFX12.1
Port changes from GFX9 to GFX12.1 mostly as-is.
Minor changes to register access code.

Assisted-by: Claude:Sonnet 4.6
Signed-off-by: David Belanger <david.belanger@amd.com>
Reviewed-by: Sreekant Somasekharan <Sreekant.Somasekharan@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:17:32 -04:00
David Belanger
fb1e65a80d drm/kfd: Add CU occupancy support to GFX12
Port changes from GFX9 to GFX12 mostly as-is.
Minor changes to register access code.

Assisted-by: Claude:Sonnet-4-6
Signed-off-by: David Belanger <david.belanger@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Reviewed-by: Sreekant Somasekharan <Sreekant.Somasekharan@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:17:29 -04:00
David Belanger
290e0be2ab drm/kfd: Add CU occupancy support to GFX11
Port changes from GFX9 to GFX11 mostly as-is.
Minor changes to register access code.

Assisted-by: Claude:Sonnet-4-6
Signed-off-by: David Belanger <david.belanger@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Reviewed-by: Sreekant Somasekharan <Sreekant.Somasekharan@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:17:21 -04:00
Bob Zhou
6760f5cb12 drm/amdgpu: avoid force-completing uninitialized UVD rings
uvd_v7_0_sw_init() does not initialize the UVD decode ring for an
SR-IOV VF. However, amdgpu_uvd_resume() unconditionally force-completes
the decode ring when restoring its fence sequence.

Skip fence completion when the fence driver is not initialized.

Fixes: 0a33b11d26 ("drm/amdgpu: mark force completed fences with -ECANCELED")
Cc: stable@vger.kernel.org
Signed-off-by: Bob Zhou <bobzhou2@amd.com>
Acked-by: Leo Liu <leo.liu@amd.com>
Acked-by: Frank Min <Frank.Min@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:14:53 -04:00
Jesse Zhang
d36fbf8218 drm/amdgpu/userq: lock and validate wptr BOs before reading their GPU offset on restore
On resume, amdgpu_userq_vm_validate_and_restore_queue() updates each queue's
wptr GPU address via amdgpu_bo_gpu_offset().

WPTR BOs are VM-mapped, but each BO has its own reservation object and is not
implicitly covered by the VM validation path here. This can leave offset reads
without proper BO locking/placement state and trigger WARN_ONs.
  ------------[ cut here ]------------
  WARNING: amdgpu_object.c:1486 at amdgpu_bo_gpu_offset+0x75/0xa0 [amdgpu], CPU#3: kworker/3:1/116
  Workqueue: events amdgpu_userq_restore_worker [amdgpu]
  RIP: 0010:amdgpu_bo_gpu_offset+0x75/0xa0 [amdgpu]
  Call Trace:
   <TASK>
   amdgpu_userq_vm_validate_and_restore_queue+0x629/0x960 [amdgpu]
   amdgpu_userq_restore_worker+0xa6/0x180 [amdgpu]
   process_scheduled_works+0xa6/0x460
   worker_thread+0x13c/0x290
   kthread+0xfb/0x140
   ret_from_fork+0x1b6/0x2b0
   ret_from_fork_asm+0x1a/0x30
   </TASK>
  ---[ end trace 0000000000000000 ]---
  ------------[ cut here ]------------
  WARNING: amdgpu_object.c:1485 at amdgpu_bo_gpu_offset+0x9a/0xa0 [amdgpu], CPU#2: kworker/2:1/127
  Workqueue: events amdgpu_userq_restore_worker [amdgpu]
  RIP: 0010:amdgpu_bo_gpu_offset+0x9a/0xa0 [amdgpu]

Add each queue's WPTR BO to the drm_exec ww context and validate it to its
allowed placement before the later offset update.

v2:
- Clarify that WPTR BOs are VM-mapped (fix incorrect "not part of VM" wording). (Christian)
- Describe both parts of the fix: lock BO reservations in drm_exec and
  validate BO placement before offset reads.

Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Jesse Zhang <Jesse.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:14:20 -04:00
Prike Liang
aef2ca9353 drm/amdgpu/mes: fix the inconsistent indenting for mes_userq_map()
Fix the inconsistent indenting warning for mes_userq_map().

Fixes: d0827dda8f ("drm/amdgpu/mes: refactor the amdgpu_mes_alloc/free_proc|gang()")
Reported-by: kernel test robot <lkp@intel.com>
Closes: https://lore.kernel.org/oe-kbuild-all/202608190252.8XCa0HqR-lkp@intel.com/
Signed-off-by: Prike Liang <Prike.Liang@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-25 18:12:53 -04:00
Lorenzo Stoakes (ARM)
51943a18ad mm: provide vma_[flags_]is_cow_mapping() and remove is_cow_mapping()
All remaining callers of is_cow_mapping() are invoking it in the form of
is_cow_mapping(vma->vm_flags) or an indirected version of this.

Therefore, provide a helper - vma_is_cow_mapping() to directly test the
VMA.

Additionally provide a new helper vma_flags_is_cow_mapping() which
performs the check using the new vma_flags_t type, and share this logic
between vma_is_cow_mapping() and vma_desc_is_cow_mapping().

With these changes, no callers of is_cow_mapping() remain, so remove it.

Also update the userland VMA tests to reflect the change.

No functional change intended.

[akpm@linux-foundation.org: fix kerneldoc comment typo, per Lorenzo]
  Link: https://lore.kernel.org/aob1goSSPH6sTN9y@gremlin
Link: https://lore.kernel.org/20260813-b4-scalable-cow-virt-pgoff-v5-2-c21581c0c3c8@kernel.org
Signed-off-by: Lorenzo Stoakes (ARM) <ljs@kernel.org>
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
Cc: Adrian Hunter <adrian.hunter@intel.com>
Cc: Alexander Deucher <alexander.deucher@amd.com>
Cc: Alexander Gordeev <agordeev@linux.ibm.com>
Cc: Alexander Shishkin <alexander.shishkin@linux.intel.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: Arnaldo Carvalho de Melo <acme@kernel.org>
Cc: Arnd Bergmann <arnd@arndb.de>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Baoquan He <baoquan.he@linux.dev>
Cc: Barry Song <baohua@kernel.org>
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Chengming Zhou <chengming.zhou@linux.dev>
Cc: Chris Li <chrisl@kernel.org>
Cc: Christan König <christian.koenig@amd.com>
Cc: Christian Borntraeger <borntraeger@linux.ibm.com>
Cc: Claudio Imbrenda <imbrenda@linux.ibm.com>
Cc: Dave Airlie <airlied@gmail.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Gerald Schaefer <gerald.schaefer@linux.ibm.com>
Cc: Greg Kroah-Hartman <gregkh@linuxfoundation.org>
Cc: Gregory Price (Meta) <gourry@gourry.net>
Cc: Harry Yoo <harry@kernel.org>
Cc: Heiko Carstens <hca@linux.ibm.com>
Cc: Huang Ray <Ray.Huang@amd.com>
Cc: "Huang, Ying" <ying.huang@linux.alibaba.com>
Cc: Ian Rogers <irogers@google.com>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: James Clark <james.clark@linaro.org>
Cc: Jan Kara <jack@suse.cz>
Cc: Jann Horn <jannh@google.com>
Cc: Janosch Frank <frankja@linux.ibm.com>
Cc: Jason Gunthorpe <jgg@ziepe.ca>
Cc: Jiri Olsa <jolsa@kernel.org>
Cc: John Hubbard <jhubbard@nvidia.com>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Kairui Song <kasong@tencent.com>
Cc: Kees Cook <kees@kernel.org>
Cc: Kemeng Shi <shikemeng@huaweicloud.com>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Liam R. Howlett <liam@infradead.org>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Maarten Lankhorst <maarten.lankhorst@linux.intel.com>
Cc: Marc Rutland <mark.rutland@arm.com>
Cc: "Masami Hiramatsu (Google)" <mhiramat@kernel.org>
Cc: Matthew Auld <matthew.auld@intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Matthew Wilcox (Oracle) <willy@infradead.org>
Cc: Maxime Ripard <mripard@kernel.org>
Cc: Miaohe Lin <linmiaohe@huawei.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Muchun Song <muchun.song@linux.dev>
Cc: Namhyung kim <namhyung@kernel.org>
Cc: Naoya Horiguchi <nao.horiguchi@gmail.com>
Cc: Nhat Pham <nphamcs@gmail.com>
Cc: Nico Pache <npache@redhat.com>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Oscar Salvador <osalvador@suse.de>
Cc: Pedro Falcato <pfalcato@suse.de>
Cc: Peter Xu <peterx@redhat.com>
Cc: Peter Zijlstra <peterz@infradead.org>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Rik van Riel <riel@surriel.com>
Cc: Rodrigo Vivi <rodrigo.vivi@intel.com>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Sven Schnelle <svens@linux.ibm.com>
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Thomas Zimemrmann <tzimmermann@suse.de>
Cc: Vasily Gorbik <gor@linux.ibm.com>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: xu xin <xu.xin16@zte.com.cn>
Cc: Zi Yan <ziy@nvidia.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
2026-08-24 18:42:50 -07:00
Dave Airlie
22e48eeaf4 amd-drm-next-7.3-2026-08-19:
amdgpu:
 - eGPU fixes
 - Runtime PM fix
 - UserQ fixes
 - Backlight fix
 - Discovery sysfs fix
 - Reset handling fixes
 - Buffer func handling fix for xgmi
 - VCN boundary check fix
 - DC lut handling fixes
 
 amdkfd:
 - Fix return value
 
 radeon:
 - iMac display fix
 -----BEGIN PGP SIGNATURE-----
 
 iHUEABYKAB0WIQQgO5Idg2tXNTSZAr293/aFa7yZ2AUCaoX2XgAKCRC93/aFa7yZ
 2McyAP45fRlBRPJeNkk5CDWFRFFJXGwmkiKcBAogxFYgzpJ7SAD7BTbXwPGugPPr
 scBCM6qcx8jBH4HSLZnV1gq5tGtcOAc=
 =ayyI
 -----END PGP SIGNATURE-----

Merge tag 'amd-drm-next-7.3-2026-08-19' of https://gitlab.freedesktop.org/agd5f/linux into drm-next

amd-drm-next-7.3-2026-08-19:

amdgpu:
- eGPU fixes
- Runtime PM fix
- UserQ fixes
- Backlight fix
- Discovery sysfs fix
- Reset handling fixes
- Buffer func handling fix for xgmi
- VCN boundary check fix
- DC lut handling fixes

amdkfd:
- Fix return value

radeon:
- iMac display fix

Signed-off-by: Dave Airlie <airlied@redhat.com>

From: Alex Deucher <alexander.deucher@amd.com>
Link: https://patch.msgid.link/20260819183622.2406038-1-alexander.deucher@amd.com
2026-08-24 09:13:44 +10:00
Dave Airlie
c44e278ce0 Linux 7.2
-----BEGIN PGP SIGNATURE-----
 
 iQFSBAABCgA8FiEEq68RxlopcLEwq+PEeb4+QwBBGIYFAmqCLGoeHHRvcnZhbGRz
 QGxpbnV4LWZvdW5kYXRpb24ub3JnAAoJEHm+PkMAQRiGJzYH/0SFjcgnk1Z3Km+3
 2kEeGAMETajW41W7+5QQkuHk83UXDxigDRoD857/d8utK90GrZAoTMS9/6zF3tra
 ht4G1yc2x7/xgVLkWii54d/sp1LEWTRDntN95fzYZwbeAXwd0AcYBlKXZYHKl4t/
 4yZCgYPmYTkewaYdbyWNPiZvCwhBUl5k1E9i/drh5IJXdgXRcqoO86FY9JX+Ks9x
 r0g+d6RIiSbDfwzgRpkBn0TRnqzh2OeBfgyrsgGZO2axwlKcA7SP0vwwT6c6nOUI
 s8F2xXrqrUI75JbSI4YbdwOSvktwbtkz83idlRAYBOdxof3LJ6i2YaxrT8iG+KUH
 l7+e18M=
 =eQMh
 -----END PGP SIGNATURE-----

BackMerge tag 'v7.2' into drm-next

Linux 7.2

There was a lot of conflicts this round between fixes and next,
and I'd like to get the merge resolutions that we have in drm-tip.

Signed-off-by: Dave Airlie <airlied@redhat.com>
2026-08-20 10:58:44 +10:00
Alex Deucher
0e4ef0ead6 drm/amdgpu: handle pipeline sync without a VM fence
If we end up emitting a VM fence keep pipeline sync
associated with that fence.  If not, emit them as
part of the IB fence.

v2: fix need_pipe_sync handling
v3: simplify the function

Cc: David Rosca <david.rosca@amd.com>
Fixes: cb1e657cca ("drm/amdgpu: handle GDS and SPM without a VM fence")
Reviewed-by: David Rosca <david.rosca@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-19 10:13:33 -04:00
David (Ming Qiang) Wu
4d73905308 drm/amdgpu/vcn: fix integer overflow in dec_msg buffer count check
If the supplied msg[2] (num_buffers) is 0x3FFFFFFF, the expression
6 + num_buffers * 4 wraps to 2 and the bounds check passes, letting
the parser loop far past the end of the message BO. Triggering it
additionally requires a ~4GiB mapping so that msg[1] survives the
earlier "header does not fit in BO" check.

Rewrite the test in division form, which is overflow-free by
construction. Also update the message to reflect that msg is invalid.

Fixes: b193019860 ("drm/amdgpu/vcn3: Prevent OOB reads when parsing dec msg")
Fixes: 0a78f2bac1 ("drm/amdgpu/vcn4: Prevent OOB reads when parsing dec msg")
Cc: stable@vger.kernel.org
Signed-off-by: David (Ming Qiang) Wu <David.Wu3@amd.com>
Reviewed-by: Leo Liu <leo.liu@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-19 10:12:34 -04:00
Zhu Lingshan
5c082f4cd1 drm/amdgpu: fix hang and race in userq destroy
When a queue is hung, the hang_detect_work is the
only way to recover it. However in amdgpu_userq_destroy(),
the hang_detect_work is cancelled too early,
resulting in amdgpu_userq_wait_for_last_fence()
may never return, leaving an uninterruptible dma_fence_wait()
hang there.

To fix this problem, this commit moves the cancelling of
hang_detect_work after amdgpu_userq_wait_for_last_fence(), and it has
to be before the unmap helper, because hang_detect_work resets the
queue, so it races with amdgpu_userq_unmap_helper() for MES operations
and queue state.

This commit splits amdgpu_userq_cleanup() into two parts:

1) amdgpu_userq_detach_doorbell(), which detaches the queue from
userq_doorbell_xa. This has to be called before the cancel, otherwise
the IRQ handlers (for example amdgpu_userq_process_fence_irq)
can re-schedule the hang_detect_work and the cancel is not final.

2) amdgpu_userq_fence_driver_free(), this has to be called after the
unmap helper, because it can release the seq64 slot that the GPU
writes fence values to.

Only one cancel_delayed_work_sync(&queue->hang_detect_work) is needed,
so other redundancies are removed.

Signed-off-by: Zhu Lingshan <lingshan.zhu@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-19 10:12:12 -04:00
Pierre-Eric Pelloux-Prayer
c675dea86a drm/amdgpu: delay ttm buffer func enablement on xgmi
When amdgpu_init_minimal_xgmi is used, SDMA engines init
is delayed so amdgpu_ttm_enable_buffer_funcs must be
called later.

Without this, the check for num_buffer_funcs_scheds will
fail and using ttm buffer funcs later will fail.

Given that amdgpu_ttm_enable_buffer_funcs is a no-op if
amdgpu_in_reset() returns true, the call has to occur
after the reset lock is dropped.

Cc: stable@vger.kernel.org
Fixes: e4029f7a94 ("drm/amdgpu: only use working sdma schedulers for ttm")
Signed-off-by: Pierre-Eric Pelloux-Prayer <pierre-eric.pelloux-prayer@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-19 10:09:46 -04:00
Yang Wang
8587d48d69 drm/amdgpu: check thunderbolt before switcheroo registration
Introduce a helper to consolidate the vga_switcheroo registration condition
used by the init and fini paths.

Keep the explicit pci_is_thunderbolt_attached() check, as dev_is_removable()
does not provide equivalent coverage for Thunderbolt-attached GPUs.
This ensures such devices remain excluded from switcheroo registration while
preserving the existing PX and Apple gmux handling.

Cc: stable@vger.kernel.org
Signed-off-by: Yang Wang <kevinyang.wang@amd.com>
Reviewed-by: Kenneth Feng <kenneth.feng@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-19 10:08:54 -04:00
Jesse Zhang
fd65d17429 drm/amdgpu: force complete the KIQ ring fences on reset
Like the MES scheduler ring, the KIQ ring sets no_scheduler = true and uses a
polling fence, so it is skipped by the force-completion loop in
amdgpu_device_pre_asic_reset(). Its hw fence value lives in wb (GTT) memory and
survives a MODE1 reset while fence_drv.sync_seq keeps advancing, so after a
reset the first KIQ submission can poll forever on a seq that is never written
back.

Force complete the KIQ ring fences too so their hw fence is realigned to
sync_seq.

Cc: stable@vger.kernel.org
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Suggested-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Jesse Zhang <Jesse.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-19 10:08:26 -04:00
Zhu Lingshan
556488b086 drm/amdgpu: validate rptr and wptr of a userq
rptr and wptr of a userq are 8 bytes aligned, and may
not placed on a page boundary.

This commit checks whether rptr and wptr are 8 bytes
aligned, and expectes 8 bytes when validates rptr/wptr VA.

With above changes, this commit fixes an regression
in amdgpu_userq_input_va_validate, where
end_addr is caculated by:
check_add_overflow(start_addr, expected_size - 1, &end_addr).
Wptr and rptr are very likely not to be page aligned,
when validating rptr and wptr, if they are located in the last
mapped page(or only one page is mapped)
and expected_size is PAGE_SIZE, end_addr will exceed the last
mapped page, means (end_addr >> AMDGPU_GPU_PAGE_SHIFT) > va_map->last,
and causing an -EINVAL, even it is a valid VA.

Signed-off-by: Zhu Lingshan <lingshan.zhu@amd.com>
Acked-by: Alex Deucher <alexander.deucher@amd.com>
Fixes: c0122bf2cc ("drm/amdgpu: fix userq VA validation for sub-page buffers")
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-19 10:06:50 -04:00
Jesse Zhang
48dc279c30 drm/amdgpu: force complete the MES ring fences on reset
The MES scheduler ring has no drm scheduler (no_scheduler = true), so it is
skipped by the force-completion loop in amdgpu_device_pre_asic_reset(). It uses
a polling fence whose hw value lives in wb (GTT) memory and survives a MODE1
reset, while fence_drv.sync_seq keeps advancing for every packet.

When the reset is triggered because MES itself stopped responding, the
timed-out packets advance sync_seq past the last hw fence value MES wrote.
After resume the first MES submission polls forever on a seq that is never
written back, failing the resume and wedging the box on a second reset:

  amdgpu: MES ring buffer is full.
  amdgpu: *ERROR* ring gfx_0.0.0 test failed (-110)
  amdgpu: resume of IP block <gfx_v11_0> failed -110
  amdgpu: GPU reset end with ret = -110

Force complete the MES scheduler ring fences together with the scheduler rings
so their hw fence is realigned to sync_seq.

v2: cover all XCCs (one scheduler ring each), not just mes.ring[0].

Cc: stable@vger.kernel.org
Signed-off-by: Jesse Zhang <Jesse.Zhang@amd.com>
Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
2026-08-19 10:06:23 -04:00