mirror of
https://github.com/torvalds/linux.git
synced 2026-09-28 12:02:03 +02:00
drm/xe/vm: nuke PTs only after unlinking contested VMAs
In xe_vm_close_and_put(), external-BO VMAs are queued on the contested
list for deferred destruction via xe_vma_destroy_unlocked(). However,
xe_vm_pt_destroy() was previously invoked before processing contested
VMAs, destroying vm->pt_root while those VMAs were still linked to their
respective buffer objects (vm_bo->list.gpuva).
If a concurrent thread evicts one of those shared buffer objects,
xe_bo_trigger_rebind() holding only bo->resv walks the BO's VMAs and, in
fault mode, calls xe_vm_invalidate_vma() -> xe_pt_zap_ptes(). Because
vm->pt_root[tile->id] is already NULL, dereferencing pt->level causes a
NULL ptr deref.
Fix this by deferring xe_vm_free_scratch() and xe_vm_pt_destroy() until
after all contested VMAs have been unlinked and destroyed.
User is reporting hitting a NULL ptr deref in xe_pt_zap_ptes(), which
could be explained by this race.
Assisted-by: LLM
Fixes: b06d47be7c ("drm/xe: Port Xe to GPUVA")
Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9290
Signed-off-by: Matthew Auld <matthew.auld@intel.com>
Cc: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: <stable@vger.kernel.org> # v6.12+
Reviewed-by: Thomas Hellström <thomas.hellstrom@linux.intel.com>
Reviewed-by: Matthew Brost <matthew.brost@intel.com>
Link: https://patch.msgid.link/20260918131034.598078-2-matthew.auld@intel.com
(cherry picked from commit c2863648959489767f08892fd6e90577d2ea0b6a)
Signed-off-by: Rodrigo Vivi <rodrigo.vivi@intel.com>
This commit is contained in:
parent
be1df8bada
commit
24a22fb3c7
|
|
@ -1947,21 +1947,13 @@ void xe_vm_close_and_put(struct xe_vm *vm)
|
|||
vma->gpuva.flags |= XE_VMA_DESTROYED;
|
||||
}
|
||||
|
||||
/*
|
||||
* All vm operations will add shared fences to resv.
|
||||
* The only exception is eviction for a shared object,
|
||||
* but even so, the unbind when evicted would still
|
||||
* install a fence to resv. Hence it's safe to
|
||||
* destroy the pagetables immediately.
|
||||
*/
|
||||
xe_vm_free_scratch(vm);
|
||||
xe_vm_pt_destroy(vm);
|
||||
xe_vm_unlock(vm);
|
||||
|
||||
/*
|
||||
* VM is now dead, cannot re-add nodes to vm->vmas if it's NULL
|
||||
* Since we hold a refcount to the bo, we can remove and free
|
||||
* the members safely without locking.
|
||||
* Unlink and destroy all contested external-BO VMAs before destroying
|
||||
* the page tables. Otherwise, concurrent eviction holding only bo->resv
|
||||
* can walk the BO's VMAs and attempt to invalidate/zap page tables that
|
||||
* have already been freed.
|
||||
*/
|
||||
list_for_each_entry_safe(vma, next_vma, &contested,
|
||||
combined_links.destroy) {
|
||||
|
|
@ -1969,6 +1961,11 @@ void xe_vm_close_and_put(struct xe_vm *vm)
|
|||
xe_vma_destroy_unlocked(vma);
|
||||
}
|
||||
|
||||
xe_vm_lock(vm, false);
|
||||
xe_vm_free_scratch(vm);
|
||||
xe_vm_pt_destroy(vm);
|
||||
xe_vm_unlock(vm);
|
||||
|
||||
xe_svm_fini(vm);
|
||||
|
||||
up_write(&vm->lock);
|
||||
|
|
|
|||
Loading…
Reference in New Issue
Block a user