From 24a22fb3c731474b68e986af6804db450fb88617 Mon Sep 17 00:00:00 2001 From: Matthew Auld Date: Fri, 18 Sep 2026 14:10:35 +0100 Subject: [PATCH] drm/xe/vm: nuke PTs only after unlinking contested VMAs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit In xe_vm_close_and_put(), external-BO VMAs are queued on the contested list for deferred destruction via xe_vma_destroy_unlocked(). However, xe_vm_pt_destroy() was previously invoked before processing contested VMAs, destroying vm->pt_root while those VMAs were still linked to their respective buffer objects (vm_bo->list.gpuva). If a concurrent thread evicts one of those shared buffer objects, xe_bo_trigger_rebind() holding only bo->resv walks the BO's VMAs and, in fault mode, calls xe_vm_invalidate_vma() -> xe_pt_zap_ptes(). Because vm->pt_root[tile->id] is already NULL, dereferencing pt->level causes a NULL ptr deref. Fix this by deferring xe_vm_free_scratch() and xe_vm_pt_destroy() until after all contested VMAs have been unlinked and destroyed. User is reporting hitting a NULL ptr deref in xe_pt_zap_ptes(), which could be explained by this race. Assisted-by: LLM Fixes: b06d47be7c83 ("drm/xe: Port Xe to GPUVA") Link: https://gitlab.freedesktop.org/drm/xe/kernel/-/work_items/9290 Signed-off-by: Matthew Auld Cc: Thomas Hellström Cc: Matthew Brost Cc: # v6.12+ Reviewed-by: Thomas Hellström Reviewed-by: Matthew Brost Link: https://patch.msgid.link/20260918131034.598078-2-matthew.auld@intel.com (cherry picked from commit c2863648959489767f08892fd6e90577d2ea0b6a) Signed-off-by: Rodrigo Vivi --- drivers/gpu/drm/xe/xe_vm.c | 21 +++++++++------------ 1 file changed, 9 insertions(+), 12 deletions(-) diff --git a/drivers/gpu/drm/xe/xe_vm.c b/drivers/gpu/drm/xe/xe_vm.c index 23952ad8951e..ef20e205a734 100644 --- a/drivers/gpu/drm/xe/xe_vm.c +++ b/drivers/gpu/drm/xe/xe_vm.c @@ -1947,21 +1947,13 @@ void xe_vm_close_and_put(struct xe_vm *vm) vma->gpuva.flags |= XE_VMA_DESTROYED; } - /* - * All vm operations will add shared fences to resv. - * The only exception is eviction for a shared object, - * but even so, the unbind when evicted would still - * install a fence to resv. Hence it's safe to - * destroy the pagetables immediately. - */ - xe_vm_free_scratch(vm); - xe_vm_pt_destroy(vm); xe_vm_unlock(vm); /* - * VM is now dead, cannot re-add nodes to vm->vmas if it's NULL - * Since we hold a refcount to the bo, we can remove and free - * the members safely without locking. + * Unlink and destroy all contested external-BO VMAs before destroying + * the page tables. Otherwise, concurrent eviction holding only bo->resv + * can walk the BO's VMAs and attempt to invalidate/zap page tables that + * have already been freed. */ list_for_each_entry_safe(vma, next_vma, &contested, combined_links.destroy) { @@ -1969,6 +1961,11 @@ void xe_vm_close_and_put(struct xe_vm *vm) xe_vma_destroy_unlocked(vma); } + xe_vm_lock(vm, false); + xe_vm_free_scratch(vm); + xe_vm_pt_destroy(vm); + xe_vm_unlock(vm); + xe_svm_fini(vm); up_write(&vm->lock);