drm/amdgpu: track faulted gfx user-queue slots

A gfx priv/bad-op fault IV carries only the HW slot (ring_id), not the
faulting user queue's doorbell. Add userq_priv_fault_slots (an atomic
bitmap of faulted slots, so concurrent faults are not dropped) and
userq_priv_fault_work to struct amdgpu_gfx; a worker drains the bitmap
and reads the doorbell back from each HQD to locate and reset the queue.

Reviewed-by: Alex Deucher <alexander.deucher@amd.com>
Signed-off-by: Jesse Zhang <Jesse.Zhang@amd.com>
Signed-off-by: Alex Deucher <alexander.deucher@amd.com>
This commit is contained in:
Jesse Zhang 2026-07-31 12:36:24 +08:00 committed by Alex Deucher
parent 70a1e9849e
commit 9ad81600b9

View File

@ -535,6 +535,9 @@ struct amdgpu_gfx {
struct mutex userq_sch_mutex;
u64 userq_sch_req_count[MAX_XCP];
bool userq_sch_inactive[MAX_XCP];
/* atomic bitmap of faulted gfx UQ slots (index = pipe | queue << 2) */
unsigned long userq_priv_fault_slots;
struct work_struct userq_priv_fault_work;
unsigned long enforce_isolation_jiffies[MAX_XCP];
unsigned long enforce_isolation_time[MAX_XCP];