From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from jeth.damsy.net (jeth.damsy.net [51.159.152.102]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3EE122E336F; Thu, 30 Oct 2025 12:09:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=51.159.152.102 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1761826190; cv=none; b=mi47mkmh+WOuUzit81hjGOF/MXAmRFg9d8XrLiDCOUDmIAC5NVKOOIuU1CxXthLfw+cy2rr+tPF9xgZYRFHU+Cvs+GuwsddexxsPwauGvDLBVwhk5qjw8pTvAWsGELAJDNbazzgEeexnk3tFyBimWPjgEQaOFIfBCiFh6omwa44= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1761826190; c=relaxed/simple; bh=Lad06SDDA/ysZfmJAwYa3FHCerdh4vO/1UsDXUWkOqQ=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=A2b8jzyGLu27h0+YT+1WXV1JuvudBdMxT2onJnkX2J2BkiR3PPSKbf6pAfGd7aU5hN15rG0fbpVjGtcaxIPO6DjlVzuar5fiIvKL2tIjRXCR0iSwieMj7ZiE5lT0ceRHhnPkipRIQpaUYfzSzRyvZVfxLpyjSgyfI6KSL2ETTTE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=damsy.net; spf=pass smtp.mailfrom=damsy.net; dkim=pass (2048-bit key) header.d=damsy.net header.i=@damsy.net header.b=ru2dQ6r/; dkim=permerror (0-bit key) header.d=damsy.net header.i=@damsy.net header.b=/T94g59U; arc=none smtp.client-ip=51.159.152.102 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=damsy.net Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=damsy.net Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=damsy.net header.i=@damsy.net header.b="ru2dQ6r/"; dkim=permerror (0-bit key) header.d=damsy.net header.i=@damsy.net header.b="/T94g59U" DKIM-Signature: v=1; a=rsa-sha256; s=202408r; d=damsy.net; c=relaxed/relaxed; h=From:To:Subject:Date:Message-ID; t=1761826013; bh=R9PAwQHeMkITXcpMrtxUNPe 014g4Qj6IEYjyr1zyJvk=; b=ru2dQ6r/nud5HOoqwBE750RvsPrgmvTNrykguX2SZ1+mpDTWiy 4ttwPDYAX22NS8J3jA4tsVpcyph9xro83VrX0ckATsG/vsCbOeaKuyktd7+9KanuzndX1YZ3nft NOnNKn6pXv64CyCxTTGzmZk/Pd1yrKhV+PWQDEdfPNb9+MLo+8LbFMg4Mz9E+pl5cHowoIcXlWg fzi/nRhXUUUVTytoIN/6eqlt5YqQOKIp2QAYkOIz111XkRyw3xp46ed1A+51krJ+w51k87TMEv2 JMKO01bbdr2IGfoHBDchwXuftllMrEdC536eRhYACi5B26/nDh329DVEq30RepHuTyQ==; DKIM-Signature: v=1; a=ed25519-sha256; s=202408e; d=damsy.net; c=relaxed/relaxed; h=From:To:Subject:Date:Message-ID; t=1761826013; bh=R9PAwQHeMkITXcpMrtxUNPe 014g4Qj6IEYjyr1zyJvk=; b=/T94g59U0g1LTiijf/QkKxE7gdalHEpNXHu/fTDmtsBrVR2R14 zCthAWVelAn2i5oU/yeb8pigXZEpZrSzexAw==; Message-ID: <442d0e70-c9e2-4bd6-a144-ea083dbf86d2@damsy.net> Date: Thu, 30 Oct 2025 13:06:52 +0100 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH v1] drm/sched: fix deadlock in drm_sched_entity_kill_jobs_cb To: phasta@kernel.org, Pierre-Eric Pelloux-Prayer , Matthew Brost , Danilo Krummrich , =?UTF-8?Q?Christian_K=C3=B6nig?= , Maarten Lankhorst , Maxime Ripard , Thomas Zimmermann , David Airlie , Simona Vetter , Sumit Semwal Cc: =?UTF-8?Q?Christian_K=C3=B6nig?= , dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org, linux-media@vger.kernel.org, linaro-mm-sig@lists.linaro.org References: <20251029091103.1159-1-pierre-eric.pelloux-prayer@amd.com> Content-Language: en-US From: Pierre-Eric Pelloux-Prayer In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit Le 30/10/2025 à 12:17, Philipp Stanner a écrit : > On Wed, 2025-10-29 at 10:11 +0100, Pierre-Eric Pelloux-Prayer wrote: >> https://gitlab.freedesktop.org/mesa/mesa/-/issues/13908 pointed out > > This link should be moved to the tag section at the bottom at a Closes: > tag. Optionally a Reported-by:, too. The bug report is about a different issue. The potential deadlock being fixed by this patch was discovered while investigating it. I'll add a Reported-by tag though. > >> a possible deadlock: >> >> [ 1231.611031]  Possible interrupt unsafe locking scenario: >> >> [ 1231.611033]        CPU0                    CPU1 >> [ 1231.611034]        ----                    ---- >> [ 1231.611035]   lock(&xa->xa_lock#17); >> [ 1231.611038]                                local_irq_disable(); >> [ 1231.611039]                                lock(&fence->lock); >> [ 1231.611041]                                lock(&xa->xa_lock#17); >> [ 1231.611044]   >> [ 1231.611045]     lock(&fence->lock); >> [ 1231.611047] >>                 *** DEADLOCK *** >> > > The commit message is lacking an explanation as to _how_ and _when_ the > deadlock comes to be. That's a prerequisite for understanding why the > below is the proper fix and solution. I copy-pasted a small chunk of the full deadlock analysis report included in the ticket because it's 300+ lines long. Copying the full log isn't useful IMO, but I can add more context. The problem is that a thread (CPU0 above) can lock the job's dependencies xa_array without disabling the interrupts. If a fence signals while CPU0 holds this lock and drm_sched_entity_kill_jobs_cb is called, it will try to grab the xa_array lock which is not possible because CPU0 holds it already. > > The issue seems to be that you cannot perform certain tasks from within > that work item? > >> My initial fix was to replace xa_erase by xa_erase_irq, but Christian >> pointed out that calling dma_fence_add_callback from a callback can >> also deadlock if the signalling fence and the one passed to >> dma_fence_add_callback share the same lock. >> >> To fix both issues, the code iterating on dependencies and re-arming them >> is moved out to drm_sched_entity_kill_jobs_work. >> >> Suggested-by: Christian König >> Signed-off-by: Pierre-Eric Pelloux-Prayer >> --- >>  drivers/gpu/drm/scheduler/sched_entity.c | 34 +++++++++++++----------- >>  1 file changed, 19 insertions(+), 15 deletions(-) >> >> diff --git a/drivers/gpu/drm/scheduler/sched_entity.c b/drivers/gpu/drm/scheduler/sched_entity.c >> index c8e949f4a568..fe174a4857be 100644 >> --- a/drivers/gpu/drm/scheduler/sched_entity.c >> +++ b/drivers/gpu/drm/scheduler/sched_entity.c >> @@ -173,26 +173,15 @@ int drm_sched_entity_error(struct drm_sched_entity *entity) >>  } >>  EXPORT_SYMBOL(drm_sched_entity_error); >> >> +static void drm_sched_entity_kill_jobs_cb(struct dma_fence *f, >> +   struct dma_fence_cb *cb); >> + >>  static void drm_sched_entity_kill_jobs_work(struct work_struct *wrk) >>  { >>   struct drm_sched_job *job = container_of(wrk, typeof(*job), work); >> - >> - drm_sched_fence_scheduled(job->s_fence, NULL); >> - drm_sched_fence_finished(job->s_fence, -ESRCH); >> - WARN_ON(job->s_fence->parent); >> - job->sched->ops->free_job(job); > > Can free_job() really not be called from within work item context? It's still called from drm_sched_entity_kill_jobs_work but the diff is slightly confusing. Pierre-Eric > > > P.