From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mout-p-202.mailbox.org (mout-p-202.mailbox.org [80.241.56.172]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 6633C2D9EF3 for ; Tue, 17 Feb 2026 10:25:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=80.241.56.172 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1771323936; cv=none; b=Gj023oRV56ASc6qEvbJg8ozerUfUqj3wlLGZpla9XxwuPoTUNpurqo/1Ee3usuigNq4Jo/NoLETwpb8YzmiGN8xihBtI5/N6nzxURSKz4JkUvvYj+vkKjGtk2e1l0meap8Y4bXIWoCaWSlw1zftqR9TztCN/Pz2D3SbyzbTCf8w= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1771323936; c=relaxed/simple; bh=RX9IUOOr9nK7eMRvUoibne4mjIsEu0xyh4TUHoS+pc0=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=SAZlS65SoisTASUUtDkf8AvKqTr7nl16PWlQy5/STVLy8LKs9VXeX0M6hKbrjqq2u6tdT0hAgPwOE+p5MdqWO9xEXZT0eL32n1H9Elj5nGOkHifYD2WubK2vQQiVWspGVgwPF24gwK3VA3Z4rfCGoPhbo3xDFTWwRFGPMBTUS14= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=mailbox.org; spf=pass smtp.mailfrom=mailbox.org; dkim=pass (2048-bit key) header.d=mailbox.org header.i=@mailbox.org header.b=ZJmOsiFU; arc=none smtp.client-ip=80.241.56.172 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=mailbox.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=mailbox.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=mailbox.org header.i=@mailbox.org header.b="ZJmOsiFU" Received: from smtp202.mailbox.org (smtp202.mailbox.org [IPv6:2001:67c:2050:b231:465::202]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (4096 bits) server-digest SHA256) (No client certificate requested) by mout-p-202.mailbox.org (Postfix) with ESMTPS id 4fFbPk6CrVz9tPF; Tue, 17 Feb 2026 11:25:30 +0100 (CET) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=mailbox.org; s=mail20150812; t=1771323930; h=from:from:reply-to:reply-to:subject:subject:date:date: message-id:message-id:to:to:cc:cc:mime-version:mime-version: content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=MYALA9bthKUi4cPtiQOjwnebj6XDjFjmpqTthtqZB7c=; b=ZJmOsiFUUuUcKJWu7XTcLVMCspZe7khN5E18Tt6tUvyHF1BFVK+TPY1vukxkMjoIMMP4y4 tiArVxMqeM0MQ4HV+rpUj7Ku7Tvb/TkuyWsZ1r681svBMlI6M+0YZ6oIq9K6iXc1yDBB6K 3QG9knlpiLMYPnJYKAnfeqSuNhR/oNYji337S40/Eajfwsvdvws81S47A7O+h3BT6hkdvL t52++fSaCnMWZLqf+PonEsczLzPi25ZJS0Z9aBdn9HISg79B/SWLzk7GjXNfS+3tH/5wRD KJkM6MKkleUKZu0bGiOVhCZ0mYi9Wns5E8Bq3vbFCiSTSDYlC6wHmU/+fjJUhg== Message-ID: <414b00dcb19c772d9a616a5aed10a5ac59dc830b.camel@mailbox.org> Subject: Re: [PATCH] drm/sched: Remove racy hack from drm_sched_fini() From: Philipp Stanner Reply-To: phasta@kernel.org To: Philipp Stanner , Matthew Brost , Danilo Krummrich , Christian =?ISO-8859-1?Q?K=F6nig?= Cc: dri-devel@lists.freedesktop.org, linux-kernel@vger.kernel.org Date: Tue, 17 Feb 2026 11:25:27 +0100 In-Reply-To: <20260108083019.63532-2-phasta@kernel.org> References: <20260108083019.63532-2-phasta@kernel.org> Content-Type: text/plain; charset="UTF-8" Content-Transfer-Encoding: quoted-printable Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MBO-RS-ID: f2e2e50f84b4298189a X-MBO-RS-META: c6mw8tb49tnn91cd395r5gtf5ya1h9rw On Thu, 2026-01-08 at 09:30 +0100, Philipp Stanner wrote: > drm_sched_fini() contained a hack to work around a race in amdgpu. > According to AMD, the hack should not be necessary anymore. In case > there should have been undetected users, >=20 > commit 975ca62a014c ("drm/sched: Add warning for removing hack in drm_sch= ed_fini()") >=20 > had added a warning one release cycle ago. >=20 > Thus, it can be derived that the hack can be savely removed by now. >=20 > Remove the hack. >=20 > Signed-off-by: Philipp Stanner > --- > As hinted at in the commit, I want to cozyly queue this one up for the > next merge window, since we're printing that warning since last merge > window already. >=20 > If someone has concerns I'm also happy to delay this patch for a few > more releases. > --- Any objections by anyone? Can I get an RB? P. > =C2=A0drivers/gpu/drm/scheduler/sched_main.c | 38 +----------------------= --- > =C2=A01 file changed, 1 insertion(+), 37 deletions(-) >=20 > diff --git a/drivers/gpu/drm/scheduler/sched_main.c b/drivers/gpu/drm/sch= eduler/sched_main.c > index 1d4f1b822e7b..381c1694a12e 100644 > --- a/drivers/gpu/drm/scheduler/sched_main.c > +++ b/drivers/gpu/drm/scheduler/sched_main.c > @@ -1416,48 +1416,12 @@ static void drm_sched_cancel_remaining_jobs(struc= t drm_gpu_scheduler *sched) > =C2=A0 */ > =C2=A0void drm_sched_fini(struct drm_gpu_scheduler *sched) > =C2=A0{ > - struct drm_sched_entity *s_entity; > =C2=A0 int i; > =C2=A0 > =C2=A0 drm_sched_wqueue_stop(sched); > =C2=A0 > - for (i =3D DRM_SCHED_PRIORITY_KERNEL; i < sched->num_rqs; i++) { > - struct drm_sched_rq *rq =3D sched->sched_rq[i]; > - > - spin_lock(&rq->lock); > - list_for_each_entry(s_entity, &rq->entities, list) { > - /* > - * Prevents reinsertion and marks job_queue as idle, > - * it will be removed from the rq in drm_sched_entity_fini() > - * eventually > - * > - * FIXME: > - * This lacks the proper spin_lock(&s_entity->lock) and > - * is, therefore, a race condition. Most notably, it > - * can race with drm_sched_entity_push_job(). The lock > - * cannot be taken here, however, because this would > - * lead to lock inversion -> deadlock. > - * > - * The best solution probably is to enforce the life > - * time rule of all entities having to be torn down > - * before their scheduler. Then, however, locking could > - * be dropped alltogether from this function. > - * > - * For now, this remains a potential race in all > - * drivers that keep entities alive for longer than > - * the scheduler. > - * > - * The READ_ONCE() is there to make the lockless read > - * (warning about the lockless write below) slightly > - * less broken... > - */ > - if (!READ_ONCE(s_entity->stopped)) > - dev_warn(sched->dev, "Tearing down scheduler with active entities!\n= "); > - s_entity->stopped =3D true; > - } > - spin_unlock(&rq->lock); > + for (i =3D DRM_SCHED_PRIORITY_KERNEL; i < sched->num_rqs; i++) > =C2=A0 kfree(sched->sched_rq[i]); > - } > =C2=A0 > =C2=A0 /* Wakeup everyone stuck in drm_sched_entity_flush for this schedu= ler */ > =C2=A0 wake_up_all(&sched->job_scheduled);