From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f49.google.com (mail-wm1-f49.google.com [209.85.128.49]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 45F623368A9 for ; Fri, 24 Jul 2026 16:33:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.49 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784910801; cv=none; b=ssNIhJFSYHab7Kg4gDy9bbyIoGMgu0IiglDfqi0yriG5QnbCRvb1j/PT/pK/p0XGDir+fHMM4JxceAlX5c/pwqjwtQ/nqswE05eYj65KikdKpq0ToqvzvG+l6M7W24lD/njyDf1HkkPo4z4M4UDl6+L8lv9TRRbICRmhOSekK7c= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784910801; c=relaxed/simple; bh=b0wibVvqA6LqIzvcm1tgZk01l9v2eVHtKm6xtaR5h8U=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=uLl01HIPclP5VemtZDKI3hUatwXoPUss620DsjF872+KXmV88ETjBKR1eIl7FvckRQuVonwkt4/hJixm5JCN/yiyWtkzS71YzP9+6NWdX18+PF+dCE7p0VS04z/zzANeGhdhQjn9UXybdxSN+jEXrA0OgZ31Eggs8istwlcDMlA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linbit.com; spf=pass smtp.mailfrom=linbit.com; dkim=pass (2048-bit key) header.d=linbit-com.20251104.gappssmtp.com header.i=@linbit-com.20251104.gappssmtp.com header.b=ZXSCdPWQ; arc=none smtp.client-ip=209.85.128.49 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linbit.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linbit.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=linbit-com.20251104.gappssmtp.com header.i=@linbit-com.20251104.gappssmtp.com header.b="ZXSCdPWQ" Received: by mail-wm1-f49.google.com with SMTP id 5b1f17b1804b1-493f75f7172so5833505e9.1 for ; Fri, 24 Jul 2026 09:33:18 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linbit-com.20251104.gappssmtp.com; s=20251104; t=1784910796; x=1785515596; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:mail-followup-to:message-id:subject:cc:to:from:date:from :to:cc:subject:date:message-id:reply-to:content-type; bh=73zYw9m2TECKXo+uu9AFah83IICDtXG569B44CjRriI=; b=ZXSCdPWQqze3Gt1FjL19HR+pts/iOLfMBPNoFoCs0wwdSBcpro5UJiRaPLKhXn/ggj 5xIfHoFPlaVVt/KEtzTR3ZpMm3HRnNYWVHCm7u8GYMy6Z+cg4sImTXi2rrOZ0EGdqa77 JxJxEQ4mKA1p3oqPPeNWLEQhzjrxq8KnX8JC7x4ZKsc0JH/wQo6NmnETgBrHQjyFFoyP hV8MfxGsRsOssPmePbKYj+Icg8UemBG9Hv205sf2ktMlj1SE64MMtZxcTZuwuJi1/qyk UC+5hy0+3KuhfWGMQ+spyePBtm0VSUut5QUri8tMoTGzBnhiJ65W3D9kKt+0gqlh5euJ xpRg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1784910796; x=1785515596; h=in-reply-to:content-disposition:content-type:mime-version :references:mail-followup-to:message-id:subject:cc:to:from:date :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=73zYw9m2TECKXo+uu9AFah83IICDtXG569B44CjRriI=; b=UVPEpxCCwN/rlTSrVb9yGHvuxx05zQxzKUenVwmtZKddCbbnOTG+f9tFwyH71imL21 Mm6GUAwSSPn88Huuetn+EXx6aGFqumgw4j7KcNj+A9PEaH8YzrafNYjZdxjiIxQNgtKZ sFVyw36Pk2SOjrkcnG7+g9+Lw9kjypSjdNLlS+9RA23RZSIZ2e23nfw9W1zQ2CVgreLB nmx56f+4Vn0HG6wd6uFFlHCyeij6QK4V4Zadg6hTlWPx1KuUTbEdpxRJOYb1m/RsRz9V ZdjJ858pzGUwpVFwH+96AAvEHuFfUFATot/lRS3F6K70wu5WZB8Wy1FIgyIft7Eo0Ebm 7r4g== X-Forwarded-Encrypted: i=1; AHgh+Ro88nkaZZzgUEM9rZaekMsVlc3jW1DFqyd6ypb2rBjbOQoc1EjvhTgaFKUlzswvKjopDXKneikB6hadQz0=@vger.kernel.org X-Gm-Message-State: AOJu0Yxnj0+bdsIBmTGzMEE1He2LPNs0e5LPNE/4EwZxMJPC670h8WQn XN5pSR/J0QM5V/67QlPUMXx5Doqtya2V6GTBqISQRXBOCHszpsH+VV6tpecUPtiCK+o= X-Gm-Gg: AR+sD10K9Tn9SeaL5OuEn/jftjA6Lvb4O5Kf9K69mcntqVZMtxVnmFpbqa0vPngx8IJ bVBli4FgzMlcfA/dxrvTnFUl+oYrDFtG94MUJHYYo7lERo4Tmy7dBUl1EFWnlrjPPjDBiauXtym LGO7a2JwyC+cf4SgJ4LzY3G/1wRyEWs0J8+ZbY96LUzM3xh0JxN0sOcy6Dj3suxBvdtWeT6m9y9 J1DCogRbaoK2cEndPIZby2gv0b218f65D5DigqBBy6N2+PzpghaZQPrK4oJX8sI0jPN7a/1XK/U wOFCuuYEYDX3E3Zh7vS8nmNzC4Q/2pYdW7ecbL4oDho6puWr3KJ36lbRbvMNv+JGacwxNiBJgEd UmPaMzcVEVSyhlk5/gxrON9X+WmuD0wYj+xORM6tMVnxECTOFfT8mVwg4uGq/Bg+8y/WqLpdvzU UCdwU9H1B+4Vtfki3nv/TvsMsCAaF4YoS0RghxQ0N0aSVckQ== X-Received: by 2002:a05:600c:529b:b0:493:e404:3727 with SMTP id 5b1f17b1804b1-49573cf6d94mr95354345e9.23.1784910796360; Fri, 24 Jul 2026 09:33:16 -0700 (PDT) Received: from localhost (h082218129081.host.wavenet.at. [82.218.129.81]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-496b48632e7sm2798035e9.7.2026.07.24.09.33.14 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jul 2026 09:33:14 -0700 (PDT) Date: Fri, 24 Jul 2026 18:33:13 +0200 From: Christoph =?utf-8?Q?B=C3=B6hmwalder?= To: Wentao Liang Cc: philipp.reisner@linbit.com, lars.ellenberg@linbit.com, axboe@kernel.dk, drbd-dev@lists.linbit.com, linux-block@vger.kernel.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: Re: [PATCH] drbd: Fix local_cnt refcount leak on ascw allocation failure in _drbd_set_state Message-ID: Mail-Followup-To: Wentao Liang , philipp.reisner@linbit.com, lars.ellenberg@linbit.com, axboe@kernel.dk, drbd-dev@lists.linbit.com, linux-block@vger.kernel.org, linux-kernel@vger.kernel.org, stable@vger.kernel.org References: <20260625151636.72599-1-vulab@iscas.ac.cn> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii; format=flowed Content-Disposition: inline In-Reply-To: <20260625151636.72599-1-vulab@iscas.ac.cn> On Thu, Jun 25, 2026 at 11:16:36PM +0800, Wentao Liang wrote: >In _drbd_set_state(), when transitioning a device to D_FAILED or >D_DISKLESS, an extra reference on local_cnt is taken via >atomic_inc(&device->local_cnt) to prevent premature destruction of >the local disk. This reference is normally released by put_ldev() >in after_state_ch(), which is called asynchronously through the >after_state_chg_work (ascw) work item. > >If the GFP_ATOMIC allocation of the ascw work item fails, the work >is never queued, after_state_ch() never runs, and the extra >local_cnt reference is permanently leaked. Additionally, the >state_change object allocated by remember_old_state() is also >leaked, along with the krefs it acquired on the resource, >connections, and devices. > >Fix both leaks in the ascw allocation failure path: > - Call put_ldev() to release the extra local_cnt reference when > the transition matches the same conditions used for the > atomic_inc. > - Call forget_state_change() to free the state_change object and > release the krefs it holds. > >Cc: stable@vger.kernel.org >Fixes: d01801710265 ("drbd: Remove the terrible DEV hack") >Signed-off-by: Wentao Liang >--- > drivers/block/drbd/drbd_state.c | 8 +++++++- > 1 file changed, 7 insertions(+), 1 deletion(-) > >diff --git a/drivers/block/drbd/drbd_state.c b/drivers/block/drbd/drbd_state.c >index adcba7f1d8ea..68e273c6d5be 100644 >--- a/drivers/block/drbd/drbd_state.c >+++ b/drivers/block/drbd/drbd_state.c >@@ -1480,7 +1480,13 @@ _drbd_set_state(struct drbd_device *device, union drbd_state ns, > drbd_queue_work(&connection->sender_work, > &ascw->w); > } else { >- drbd_err(device, "Could not kmalloc an ascw\n"); >+ if ((os.disk != D_FAILED && ns.disk == D_FAILED) || >+ (os.disk != D_DISKLESS && ns.disk == D_DISKLESS)) >+ put_ldev(device); Thanks for the patch. The logic itself looks correct to me. >+ >+ forget_state_change(state_change); >+ drbd_err(device, "Could not kmalloc an ascw, state change %p -> %p leaked\n", >+ &os, &ns); However, this error message is nonsensical. If anything, we should print some halfway human-readable identifier for the state values here, not the pointer. Also, the state change is precisely *not* leaked at the point this message triggers, since we free it here. What actually gets lost is the effects of the state change, so if anything we should point that out here. But I think just keeping the original message is fine. > } > > return rv; >-- >2.39.5 (Apple Git-154) Also, the Fixes tag points to the wrong commit, that was just a mechanical change. The actual breakage was introduced in commit 82f59cc63538 ("drbd: fix potential deadlock on detach").