From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-oa1-f45.google.com (mail-oa1-f45.google.com [209.85.160.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7A3D6381B15 for ; Wed, 27 May 2026 16:03:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.45 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779897812; cv=none; b=iUOANW3UfpWAvOJb85qJZ9L5+znVUBLUrmVaA8BZDmuUStvI2sfVuNq5COo4IsiBBo1SrEy5KUxWDXpzjG1T2e3s0zrWr5wM3Mk6Ie7dlgZrCuI+Hdq+N7KYKrdV8lgNh40eiR6E8OJXRFhYjKBxgRuvUduYEFxsXEsFNoD4ZZU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779897812; c=relaxed/simple; bh=jfHYVvAv2KLPbWszFU9UWn6gSBGjDMujZa9kYNQiUyw=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=oSdTIApOTihn90vYiCTOTZeT6GxEbD/UHKc5cVnogTgCM3oSOi+Jh60wYPjB4iMb5+rU3cZJfMYk2HJTdFXrSCQeCixQxVQiDgrsRfv5DyZX+EtzNkw5iqBPospbmVW6fbXGbFhtrcVyzZDub+dLaQMU9mBzeP98F9s4V0sk3f0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b=p3PoEzqv; arc=none smtp.client-ip=209.85.160.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20251104.gappssmtp.com header.i=@kernel-dk.20251104.gappssmtp.com header.b="p3PoEzqv" Received: by mail-oa1-f45.google.com with SMTP id 586e51a60fabf-43b7e186a0cso1580951fac.0 for ; Wed, 27 May 2026 09:03:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20251104.gappssmtp.com; s=20251104; t=1779897809; x=1780502609; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=gntijPqsvuiGRryRGaewxCELZ5a6Qr0krHV7SCIDMB4=; b=p3PoEzqvIwCDKqu1d2DE3TstqFbMSY6eCzEUZ4nAjFKTII7VwecNYGybvd8CjXKkC3 tsupMB/EID6Aa2HQeld19Q3tlb2tUw8wfiMKcODRX1Y6d6GEKvv4g6tprpxzq/YzwO2B gXeMDo5zxBSGLho53SSlLAn0B2rDXwcFzs+S5K939WEcgO5+DE0T0oJUZO4CSqnf7tA9 8BxHBE+WPwFLpQd1aVCL3T89ed8GUP79mY17chmOQWYDx28TS7o2ju0jadAWemuyQJ3z PqJp2cEqb5XVt4hcD13u5AqmxVszUsy8oviQCzXSJbpOSiS1/WwddFRRuBDymuKCC6EI 6hOA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1779897809; x=1780502609; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=gntijPqsvuiGRryRGaewxCELZ5a6Qr0krHV7SCIDMB4=; b=EvQdKjPmSNqMWKK+xo1rPE4OUCtx4w4fiTaG24jQU/46+yEGSWjdVJawYbJWtt3oB4 QeVYslsNpvqqyONZ4bLmlvNkqm9l/NS/lJrshJfsBJXr+8vrFDqBSLRvklNSx2Rrs61S sv9+Bjq+RaQ5NE8XoQwPj0aHSPcecwT9ugbj4D8/CRmDWOnS71RBbvk7TwgiFamgZZhW aWI1lSdvyJIBi75bIMDCmk+mIEH1Gisa7JTKE1pLkQSJ1zk5DwqWvVTLStWfBrAO80KJ 3ztyt/9LnXg6NHiZGI0bUg/s2N1toa63eBpu6JZqHvkhdN5D7UPC6fqkMNK9BB8ZoX6F 366g== X-Gm-Message-State: AOJu0YwV464NTnFIDxLz0IXoFvaEpMcLZBO5utbfmNcs1E6iSqtGCTVX ZSA8kboEey53uoOoe/PmMcOPE+FIHFVY0pPLaGqmuK/07SNnqoFbe7bXMFlJKHr/JRMOyWHczWC FYNNrxes= X-Gm-Gg: Acq92OGWQ9aw4SYsQzbB8wFnlhLX5B8DY2aWOfFvcnkTdeXbgTJBugE5C3uI2+Ms8AS tt+mYExLVC3OXeiEfriJF0Qf3mM/gfWMHAHDGz94Oy9jFFX5x1G6L+stCzhUYCQONQh2ouCWRGD Hn+P0RtDWux8U1C7V75RYouF0USRxWktvEONAC4QD7Vox2dQkmfPvrkE7beZTvZn/1omzoScaFx x6gzTR1/EzR+ny+1TkP1qwWptTFSIm24wFmOQKWiOhh4MlXWzAoiuTwPtCs0wrw+nQKgQws0ZlL fL0mpJ74JvTQXKhFTWAcjZPy5rqiAB4oqeLHx7AaPos6AbZg35MgbmShfMZqCDdf4UFNkEGi0su mNSTp5MfcBLaXLDSc2HHCSighTowfFS+dkg+4zvE8Pb/bkgixrd+1QupBc8opwZF03nWNaNd6P+ DKPZbHSEH7oJ7+3PDxzY+iBx/g9xIawDZihM3C4ii2oa4RuGjmnxEl2EPuZxoeaPsAJbMUgom2c 7dcAFpuqCAj7vUslFM= X-Received: by 2002:a05:6871:4390:b0:43a:f95e:cf14 with SMTP id 586e51a60fabf-43b5aaed6f2mr14617454fac.12.1779897808999; Wed, 27 May 2026 09:03:28 -0700 (PDT) Received: from [192.168.1.102] ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-43b639f3609sm16097799fac.13.2026.05.27.09.03.27 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Wed, 27 May 2026 09:03:27 -0700 (PDT) Message-ID: <919d86f3-1164-4084-9f72-d3ead0522c5e@kernel.dk> Date: Wed, 27 May 2026 10:03:26 -0600 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH] io_uring/io-wq: re-check IO_WQ_BIT_EXIT for each linked work item To: Runyu Xiao , io-uring@vger.kernel.org Cc: linux-kernel@vger.kernel.org, gregkh@linuxfoundation.org, jianhao.xu@seu.edu.cn, stable@vger.kernel.org References: <20260527143726.1272269-1-runyu.xiao@seu.edu.cn> Content-Language: en-US From: Jens Axboe In-Reply-To: <20260527143726.1272269-1-runyu.xiao@seu.edu.cn> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 5/27/26 8:37 AM, Runyu Xiao wrote: > Commit bdf0bf73006e ("io_uring/io-wq: check IO_WQ_BIT_EXIT inside work > run loop") fixed the obvious case where io_worker_handle_work() took one > exit-bit snapshot before draining pending work, but the fix stops one > level too early. > > io_worker_handle_work() now re-checks IO_WQ_BIT_EXIT in its outer work > run loop, yet it still snapshots that bit once before processing a > whole dependent linked-work chain. If io_wq_exit_start() sets > IO_WQ_BIT_EXIT after the first linked item has started, the remaining > linked items can still reuse stale do_kill = false, skip > IO_WQ_WORK_CANCEL, and continue running after exit has begun. > > That means the previous fix did not fully eliminate the exit-latency > problem; it only narrowed it to linked chains. A long or slow linked > chain can still keep io-wq exit waiting for work that should already > have been canceled. > > The issue was found on Linux v6.18.21 by our static-analysis tool, > which flagged linked-work loops that snapshot shared exit state > outside per-item cancel decisions, and was then confirmed by manual > auditing of io_worker_handle_work(). It was later reproduced with a > QEMU no-device validation selftest that preserved the same contract: > a three-node unbound linked chain, an exit actor setting > IO_WQ_BIT_EXIT after work1, and slow post-exit linked work. With a > 3000 ms delay injected into each post-exit item, the buggy path > spends about 6066 ms after exit running work2/work3, while the fixed > path cancels both and finishes in about 2 ms. > > Re-check test_bit(IO_WQ_BIT_EXIT, &wq->state) for each iteration of the > dependent-link loop, right before deciding whether to cancel the > current work item. That closes the remaining stale-snapshot window and > prevents linked post-exit work from stretching shutdown latency. I think this change makes sense to further cut down on the time, but you need to send it in for the _upstream_ kernel, stable only does backports of those. Eg if you send this one for current -git and mark it fixing the correct upstream commit (not the stable one) and add CC stable, then it'll wind up in stable as well. -- Jens Axboe