From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 83E062F25F5 for ; Mon, 2 Feb 2026 06:09:41 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770012581; cv=none; b=Rp2tRpGGQ3A9X+gne77cPDHX4QZE6Z/mtJho5as9YxePq0lJjiqcsxKGsNS1DJhaQlC0Fbv5LTRYKITfneQMxs+cFGTvlNRn5/DqaCbvq+f8GOlf7Fg3PBxWZWQjaOHd1HHX5PWeteNc1Zk5t6wq0Z9b593r649kAYA72wg2Zn8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770012581; c=relaxed/simple; bh=qgolyN5UP4lvVIaa+M7uL8rq21flrl+ZcqMblfNbGaI=; h=Date:From:To:Cc:Subject:Message-Id:In-Reply-To:References: Mime-Version:Content-Type; b=WZYuPkdeaEkwDaitlCQdJW1mX4DOZc5FrXAap3CExZjKvHTruse3jBIOnMLUNOvP/ow8IkCpbkxT5vmRU0iwOcXzrm1efXR6YrXxa7DiDQ9NAMgmb721wgukrrs62wLbIwdOW0xw3nRTugixlAwyhaylZuQr8RS/VjuEXGBPiME= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=Wfh7P3qr; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="Wfh7P3qr" Received: by smtp.kernel.org (Postfix) with ESMTPSA id 3FBA3C19421; Mon, 2 Feb 2026 06:09:38 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1770012581; bh=qgolyN5UP4lvVIaa+M7uL8rq21flrl+ZcqMblfNbGaI=; h=Date:From:To:Cc:Subject:In-Reply-To:References:From; b=Wfh7P3qreD6JpAZHuJfgztoilFQtDnLbEi77/JffbEauQlwTR5MqwGU5plck75EF9 GI9t62O8vnSeagRvMH1j2i9XYbjWKxNQ4HpydqM/jOb8nz5lH/N6Yqs62qH1phjP4Y Y4ybKkSkYWcBzXKk/LZ3nSCee6SXs8HM5VLbp0WC4wbRBdPE5IrrrN4LVu3E5vybPJ pgfVN9Ce8R/wjnXNPnmMYRjsIWpKUi/mL2A1JBepGWXTrfoccI6p8sdIbYB/v18ICt f0N8GkgZ1LBrt8anFJH5KeX43/qtcf82wLNX/aGiwUAXtahJzw4LBKshr5ELwkGuaf kgm1aVeW+QFQw== Date: Mon, 2 Feb 2026 15:09:35 +0900 From: Masami Hiramatsu (Google) To: Aaron Tomlin Cc: akpm@linux-foundation.org, lance.yang@linux.dev, gregkh@linuxfoundation.org, pmladek@suse.com, joel.granados@kernel.org, neelx@suse.com, sean@ashe.io, mproche@gmail.com, chjohnst@gmail.com, nick.lange@gmail.com, linux-kernel@vger.kernel.org Subject: Re: [v7 PATCH 2/2] hung_task: Enable runtime reset of hung_task_detect_count Message-Id: <20260202150935.63878390c9e8c3d238d143a1@kernel.org> In-Reply-To: <20260125135848.3356585-3-atomlin@atomlin.com> References: <20260125135848.3356585-1-atomlin@atomlin.com> <20260125135848.3356585-3-atomlin@atomlin.com> X-Mailer: Sylpheed 3.8.0beta1 (GTK+ 2.24.33; x86_64-pc-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Sun, 25 Jan 2026 08:58:48 -0500 Aaron Tomlin wrote: > Currently, the hung_task_detect_count sysctl provides a cumulative count > of hung tasks since boot. In long-running, high-availability > environments, this counter may lose its utility if it cannot be reset > once an incident has been resolved. Furthermore, the previous > implementation relied upon implicit ordering, which could not strictly > guarantee that diagnostic metadata published by one CPU was visible to > the panic logic on another. > > This patch introduces the capability to reset the detection count by > writing "0" to the hung_task_detect_count sysctl. The proc_handler logic > has been updated to validate this input and atomically reset the > counter. > > The synchronisation of sysctl_hung_task_detect_count relies upon a > transactional model to ensure the integrity of the detection counter > against concurrent resets from userspace. The application of > atomic_long_read_acquire() and atomic_long_cmpxchg_release() is correct > and provides the following guarantees: > > 1. Prevention of Load-Store Reordering via Acquire Semantics By > utilising atomic_long_read_acquire() to snapshot the counter > before initiating the task traversal, we establish a strict > memory barrier. This prevents the compiler or hardware from > reordering the initial load to a point later in the scan. Without > this "acquire" barrier, a delayed load could potentially read a > "0" value resulting from a userspace reset that occurred > mid-scan. This would lead to the subsequent cmpxchg succeeding > erroneously, thereby overwriting the user's reset with stale > increment data. > > 2. Atomicity of the "Commit" Phase via Release Semantics The > atomic_long_cmpxchg_release() serves as the transaction's commit > point. The "release" barrier ensures that all diagnostic > recordings and task-state observations made during the scan are > globally visible before the counter is incremented. > > 3. Race Condition Resolution This pairing effectively detects any > "out-of-band" reset of the counter. If > sysctl_hung_task_detect_count is modified via the procfs > interface during the scan, the final cmpxchg will detect the > discrepancy between the current value and the "acquire" snapshot. > Consequently, the update will fail, ensuring that a reset command > from the administrator is prioritised over a scan that may have > been invalidated by that very reset. > > Signed-off-by: Aaron Tomlin Looks good to me. Reviewed-by: Masami Hiramatsu (Google) Thanks, > --- > Documentation/admin-guide/sysctl/kernel.rst | 3 +- > kernel/hung_task.c | 58 ++++++++++++++++++--- > 2 files changed, 53 insertions(+), 8 deletions(-) > > diff --git a/Documentation/admin-guide/sysctl/kernel.rst b/Documentation/admin-guide/sysctl/kernel.rst > index 239da22c4e28..68da4235225a 100644 > --- a/Documentation/admin-guide/sysctl/kernel.rst > +++ b/Documentation/admin-guide/sysctl/kernel.rst > @@ -418,7 +418,8 @@ hung_task_detect_count > ====================== > > Indicates the total number of tasks that have been detected as hung since > -the system boot. > +the system boot or since the counter was reset. The counter is zeroed when > +a value of 0 is written. > > This file shows up if ``CONFIG_DETECT_HUNG_TASK`` is enabled. > > diff --git a/kernel/hung_task.c b/kernel/hung_task.c > index df10830ed9ef..350093de0535 100644 > --- a/kernel/hung_task.c > +++ b/kernel/hung_task.c > @@ -306,7 +306,11 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout) > int need_warning = sysctl_hung_task_warnings; > unsigned long si_mask = hung_task_si_mask; > > - total_count = atomic_long_read(&sysctl_hung_task_detect_count); > + /* > + * The counter might get reset. Remember the initial value. > + * Acquire prevents reordering task checks before this point. > + */ > + total_count = atomic_long_read_acquire(&sysctl_hung_task_detect_count); > /* > * If the system crashed already then all bets are off, > * do not report extra hung tasks: > @@ -337,10 +341,11 @@ static void check_hung_uninterruptible_tasks(unsigned long timeout) > return; > > /* > - * This counter tracks the total number of tasks detected as hung > - * since boot. > + * Do not count this round when the global counter has been reset > + * during this check. Release ensures we see all hang details > + * recorded during the scan. > */ > - atomic_long_cmpxchg_relaxed(&sysctl_hung_task_detect_count, > + atomic_long_cmpxchg_release(&sysctl_hung_task_detect_count, > total_count, total_count + > this_round_count); > > @@ -366,6 +371,46 @@ static long hung_timeout_jiffies(unsigned long last_checked, > } > > #ifdef CONFIG_SYSCTL > + > +/** > + * proc_dohung_task_detect_count - proc handler for hung_task_detect_count > + * @table: Pointer to the struct ctl_table definition for this proc entry > + * @dir: Flag indicating the operation > + * @buffer: User space buffer for data transfer > + * @lenp: Pointer to the length of the data being transferred > + * @ppos: Pointer to the current file offset > + * > + * This handler is used for reading the current hung task detection count > + * and for resetting it to zero when a write operation is performed using a > + * zero value only. > + * Return: 0 on success, or a negative error code on failure. > + */ > +static int proc_dohung_task_detect_count(const struct ctl_table *table, int dir, > + void *buffer, size_t *lenp, loff_t *ppos) > +{ > + unsigned long detect_count; > + struct ctl_table proxy_table; > + int err; > + > + proxy_table = *table; > + proxy_table.data = &detect_count; > + > + if (SYSCTL_KERN_TO_USER(dir)) > + detect_count = atomic_long_read(&sysctl_hung_task_detect_count); > + > + err = proc_doulongvec_minmax(&proxy_table, dir, buffer, lenp, ppos); > + if (err < 0) > + return err; > + > + if (SYSCTL_USER_TO_KERN(dir)) { > + if (detect_count) > + return -EINVAL; > + atomic_long_set(&sysctl_hung_task_detect_count, 0); > + } > + > + return 0; > +} > + > /* > * Process updating of timeout sysctl > */ > @@ -446,10 +491,9 @@ static const struct ctl_table hung_task_sysctls[] = { > }, > { > .procname = "hung_task_detect_count", > - .data = &sysctl_hung_task_detect_count, > .maxlen = sizeof(unsigned long), > - .mode = 0444, > - .proc_handler = proc_doulongvec_minmax, > + .mode = 0644, > + .proc_handler = proc_dohung_task_detect_count, > }, > { > .procname = "hung_task_sys_info", > -- > 2.51.0 > -- Masami Hiramatsu (Google)