From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S964958AbbLQWTA (ORCPT ); Thu, 17 Dec 2015 17:19:00 -0500 Received: from mail.linuxfoundation.org ([140.211.169.12]:38545 "EHLO mail.linuxfoundation.org" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S934305AbbLQWSG (ORCPT ); Thu, 17 Dec 2015 17:18:06 -0500 Date: Thu, 17 Dec 2015 14:18:05 -0800 From: Andrew Morton To: Tetsuo Handa Cc: linux-kernel@vger.kernel.org, oleg@redhat.com, atomlin@redhat.com Subject: Re: [PATCH] kernel/hung_task.c: use timeout diff when timeout is updated Message-Id: <20151217141805.f418cf9b137da08656504001@linux-foundation.org> In-Reply-To: <201512172123.DFJ69220.SFFOLOJtVHOQMF@I-love.SAKURA.ne.jp> References: <201512172123.DFJ69220.SFFOLOJtVHOQMF@I-love.SAKURA.ne.jp> X-Mailer: Sylpheed 3.4.1 (GTK+ 2.24.23; x86_64-pc-linux-gnu) Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Thu, 17 Dec 2015 21:23:03 +0900 Tetsuo Handa wrote: > >From 529ff00b556e110c6e801c39e94b06f559307136 Mon Sep 17 00:00:00 2001 > From: Tetsuo Handa > Date: Thu, 17 Dec 2015 16:27:08 +0900 > Subject: [PATCH] kernel/hung_task.c: use timeout diff when timeout is updated > > When new timeout is written to /proc/sys/kernel/hung_task_timeout_secs, > khungtaskd is interrupted and again sleeps for full timeout duration. > > This means that hang task will not be checked if new timeout is written > periodically within old timeout duration and/or checking of hang task > will be delayed for up to previous timeout duration. > Fix this by remembering last time khungtaskd checked hang task. > > This change will allow other watchdog tasks (if any) to share khungtaskd > by sleeping for minimal timeout diff of all watchdog tasks. Doing more > watchdog tasks from khungtaskd will reduce the possibility of printk() > collisions by multiple watchdog threads. This seems like reasonable behaviour, but it is a non-backward-compatible change. I don't know how important that is - probably "not very". Please let's fully describe this behaviour in the documentation. Documentation/sysctl/kernel.txt. It appears that if userspace writes a new timeout which has already expired, a check is triggered immediately, correct? Let's ensure that this is documented as well. And tested! And it would be helpful to add a comment to hung_timeout_jiffies() which describes the behaviour and explains the reasons for it.