From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752539AbdLCOsU (ORCPT ); Sun, 3 Dec 2017 09:48:20 -0500 Received: from Galois.linutronix.de ([146.0.238.70]:55600 "EHLO Galois.linutronix.de" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1752013AbdLCOsT (ORCPT ); Sun, 3 Dec 2017 09:48:19 -0500 Date: Sun, 3 Dec 2017 15:48:14 +0100 (CET) From: Thomas Gleixner To: Dmitry Vyukov cc: syzbot , Greg Kroah-Hartman , Kate Stewart , LKML , Linux-MM , Philippe Ombredanne , syzkaller-bugs@googlegroups.com Subject: Re: BUG: workqueue lockup (2) In-Reply-To: Message-ID: References: <94eb2c03c9bc75aff2055f70734c@google.com> User-Agent: Alpine 2.20 (DEB 67 2015-01-07) MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII X-Linutronix-Spam-Score: -1.0 X-Linutronix-Spam-Level: - X-Linutronix-Spam-Status: No , -1.0 points, 5.0 required, ALL_TRUSTED=-1,SHORTCIRCUIT=-0.0001 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Sun, 3 Dec 2017, Dmitry Vyukov wrote: > On Sun, Dec 3, 2017 at 3:31 PM, syzbot > > wrote: > > Hello, > > > > syzkaller hit the following crash on > > 2db767d9889cef087149a5eaa35c1497671fa40f > > git://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/master > > compiler: gcc (GCC) 7.1.1 20170620 > > .config is attached > > Raw console output is attached. > > > > Unfortunately, I don't have any reproducer for this bug yet. > > > > > > BUG: workqueue lockup - pool cpus=0 node=0 flags=0x0 nice=0 stuck for 48s! > > BUG: workqueue lockup - pool cpus=0-1 flags=0x4 nice=0 stuck for 47s! > > Showing busy workqueues and worker pools: > > workqueue events: flags=0x0 > > pwq 0: cpus=0 node=0 flags=0x0 nice=0 active=4/256 > > pending: perf_sched_delayed, vmstat_shepherd, jump_label_update_timeout, > > cache_reap > > workqueue events_power_efficient: flags=0x80 > > pwq 0: cpus=0 node=0 flags=0x0 nice=0 active=4/256 > > pending: neigh_periodic_work, neigh_periodic_work, do_cache_clean, > > reg_check_chans_work > > workqueue mm_percpu_wq: flags=0x8 > > pwq 0: cpus=0 node=0 flags=0x0 nice=0 active=1/256 > > pending: vmstat_update > > workqueue writeback: flags=0x4e > > pwq 4: cpus=0-1 flags=0x4 nice=0 active=1/256 > > in-flight: 3401:wb_workfn > > workqueue kblockd: flags=0x18 > > pwq 1: cpus=0 node=0 flags=0x0 nice=-20 active=1/256 > > pending: blk_mq_timeout_work > > pool 4: cpus=0-1 flags=0x4 nice=0 hung=0s workers=11 idle: 3423 4249 92 21 > > > This error report does not look actionable. Perhaps if code that > detect it would dump cpu/task stacks, it would be actionable. That might be related to the RCU stall issue we are chasing, where a timer does not fire for yet unknown reasons. We have a reproducer now and hopefully a solution in the next days. Thanks, tglx