From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752694AbZHSS2n (ORCPT ); Wed, 19 Aug 2009 14:28:43 -0400 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S1751515AbZHSS2n (ORCPT ); Wed, 19 Aug 2009 14:28:43 -0400 Received: from newpeace.netnation.com ([204.174.223.7]:47030 "EHLO peace.netnation.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S1751491AbZHSS2m (ORCPT ); Wed, 19 Aug 2009 14:28:42 -0400 X-Greylist: delayed 1249 seconds by postgrey-1.27 at vger.kernel.org; Wed, 19 Aug 2009 14:28:42 EDT Date: Wed, 19 Aug 2009 11:07:54 -0700 From: Simon Kirby To: linux-kernel@vger.kernel.org Subject: Processes hanging under heavy write loads Message-ID: <20090819180754.GB7068@hostway.ca> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline User-Agent: Mutt/1.5.13 (2006-08-11) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi all, On an storage head box running 2.6.30, it's easy to see even sshd hang when allocating memory to send a packet (eg: while watching "top"), sometimes for several seconds. The hung process detector, with the timeout lowered a bit, spits out a backtrace such as: INFO: task sshd:31015 blocked for more than 4 seconds. "echo 0 > /proc/sys/kernel/hung_task_timeout_secs" disables this message. sshd D ffffffff8087b144 0 31015 3378 ffff8801c5afd918 0000000000000086 0000000000000000 ffff880100b3e070 ffff880100b3ddc0 ffff880183757080 ffff880100b3e070 ffffe2000dd81780 ffff8801c5afd8f8 ffffffff8028c235 ffffe2000e578bd0 ffffffffffffffff Call Trace: [] ? determine_dirtyable_memory+0x15/0x30 [] __mutex_lock_slowpath+0xd1/0x150 [] mutex_lock+0x1e/0x40 [] shrink_icache_memory+0x7d/0x2b0 [] shrink_slab+0x125/0x180 [] try_to_free_pages+0x26a/0x3e0 [] ? isolate_pages_global+0x0/0x290 [] __alloc_pages_internal+0x19f/0x440 [] ? pollwake+0x0/0x60 [] __slab_alloc+0x151/0x570 [] ? __alloc_skb+0x46/0x170 [] kmem_cache_alloc+0xb9/0x110 [] __alloc_skb+0x46/0x170 [] sk_stream_alloc_skb+0x41/0x110 [] tcp_sendmsg+0x2f0/0xad0 [] sock_aio_write+0xf0/0x100 [] do_sync_write+0xf1/0x130 [] ? autoremove_wake_function+0x0/0x40 [] ? current_fs_time+0x22/0x30 [] ? tty_ldisc_deref+0x58/0x70 [] vfs_write+0x175/0x180 [] sys_write+0x50/0x90 [] system_call_fastpath+0x16/0x1b ...This mutex appears to be iprune_mutex, called from prune_icache in fs/inode.c. I watched this for a while, and all of the backtraces seem to be the same. Would it be a reasonable idea to convert this to a mutex_trylock since a holder of it is trying to do the same work anyway? I'm not sure what is taking so long during heavy write sessions, but it has to be either invalidate_inodes() or prune_icache(). The current behaviour is horrible to work with when non-guilty processes, such as sshd, happen to get stuck on it... Simon-