From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1755688AbcJRM2A (ORCPT ); Tue, 18 Oct 2016 08:28:00 -0400 Received: from mail-lf0-f65.google.com ([209.85.215.65]:33176 "EHLO mail-lf0-f65.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1753682AbcJRM1x (ORCPT ); Tue, 18 Oct 2016 08:27:53 -0400 Date: Tue, 18 Oct 2016 14:27:49 +0200 From: Michal Hocko To: Tetsuo Handa Cc: akpm@linux-foundation.org, hannes@cmpxchg.org, mgorman@suse.de, dave.hansen@intel.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: How to make warn_alloc() reliable? Message-ID: <20161018122749.GE12092@dhcp22.suse.cz> References: <201610182004.AEF87559.FOOHVLJOQFFtSM@I-love.SAKURA.ne.jp> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <201610182004.AEF87559.FOOHVLJOQFFtSM@I-love.SAKURA.ne.jp> User-Agent: Mutt/1.6.0 (2016-04-01) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Tue 18-10-16 20:04:20, Tetsuo Handa wrote: [...] > @@ -1697,11 +1697,25 @@ static bool inactive_reclaimable_pages(struct lruvec *lruvec, > int file = is_file_lru(lru); > struct pglist_data *pgdat = lruvec_pgdat(lruvec); > struct zone_reclaim_stat *reclaim_stat = &lruvec->reclaim_stat; > + unsigned long wait_start = jiffies; > + unsigned int wait_timeout = 10 * HZ; > + long last_diff = 0; > + long diff; > > if (!inactive_reclaimable_pages(lruvec, sc, lru)) > return 0; > > - while (unlikely(too_many_isolated(pgdat, file, sc))) { > + while (unlikely((diff = too_many_isolated(pgdat, file, sc)) > 0)) { > + if (diff < last_diff) { > + wait_start = jiffies; > + wait_timeout = 10 * HZ; > + } else if (time_after(jiffies, wait_start + wait_timeout)) { > + warn_alloc(sc->gfp_mask, > + "shrink_inactive_list() stalls for %ums", > + jiffies_to_msecs(jiffies - wait_start)); > + wait_timeout += 10 * HZ; > + } > + last_diff = diff; > congestion_wait(BLK_RW_ASYNC, HZ/10); > > /* We are about to die and free our memory. Return now. */ > ---------- [...] > So, how can we make warn_alloc() reliable? This is not about warn_alloc reliability but more about too_many_isolated waiting for an unbounded amount of time. And that should be fixed. I do not have a good idea how right now. -- Michal Hocko SUSE Labs