mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: KOSAKI Motohiro <kosaki.motohiro@jp.fujitsu.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: kosaki.motohiro@jp.fujitsu.com,
	David Rientjes <rientjes@google.com>,
	Linus Torvalds <torvalds@linux-foundation.org>,
	linux-kernel@vger.kernel.org, Rik van Riel <riel@redhat.com>,
	Oleg Nesterov <oleg@redhat.com>
Subject: Re: Linux 2.6.38
Date: Wed, 16 Mar 2011 18:09:57 +0900 (JST)	[thread overview]
Message-ID: <20110315153801.3526.A69D9226@jp.fujitsu.com> (raw)
In-Reply-To: <20110314232156.0c363813.akpm@linux-foundation.org>


[Yesterdays earthquake was announced magnitude 6.5, but M6 quake is
 no longer treated significant news in this country. We are living 
 slightly in a floating mood.]


> So we're talking about three patches:
> 
> oom-prevent-unnecessary-oom-kills-or-kernel-panics.patch
> oom-skip-zombies-when-iterating-tasklist.patch
> oom-avoid-deferring-oom-killer-if-exiting-task-is-being-traced.patch
> 
> all appended below.
> 
> About all of which Oleg had serious complaints, some of which haven't
> yet been addressed.
> 
> And that's OK.  As I said, please let's work through it and get it right.

I haven't understand what is "OK" and what do you want talk. probably
the reason is in my language skill or I haven't catch up Oleg and David
discussion. then instead, I'll post my debugging progressing condition.


 o vmscan.c#all_unreclaimable() might return false negative and lead
   to prevent oom-killer by mistaken. Why? zone->pages_scanned is not
   protected by lock, in other words, it's unstable value. in the other
   hands, x86 ZONE_DMA has only a very little memory, then usually
   never recover all_unreclaimable=no if once become all_unreclaimable=yes.
   then, if zone state become unmatched (eg pages_scanned=0 and all_unreclaimable=yes)
   it can't be recovered never. I mean I could reproduced Andrey reported issue.

 o oom_kill.c#boost_dying_task_prio() makes kernel hang-up if user
   are using cpu cgroups. because cpu cgroup has inadequate default
   RT rt_runtime_us (0 by default. 0 mean RT tasks can't run at all).

 o oom_kill.c#TIF_MEMDIE check makes kernel hang-up. I haven't catch
   the exact reason of a oom killed process sticking even though zone has
   enough memory. 

My dislikeness is, Many people in the list fun to make flamewar but 
nobody except really a few developers run the real code nor join to 
debug real and actual reported issue. In fact, Andrey made testcase and
reported his test environment and help we made reproduce envronemnt.

I also dislike some developer say they haven't seen oom livelock case yet.
It indicate they haven't tested stress workload oom scenario. How do i
trust an untested patch, an untested guys? All developer have to test
until seen oom livelock.

I know oom debugging is very painful and need to take a lot of time.
much false positive, much unfixable live lock, need mililion reset. 
But, I don't think this is good reason to take untested.

Now I'm only access a three years old PC. Therefore, I have no reason
anyone can't debug the issue.



  reply	other threads:[~2011-03-16  9:10 UTC|newest]

Thread overview: 74+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2011-03-15  1:49 Linus Torvalds
2011-03-15  3:13 ` David Rientjes
2011-03-15  4:06   ` Steven Rostedt
2011-03-15  4:14   ` Linus Torvalds
2011-03-15  4:29     ` David Rientjes
2011-03-15  4:33   ` Andrew Morton
2011-03-15  4:50     ` David Rientjes
2011-03-15  6:21       ` Andrew Morton
2011-03-16  9:09         ` KOSAKI Motohiro [this message]
2011-03-22 11:04           ` [patch 0/5] oom: a few anti fork bomb patches KOSAKI Motohiro
2011-03-22 11:05             ` [PATCH 1/5] vmscan: remove all_unreclaimable check from direct reclaim path completely KOSAKI Motohiro
2011-03-22 14:49               ` Minchan Kim
2011-03-23  5:21                 ` KOSAKI Motohiro
2011-03-23  6:59                   ` Minchan Kim
2011-03-23  7:13                     ` KOSAKI Motohiro
2011-03-23  8:24                       ` Minchan Kim
2011-03-23  8:44                         ` KOSAKI Motohiro
2011-03-23  9:02                           ` Minchan Kim
2011-03-24  2:11                             ` KOSAKI Motohiro
2011-03-24  2:21                               ` Andrew Morton
2011-03-24  2:48                                 ` KOSAKI Motohiro
2011-03-24  3:04                                   ` Andrew Morton
2011-03-24  5:35                                     ` KOSAKI Motohiro
2011-03-24  4:19                               ` Minchan Kim
2011-03-24  5:35                                 ` KOSAKI Motohiro
2011-03-24  5:53                                   ` Minchan Kim
2011-03-24  6:16                                     ` KOSAKI Motohiro
2011-03-24  6:32                                       ` Minchan Kim
2011-03-24  7:03                                         ` KOSAKI Motohiro
2011-03-24  7:25                                           ` Minchan Kim
2011-03-24  7:28                                             ` KOSAKI Motohiro
2011-03-24  7:34                                               ` Minchan Kim
2011-03-24  7:41                                                 ` Minchan Kim
2011-03-24  7:43                                                 ` KOSAKI Motohiro
2011-03-24  7:43                                   ` Minchan Kim
2011-03-23  7:41               ` KAMEZAWA Hiroyuki
2011-03-23  7:55                 ` KOSAKI Motohiro
2011-03-22 11:08             ` [PATCH 3/5] oom: create oom autogroup KOSAKI Motohiro
2011-03-22 23:21               ` Minchan Kim
2011-03-23  1:27                 ` KOSAKI Motohiro
2011-03-23  2:41                   ` Mike Galbraith
2011-03-22 11:08             ` [PATCH 4/5] mm: introduce wait_on_page_locked_killable KOSAKI Motohiro
2011-03-23  7:44               ` KAMEZAWA Hiroyuki
2011-03-24 15:04               ` Minchan Kim
2011-03-22 11:09             ` [PATCH 5/5] x86,mm: make pagefault killable KOSAKI Motohiro
2011-03-23  7:49               ` KAMEZAWA Hiroyuki
2011-03-23  8:09                 ` KOSAKI Motohiro
2011-03-23 14:34                   ` Linus Torvalds
2011-03-24 15:10               ` Minchan Kim
2011-03-24 17:13               ` Oleg Nesterov
2011-03-24 17:34                 ` Linus Torvalds
2011-03-28  7:00                   ` KOSAKI Motohiro
     [not found]             ` <20110322200657.B064.A69D9226@jp.fujitsu.com>
     [not found]               ` <20110323164229.6b647004.kamezawa.hiroyu@jp.fujitsu.com>
2011-03-23 13:40                 ` [PATCH 2/5] Revert "oom: give the dying task a higher priority" Luis Claudio R. Goncalves
2011-03-24  0:06                   ` KOSAKI Motohiro
2011-03-24 15:27               ` Minchan Kim
2011-03-28  9:48                 ` KOSAKI Motohiro
2011-03-28 12:28                   ` Minchan Kim
2011-03-28  9:51                 ` Peter Zijlstra
2011-03-28 12:21                   ` Minchan Kim
2011-03-28 12:28                     ` Peter Zijlstra
2011-03-28 12:40                       ` Minchan Kim
2011-03-28 13:10                         ` Luis Claudio R. Goncalves
2011-03-28 13:18                           ` Peter Zijlstra
2011-03-28 13:56                             ` Luis Claudio R. Goncalves
2011-03-29  2:46                             ` KOSAKI Motohiro
2011-03-28 13:48                           ` Minchan Kim
2011-03-15 21:08       ` Linux 2.6.38 Oleg Nesterov
2011-03-15  3:14 ` Steven Rostedt
2011-03-15  4:15   ` Linus Torvalds
2011-03-16 17:30 ` i915/kms regression after 2.6.38-rc8 (was: Re: Linux 2.6.38) Melchior FRANZ
2011-03-16 19:22   ` i915/kms regression after 2.6.38-rc8 Jiri Slaby
2011-03-16 19:43   ` i915/kms regression after 2.6.38-rc8 (was: Re: Linux 2.6.38) Chris Wilson
2011-03-16 21:09     ` i915/kms regression after 2.6.38-rc8 Melchior FRANZ
2011-03-20 18:30   ` i915/kms regression after 2.6.38-rc8 (was: Re: Linux 2.6.38) Maciej Rutecki

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20110315153801.3526.A69D9226@jp.fujitsu.com \
    --to=kosaki.motohiro@jp.fujitsu.com \
    --cc=akpm@linux-foundation.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=oleg@redhat.com \
    --cc=riel@redhat.com \
    --cc=rientjes@google.com \
    --cc=torvalds@linux-foundation.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®