* oom killer and its superior braindamage in 2.4
@ 2003-02-23 16:06 David Mansfield
2003-02-23 16:25 ` Faik Uygur
2003-02-23 16:25 ` Rik van Riel
0 siblings, 2 replies; 15+ messages in thread
From: David Mansfield @ 2003-02-23 16:06 UTC (permalink / raw)
To: linux-kernel, Rik van Riel, Marc-Christian Petersen
Marc, Rik,
> - Feb 21 10:04:57 codeman kernel: Out of Memory: Killed process 2657
> (apache).
>
> The above log entry (apache) appeared for about 4 hours every some
> seconds (same PID) until I thought about sysrq-b to get out of this
> braindead behaviour. The machine was somewhat dead for me because I was
> not able to do anything but sysrq. The system itself was _not_ dead,
> there was massive disk i/o. This is 2.4.20 vanilla.
This exact thing happened to me as well, on a 2.4.20-pre that hasn't been
upgraded to 2.4.20 yet. The thing that concerns me most is:
Why won't the system kill the process it claims to be killing?
If, in Marc's case, the system wants to kill PID 2657, a lowly sleeping
apache process, why can't it? This is a bug for sure.
For me, there was some python process chosen as the one for killing and it
repeated the 'Out of Memory: Killed process xxxxx (python)' for hours
while making no progress. The machine was still routing packets but I
couldn't log in. Sys-rq was disabled, so I was forced to use the big red
button.
Rik, any ideas?
David
--
/==============================\
| David Mansfield |
| lkml@dm.cobite.com |
\==============================/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-23 16:06 oom killer and its superior braindamage in 2.4 David Mansfield
@ 2003-02-23 16:25 ` Faik Uygur
2003-02-23 16:25 ` Rik van Riel
1 sibling, 0 replies; 15+ messages in thread
From: Faik Uygur @ 2003-02-23 16:25 UTC (permalink / raw)
To: David Mansfield; +Cc: linux-kernel
> This exact thing happened to me as well, on a 2.4.20-pre that hasn't been
> upgraded to 2.4.20 yet. The thing that concerns me most is:
>
> Why won't the system kill the process it claims to be killing?
>
> If, in Marc's case, the system wants to kill PID 2657, a lowly sleeping
> apache process, why can't it? This is a bug for sure.
>
> For me, there was some python process chosen as the one for killing and it
> repeated the 'Out of Memory: Killed process xxxxx (python)' for hours
> while making no progress. The machine was still routing packets but I
> couldn't log in. Sys-rq was disabled, so I was forced to use the big red
> button.
>
> Rik, any ideas?
But, did you follow that thread? Rik van Riel, already suggested a solution for
the problem.
http://marc.theaimsgroup.com/?l=linux-kernel&m=104594301523518&w=2
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-23 16:06 oom killer and its superior braindamage in 2.4 David Mansfield
2003-02-23 16:25 ` Faik Uygur
@ 2003-02-23 16:25 ` Rik van Riel
2003-02-23 18:07 ` David Mansfield
1 sibling, 1 reply; 15+ messages in thread
From: Rik van Riel @ 2003-02-23 16:25 UTC (permalink / raw)
To: David Mansfield; +Cc: linux-kernel, Marc-Christian Petersen
On Sun, 23 Feb 2003, David Mansfield wrote:
> Rik, any ideas?
You could try the patch I sent to Marc and linux-kernel
yesterday afternoon ;)
Rik
--
Engineers don't grow up, they grow sideways.
http://www.surriel.com/ http://kernelnewbies.org/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-23 16:25 ` Rik van Riel
@ 2003-02-23 18:07 ` David Mansfield
2003-02-23 20:14 ` Rik van Riel
0 siblings, 1 reply; 15+ messages in thread
From: David Mansfield @ 2003-02-23 18:07 UTC (permalink / raw)
To: Rik van Riel; +Cc: linux-kernel, Marc-Christian Petersen
> On Sun, 23 Feb 2003, David Mansfield wrote:
>
> > Rik, any ideas?
>
> You could try the patch I sent to Marc and linux-kernel
> yesterday afternoon ;)
>
You miss my point completely. The kernel has ALREADY chosen a task to
kill. I don't care to adjust the 'badness' function. The kernel has
already chosen a bad task.
If you read my post, the bug is that the kernel CANNOT kill that process?
Why? If it's really a bad process, shouldn't it be the one that gets
killed?
With you patch we have:
1) Kernel goes OOM
2) Kernel picks the worst task to kill using badness()
3) Kernel attempts to kill this task but fails due to some {reason|bug}.
4) Kernel now picks some other task to kill even though the 'baddest' one
is allowed to hang out.
This is my question, and I don't see how the patch addresses it.
David
--
/==============================\
| David Mansfield |
| lkml@dm.cobite.com |
\==============================/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-23 18:07 ` David Mansfield
@ 2003-02-23 20:14 ` Rik van Riel
2003-02-23 20:22 ` David Mansfield
0 siblings, 1 reply; 15+ messages in thread
From: Rik van Riel @ 2003-02-23 20:14 UTC (permalink / raw)
To: David Mansfield; +Cc: linux-kernel, Marc-Christian Petersen
On Sun, 23 Feb 2003, David Mansfield wrote:
> If you read my post, the bug is that the kernel CANNOT kill that
> process? Why? If it's really a bad process, shouldn't it be the one
> that gets killed?
> This is my question, and I don't see how the patch addresses it.
And you won't see one, either. You cannot change the
semantics of uninterruptible sleep, nor can the OOM
killer change other device driver things.
This means the OOM killer has little choice but to
"hope for the best" and pick another process if the
first process chosen can't exit.
If you think you can fix all drivers to work fine
when tasks suddenly disappear, I guess you might
wnat to create such a patch ...
regards,
Rik
--
Engineers don't grow up, they grow sideways.
http://www.surriel.com/ http://kernelnewbies.org/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-23 20:14 ` Rik van Riel
@ 2003-02-23 20:22 ` David Mansfield
2003-02-23 20:53 ` Rik van Riel
0 siblings, 1 reply; 15+ messages in thread
From: David Mansfield @ 2003-02-23 20:22 UTC (permalink / raw)
To: Rik van Riel; +Cc: David Mansfield, linux-kernel, Marc-Christian Petersen
>
> > If you read my post, the bug is that the kernel CANNOT kill that
> > process? Why? If it's really a bad process, shouldn't it be the one
> > that gets killed?
>
> > This is my question, and I don't see how the patch addresses it.
>
> And you won't see one, either. You cannot change the
> semantics of uninterruptible sleep, nor can the OOM
> killer change other device driver things.
So you're saying that a process can stay in the D state, without ever
getting enough resources to complete a single Uninteruptible wait, for
hours at a time?
Ok. Now I understand your patch. Thanks for the info.
You should push your patch to Marcelo.
Thanks,
David
--
/==============================\
| David Mansfield |
| david@cobite.com |
\==============================/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-23 20:22 ` David Mansfield
@ 2003-02-23 20:53 ` Rik van Riel
0 siblings, 0 replies; 15+ messages in thread
From: Rik van Riel @ 2003-02-23 20:53 UTC (permalink / raw)
To: David Mansfield; +Cc: David Mansfield, linux-kernel, Marc-Christian Petersen
On Sun, 23 Feb 2003, David Mansfield wrote:
> So you're saying that a process can stay in the D state, without ever
> getting enough resources to complete a single Uninteruptible wait, for
> hours at a time?
Or even in the R state, but that would only happen when there
is a kernel bug. The OOM killer can do nothing but hope for
the best and try another process if the first one doesn't want
to exit.
> Ok. Now I understand your patch. Thanks for the info.
>
> You should push your patch to Marcelo.
Will do.
cheers,
Rik
--
Engineers don't grow up, they grow sideways.
http://www.surriel.com/ http://kernelnewbies.org/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-24 9:13 Mikael Starvik
@ 2003-02-24 9:27 ` Marc-Christian Petersen
0 siblings, 0 replies; 15+ messages in thread
From: Marc-Christian Petersen @ 2003-02-24 9:27 UTC (permalink / raw)
To: Mikael Starvik, 'linux-kernel@vger.kernel.org'
Cc: Jonas Holmberg, Sebastian Sjoberg
On Monday 24 February 2003 10:13, Mikael Starvik wrote:
Hi Mikael,
> Does everyone agree that killing a process is always the best approach
> to resolve an OOM? If the OOM is caused by e.g. a growing tmpfs or
> memory leaks in the kernel it won't help much to kill processes that
> may respawn.
Well, I don't agree that it's always the best approach. Other bad things, you
metioned it, can happen.
> Would it be useful if it was possible to register another oom-handler?
> Some architectures could then choose to e.g. reboot the system instead.
I'd like to see _an option_ (read: not default but an option, e.g. boot
parameter) that will reboot the machine after $specified_time if an OOM
killing action does not stop.
ciao, Marc
^ permalink raw reply [flat|nested] 15+ messages in thread
* RE: oom killer and its superior braindamage in 2.4
@ 2003-02-24 9:13 Mikael Starvik
2003-02-24 9:27 ` Marc-Christian Petersen
0 siblings, 1 reply; 15+ messages in thread
From: Mikael Starvik @ 2003-02-24 9:13 UTC (permalink / raw)
To: 'Marc-Christian Petersen',
'linux-kernel@vger.kernel.org'
Cc: Jonas Holmberg, Sebastian Sjoberg
Does everyone agree that killing a process is always the best approach
to resolve an OOM? If the OOM is caused by e.g. a growing tmpfs or
memory leaks in the kernel it won't help much to kill processes that
may respawn.
Would it be useful if it was possible to register another oom-handler?
Some architectures could then choose to e.g. reboot the system instead.
/Mikael
-----Original Message-----
From: linux-kernel-owner@vger.kernel.org
[mailto:linux-kernel-owner@vger.kernel.org]On Behalf Of Marc-Christian
Petersen
Sent: Saturday, February 22, 2003 8:35 PM
To: linux-kernel@vger.kernel.org
Subject: oom killer and its superior braindamage in 2.4
Hi all,
I just thought (ok it was yesterday) about stress testing my mysql db.
I used this:
- mystress.pl localhost mysql root test 600 300 60 "select * from user"
It worked like a charme. So I tried:
- mystress.pl localhost mysql root test 1800 900 60 "select * from user"
My machine has 512MB RAM and 512MB SWAP.
I expected that the 2nd run will OOM my machine but I did not expect this
silly behaviour.
The following log entry appeared only _once_ (there were ~700 mysqld running)
- Feb 21 10:03:22 codeman kernel: Out of Memory: Killed process 1463 (mysqld).
Instead of really killing either mysqld or mystress.pl the OOM killer decided
to kill apache (apache did nothing but had 5 threads sleeping)
- Feb 21 10:04:57 codeman kernel: Out of Memory: Killed process 2657 (apache).
The above log entry (apache) appeared for about 4 hours every some seconds
(same PID) until I thought about sysrq-b to get out of this braindead
behaviour. The machine was somewhat dead for me because I was not able to do
anything but sysrq. The system itself was _not_ dead, there was massive disk
i/o. This is 2.4.20 vanilla.
Is there any chance we can fix this up?
ciao, Marc
-
To unsubscribe from this list: send the line "unsubscribe linux-kernel" in
the body of a message to majordomo@vger.kernel.org
More majordomo info at http://vger.kernel.org/majordomo-info.html
Please read the FAQ at http://www.tux.org/lkml/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-23 20:18 ` Rik van Riel
@ 2003-02-23 20:29 ` Marc-Christian Petersen
0 siblings, 0 replies; 15+ messages in thread
From: Marc-Christian Petersen @ 2003-02-23 20:29 UTC (permalink / raw)
To: Rik van Riel; +Cc: linux-kernel, Andrew Morton, David Mansfield
On Sunday 23 February 2003 21:18, Rik van Riel wrote:
Hi Rik,
> It'd be interesting to know where these processes are spending
> their CPU time and why they're not catching their signals.
I'll look into it again when I do the next run.
> > Sysrq-i gave me the chance to get out of the OOM killing process and
> > only kernel threads were left + getty's so I was able to log in again.
> Strange, so sysrq-i manages to kill the processes, but the OOM
> killer doesn't kill the processes ?
yep, so it is.
> This is very suspect because the OOM killer uses force_sig in
> the same way the sysrq-i handler does...
indeed. Well, sysrq-i need about 5 seconds to give me my getty back.
Anyway, your patch should go into -BK. Your patch does _not_ introduce this
behaviour, it's present even w/o your patch but your approach makes things
better :)
ciao, Marc
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-23 17:35 ` Marc-Christian Petersen
@ 2003-02-23 20:18 ` Rik van Riel
2003-02-23 20:29 ` Marc-Christian Petersen
0 siblings, 1 reply; 15+ messages in thread
From: Rik van Riel @ 2003-02-23 20:18 UTC (permalink / raw)
To: Marc-Christian Petersen; +Cc: linux-kernel, Andrew Morton
On Sun, 23 Feb 2003, Marc-Christian Petersen wrote:
> > Does the below patch fix your problem ?
> With your patch, mystress.pl was marked to get killed, every PID only
> once, no apache or similar (good). ... But the strange thing is, that it
> seems none of the processes, which are marked to be killed, get killed.
> So sysrq-t tells me.
It'd be interesting to know where these processes are spending
their CPU time and why they're not catching their signals.
> Sysrq-i gave me the chance to get out of the OOM killing process and
> only kernel threads were left + getty's so I was able to log in again.
Strange, so sysrq-i manages to kill the processes, but the OOM
killer doesn't kill the processes ?
This is very suspect because the OOM killer uses force_sig in
the same way the sysrq-i handler does...
regards,
Rik
--
Engineers don't grow up, they grow sideways.
http://www.surriel.com/ http://kernelnewbies.org/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-22 20:32 ` Rik van Riel
@ 2003-02-23 17:35 ` Marc-Christian Petersen
2003-02-23 20:18 ` Rik van Riel
0 siblings, 1 reply; 15+ messages in thread
From: Marc-Christian Petersen @ 2003-02-23 17:35 UTC (permalink / raw)
To: Rik van Riel; +Cc: linux-kernel, Andrew Morton
On Saturday 22 February 2003 21:32, Rik van Riel wrote:
Hi Rik,
> > > - Feb 21 10:04:57 codeman kernel: Out of Memory: Killed process 2657
> > > (apache).
> > > The above log entry (apache) appeared for about 4 hours every some
> > > seconds (same PID) until I thought about sysrq-b
> > > Is there any chance we can fix this up?
> > Yes.
> Never mind my last idea, it can be done much simpler ;)
hehe :)
> Does the below patch fix your problem ?
Well, this makes a difference. I filled up my memory with something else
before starting mystress.pl because of top's|ps' slowness with many processes.
I had about 400 processes. The test from yesterday had ~ 1800.
With your patch, mystress.pl was marked to get killed, every PID only once, no
apache or similar (good). ... But the strange thing is, that it seems none of
the processes, which are marked to be killed, get killed. So sysrq-t tells
me. Sysrq-i gave me the chance to get out of the OOM killing process and only
kernel threads were left + getty's so I was able to log in again.
ciao, Marc
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-22 20:14 ` Rik van Riel
@ 2003-02-22 20:32 ` Rik van Riel
2003-02-23 17:35 ` Marc-Christian Petersen
0 siblings, 1 reply; 15+ messages in thread
From: Rik van Riel @ 2003-02-22 20:32 UTC (permalink / raw)
To: Marc-Christian Petersen; +Cc: linux-kernel
On Sat, 22 Feb 2003, Rik van Riel wrote:
> On Sat, 22 Feb 2003, Marc-Christian Petersen wrote:
>
> > - Feb 21 10:04:57 codeman kernel: Out of Memory: Killed process 2657 (apache).
> >
> > The above log entry (apache) appeared for about 4 hours every some
> > seconds (same PID) until I thought about sysrq-b
>
> > Is there any chance we can fix this up?
>
> Yes.
Never mind my last idea, it can be done much simpler ;)
Does the below patch fix your problem ?
Rik
--
Engineers don't grow up, they grow sideways.
http://www.surriel.com/ http://kernelnewbies.org/
===== mm/oom_kill.c 1.11 vs edited =====
--- 1.11/mm/oom_kill.c Fri Aug 16 10:59:46 2002
+++ edited/mm/oom_kill.c Sat Feb 22 17:31:49 2003
@@ -61,6 +61,9 @@
if (!p->mm)
return 0;
+
+ if (p->flags & PF_MEMDIE)
+ return 0;
/*
* The memory size of the process is the basis for the badness.
*/
^ permalink raw reply [flat|nested] 15+ messages in thread
* Re: oom killer and its superior braindamage in 2.4
2003-02-22 19:35 Marc-Christian Petersen
@ 2003-02-22 20:14 ` Rik van Riel
2003-02-22 20:32 ` Rik van Riel
0 siblings, 1 reply; 15+ messages in thread
From: Rik van Riel @ 2003-02-22 20:14 UTC (permalink / raw)
To: Marc-Christian Petersen; +Cc: linux-kernel
On Sat, 22 Feb 2003, Marc-Christian Petersen wrote:
> - Feb 21 10:04:57 codeman kernel: Out of Memory: Killed process 2657 (apache).
>
> The above log entry (apache) appeared for about 4 hours every some
> seconds (same PID) until I thought about sysrq-b
> Is there any chance we can fix this up?
Yes.
1) add a VM_KILLED flag
2) set this flag on the p->mm->def_flags when you kill a
process/thread from oom_kill.c
3) clear the flag on process exit
4) on a new call to oom_kill, skip processes when
(p->mm->def_flags & VM_KILLED) by returning 0
points for such a process
cheers,
Rik
--
Engineers don't grow up, they grow sideways.
http://www.surriel.com/ http://kernelnewbies.org/
^ permalink raw reply [flat|nested] 15+ messages in thread
* oom killer and its superior braindamage in 2.4
@ 2003-02-22 19:35 Marc-Christian Petersen
2003-02-22 20:14 ` Rik van Riel
0 siblings, 1 reply; 15+ messages in thread
From: Marc-Christian Petersen @ 2003-02-22 19:35 UTC (permalink / raw)
To: linux-kernel
Hi all,
I just thought (ok it was yesterday) about stress testing my mysql db.
I used this:
- mystress.pl localhost mysql root test 600 300 60 "select * from user"
It worked like a charme. So I tried:
- mystress.pl localhost mysql root test 1800 900 60 "select * from user"
My machine has 512MB RAM and 512MB SWAP.
I expected that the 2nd run will OOM my machine but I did not expect this
silly behaviour.
The following log entry appeared only _once_ (there were ~700 mysqld running)
- Feb 21 10:03:22 codeman kernel: Out of Memory: Killed process 1463 (mysqld).
Instead of really killing either mysqld or mystress.pl the OOM killer decided
to kill apache (apache did nothing but had 5 threads sleeping)
- Feb 21 10:04:57 codeman kernel: Out of Memory: Killed process 2657 (apache).
The above log entry (apache) appeared for about 4 hours every some seconds
(same PID) until I thought about sysrq-b to get out of this braindead
behaviour. The machine was somewhat dead for me because I was not able to do
anything but sysrq. The system itself was _not_ dead, there was massive disk
i/o. This is 2.4.20 vanilla.
Is there any chance we can fix this up?
ciao, Marc
^ permalink raw reply [flat|nested] 15+ messages in thread
end of thread, other threads:[~2003-02-24 9:51 UTC | newest]
Thread overview: 15+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2003-02-23 16:06 oom killer and its superior braindamage in 2.4 David Mansfield
2003-02-23 16:25 ` Faik Uygur
2003-02-23 16:25 ` Rik van Riel
2003-02-23 18:07 ` David Mansfield
2003-02-23 20:14 ` Rik van Riel
2003-02-23 20:22 ` David Mansfield
2003-02-23 20:53 ` Rik van Riel
-- strict thread matches above, loose matches on Subject: below --
2003-02-24 9:13 Mikael Starvik
2003-02-24 9:27 ` Marc-Christian Petersen
2003-02-22 19:35 Marc-Christian Petersen
2003-02-22 20:14 ` Rik van Riel
2003-02-22 20:32 ` Rik van Riel
2003-02-23 17:35 ` Marc-Christian Petersen
2003-02-23 20:18 ` Rik van Riel
2003-02-23 20:29 ` Marc-Christian Petersen
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®