mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* 2.4.10 VM, active cache pages, and OOM
@ 2001-09-29 16:19 Tobias Ringstrom
  2001-09-29 17:16 ` Sebastian Benoit
  2001-09-29 18:36 ` Linus Torvalds
  0 siblings, 2 replies; 6+ messages in thread
From: Tobias Ringstrom @ 2001-09-29 16:19 UTC (permalink / raw)
  To: Kernel Mailing List

First I'd like to say that the 2.4.10 VM works great for my desktop and
home server, much better than previous versions.  I have not tried Alan's
kernels.

I do have one problem, though, and it is illustrated by the following very
simple program:

	#include <unistd.h>
	int main()
	{
		char buf[512];
		while (read(0, buf, sizeof(buf)) == sizeof(buf))
			;
		return 0;
	}

The program should be reading a block device, but a big file probably does
the trick as well.

	./a.out < /dev/hde1

When the program is running, all cached pages pop up in the active list,
and when the memory is full of active pages, the computer starts to page
out stuff, becomes VERY unresponsive, and after half a minute or so it
goes OOM and starts killing processes.  There are lots and lots of free
swap at this time.  I also get a bunch of 0-order allocation failures in
the log.

(I'd say that the OOM killer does seem to kill the most memory-hog-like
processes, but the problem is that it is not the processes that use up all
the memory, it is the active cache pages.)

If the buf size is changed to a multiple of the page size, such as 4096,
the cache pages are instead added to the inactive list, and the system is
very responsive, no paging occurs, and it does not go OOM.  In other
words, it works perfectly.

I assume that the difference between a buf size of 512 and 4096 is that
for the 512-byte case, each page is touched more than once, and that's why
the system think the pages are active.  This is a very wrong decision,
since I'm doing a sequential read.

Fixing that particular problem will get rid of my problem, but I'm
guessing that it would only hide another real problem, which is that
2.4.10 has a huge problem freeing pages from the list of active pages,
even if they are clean, and thus making a wrong decision on the
availibility of free(able) pages.

Am I right to assume that if I would make the program do random seeks, or
read each page twice, the pages would again be added to the active list,
even if I would read whole pages at a time?

I also wonder why the system get so unresponsive before it goes OOM.
Perhaps there is a kernel process running, scanning lists trying to free
memory, but not finding any, wasting all CPU cycles.

What do you think?

/Tobias


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: 2.4.10 VM, active cache pages, and OOM
  2001-09-29 16:19 2.4.10 VM, active cache pages, and OOM Tobias Ringstrom
@ 2001-09-29 17:16 ` Sebastian Benoit
  2001-09-29 18:36 ` Linus Torvalds
  1 sibling, 0 replies; 6+ messages in thread
From: Sebastian Benoit @ 2001-09-29 17:16 UTC (permalink / raw)
  To: Tobias Ringstrom; +Cc: Kernel Mailing List

[-- Attachment #1: Type: text/plain, Size: 698 bytes --]

Tobias Ringstrom(tori@ringstrom.mine.nu)@2001.09.29 18:19:13 +0000:
> First I'd like to say that the 2.4.10 VM works great for my desktop and
> home server, much better than previous versions.  I have not tried Alan's
> kernels.

[snip] 

> What do you think?

Try the vm-tweaks-2 patch, see

 http://uwsg.iu.edu/hypermail/linux/kernel/0109.3/0626.html

with that I ran 2.4.10 for 2 days w/o problems. best 2.4.x ever ;) 

Right now i have switched over to 2.4.9-ac17 and i must say it works just as
well.

/B.

-- 
Sebastian Benoit <ben-lists@andastra.de>
OpenPGP-Key ID 0x82AE75E4                             
fingerprint 0BDA 0CB7 9BCA AF77 28EE  D91A 396D 93BC 82AE 75E

[-- Attachment #2: Type: application/pgp-signature, Size: 232 bytes --]

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: 2.4.10 VM, active cache pages, and OOM
  2001-09-29 16:19 2.4.10 VM, active cache pages, and OOM Tobias Ringstrom
  2001-09-29 17:16 ` Sebastian Benoit
@ 2001-09-29 18:36 ` Linus Torvalds
  2001-09-29 18:46   ` Rik van Riel
  2001-09-29 19:43   ` Tobias Ringstrom
  1 sibling, 2 replies; 6+ messages in thread
From: Linus Torvalds @ 2001-09-29 18:36 UTC (permalink / raw)
  To: linux-kernel

In article <Pine.LNX.4.33.0109291645260.16885-100000@boris.prodako.se>,
Tobias Ringstrom  <tori@ringstrom.mine.nu> wrote:
>
>I assume that the difference between a buf size of 512 and 4096 is that
>for the 512-byte case, each page is touched more than once, and that's why
>the system think the pages are active.  This is a very wrong decision,
>since I'm doing a sequential read.
>
>Fixing that particular problem will get rid of my problem, but I'm
>guessing that it would only hide another real problem, which is that
>2.4.10 has a huge problem freeing pages from the list of active pages,
>even if they are clean, and thus making a wrong decision on the
>availibility of free(able) pages.

Absolutely right.

It's probably worth fixing the "sequential accesses of < pagesize count
as 'active'" problem too, but the real issue is that if you get into a
situation with _many_ more active pages than inactive, the plain 2.4.10
VM doesn't age the active list nearly fast enough.

That's fixed in Andrea's VM tweaks, but if you want to look into this,
the basic problem is in mm/vmscan.c, shrink_caches(), which in plain
2.4.10 does

        /* Do we want to age the active list? */
        if (nr_inactive_pages < nr_active_pages*2)
                refill_inactive(nr_pages);

which doesn't take into account just _how_ imbalanced the active list
is. So if the active list is huge, it will still just scan a small fixed
percentage of it (and to make matters worse, the small part of it is
proportional to the size of the _inactive_ list, so if the inactive list
is small, that just makes the problem worse.

What Andreas fix does is to make the refill rate be proportional to the
sizes of the lists, which should fix this problem for you. 

However, I'd also like to fix generic_file_read() to only mark the page
accessed when we're touching it for the first time, and notice
sequential accesses automatically. That way the use-once logic doesn't
depend on the read size - which is a totally independent problem.

If you want to test, the fix for _that_ is in mm/filemap.c:
do_generic_file_read(), where the code does:

		...
                ret = actor(desc, page, offset, nr);
                offset += ret;
                index += offset >> PAGE_CACHE_SHIFT;
                offset &= ~PAGE_CACHE_MASK;

                mark_page_accessed(page);
		...

and it would be interesting to hear if the behaviour improves with the
above mark_page_accessed() logic moved a bit and changed to:

		...
                ret = actor(desc, page, offset, nr);
		if (!offset || !file->f_reada)
			mark_page_accessed(page);
                offset += ret;
                index += offset >> PAGE_CACHE_SHIFT;
                offset &= ~PAGE_CACHE_MASK;
		...

(which basically says: we only mark the page accessed if we read the
_beginning_ of the page, or if we just did a seek to it)

Btw, if you test the above change out and confirm that it fixes te
behaviour, please send me an acknowledgement email - I've not done it in
my own tree yet, and unless I get a "yes, that works well" email I won't
be doing it..

		Linus

^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: 2.4.10 VM, active cache pages, and OOM
  2001-09-29 18:36 ` Linus Torvalds
@ 2001-09-29 18:46   ` Rik van Riel
  2001-09-29 19:21     ` Linus Torvalds
  2001-09-29 19:43   ` Tobias Ringstrom
  1 sibling, 1 reply; 6+ messages in thread
From: Rik van Riel @ 2001-09-29 18:46 UTC (permalink / raw)
  To: Linus Torvalds; +Cc: linux-kernel

On Sat, 29 Sep 2001, Linus Torvalds wrote:

> (which basically says: we only mark the page accessed if we read the
> _beginning_ of the page, or if we just did a seek to it)

That should work for linear IO, but I fear what influence
such a thing would have on eg. database indexes ;)

Rik
-- 
IA64: a worthy successor to i860.

http://www.surriel.com/		http://distro.conectiva.com/

Send all your spam to aardvark@nl.linux.org (spam digging piggy)


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: 2.4.10 VM, active cache pages, and OOM
  2001-09-29 18:46   ` Rik van Riel
@ 2001-09-29 19:21     ` Linus Torvalds
  0 siblings, 0 replies; 6+ messages in thread
From: Linus Torvalds @ 2001-09-29 19:21 UTC (permalink / raw)
  To: Rik van Riel; +Cc: linux-kernel


On Sat, 29 Sep 2001, Rik van Riel wrote:
> On Sat, 29 Sep 2001, Linus Torvalds wrote:
>
> > (which basically says: we only mark the page accessed if we read the
> > _beginning_ of the page, or if we just did a seek to it)
>
> That should work for linear IO, but I fear what influence
> such a thing would have on eg. database indexes ;)

Well, for things that seek, the behaviour will be the same as it was
before: it will always mark the page accessed, because "file->f_reada"
will always be zero for the first read after a lseek.

That's why we have the "or if we just did a seek to it". You cannot _just_
test for "did we read the beginning of a page", because that fails for
seekers, whether database or otherwise.

		Linus


^ permalink raw reply	[flat|nested] 6+ messages in thread

* Re: 2.4.10 VM, active cache pages, and OOM
  2001-09-29 18:36 ` Linus Torvalds
  2001-09-29 18:46   ` Rik van Riel
@ 2001-09-29 19:43   ` Tobias Ringstrom
  1 sibling, 0 replies; 6+ messages in thread
From: Tobias Ringstrom @ 2001-09-29 19:43 UTC (permalink / raw)
  To: Linus Torvalds; +Cc: linux-kernel

On Sat, 29 Sep 2001, Linus Torvalds wrote:

> However, I'd also like to fix generic_file_read() to only mark the page
> accessed when we're touching it for the first time, and notice
> sequential accesses automatically. That way the use-once logic doesn't
> depend on the read size - which is a totally independent problem.
>
> If you want to test, the fix for _that_ is in mm/filemap.c:
> do_generic_file_read(), where the code does:
>
> 		...
>                 ret = actor(desc, page, offset, nr);
>                 offset += ret;
>                 index += offset >> PAGE_CACHE_SHIFT;
>                 offset &= ~PAGE_CACHE_MASK;
>
>                 mark_page_accessed(page);
> 		...
>
> and it would be interesting to hear if the behaviour improves with the
> above mark_page_accessed() logic moved a bit and changed to:
>
> 		...
>                 ret = actor(desc, page, offset, nr);
> 		if (!offset || !file->f_reada)
> 			mark_page_accessed(page);
>                 offset += ret;
>                 index += offset >> PAGE_CACHE_SHIFT;
>                 offset &= ~PAGE_CACHE_MASK;
> 		...
>
> (which basically says: we only mark the page accessed if we read the
> _beginning_ of the page, or if we just did a seek to it)
>
> Btw, if you test the above change out and confirm that it fixes te
> behaviour, please send me an acknowledgement email - I've not done it in
> my own tree yet, and unless I get a "yes, that works well" email I won't
> be doing it..

Yes, that works well, and I tried with a block sizes of 1, 512, 4095 and
4096.  The cache pages are not beeing activated now.  When reading the
same buf twice with a seek between the reads, the pages are activated, as
expected.

I'll have a look at the other problem, and Andrea's solution, later.

/Tobias


^ permalink raw reply	[flat|nested] 6+ messages in thread

end of thread, other threads:[~2001-09-29 19:44 UTC | newest]

Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2001-09-29 16:19 2.4.10 VM, active cache pages, and OOM Tobias Ringstrom
2001-09-29 17:16 ` Sebastian Benoit
2001-09-29 18:36 ` Linus Torvalds
2001-09-29 18:46   ` Rik van Riel
2001-09-29 19:21     ` Linus Torvalds
2001-09-29 19:43   ` Tobias Ringstrom

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®