mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* problem with mmap over nfs
@ 2005-05-06  9:50 Fabio Brugnara
  2005-05-06 11:54 ` Andrew Morton
  0 siblings, 1 reply; 4+ messages in thread
From: Fabio Brugnara @ 2005-05-06  9:50 UTC (permalink / raw)
  To: linux-kernel

Hi all,

	I hope I will be able to find some advice for solving a problem I
observed from going to linux 2.4 to linux 2.6. I'm not a kernel expert, and
was not able to get help from those I know. I tried to search the archives
for something related, but was also unsuccessful. So I try here hoping that
someone will have the patience to read the message and maybe give some
help.

	If anyone answers, please post also the message in CC directly to
me, as I'm not subscribed to the list (it doesn't make much sense at my
level of expertise).

	The problem is related to the use of memory mapped files over a nfs
mounted filesystem. In a rather complex system (speech recognition) that we
developed and use, we need to share large read-only data structures between
different processes, also on several machines. Until a few months ago, we
observed that it was perfectly adequate to just mmap() a file residing on a
particular disk, that machines other as the owner mounted via nfs. This was
very convenient, as we had only a single physical copy of the data
structure, and it did not introduce any significant performance penalty. We
have used this method for years.
	Now, after the machines have been upgraded to kernel 2.6.10 from
kernel 2.4.20, something disappointing happens. Everything still works
correctly, but somehow it introduces a massive slowdown of the machines.
While the processes are running, the machines that map the shared file via
nfs (not the one that owns it) report (with "top") a very high usage of
system time (e.g. 50% or more), and also become very unresponsive at the
shell prompt.
	Another strange thing that can be seen is that in this situation
the "top" process itself, that normally reports something like 0.3% CPU
time), starts using percentages of 30% CPU time or more.
	Beside observing the behavior with "top", we also measure the time
processes take in the program itself, using the "times()" primitive at
the end of the processing. A typical short run of the system on the old
kernel reported numbers such as:

Times: user: 1125.50, system: 6.28

that is, with a proportion always much lower that 1% between system time
and user time. Now, we see things like:

Times: user: 660.29, system: 326.64

i.e., a proportion of 50% between the two (Absolute values cannot
be compared, as the hardware is different).
Typically we run two of these processes on each machine, which
are 2-CPU SMP Xeon systems with 2 or 4 Gb of memory.

	Notice that, if we take care to copy the shared file on local
disks on each machine, and perform the mmap() on the local copy for each
of them, the problem disappears, i.e. the machine do not become overloaded,
and all timings are perfectly reasonable as before. So it's not a problem
of mmap() in itself, but something that comes from the combination of
mmap() and nfs, and maybe sychronization in the access of shared pages
on an SMP system.
	I'm aware that I should present some simple program that triggers
the misbehavior, but unfortunately I was not able to reproduce the problem
in a simple context. The system in question is rather complex,
computationally expensive and makes an intensive use both of these static
shared structures, and of large amounts of dynamically allocated memory
(process size is typically of several hundreds Mb, and CPU usage as high
as possible).

If all this story flashes some lamp in the knowledgeable heads that
populate this list, please let me know

best regards,
thank you for your patience
Fabio

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: problem with mmap over nfs
  2005-05-06  9:50 problem with mmap over nfs Fabio Brugnara
@ 2005-05-06 11:54 ` Andrew Morton
  2005-05-06 12:20   ` Fabio Brugnara
  2005-05-12  8:00   ` Fabio Brugnara
  0 siblings, 2 replies; 4+ messages in thread
From: Andrew Morton @ 2005-05-06 11:54 UTC (permalink / raw)
  To: Fabio Brugnara; +Cc: linux-kernel

Fabio Brugnara <brugnara@itc.it> wrote:
>
> 	The problem is related to the use of memory mapped files over a nfs
>  mounted filesystem. In a rather complex system (speech recognition) that we
>  developed and use, we need to share large read-only data structures between
>  different processes, also on several machines. Until a few months ago, we
>  observed that it was perfectly adequate to just mmap() a file residing on a
>  particular disk, that machines other as the owner mounted via nfs. This was
>  very convenient, as we had only a single physical copy of the data
>  structure, and it did not introduce any significant performance penalty. We
>  have used this method for years.
>  	Now, after the machines have been upgraded to kernel 2.6.10 from
>  kernel 2.4.20, something disappointing happens. Everything still works
>  correctly, but somehow it introduces a massive slowdown of the machines.
>  While the processes are running, the machines that map the shared file via
>  nfs (not the one that owns it) report (with "top") a very high usage of
>  system time (e.g. 50% or more), and also become very unresponsive at the
>  shell prompt.

Could you please generate a kernel profile?

- Compile with CONFIG_PROFILING

- Start the workload, wait for steady state.

- As root, run:

#!/bin/sh

SM=/boot/System.map
TIMEFILE=/tmp/prof.time
readprofile -r
sleep 10
readprofile -n -v -m $SM | sort -n +2 | tail -40 | tee $TIMEFILE >&2

(make sure that /boot/System.map is from the currently-running kernel)

More in Documentation/basic_profiling.txt



Even better, learn to drive oprofile.  Once it's running properly I usually
use this silly script:

#!/bin/sh
opcontrol --stop
opcontrol --shutdown
rm -rf /var/lib/oprofile
opcontrol --vmlinux=/boot/vmlinux-$(uname -r)
opcontrol --start-daemon
opcontrol --start
sleep 10
opcontrol --stop
opcontrol --shutdown
opreport -l /boot/vmlinux-$(uname -r) | head -50


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: problem with mmap over nfs
  2005-05-06 11:54 ` Andrew Morton
@ 2005-05-06 12:20   ` Fabio Brugnara
  2005-05-12  8:00   ` Fabio Brugnara
  1 sibling, 0 replies; 4+ messages in thread
From: Fabio Brugnara @ 2005-05-06 12:20 UTC (permalink / raw)
  To: Andrew Morton; +Cc: Fabio Brugnara, linux-kernel

> Could you please generate a kernel profile?
> 
> - Compile with CONFIG_PROFILING
> 
> - Start the workload, wait for steady state.
> 
> - As root, run:
> 
> #!/bin/sh
> 
> SM=/boot/System.map
> TIMEFILE=/tmp/prof.time
> readprofile -r
> sleep 10
> readprofile -n -v -m $SM | sort -n +2 | tail -40 | tee $TIMEFILE >&2
> 
> (make sure that /boot/System.map is from the currently-running kernel)
> 
> More in Documentation/basic_profiling.txt
> 
> 
> 
> Even better, learn to drive oprofile.  Once it's running properly I usually
> use this silly script:
> 
> #!/bin/sh
> opcontrol --stop
> opcontrol --shutdown
> rm -rf /var/lib/oprofile
> opcontrol --vmlinux=/boot/vmlinux-$(uname -r)
> opcontrol --start-daemon
> opcontrol --start
> sleep 10
> opcontrol --stop
> opcontrol --shutdown
> opreport -l /boot/vmlinux-$(uname -r) | head -50


Wow! Thank you, Andrew.
If we have to consider me alone, this sounds like a dentist
that says "please extract your bad teeth and send it for the fix".
But I hope I'm able to do it with the help of our system administrators.

best regards,
thank you again for your attention,
Fabio


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: problem with mmap over nfs
  2005-05-06 11:54 ` Andrew Morton
  2005-05-06 12:20   ` Fabio Brugnara
@ 2005-05-12  8:00   ` Fabio Brugnara
  1 sibling, 0 replies; 4+ messages in thread
From: Fabio Brugnara @ 2005-05-12  8:00 UTC (permalink / raw)
  To: Andrew Morton; +Cc: Fabio Brugnara, linux-kernel

On Fri, May 06, 2005 at 04:54:46AM -0700, Andrew Morton wrote:
>
> Could you please generate a kernel profile?
>
> - Compile with CONFIG_PROFILING
>
> - Start the workload, wait for steady state.
>
> - As root, run:
>
> #!/bin/sh
>
> SM=/boot/System.map
> TIMEFILE=/tmp/prof.time
> readprofile -r
> sleep 10
> readprofile -n -v -m $SM | sort -n +2 | tail -40 | tee $TIMEFILE >&2
>
> (make sure that /boot/System.map is from the currently-running kernel)
>
> More in Documentation/basic_profiling.txt
>
>
>
> Even better, learn to drive oprofile.  Once it's running properly I usually
> use this silly script:
>
> #!/bin/sh
> opcontrol --stop
> opcontrol --shutdown
> rm -rf /var/lib/oprofile
> opcontrol --vmlinux=/boot/vmlinux-$(uname -r)
> opcontrol --start-daemon
> opcontrol --start
> sleep 10
> opcontrol --stop
> opcontrol --shutdown
> opreport -l /boot/vmlinux-$(uname -r) | head -50

Hi Andrew,

The short story:

Alarm is over. Everything works perfectly with 2.6.11.

The long story:

I could not directly do what was suggested, so I went for help to one of
our sysadmin. He installed oprofile and everything, but discovered that we
could not use it, because we didn't have the vmlinux image of the running
kernel, only vmlinuz.  After trying to recover vmlinux from vmlinuz and
discovering that it's not possible, he decided to recompile the kernel.
But, since now 2.6.11 is released, he decided also to try the latest
version.  Well, when running under 2.6.11 the strange phenomenon of
excessive system usage does not appear anymore. What we observed was
therefore just another symptom of some bug (who knows what ... ) that was
already known and fixed.

best regards,
thank you again for your attention,
Fabio

PS: just to be precise:

the kernel that had the problem was:

$ uname -r
2.6.10-1.741_FC3smp

the kernel that is now working well is:

$ uname -r
2.6.11-1.14.RH9smp


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2005-05-12  7:57 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2005-05-06  9:50 problem with mmap over nfs Fabio Brugnara
2005-05-06 11:54 ` Andrew Morton
2005-05-06 12:20   ` Fabio Brugnara
2005-05-12  8:00   ` Fabio Brugnara

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®