mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* Re: NUMA allocator on Opteron systems does non-local allocation on node0
@ 2008-10-13 10:46 Andi Kleen
  2008-10-13 15:27 ` Oliver Weihe
  0 siblings, 1 reply; 3+ messages in thread
From: Andi Kleen @ 2008-10-13 10:46 UTC (permalink / raw)
  To: linux-kernel; +Cc: o.weihe

[Another copy of the reply with linux-kernel added this time]

> In my setup I'm allocating an array of ~7GiB memory size in a
> singlethreaded application.
> Startup: numactl --cpunodebind=X ./app
> For X=1,2,3 it works as expected, all memory is allocated on the local
> node.
> For X=0 I can see the memory beeing allocated on node0 as long as ~3GiB
> are "free" on node0. At this point the kernel starts using memory from
> node1 for the app!

Hmm, that sounds like it doesn't want to use the 4GB DMA zone.

Normally there should be no protection on it, but perhaps something 
broke.

What does cat /proc/sys/owmem_reserve_ratio say?

> 
> For parallel realworld apps I've seen a performance penalty of 30%
> compared to older kernels!

Compared to what older kernels? When did it start?

-Andi

-- 
ak@linux.intel.com

^ permalink raw reply	[flat|nested] 3+ messages in thread
* NUMA allocator on Opteron systems does non-local allocation on node0
@ 2008-10-13 10:15 Oliver Weihe
  0 siblings, 0 replies; 3+ messages in thread
From: Oliver Weihe @ 2008-10-13 10:15 UTC (permalink / raw)
  To: Andi Kleen; +Cc: linux-kernel

Hi Andi,

I'm not sure if you're the right person for this but I hope you are!

I've notived that the memory allocation on NUMA systems (Opterons) does
memory allocation on non-local nodes for processes running node0 even if
local memory is available. (Kernel 2.6.25 and above)

Currently I'm playing around with a quadsocket quadcore Opteron but I've
observed this behavior on other Opteron systems aswell.

Hardware specs:
1x Supermicro H8QM3-2
4x Quadcore Opteron
16x 2GiB (8 GiB memory per node)

OS:
currently openSUSE 10.3 but I've observed this on other distros aswell
Kernel: 2.6.22.* (openSUSE) / 2.6.25.4 / 2.6.25.5 / 2.6.27 (vanilla
config)

Steps to reproduce:
Start an application which needs alot of memory and watch the memory
usage per node (I'm using "watch -n 1 numastat --hardware" to watch the
memory usage per node)
A quick&dirty code which allocates a big array and writes data into the
array is enough!

In my setup I'm allocating an array of ~7GiB memory size in a
singlethreaded application.
Startup: numactl --cpunodebind=X ./app
For X=1,2,3 it works as expected, all memory is allocated on the local
node.
For X=0 I can see the memory beeing allocated on node0 as long as ~3GiB
are "free" on node0. At this point the kernel starts using memory from
node1 for the app!

For parallel realworld apps I've seen a performance penalty of 30%
compared to older kernels!

numactl --cpunodebind=0 --membind=0 ./app "solves" the problem in this
case but thats not the point!

-- 

Regards,
Oliver Weihe


^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2008-10-13 15:27 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2008-10-13 10:46 NUMA allocator on Opteron systems does non-local allocation on node0 Andi Kleen
2008-10-13 15:27 ` Oliver Weihe
  -- strict thread matches above, loose matches on Subject: below --
2008-10-13 10:15 Oliver Weihe

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®