* Lmbench performance drop 2.6.18-->2.6.27
@ 2009-11-19 21:13 Ajay Patel
2009-11-20 10:20 ` Mike Galbraith
0 siblings, 1 reply; 2+ messages in thread
From: Ajay Patel @ 2009-11-19 21:13 UTC (permalink / raw)
To: linux-kernel
[-- Attachment #1: Type: text/plain, Size: 616 bytes --]
Hi all,
Part of our evaluation to upgrade kernel we
ran lmbench. The lmbench results shows significant performance
drop from 2.6.18 to 2.6.27. (Results attached)
The benchmark was performed on same hardware with
different distro. (Quad-Core AMD Opteron(tm) Processor 2346 HE,
cpu MHz : 1795.597, cache size : 512 KB, x86_64 kernel)
The 2.6.18 based distribution was from CentOS release 5.4.
(Linux cento-5.4 2.6.18-164.el5).
The 2.6.27 based distribution was from FC10.
(Fedora Core release 10 2.6.27.5-117.fc10-x86_64).
Does this results make sense? Is this expected?
Am I doing something wrong?
Thanks
Ajay
[-- Attachment #2: lmbench-amd-distro.txt --]
[-- Type: text/plain, Size: 8445 bytes --]
L M B E N C H 3 . 0 S U M M A R Y
------------------------------------
(Alpha software, do not distribute)
Basic system parameters
------------------------------------------------------------------------------
Host OS Description Mhz tlb cache mem scal
pages line par load
bytes
--------- ------------- ----------------------- ---- ----- ----- ------ ----
cento-5.4 Linux 2.6.18- x86_64-linux-gnu 1797 48 64 1.0200 1
cento-5.4 Linux 2.6.18- x86_64-linux-gnu 1797 48 64 1.0200 1
cento-5.4 Linux 2.6.18- x86_64-linux-gnu 1797 48 64 1.0200 1
fedora-10 Linux 2.6.27. x86_64-linux-gnu 1791 48 64 5.4100 1
fedora-10 Linux 2.6.27. x86_64-linux-gnu 1791 48 64 5.4400 1
fedora-10 Linux 2.6.27. x86_64-linux-gnu 1791 48 64 5.4400 1
Processor, Processes - times in microseconds - smaller is better
------------------------------------------------------------------------------
Host OS Mhz null null open slct sig sig fork exec sh
call I/O stat clos TCP inst hndl proc proc proc
--------- ------------- ---- ---- ---- ---- ---- ---- ---- ---- ---- ---- ----
cento-5.4 Linux 2.6.18- 1797 0.24 0.33 2.31 3.01 5.14 0.36 1.36 195. 611. 2444
cento-5.4 Linux 2.6.18- 1797 0.24 0.34 2.17 3.21 5.27 0.36 1.46 186. 616. 2466
cento-5.4 Linux 2.6.18- 1797 0.24 0.33 2.26 3.25 5.08 0.38 1.36 195. 595. 2294
fedora-10 Linux 2.6.27. 1791 0.24 0.39 3.60 5.08 4.37 0.37 1.66 228. 847. 3077
fedora-10 Linux 2.6.27. 1791 0.24 0.39 3.40 5.06 4.38 0.37 1.64 227. 843. 2977
fedora-10 Linux 2.6.27. 1791 0.24 0.40 3.50 4.98 4.33 0.37 1.61 218. 841. 2996
Basic integer operations - times in nanoseconds - smaller is better
-------------------------------------------------------------------
Host OS intgr intgr intgr intgr intgr
bit add mul div mod
--------- ------------- ------ ------ ------ ------ ------
cento-5.4 Linux 2.6.18- 0.5700 0.5600 0.2000 28.7 16.4
cento-5.4 Linux 2.6.18- 0.5600 0.5600 0.2000 28.5 16.3
cento-5.4 Linux 2.6.18- 0.5700 0.5600 0.2000 28.7 16.3
fedora-10 Linux 2.6.27. 0.5700 0.5700 0.2300 28.8 16.4
fedora-10 Linux 2.6.27. 0.5700 0.5600 0.2300 28.8 16.4
fedora-10 Linux 2.6.27. 0.5700 0.5600 0.2300 28.8 16.4
Basic uint64 operations - times in nanoseconds - smaller is better
------------------------------------------------------------------
Host OS int64 int64 int64 int64 int64
bit add mul div mod
--------- ------------- ------ ------ ------ ------ ------
cento-5.4 Linux 2.6.18- 0.570 0.2500 43.4 45.1
cento-5.4 Linux 2.6.18- 0.560 0.2500 43.4 44.8
cento-5.4 Linux 2.6.18- 0.560 0.2500 43.4 45.0
fedora-10 Linux 2.6.27. 0.570 0.2500 43.5 45.2
fedora-10 Linux 2.6.27. 0.570 0.2500 43.5 45.2
fedora-10 Linux 2.6.27. 0.570 0.2500 43.5 45.2
Basic float operations - times in nanoseconds - smaller is better
-----------------------------------------------------------------
Host OS float float float float
add mul div bogo
--------- ------------- ------ ------ ------ ------
cento-5.4 Linux 2.6.18- 2.2500 2.2800 10.6 7.3500
cento-5.4 Linux 2.6.18- 2.2500 2.2700 10.5 7.3300
cento-5.4 Linux 2.6.18- 2.2500 2.2800 10.5 7.3400
fedora-10 Linux 2.6.27. 2.2600 2.2900 10.6 7.3600
fedora-10 Linux 2.6.27. 2.2600 2.2900 10.6 7.3600
fedora-10 Linux 2.6.27. 2.2600 2.2900 10.6 7.3700
Basic double operations - times in nanoseconds - smaller is better
------------------------------------------------------------------
Host OS double double double double
add mul div bogo
--------- ------------- ------ ------ ------ ------
cento-5.4 Linux 2.6.18- 2.2500 2.2800 12.8 9.6100
cento-5.4 Linux 2.6.18- 2.2400 2.2800 12.7 9.5300
cento-5.4 Linux 2.6.18- 2.2500 2.2800 12.8 9.5900
fedora-10 Linux 2.6.27. 2.2600 2.2900 12.8 9.6200
fedora-10 Linux 2.6.27. 2.2600 2.2900 12.8 9.6200
fedora-10 Linux 2.6.27. 2.2600 2.2900 12.8 9.6200
Context switching - times in microseconds - smaller is better
-------------------------------------------------------------------------
Host OS 2p/0K 2p/16K 2p/64K 8p/16K 8p/64K 16p/16K 16p/64K
ctxsw ctxsw ctxsw ctxsw ctxsw ctxsw ctxsw
--------- ------------- ------ ------ ------ ------ ------ ------- -------
cento-5.4 Linux 2.6.18- 0.9500 4.7300 8.3600 4.3700 8.2400 5.13000 12.6
cento-5.4 Linux 2.6.18- 0.8700 3.0300 5.3900 4.9400 8.4100 4.82000 12.8
cento-5.4 Linux 2.6.18- 0.9400 3.7100 7.1400 4.3200 7.5600 4.66000 10.7
fedora-10 Linux 2.6.27. 7.5300 2.5800 11.2 5.9300 9.9500 5.10000 15.0
fedora-10 Linux 2.6.27. 7.2400 2.4400 3.4600 6.1300 8.9000 7.19000 13.1
fedora-10 Linux 2.6.27. 7.4200 7.4500 11.4 4.8600 9.6700 5.43000 9.42000
*Local* Communication latencies in microseconds - smaller is better
---------------------------------------------------------------------
Host OS 2p/0K Pipe AF UDP RPC/ TCP RPC/ TCP
ctxsw UNIX UDP TCP conn
--------- ------------- ----- ----- ---- ----- ----- ----- ----- ----
cento-5.4 Linux 2.6.18- 0.950 6.480 6.87 20.0 21.1 19.3 23.8 27.
cento-5.4 Linux 2.6.18- 0.870 9.939 9.61 20.3 24.0 18.8 26.8 27.
cento-5.4 Linux 2.6.18- 0.940 11.7 6.96 20.4 21.3 21.7 25.0 29.
fedora-10 Linux 2.6.27. 7.530 7.751 9.56 34.6 39.5 36.1 47.6 59.
fedora-10 Linux 2.6.27. 7.240 7.733 9.16 34.5 24.6 41.0 32.5 110.
fedora-10 Linux 2.6.27. 7.420 7.648 9.35 18.2 39.9 40.3 47.9 58.
File & VM system latencies in microseconds - smaller is better
-------------------------------------------------------------------------------
Host OS 0K File 10K File Mmap Prot Page 100fd
Create Delete Create Delete Latency Fault Fault selct
--------- ------------- ------ ------ ------ ------ ------- ----- ------- -----
cento-5.4 Linux 2.6.18- 15.8 9.5580 47.9 21.2 5150.0 0.416 1.60530 2.484
cento-5.4 Linux 2.6.18- 15.7 9.8031 48.5 21.4 4916.0 0.281 1.65810 2.503
cento-5.4 Linux 2.6.18- 15.9 9.4592 48.9 21.1 5139.0 0.394 1.47310 2.435
fedora-10 Linux 2.6.27. 61.5 13.9 103.5 28.1 6162.0 0.477 1.66810 2.297
fedora-10 Linux 2.6.27. 63.5 13.9 111.6 29.0 6103.0 0.384 1.75870 2.305
fedora-10 Linux 2.6.27. 62.2 14.1 110.6 29.2 6344.0 0.459 1.79540 2.297
*Local* Communication bandwidths in MB/s - bigger is better
-----------------------------------------------------------------------------
Host OS Pipe AF TCP File Mmap Bcopy Bcopy Mem Mem
UNIX reread reread (libc) (hand) read write
--------- ------------- ---- ---- ---- ------ ------ ------ ------ ---- -----
cento-5.4 Linux 2.6.18- 1887 1202 946. 1678.9 2952.1 1473.7 1400.5 2425 1409.
cento-5.4 Linux 2.6.18- 1337 1195 940. 1700.9 3027.9 1477.5 1400.7 2423 1368.
cento-5.4 Linux 2.6.18- 781. 1098 864. 1715.8 3011.9 1202.8 1408.5 2190 1410.
fedora-10 Linux 2.6.27. 1604 1181 923. 1541.7 2666.6 1447.8 1342.4 2429 1408.
fedora-10 Linux 2.6.27. 1592 1782 925. 1590.7 2886.1 1432.6 1391.9 2430 1401.
fedora-10 Linux 2.6.27. 1606 1170 926. 1550.8 2640.2 1432.3 1354.4 2429 1406.
Memory latencies in nanoseconds - smaller is better
(WARNING - may not be correct, check graphs)
------------------------------------------------------------------------------
Host OS Mhz L1 $ L2 $ Main mem Rand mem Guesses
--------- ------------- --- ---- ---- -------- -------- -------
cento-5.4 Linux 2.6.18- 1797 1.6860 8.6140 99.3 159.2
cento-5.4 Linux 2.6.18- 1797 1.6870 8.6210 99.3 155.0
cento-5.4 Linux 2.6.18- 1797 1.6810 8.6140 106.0 155.9
fedora-10 Linux 2.6.27. 1791 1.6920 8.6470 99.5 153.2
fedora-10 Linux 2.6.27. 1791 1.6920 8.6500 99.5 153.5
fedora-10 Linux 2.6.27. 1791 1.6920 8.6500 99.5 154.4
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: Lmbench performance drop 2.6.18-->2.6.27
2009-11-19 21:13 Lmbench performance drop 2.6.18-->2.6.27 Ajay Patel
@ 2009-11-20 10:20 ` Mike Galbraith
0 siblings, 0 replies; 2+ messages in thread
From: Mike Galbraith @ 2009-11-20 10:20 UTC (permalink / raw)
To: Ajay Patel; +Cc: linux-kernel
On Thu, 2009-11-19 at 13:13 -0800, Ajay Patel wrote:
> Hi all,
>
> Part of our evaluation to upgrade kernel we
> ran lmbench. The lmbench results shows significant performance
> drop from 2.6.18 to 2.6.27. (Results attached)
>
> The benchmark was performed on same hardware with
> different distro. (Quad-Core AMD Opteron(tm) Processor 2346 HE,
> cpu MHz : 1795.597, cache size : 512 KB, x86_64 kernel)
>
> The 2.6.18 based distribution was from CentOS release 5.4.
> (Linux cento-5.4 2.6.18-164.el5).
> The 2.6.27 based distribution was from FC10.
> (Fedora Core release 10 2.6.27.5-117.fc10-x86_64).
>
> Does this results make sense? Is this expected?
> Am I doing something wrong?
I think most of what you're seeing is config differences and the fact
that microbenchmarks are excellent at showing how horribly expensive
cache misses are. You see radically different numbers when two halves
of a microbenchmark land in the same cache vs landing across the fence
from one another.
Sometimes affinity is great, sometimes not so great. For evaluating,
wide spectrum testing is much safer than microbenchmarks, they can be
very misleading. lmbench is really good at showing us that affinity
logic has always been a sore spot, and I'm very sure that's what you're
seeing.
Below are some numbers from my supermarket Q6600 box. Notice the wild
swings when affinity goes wrong/right, same as your numbers. Check out
the UNIX socket numbers in the last three lines. Wake affine to cache
rather than to CPU, and that's what you get for that microbenchmark.
Latency numbers don't look very appetizing, but throughput goes through
the roof.
*Local* Communication latencies in microseconds - smaller is better
---------------------------------------------------------------------
Host OS 2p/0K Pipe AF UDP RPC/ TCP RPC/ TCP
ctxsw UNIX UDP TCP conn
--------- ------------- ----- ----- ---- ----- ----- ----- ----- ----
marge Linux 2.6.27. 1.070 3.757 4.87 8.305 12.7 12.9 16.1 34.
marge Linux 2.6.27. 0.880 3.364 5.72 7.950 11.8 12.8 15.3 34.
marge Linux 2.6.27. 0.810 3.458 4.77 7.978 11.7 12.9 15.3 34.
marge Linux 2.6.27. 0.860 3.408 5.80 7.981 11.9 30.4 15.3 34.
marge Linux 2.6.31. 0.820 3.127 5.63 8.060 13.0 29.6 16.6 34.
marge Linux 2.6.31. 0.760 3.119 5.64 8.105 13.1 29.6 16.8 36.
marge Linux 2.6.22. 4.420 10.9 20.7 7.705 10.7 22.5 29.9 27.
marge Linux 2.6.22. 4.460 3.298 5.16 7.648 10.7 22.4 30.0 26.
marge Linux 2.6.32- 0.830 3.069 4.92 8.337 13.0 13.1 17.6 37.
marge Linux 2.6.32- 0.810 3.041 5.72 9.902 13.4 13.2 17.3 36.
marge Linux 2.6.32- 0.800 3.050 5.65 8.312 13.4 13.0 17.2 36.
marge Linux 2.6.32- 0.790 3.113 5.55 9.925 13.4 13.0 17.4 36.
marge Linux 2.6.32- 0.800 3.082 4.78 8.353 13.4 13.0 17.2 36.
marge Linux 2.6.32- 0.780 3.086 5.59 8.332 13.4 13.1 17.2 36.
marge Linux 2.6.32- 1.470 4.758 7.32 10.1 13.4 12.9 16.7 21.
marge Linux 2.6.32- 1.460 4.736 7.60 10.1 13.4 13.0 16.5 21.
marge Linux 2.6.32- 1.290 4.571 7.69 10.1 13.3 12.9 16.4 21.
*Local* Communication bandwidths in MB/s - bigger is better
-----------------------------------------------------------------------------
Host OS Pipe AF TCP File Mmap Bcopy Bcopy Mem Mem
UNIX reread reread (libc) (hand) read write
--------- ------------- ---- ---- ---- ------ ------ ------ ------ ---- -----
marge Linux 2.6.27. 2706 2535 1142 2793.2 4786.3 1285.4 1235.7 4456 1685.
marge Linux 2.6.27. 2737 2811 749. 2773.3 4798.8 1245.3 1235.9 4387 1685.
marge Linux 2.6.27. 2746 2838 2617 2785.0 4763.6 1239.6 1237.8 4423 1688.
marge Linux 2.6.27. 2711 2776 2653 2772.8 4842.4 1404.5 1394.1 4492 1765.
marge Linux 2.6.31. 2759 2847 735. 2808.2 4800.2 1238.4 1231.8 4485 1683.
marge Linux 2.6.31. 2756 2842 1125 2792.1 4788.9 1239.8 1235.0 4475 1681.
marge Linux 2.6.22. 2692 1843 1017 2746.4 4763.6 1293.3 1278.4 4439 1678.
marge Linux 2.6.22. 2865 1872 1015 2776.4 4803.4 1291.7 1320.8 4421 1679.
marge Linux 2.6.32- 2780 2889 2812 2802.0 4833.7 1239.2 1233.6 4510 1683.
marge Linux 2.6.32- 2782 2892 2808 2808.4 4819.8 1240.5 1232.3 4438 1682.
marge Linux 2.6.32- 2786 2877 2810 2824.7 4802.2 1237.9 1234.8 4433 1601.
marge Linux 2.6.32- 2767 2873 1129 2801.7 4803.4 1238.3 1234.6 4467 1688.
marge Linux 2.6.32- 2761 2890 736. 2814.8 4793.0 1245.4 1233.4 4483 1687.
marge Linux 2.6.32- 2748 2883 755. 2805.1 4823.6 1243.0 1233.9 4353 1683.
marge Linux 2.6.32- 1776 5116 2814 2809.7 4759.5 1245.4 1232.3 4451 1681.
marge Linux 2.6.32- 2997 5119 1147 2809.7 4787.3 1242.2 1236.4 4462 1682.
marge Linux 2.6.32- 2999 5122 1138 2808.3 4803.8 1242.7 1236.5 4498 1685.
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2009-11-20 10:20 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2009-11-19 21:13 Lmbench performance drop 2.6.18-->2.6.27 Ajay Patel
2009-11-20 10:20 ` Mike Galbraith
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®