From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S932523AbbJINA5 (ORCPT ); Fri, 9 Oct 2015 09:00:57 -0400 Received: from mailout3.samsung.com ([203.254.224.33]:56671 "EHLO mailout3.samsung.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S932251AbbJINAn (ORCPT ); Fri, 9 Oct 2015 09:00:43 -0400 X-AuditID: cbfee68f-f796f6d0000014a4-79-5617ba787920 From: PINTU KUMAR To: "'Michal Hocko'" Cc: akpm@linux-foundation.org, minchan@kernel.org, dave@stgolabs.net, koct9i@gmail.com, rientjes@google.com, hannes@cmpxchg.org, penguin-kernel@i-love.sakura.ne.jp, bywxiaobai@163.com, mgorman@suse.de, vbabka@suse.cz, js1304@gmail.com, kirill.shutemov@linux.intel.com, alexander.h.duyck@redhat.com, sasha.levin@oracle.com, cl@linux.com, fengguang.wu@intel.com, linux-kernel@vger.kernel.org, linux-mm@kvack.org, cpgs@samsung.com, pintu_agarwal@yahoo.com, pintu.ping@gmail.com, vishnu.ps@samsung.com, rohit.kr@samsung.com, c.rajkumar@samsung.com, sreenathd@samsung.com References: <1443696523-27262-1-git-send-email-pintu.k@samsung.com> <20151001133843.GG24077@dhcp22.suse.cz> <010401d0ff34$f48e8eb0$ddabac10$@samsung.com> <20151005122258.GA7023@dhcp22.suse.cz> <014e01d10004$c45bba30$4d132e90$@samsung.com> <20151006154152.GC20600@dhcp22.suse.cz> <023601d1010f$787696b0$6963c410$@samsung.com> <20151008141851.GD426@dhcp22.suse.cz> <032501d101e3$82588ba0$8709a2e0$@samsung.com> <20151008163049.GJ426@dhcp22.suse.cz> In-reply-to: <20151008163049.GJ426@dhcp22.suse.cz> Subject: RE: [PATCH 1/1] mm: vmstat: Add OOM kill count in vmstat counter Date: Fri, 09 Oct 2015 18:29:49 +0530 Message-id: <03ad01d10292$a45c53d0$ed14fb70$@samsung.com> MIME-version: 1.0 Content-type: text/plain; charset=US-ASCII Content-transfer-encoding: 7bit X-Mailer: Microsoft Outlook 14.0 Thread-index: AQHFyC/Uoy22OwXqBhG9QnqJ8jAPkwHWGZsAAeikRuMArEHMsgIq7xL1AZSJcYQCrtP9xgIqnS+7AoRz86MCrN7LU53nd8UQ Content-language: en-us X-Brightmail-Tracker: H4sIAAAAAAAAA02Sa0hTYRjHeXfec86UhJNlvk3TWFYodPHaK12Y9cETEXSjyC+59GAXnWtH LfFDmqWoMU1X1nHFwgtmi+EMK7O0WXYzM7SWliNTJ6ZTC1oXzdppBX778ef/PP//A4+U8O4m ZdLDqjROo1ImyylPaFwYkbXqRJPv3rUDpvVYbzJSeGa2HGKtfiUu0pVJsNU5DvCoJRhf7zUC PDliIvB183Z86d5lEveO6CG+VvCBxN1NegrbjL9JXDZhB3jMWUPgmq+TNK44PUFi6+gFiJ2O dzTOq6qX4JHcMxBXPnxH4KLBXIgrcrQA67T9QCFjK2t1JNs2Pkmwd4R+mjWY09k7N+okbGXz qIQ11xVQrPlLKc0+uTgN2aGecgl75elO9vNwH2Qn77+mWO3NOsCerzjJdhge0jsWxHluSOSS D2dwmjWb4j0PNQpaqL676kT7+BCRDZzLCoFUipgIpDsVUwg8XLgIddlMlMjeTC1Ave/Rf0tn 93G3LABkP5tYCDxd7ACorX5IInooZiVqb/USPQuZFSi76RUtegimHaLhjlzaPdxFoJIOhej3 YMLQORsvyguYrahq+iUQZcgsR7qfyaLsxUSjW1WdlJvno+9lNigywYSghsZTpJsDUYPRQbjb L0W3X4wBdwU1qv5koNweX1T6YeBvHcS0eaBW682/A5BhkLPMAt0nLkHm1n97FqMHtW9hCUDC nGhhTrQwJ1qYE2EAsA74cOoENX8wSRO+mlem8OmqpNUJqSlm4Pq157P24tugv3W9BTBSIJ/n heN893qTygw+M8UCIl2NzhEyn4RU13uq0g6EhkeF4ciIyPCwddFRcl+vq7Ifu72ZJGUad5Tj 1JzmgCY9meMtQCL1kGWDgBsKR8ejmAG9SZgNeXysodj/WfeG/Ys3dQ7uih+eeDPoDJZL6ax9 JTMhjnuNgap8uTVoc7PfVGxfTY1PT3SOHzNO9tVntOXNTFn9A4WBkj0bp1vWZTU8rc3nlc3G WFuBISso6ldXasWWzIChI9qlLWtphf3YttJvH+N+jiqqB+WQP6QMDSE0vPIPWYVpjWYDAAA= X-Brightmail-Tracker: H4sIAAAAAAAAA2WSe2xLYRjGfT2np92onNVlXzqijpi4lNZWvrmFEI64za0RCXPWHt1oz5qe FSOYjcVGGmyMbmNsFqYsKxnqEoq5zdzNzGbW6TI2i9siI5seDVn4/vrlyfM8ed8vrxSTHyMU 0ngukbVyjIkigvH7nd4Q1QZ3qE792qNGuSVOAv3szMaRPXcY2pWVKUJV7S0ANXuGo1PVToDa mkowdMo1Dx26kidG1U25ODqZXi9GT925BKpzdolR5kcfQB/aizBU9K1NgnK2fxSjquYDOGpv rZGgtMJSEWpK3YGjgps1GNrlTcVRzjY7QFn2WjBVQRecyBLTN1raMPqio1ZC57ts9MXTxSK6 4HKziHYVpxO06/M+CX3n4A+cbnyWLaIP311If3r3Cqfbrj4naPu5YkDvz9lKV+TflET3WZ4M JsWxjIG1KllOn2CI54yTqTmLY6bHaMepNSpNFBpPKTnGzE6mZsyNVs2MN/k/hlKuY0w2vxTN 8Dw1Zsr/DUuXzFShP8FlqnlLFv3NjFX/81Y5QVyZw45bLqk2lLc0YsmgfUgGkEohGQkrn67P AEF+7A8f1ZUQAstJB4C+3YYMEOznVgBvlDaKBD9BDoPl12SCpy8ZDpPdjyWCByPLcfiuIlUS CD/C4J6KqYI/iBwL99bxgtyHnA0LfzwEgoyTQ2FWh0mQZWQUPF9YSQQ4BH7PrMMFxsgR8GxZ ijjAg+BZZysWGFMJLzz4AAIjWODx9/lEwBMK99W/lewBcke3Kke3Kke3Kke3SD7AiwFkLXoL H2s0azh2/WieMfM2zjhan2B2gd/36VNcANfdszyAlAKqlwwtD9XJxcw6PsnsAVCKUX1lT7b4 JZmBSdrIWhNirDYTy3uA1r/rXkzRT5/gv3YuMUYTETleG6GNikSR46KoUNlLTw+dnDQyiexa lrWw1j85kTRIkQySvl7/2lD23XtlW6dc5n2FVjZUvTG8iXV3gNouw0Td7bCj0ZdGpXxZXRpL kIaGFPOasEPumpdHhoakzwnJ7q2/N8Gb9qlJ2RjxeWf5lp5aR9Hx87oF4el3Ntk6mJEd09Jm 6CrX+sJerJqftqK63ubbfOZxUdfAAXnGja+5wSm31BODKZyPYzQjMCvP/AIvI8uutQMAAA== DLP-Filter: Pass X-MTR: 20000000000000000@CPGS X-CFilter-Loop: Reflected Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org > -----Original Message----- > From: Michal Hocko [mailto:mhocko@kernel.org] > Sent: Thursday, October 08, 2015 10:01 PM > To: PINTU KUMAR > Cc: akpm@linux-foundation.org; minchan@kernel.org; dave@stgolabs.net; > koct9i@gmail.com; rientjes@google.com; hannes@cmpxchg.org; penguin- > kernel@i-love.sakura.ne.jp; bywxiaobai@163.com; mgorman@suse.de; > vbabka@suse.cz; js1304@gmail.com; kirill.shutemov@linux.intel.com; > alexander.h.duyck@redhat.com; sasha.levin@oracle.com; cl@linux.com; > fengguang.wu@intel.com; linux-kernel@vger.kernel.org; linux-mm@kvack.org; > cpgs@samsung.com; pintu_agarwal@yahoo.com; pintu.ping@gmail.com; > vishnu.ps@samsung.com; rohit.kr@samsung.com; c.rajkumar@samsung.com; > sreenathd@samsung.com > Subject: Re: [PATCH 1/1] mm: vmstat: Add OOM kill count in vmstat counter > > On Thu 08-10-15 21:36:24, PINTU KUMAR wrote: > [...] > > Whereas, these OOM logs were not found in /var/log/messages. > > May be we do heavy logging because in ageing test we enable maximum > > functionality (Wifi, BT, GPS, fully loaded system). > > If you swamp your logs so heavily that even critical messages won't make it into > the log files then your logging is basically useless for anything serious. But that is > not really that important. > > > Hope, it is clear now. If not, please ask me for more information. > > > > > > > > > Now, every time this dumping is not feasible. And instead of > > > > counting manually in log file, we wanted to know number of oom > > > > kills happened during > > > this tests. > > > > So we decided to add a counter in /proc/vmstat to track the kernel > > > > oom_kill, and monitor it during our ageing test. > > > > > > > > Basically, we wanted to tune our user space LMK killer for > > > > different threshold values, so that we can completely avoid the kernel oom > kill. > > > > So, just by looking into this counter, we could able to tune the > > > > LMK threshold values without depending on the kernel log messages. > > > > > > Wouldn't a trace point suit you better for this particular use case > > > considering this is a testing environment? > > > > > Tracing for oom_kill count? > > Actually, tracing related configs will be normally disabled in release binary. > > Yes but your use case described a testing environment. > > > And it is not always feasible to perform tracing for such long duration tests. > > I do not see why long duration would be a problem. Each tracepoint can be > enabled separatelly. > > > Then it should be valid for other counters as well. > > > > > > Also, in most of the system /var/log/messages are not present and > > > > we just depends on kernel dmesg output, which is petty small for longer > run. > > > > Even if we reduce the loglevel to 4, it may not be suitable to > > > > capture all > > logs. > > > > > > Hmm, I would consider a logless system considerably crippled but I > > > see your point and I can imagine that especially small devices might > > > try to save every single B of the storage. Such a system is > > > basically undebugable IMO but it > > still > > > might be interesting to see OOM killer traces. > > > > > Exactly, some of the small embedded systems might be having 512MB, > > 256MB, 128MB, or even lesser. > > Also, the storage space will be 8GB or below. > > In such a system we cannot afford heavy log files and exact tuning and > > stability is most important. > > And that is what log level is for. If your logs are heavy with error levels then you > are far from being production ready... ;) > > > Even all tracing / profiling configs will be disabled to lowest level > > for reducing kernel code size as well. > > What level is that? crit? Is err really that noisy? > No. I was talking about kernel configs. Normally we keep some profiling/tracing related configs disabled for low memory system, to save some kernel code size. The point is that it's always not easy for all systems to heavily depends on logging and tracing. Else, the other counters would also not be required. We thought that the /proc/vmstat output (which is ideally available in all systems, small or big, embedded or none embedded), it can quickly tell us what has happened really. > [...] > > > > Ok, you are suggesting to divide the oom_kill counter into 2 parts > > > > (global & > > > > memcg) ? > > > > May be something like: > > > > nr_oom_victims > > > > nr_memcg_oom_victims > > > > > > You do not need the later. Memcg interface already provides you with > > > a notification API and if a counter is _really_ needed then it > > > should be per-memcg not a global cumulative number. > > > > Ok, for memory cgroups, you mean to say this one? > > sh-3.2# cat /sys/fs/cgroup/memory/memory.oom_control > > oom_kill_disable 0 > > under_oom 0 > > Yes this is the notification API. > > > I am actually confused here what to do next? > > Shall I push a new patch set with just: > > nr_oom_victims counter ? > > Yes you can repost with a better description about a typical usage scenarios. I > cannot say I would be completely sold to this because the only relevant usecase > I've heard so far is the logless system which is pretty much a corner case. This is > not a reason to nack it though. It is definitely better than the original oom_stall > suggestion because it has a clear semantic at least. Ok, thank you very much for your suggestions. I agree, oom_stall is not so important. I will try to submit a new patch set with only _nr_oom_victims_ with the descriptions about the usefulness that I came across. If anybody else can point out other use cases, please let me know. I will be happy to try that and share the results. > -- > Michal Hocko > SUSE Labs