From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1755224AbdCGMfE (ORCPT ); Tue, 7 Mar 2017 07:35:04 -0500 Received: from mailout2.samsung.com ([203.254.224.25]:53679 "EHLO mailout2.samsung.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1754860AbdCGMew (ORCPT ); Tue, 7 Mar 2017 07:34:52 -0500 X-AuditID: b6c32a35-f79d66d000001a37-39-58be975dfab8 Subject: Re: counting file descriptors with a cgroup controller To: Tejun Heo Cc: lizefan@huawei.com, hannes@cmpxchg.org, =?UTF-8?Q?=c5=81ukasz_Stelmach?= , linux-kernel@vger.kernel.org, Karol Lewandowski , cgroups@vger.kernel.org From: Krzysztof Opasiak Message-id: <7fbd9c4c-76ca-4073-9afa-1ab54364ec79@samsung.com> Date: Tue, 07 Mar 2017 12:19:52 +0100 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:45.0) Gecko/20100101 Thunderbird/45.3.0 MIME-version: 1.0 In-reply-to: <20170306185820.GA19696@htj.duckdns.org> Content-type: text/plain; charset=windows-1252; format=flowed Content-transfer-encoding: 7bit X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFupnleLIzCtJLcpLzFFi42LZdljTQDdu+r4Ig3fXGC1uLJ/BYrF6k69F 46e5zBY3D61gtLi8aw6bxaSOXnaLX8uPMjqwexx+857Zo+XIW1aPTas62Tz6tqxi9Pi8SS6A NYrLJiU1J7MstUjfLoEro2/LaaaCPVIVV2/fZGpg7BftYuTkkBAwkXiw7DcbhC0mceHeeiCb i0NIYAejxNGGf6wQTjuTxOeuJUwwHdP+nGeHSCxnlPi4+QoLhHOfUaL10i5mkCphAXuJ8+cO soDYIgKyElemPWQEKWIWOMcosfbfKqB2Dg42AX2JebtEQUxeATuJJZ1hICaLgKrE7B2SIKao QIRE/xl1kCG8AoISPybfAxvIKWAq8XXvaVYQm1nAUeLBop1QtrzE5jVvmUEWSQisYpfY93Ex K8gcCaALNh1ghjBdJFq+h0J8Iizx6vgWdghbWmLVv1tMEK3NjBIde56xQDgTGCW2rTsEVWUt 8WfVRDaIZXwS7772QM3nlehoE4Io8ZC4PGEzC4TtKLH63xwwW0jgOaPExIMGExjlZyF5ZxaS F2YheWEBI/MqRrHUguLc9NRiwwJDveLE3OLSvHS95PzcTYzgZKJluoNxyjmfQ4wCHIxKPLw7 cvdGCLEmlhVX5h5ilOBgVhLhFZ+6L0KINyWxsiq1KD++qDQntfgQozQHi5I4L6vBxAghgfTE ktTs1NSC1CKYLBMHp1QDY65kW4C3XlM6q8Cl9m+5X6wip+iYCivN/bD61U6bj3eeKq0MD5oo mql+yejOo+MPF3mUyD9S6U/oeO1nzL4+226dXvLc6LlLPz6+WimyUuHeI6nXR2bcv3XsguKD GzdqZ1z2Oy5zRWOvzsJn4Qs2Jaz6beVapbfHU3Nt0L1ezpP60sXvr8jnCiixFGckGmoxFxUn AgDTs1yuIgMAAA== X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFtrEIsWRmVeSWpSXmKPExsVy+t9jAd3Y6fsiDOb80Le4sXwGi8XqTb4W jZ/mMlvcPLSC0eLyrjlsFpM6etktfi0/yujA7nH4zXtmj5Yjb1k9Nq3qZPPo27KK0ePzJrkA 1ig3m4zUxJTUIoXUvOT8lMy8dFul0BA3XQslhbzE3FRbpQhd35AgJYWyxJxSIM/IAA04OAe4 Byvp2yW4ZfRtOc1UsEeq4urtm0wNjP2iXYycHBICJhLT/pxnh7DFJC7cW88GYgsJLGWUePjX oYuRC8h+yCjRvvMOC0hCWMBe4vy5g2C2iICsxJVpDxkhil4yStw8cIwFxGEWuMAosfP/UuYu Rg4ONgF9iXm7REFMXgE7iSWdYSAmi4CqxOwdkiBjRAUiJG497AAbySsgKPFj8j0wm1PAVOLr 3tOsIDazgK3EgvfrWCBseYnNa94yT2AUmIWkZRaSsllIyhYwMq9ilEgtSC4oTkrPNcxLLdcr TswtLs1L10vOz93ECI6vZ1I7GA/ucj/EKMDBqMTDm5C9N0KINbGsuDL3EKMEB7OSCK/41H0R QrwpiZVVqUX58UWlOanFhxhNgf6YyCwlmpwPjP28knhDE3MTc2MDC3NLSxMjJXHextnPwoUE 0hNLUrNTUwtSi2D6mDg4pRoYl36T3B4hxqYUdqbm+YGsL9X7FauUFuY83PSv9p0r467wK+wC FfrzjjH3aolpx2T0LnnKMydi1YeKlWqmPrv0ey+xXfeq/VKxR1zE/hd3SPTRdewbPjgtb5F/ NSX9QK/cnkamKZfuBG6o1bbu8zTau+WQYVqPbO6RXVz37Z346oy8wv5eOSasxFKckWioxVxU nAgAsRIem8UCAAA= X-MTR: 20000000000000000@CPGS X-CMS-MailID: 20170307111957epcas1p3a2fe0184652fcf7b74268ef74d293a8e X-Msg-Generator: CA X-Sender-IP: 203.254.230.26 X-Local-Sender: =?UTF-8?B?S3J6eXN6dG9mIE9wYXNpYWsbU1JQT0wtU3lzdGVtIChUUCkb?= =?UTF-8?B?7IK87ISx7KCE7J6QG1NvZnR3YXJlIEVuZ2luZWVy?= X-Global-Sender: =?UTF-8?B?S3J6eXN6dG9mIE9wYXNpYWsbU1JQT0wtU3lzdGVtIChUUCkb?= =?UTF-8?B?U2Ftc3VuZ8KgRWxlY3Ryb25pY3MbU29mdHdhcmUgRW5naW5lZXI=?= X-Sender-Code: =?UTF-8?B?QzEwG0VIURtDMTBDRDAyQ0QwMjczOTY=?= CMS-TYPE: 101P X-HopCount: 7 X-CMS-RootMailID: 20170217093725eucas1p12478baf297d25303f3020f4973fbf3b0 X-RootMTR: 20170217093725eucas1p12478baf297d25303f3020f4973fbf3b0 References: <87poihtaya.fsf%l.stelmach@samsung.com> <9a57890c-d9e9-5719-e155-ce1161795a02@samsung.com> <20170306185820.GA19696@htj.duckdns.org> Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi On 03/06/2017 07:58 PM, Tejun Heo wrote: > Hello, > > On Fri, Feb 17, 2017 at 12:37:11PM +0100, Krzysztof Opasiak wrote: >>> We need to limit and monitor the number of file descriptors processes >>> keep open. If a process exceeds certain limit we'd like to terminate it >>> and restart it or reboot the whole system. Currently the RLIMIT API >>> allows limiting the number of file descriptors but to achieve our goals >>> we'd need to make sure all programmes we run handle EMFILE errno >>> properly. That is why we consider developing a cgroup controller that >>> limits the number of open file descriptors of its members (similar to >>> memory controler). >>> >>> Any comments? Is there any alternative that: >>> >>> + does not require modifications of user-land code, >>> + enables other process (e.g. init) to be notified and apply policy. > > Hmm... I'm not quite sure fds qualify as an independent system-wide > resource. We did that for pids because pids are globally limited and > can run out way earlier than memory backing it. I don't think we have > similar restructions for fds, do we? Well I'm not aware of such restrictions... So maybe let me clarify our use case so we can have some more discussion about this. We are dealing with task of monitoring system services on an IoT system. So this system needs to run as long as possible without reboot just like server. In server world almost whole system state is being monitored by services like nagios. They measure each parameter (like cpu, memory etc) with some interval. Unfortunately we cannot use this it in an embedded system due to power consumption. So generally now we consider two approaches: 1) Use rlimits when possible to limit resources for each process. The problem here is that this creates an implicit requirement that all system services are well written and able to detect that they for example run out of fd and they will just exit with a suitable error code instead of hanging forever and responding to clients that they are unable to handle their request due to lack of fd. This is hard specially when service use a lot of libraries under the hood because they also need to return this error code from each functions which opens some files. This is especially hard when using some proprietary services or libraries for we don't have access to source code. 2) Use cgroups to limit and monitor resources usage Generally systemd creates a cgroup for each service. cgroups like memory cgroup has an ability to notify userspace when memory usage reaches some level. So for example systemd could get notification that one of cgroups is using more memory than it should but as long as it's not a hard limit of the cgroup this service is not going to even notice this. So instead of returning error from for example malloc() in service, systemd could just send signal to that service and ask it to exit gracefully and the restart it. The disadvantage of this solution is the need of having cgroup for each resource we would like to monitor. For now we have suitable cgroups for everything we need apart from file descriptors. What do you think about this? Maybe you have some other ideas how we could achieve this? Best regards, -- Krzysztof Opasiak Samsung R&D Institute Poland Samsung Electronics