From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S934066AbdKPJN4 (ORCPT ); Thu, 16 Nov 2017 04:13:56 -0500 Received: from mail-eopbgr50116.outbound.protection.outlook.com ([40.107.5.116]:45082 "EHLO EUR03-VE1-obe.outbound.protection.outlook.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S1756730AbdKPJNr (ORCPT ); Thu, 16 Nov 2017 04:13:47 -0500 Authentication-Results: spf=none (sender IP is ) smtp.mailfrom=ktkhai@virtuozzo.com; Subject: Re: [PATCH] net: Convert net_mutex into rw_semaphore and down read it on net->init/->exit To: "Eric W. Biederman" Cc: davem@davemloft.net, vyasevic@redhat.com, kstewart@linuxfoundation.org, pombredanne@nexb.com, vyasevich@gmail.com, mark.rutland@arm.com, gregkh@linuxfoundation.org, adobriyan@gmail.com, fw@strlen.de, nicolas.dichtel@6wind.com, xiyou.wangcong@gmail.com, roman.kapl@sysgo.com, paul@paul-moore.com, dsahern@gmail.com, daniel@iogearbox.net, lucien.xin@gmail.com, mschiffer@universe-factory.net, rshearma@brocade.com, linux-kernel@vger.kernel.org, netdev@vger.kernel.org, avagin@virtuozzo.com, gorcunov@virtuozzo.com References: <151066759055.14465.9783879083192000862.stgit@localhost.localdomain> <87tvxw3vpe.fsf@xmission.com> <8ee39a83-28c5-2f51-d711-4d0ca56daf9a@virtuozzo.com> <87wp2r33pz.fsf@xmission.com> From: Kirill Tkhai Message-ID: Date: Thu, 16 Nov 2017 12:13:37 +0300 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101 Thunderbird/52.4.0 MIME-Version: 1.0 In-Reply-To: <87wp2r33pz.fsf@xmission.com> Content-Type: text/plain; charset=utf-8 Content-Language: en-US Content-Transfer-Encoding: 7bit X-Originating-IP: [195.214.232.6] X-ClientProxiedBy: HE1PR07CA0020.eurprd07.prod.outlook.com (2603:10a6:7:67::30) To HE1PR0801MB1338.eurprd08.prod.outlook.com (2603:10a6:3:39::28) X-MS-PublicTrafficType: Email X-MS-Office365-Filtering-Correlation-Id: 78d167ed-7384-49f7-f970-08d52cd25486 X-Microsoft-Antispam: UriScan:;BCL:0;PCL:0;RULEID:(22001)(4534020)(4602075)(4627115)(201703031133081)(201702281549075)(2017052603199);SRVR:HE1PR0801MB1338; X-Microsoft-Exchange-Diagnostics: 1;HE1PR0801MB1338;3:pkN9xBl95ciZtf6ozCZ45HivV/30pQMeBw/AMZF5wIlus6u51tJX+rbWI5xTtZZPMvNj2looac7pFT/9WABYoEr8QmMwg3lQ8odlVE7uZ+udhAbbx9MMpLo2gBtGEWTACT6F3y9he7KXhlMTuZX5Bm1QDA7q14Fn3odW5lb5p+HyVUmlg0QqtCo9QWV0XEYnQB+ZUdbxv8K4fTvJ9zhXsNGFDdPAdWTeJaJiO0CxKMKcbwclIA8KhX2XGW/ZoSPu;25:Xzgs3TttQW5BqvTGVLX+oVB7GXfvsxWH/rslzvpcju1hgyTRXr6xJsFU50J76AmbsZqcDde3PYMyLH9b6vVHQAOPZ8H3phNMLhwg+6bp2s4XFhuN68LAj3PQRCAqYP4ownhj/itX13nXpAVuSdN48I77JP2o42DV9Fz3c+HA5ZPy/ZNvnMPg2WcAOCEScu1oxsZWEWRDd1kOTBS8P7g00j42WzWGr4Sl68xRzCEULvV71D++Amg5D9q1uSdLj7puIMIoZowrqHd7uHxpROOOlRaCE6nazsKJAkBSagbUqgTW9q9K3zReoqjLLePtm0NrMUtBWtbigb1I8wSzOpA39Q==;31:xnsKswG1KDlvEMXRF+bJsgtGz5qPxl02xrfR8vxk0JpS4+vtT8CgB3Sm3yLHW+zQHWH5A6wDbJ2qxKVC7ZpVEfPWwJR8wj8oGPHDSd4c5x/5UxCcxXqTCAipnYi64ymDEQEtUB6AbRMh8g5PqkvJFtIi3h4J+fNHwrbenaiDhYkK0ZsFBcyzuBLRndH7thILu8vIklRk+HGU4nMzooMsTxFI+vxs3iGZ328It9pHd4g= X-MS-TrafficTypeDiagnostic: HE1PR0801MB1338: X-Microsoft-Exchange-Diagnostics: 1;HE1PR0801MB1338;20:DA2WSyIBUpvj9t04h4ZYHYYy/YtlHk5qzXZrL8AEWERLU7aQU7tclMMJaIYu+/cd2PL16Or0TMqMo72kcZaJv3kTkO39T4DSCvqCgwm9GeqDdn3+gf+dHSfZ8y/CtVbgv8bXvKaEtC9a3pAGfv7GxtsDW2d/9YykCuqZODpRmTzCmnbLDviQB1dD63lhrFgaxqhZ7N2+722tQWvy0MBUKsWexn0zNoHom7D99R8gR8demLLGVQ0tPAGjLSnhQagFsVQ1nTE4asvcb+W+try1jA+rkcSLwKaWs9wg8QNmvggoaWUWFFtBvBVsG8p3jGogjOgCxYzMI9GH7Am6X72i8D59QgFSruzVxQK+NQCwPgVYNfDGDJ1MVWvz1zrJ62yZvWaLX6KLBVWJCRtmrQ9IWe7U9vDSa8E5TQYHW23tO5A=;4:ROb0Wg5RO9VmyCnbyrOMHftlhFhFh8LDXpH/smDSdf7GzB4LcBqo6KnuRF0PuNNMVU94RxNHSZsnu7kZf3h2cUuIlC5NrONJO0+kxw4mL+U7YG0H04UeuZhu1hCZWT/sA6+QGjfVr6q7x2khJcyis9iM8dYCjjbc07qPFAqW/xKxz6mpsJtcNLovWTDOcG1geJLp1E/SZCGMQLqMk1uFhgGzyl+nKZUbOduofhD0JCJ2FyaqgLRs7nbLirt5ZySiZ0/52fmZoz5FLza2nhnnOQ== X-Microsoft-Antispam-PRVS: X-Exchange-Antispam-Report-Test: UriScan:; X-Exchange-Antispam-Report-CFA-Test: BCL:0;PCL:0;RULEID:(100000700101)(100105000095)(100000701101)(100105300095)(100000702101)(100105100095)(6040450)(2401047)(8121501046)(5005006)(100000703101)(100105400095)(93006095)(93001095)(3002001)(3231022)(10201501046)(6041248)(20161123555025)(20161123564025)(20161123562025)(201703131423075)(201702281528075)(201703061421075)(201703061406153)(20161123560025)(20161123558100)(6072148)(201708071742011)(100000704101)(100105200095)(100000705101)(100105500095);SRVR:HE1PR0801MB1338;BCL:0;PCL:0;RULEID:(100000800101)(100110000095)(100000801101)(100110300095)(100000802101)(100110100095)(100000803101)(100110400095)(100000804101)(100110200095)(100000805101)(100110500095);SRVR:HE1PR0801MB1338; X-Forefront-PRVS: 0493852DA9 X-Forefront-Antispam-Report: SFV:NSPM;SFS:(10019020)(6049001)(6009001)(346002)(376002)(51914003)(199003)(51444003)(189002)(24454002)(16576012)(316002)(68736007)(64126003)(58126008)(93886005)(25786009)(65806001)(66066001)(7736002)(65956001)(8676002)(2906002)(81166006)(33646002)(305945005)(16526018)(81156014)(53936002)(107886003)(31686004)(6116002)(6916009)(2950100002)(6666003)(3846002)(65826007)(7416002)(5660300001)(50986999)(76176999)(189998001)(54356999)(4326008)(105586002)(230700001)(83506002)(50466002)(39060400002)(47776003)(6486002)(77096006)(31696002)(86362001)(6246003)(106356001)(53546010)(97736004)(36756003)(229853002)(478600001)(101416001)(23676003)(8936002);DIR:OUT;SFP:1102;SCL:1;SRVR:HE1PR0801MB1338;H:[172.16.25.196];FPR:;SPF:None;PTR:InfoNoRecords;A:1;MX:1;LANG:en; X-Microsoft-Exchange-Diagnostics: =?utf-8?B?MTtIRTFQUjA4MDFNQjEzMzg7MjM6ejNvOUYwYWR6aVRTakZNKy9ieHdLNTM5?= =?utf-8?B?QUNMbUQ1UnNOakZKZjJRZng5TExZeUgrU0V4T0xXcUtLQkxudjlHMk5mV2hX?= =?utf-8?B?ai9Nd3Z0N1Q5bWgxM3RHN3M4N2FOWHRQU3pNeUw1V2V4ZGJpVkpYMDAyK0Iy?= =?utf-8?B?TkpwYUNiVS9yanNlN3F0b3NUSy82bVU5NGw2M2NlK3lVMjJnaEIzUXhtNWkw?= =?utf-8?B?RXNSRlpGMi8vY3RrbnJaejh2UE5vWCswdVlEd0tiZk03OCtWQ3MxbVdkbW0v?= =?utf-8?B?TytJOUx3OEovNEVJNWJES3U0WEhwQWVkSzQyLzFBMllpMWtSQ2ZpSDRhU2xC?= =?utf-8?B?QVhFeTJjY3V2Uk1PdGxLWTB1RmR0RjBmTHpKa2o3blNSUWd1d0xnNGRaYXlj?= =?utf-8?B?UnltRldtTTFRT2J5Uk5kVnRJVFpDUFl6Zno2djF0SVZlaXNzWjZkejZmVTB3?= =?utf-8?B?MU5pbUhoSU5ySHhZSzUvY2NUenZ0dy9MbFpIamNTazQzMSttZkx4amF5dGMz?= =?utf-8?B?eCtLTi82cDBIcjNzb29paktJb29vM290WXFCQWtCM3NBY012REtWRlQ2MWFD?= =?utf-8?B?NjJsYXRxMi9NUElkZkxrK1ZHT0J3ZXZhL0FxWVpVcWlWeVh3TXVtQ3Q3V3Qw?= =?utf-8?B?UTV0cXV0c1JoS09RY1MxYUp6ODJtc0FhbmFkOFF0V2ZTUXM3azVLdFZxQUhF?= =?utf-8?B?NHN3TnFPRjB6RzBxalEwYi9JVVBub2lVbi9xbmpTZXZhRUhnQ05NQ0NLemp6?= =?utf-8?B?YkFHSVp4VEphSS9QaGRxTkFCOXJpRGpVcC9BcGhhVlVwWXRhTHk4ODN5di82?= =?utf-8?B?VjczZ1piNEdxellpWTlGMHdwZFJVWWVmc1VGQXR0Mmt1aUJySmNRQTFjS20w?= =?utf-8?B?bEdqdi9OV294QU01NEtVdXNSc3BmRTN0SkpSc1RiUVNkU0ZBY3BwaEZ3UFZY?= =?utf-8?B?aXdSVU9RT3RaYjVhL3dLUFhCZFVWdzI2WXNiNmxYRUFCMmhZQ24wZ1FpV3ht?= =?utf-8?B?ZU1IZTRUMjlxSlhMVTdqcW9GcVpEMzFjMkd6RDdOcG1KRS9oQm5rK1FyWFpE?= =?utf-8?B?SVpOSUJBdC8zUnB1RDV5Qm5wdUx4dlRsYkxtci9vc1NZZHhaR0dQK2lUbGds?= =?utf-8?B?YkNkWUpVSXpmN0R3VGxWUEkyRzNoM2UvZ3FlbWEvV2tKSXNOZGE4WjBLeHRL?= =?utf-8?B?Z1QzWjlTc0hpaFpocXQ3R0FOMEhWVU5SRXBtdVNmZHp2Ly82MzdkWXJzZkJh?= =?utf-8?B?MFA1L3V6QU9ld2dVVDFHWWpJSTRibUx1R1RjU0xPY0dKWVdRQ24xZkh2MDZY?= =?utf-8?B?MHZGUHVHR2dvM1l1UzUwSXpsQ2huelFWSDlzYWhQaHJsQnd6VVkyakF1dDFB?= =?utf-8?B?bFJ6TlBqUzE2M1JqRXpPeFB1RUt0d3BNdmtFcHloODY5c1dUSEdEd24xSVpI?= =?utf-8?B?YUlNN2dLQnZlMUpPRjdLVExndzZkeVdmZUx2eTMwTnM0RmxWSGVQZzA4ZWxV?= =?utf-8?B?cjRtYXJGK1diWVpoWXp2YkphS3NFT1VReURDL0k4UExIR0pzOEhqYW1LVTE0?= =?utf-8?B?SzBySFBpZ1dGMFBUZ3lyMWdOTGtvanlGamxiaVgwdGNsTjNZenlSSTA4UkF5?= =?utf-8?B?bFJ5NEhVc3JJZnVuVHd0K0pkMWJtYnNCaVdrd2dGTGplZ0hlTmptdHpQQVkr?= =?utf-8?B?cWJyYnlLdk41R2owYVNiNUl1WVNpaHBrdHI5N25WU2JBY3ZoaisvUXZCbHMw?= =?utf-8?B?VmpPbU1oMlBJQTlua2RqNTZnSisrL2lJajFwYWI0OXlrOUxLYlZlbk84WjMz?= =?utf-8?B?NDdabmZ4R0RwQ3g5OUZneXV6MWxDclRFSDlRYU9YWEdEbkRaTnByUWhodkdE?= =?utf-8?B?WWFjVXpXODdNc1k3UDNnV0dhb1Jmd1hFSUF3Ukt2RHROZXF6U3NxblJkbXE1?= =?utf-8?Q?G4ldoVpg8kvrJgVvopxDWFxAE/p72zls=3D?= X-Microsoft-Exchange-Diagnostics: 1;HE1PR0801MB1338;6:ct7SOwtVyswTaFT71dpbMZNUsOYdS5Ui+IYVz22XW0WtDOiMnkfaqrj+Xs1jU+TtRdm34cyq8NhJ75L9R0Nh9sWRO7t/UZrnFAi1Rr0lQVjmww+a+LgsvW04sciF/wFohePPkRAB9cDYdrO+YPiZaDQD6y7AIC+ZP1t8Md6VBWZ60p6tSP4wLQsvJLPNwj6g6utS3GfHe1VJhmE9KWqXF09xwRthss9fUMZabeXHgcmZrKl8x6MkbNfNYfq0yIeSQDezUWvvufzDLQevw5nVzvVgfrIdv3WEYf83iEB676sSCLYJjrhWy9D5yvoWZ+0j9lNIw5zVpd/86GImaouYpXfquUtKLPtIpQyFhONr2Ic=;5:fNErsotQjh7+28k8rm/VjOrlJ8TiJTar6/zCxrFS0rypvaYAsUxbBlENYhqsH0qCdT8deTaSsQ7HTmA4b0maq7I3D56q1XCurgNYTwyYQGYvKpEaXyLnAnAdI7Cmuvd013ikMFfmhAd95HkzV3OWB6GD/nhXWiT0CGmn5yhYGUQ=;24:n1IjKkQQIgSoKC7xq2zVHkhk8IZf7B3EYqXnPgqxBel3YgYjhLbohRx+aUdHXi8X5Y3+Zx/uMGAXD+D68tsg7XvE+Dt7qDXQsBGS4JvvwRo=;7:7MlEeh9FKLNEUnTwpYJghbOcE8Ioq7YEbca5D+MXQwG9n9/1xQM2TF9VsRs93tSW0fKYGfY5yCiMu2ev8GyvpZrvZJcP03Gkr1dQj9StukJT0Q8oHysF2yZtjup4GVQrOuCWYH2SFW3i5A5q2wmSQmjB2J2EwqapLqRVsFulflIYJSVhBBpVHBvHI83LRbwalCQt/BDqYzTlVVPxEIjsCRFV0KCY6v8GE7r+gmR53V/gCtc/AG0v+fI4xi6n5mDK SpamDiagnosticOutput: 1:99 SpamDiagnosticMetadata: NSPM X-Microsoft-Exchange-Diagnostics: 1;HE1PR0801MB1338;20:U6OIDSEQUyafp3PJUx3r96GuVJIj0tIRz3nM1cqM7O5W80KhuiE4QUX/1cf9+9ml/c15ARcUJvyEbitONwZH2Olv6VhuGE7FFLdXDdUHGcXDxCNfrHVpKtlyObHy5nsvXJ5BaUMbDL9m3Es2PAuMAmXdkFd0K19LTnqdy5d0sBg= X-OriginatorOrg: virtuozzo.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 16 Nov 2017 09:13:40.2397 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 78d167ed-7384-49f7-f970-08d52cd25486 X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: 0bc7f26d-0264-416e-a6fc-8352af79c58f X-MS-Exchange-Transport-CrossTenantHeadersStamped: HE1PR0801MB1338 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 15.11.2017 19:29, Eric W. Biederman wrote: > Kirill Tkhai writes: > >> On 15.11.2017 09:25, Eric W. Biederman wrote: >>> Kirill Tkhai writes: >>> >>>> Curently mutex is used to protect pernet operations list. It makes >>>> cleanup_net() to execute ->exit methods of the same operations set, >>>> which was used on the time of ->init, even after net namespace is >>>> unlinked from net_namespace_list. >>>> >>>> But the problem is it's need to synchronize_rcu() after net is removed >>>> from net_namespace_list(): >>>> >>>> Destroy net_ns: >>>> cleanup_net() >>>> mutex_lock(&net_mutex) >>>> list_del_rcu(&net->list) >>>> synchronize_rcu() <--- Sleep there for ages >>>> list_for_each_entry_reverse(ops, &pernet_list, list) >>>> ops_exit_list(ops, &net_exit_list) >>>> list_for_each_entry_reverse(ops, &pernet_list, list) >>>> ops_free_list(ops, &net_exit_list) >>>> mutex_unlock(&net_mutex) >>>> >>>> This primitive is not fast, especially on the systems with many processors >>>> and/or when preemptible RCU is enabled in config. So, all the time, while >>>> cleanup_net() is waiting for RCU grace period, creation of new net namespaces >>>> is not possible, the tasks, who makes it, are sleeping on the same mutex: >>>> >>>> Create net_ns: >>>> copy_net_ns() >>>> mutex_lock_killable(&net_mutex) <--- Sleep there for ages >>>> >>>> The solution is to convert net_mutex to the rw_semaphore. Then, >>>> pernet_operations::init/::exit methods, modifying the net-related data, >>>> will require down_read() locking only, while down_write() will be used >>>> for changing pernet_list. >>>> >>>> This gives signify performance increase, like you may see below. There >>>> is measured sequential net namespace creation in a cycle, in single >>>> thread, without other tasks (single user mode): >>>> >>>> 1)int main(int argc, char *argv[]) >>>> { >>>> unsigned nr; >>>> if (argc < 2) { >>>> fprintf(stderr, "Provide nr iterations arg\n"); >>>> return 1; >>>> } >>>> nr = atoi(argv[1]); >>>> while (nr-- > 0) { >>>> if (unshare(CLONE_NEWNET)) { >>>> perror("Can't unshare"); >>>> return 1; >>>> } >>>> } >>>> return 0; >>>> } >>>> >>>> Origin, 100000 unshare(): >>>> 0.03user 23.14system 1:39.85elapsed 23%CPU >>>> >>>> Patched, 100000 unshare(): >>>> 0.03user 67.49system 1:08.34elapsed 98%CPU >>>> >>>> 2)for i in {1..10000}; do unshare -n bash -c exit; done >>>> >>>> Origin: >>>> real 1m24,190s >>>> user 0m6,225s >>>> sys 0m15,132s >>>> >>>> Patched: >>>> real 0m18,235s (4.6 times faster) >>>> user 0m4,544s >>>> sys 0m13,796s >>>> >>>> This patch requires commit 76f8507f7a64 "locking/rwsem: Add down_read_killable()" >>>> from Linus tree (not in net-next yet). >>> >>> Using a rwsem to protect the list of operations makes sense. >>> >>> That should allow removing the sing >>> >>> I am not wild about taking a the rwsem down_write in >>> rtnl_link_unregister, and net_ns_barrier. I think that works but it >>> goes from being a mild hack to being a pretty bad hack and something >>> else that can kill the parallelism you are seeking it add. >>> >>> There are about 204 instances of struct pernet_operations. That is a >>> lot of code to have carefully audited to ensure it can in parallel all >>> at once. The existence of the exit_batch method, net_ns_barrier, >>> for_each_net and taking of net_mutex in rtnl_link_unregister all testify >>> to the fact that there are data structures accessed by multiple network >>> namespaces. >>> >>> My preference would be to: >>> >>> - Add the net_sem in addition to net_mutex with down_write only held in >>> register and unregister, and maybe net_ns_barrier and >>> rtnl_link_unregister. >>> >>> - Factor out struct pernet_ops out of struct pernet_operations. With >>> struct pernet_ops not having the exit_batch method. With pernet_ops >>> being embedded an anonymous member of the old struct pernet_operations. >>> >>> - Add [un]register_pernet_{sys,dev} functions that take a struct >>> pernet_ops, that don't take net_mutex. Have them order the >>> pernet_list as: >>> >>> pernet_sys >>> pernet_subsys >>> pernet_device >>> pernet_dev >>> >>> With the chunk in the middle taking the net_mutex. >> >> I think this approach will work. Thanks for the suggestion. Some more >> thoughts to the plan below. >> >> The only difficult thing there will be to choose the right order >> to move ops from pernet_subsys to pernet_sys and from pernet_device >> to pernet_dev one by one. >> >> This is rather easy in case of tristate drivers, as modules may be loaded >> at any time, and the only important order is dependences between them. >> So, it's possible to start from a module, who has no dependences, >> and move it to pernet_sys, and then continue with modules, >> who have no more dependences in pernet_subsys. For pernet_device >> it's vise versa. >> >> In case of bool drivers, the situation may be worse, because >> the order is not so clear there. The same priority initcalls >> (for example, initcalls, registered via core_initcall()) may require >> the certain order, driven by linking order. I know one example from >> device mapper code, which lives here: drivers/md/Makefile. >> This problem is also solvable, even if such places do not contain >> comments about linking order. It's just need to respect Makefile >> order, when choosing a new candidate to move. >> >>> I think I would enforce the ordering with a failure to register >>> if a subsystem or a device tried to register out of order. >>> >>> - Disable use of the single threaded workqueue if nothing needs the >>> net_mutex. >> >> We may use per-cpu worqueues in the future. The idea to refuse using >> worqueue doesn't seem good for me, because asynchronous net destroying >> looks very useful. > > per-cpu workqueues are fine, and definitely what I am expecting. If we > are doing this I want to get us off the single threaded workqueue that > serializes all of the cleanup. That has a huge potential for > simplifying things and reducing maintenance if running everything in > parallel is actually safe. > > I forget how the modern per-cpu workqueues work with respect to sleeps > and locking (I don't remember if when piece of work sleeps for a long > time we spin up another thread per-cpu workqueue thread, and thus avoid > priority inversion problems). As I know, there is no a problem for worker thread to sleep, as there is a pool. And right after one leaves cpu runqueue, another worker wakes up. > If locks between workqueues are not a problem we could start the > transition off of the single-threaded serializing workqueue sooner > rather than later. > >>> - Add a test mode that deliberartely spawns threads on multiple >>> processors and deliberately creates multiple network namespaces >>> at the same time. >>> >>> - Add a test mode that deliberately spawns threads on multiple >>> processors and delibearate destrosy multiple network namespaces >>> at the same time.> >>> - Convert the code to unlocked operation one pernet_operations to at a >>> time. Being careful with the loopback device it's order in the list >>> strongly matters. >>> >>> - Finally remove the unnecessary code. >>> >>> >>> At the end of the day because all of the operations for one network >>> namespace will run in parallel with all of the operations for another >>> network namespace all of the sophistication that goes into batching the >>> cleanup of multiple network namespaces can be removed. As different >>> tasks (not sharing a lock) can wait in syncrhonize_rcu at the same time >>> without slowing each other down. >>> >>> I think we can remove the batching but I am afraid that will lead us into >>> rtnl_lock contention. >> >> I've looked into this lock. It used in many places for many reasons. >> It's a little strange, nobody tried to break it up in several small locks.. > > It gets tricky. At this point getting net_mutex is enough to start > with. Mostly the rtnl_lock covers the slow path which tends to keep it > from rising to the top of the priority list. > >>> I am definitely not comfortable with changing the locking on so much >>> code without an explanation at all why it is safe in the commit comments >>> in all 204 instances. Which combined equal most of the well maintained >>> and regularly used part of the networking stack. >> >> David, >> >> could you please check the plan above and say whether 1)it's OK for you, >> and 2)if so, will you expect all the changes are made in one big 200+ patch set >> or we may go sequentially(50 patches a time, for example)? > > I think I would start with the something covering the core networking > pieces so the improvements can be seen. > > > Eric > >