From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S965914AbdKQQrN (ORCPT ); Fri, 17 Nov 2017 11:47:13 -0500 Received: from mail-eopbgr20105.outbound.protection.outlook.com ([40.107.2.105]:23915 "EHLO EUR02-VE1-obe.outbound.protection.outlook.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S965841AbdKQQrE (ORCPT ); Fri, 17 Nov 2017 11:47:04 -0500 Subject: Re: [PATCH] net: Convert net_mutex into rw_semaphore and down read it on net->init/->exit To: "Eric W. Biederman" Cc: davem@davemloft.net, vyasevic@redhat.com, kstewart@linuxfoundation.org, pombredanne@nexb.com, vyasevich@gmail.com, mark.rutland@arm.com, gregkh@linuxfoundation.org, adobriyan@gmail.com, fw@strlen.de, nicolas.dichtel@6wind.com, xiyou.wangcong@gmail.com, roman.kapl@sysgo.com, paul@paul-moore.com, dsahern@gmail.com, daniel@iogearbox.net, lucien.xin@gmail.com, mschiffer@universe-factory.net, rshearma@brocade.com, linux-kernel@vger.kernel.org, netdev@vger.kernel.org, avagin@virtuozzo.com, gorcunov@virtuozzo.com References: <151066759055.14465.9783879083192000862.stgit@localhost.localdomain> <87tvxw3vpe.fsf@xmission.com> <8ee39a83-28c5-2f51-d711-4d0ca56daf9a@virtuozzo.com> <87wp2r33pz.fsf@xmission.com> From: Kirill Tkhai Message-ID: Date: Fri, 17 Nov 2017 19:46:55 +0300 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:52.0) Gecko/20100101 Thunderbird/52.4.0 MIME-Version: 1.0 In-Reply-To: <87wp2r33pz.fsf@xmission.com> Content-Type: text/plain; charset=utf-8 Content-Language: en-US Content-Transfer-Encoding: 7bit X-Originating-IP: [195.214.232.6] X-ClientProxiedBy: HE1PR0701CA0064.eurprd07.prod.outlook.com (2603:10a6:3:9e::32) To AM5PR0801MB1332.eurprd08.prod.outlook.com (2603:10a6:203:1f::10) X-MS-PublicTrafficType: Email X-MS-Office365-Filtering-Correlation-Id: aa861412-6242-4822-9076-08d52ddad236 X-Microsoft-Antispam: UriScan:;BCL:0;PCL:0;RULEID:(22001)(4534020)(4602075)(7168020)(4627115)(201703031133081)(201702281549075)(2017052603258);SRVR:AM5PR0801MB1332; X-Microsoft-Exchange-Diagnostics: 1;AM5PR0801MB1332;3:J5S+pcCScFFfZVJBARwzCPH7msGluRbjw1TvbzgJQC1h9ZkgoTmMX7LBEoVcovhM0CEx6CdcXDWgBZTDeCbgZWmWmV7XIs5Cd8ZbIwu29YuGZRy1Jjzdu1lv2taici6Om6dtNK7eMmSM5u1uwXE8K7qUbtZpP+yXrt0hChJ13xky4G6H1c+ZZZK/E7Z3wPLSvBrMxyS4e3WDL67fInDVWNi70B+DfdrUj2vlmCCH3ImEBAhDOGl6IEaP17Jv2mHe;25:ihK67Ct75Qr8PJh8WF6NwDjUjEl8ogWMwZPOnpJzA6p+xH3vuHmT3O5oCpZJQ8BCvLl7eOeX3nzL5JE4q+mIeUP7mW00DTed106W+/lpkHnEqI/n7uSbNQEVAPJ2Jzw8r+UvP677vqRAPJf2PIefMTw7tllW+FWrCwvRmFtXQhBD1tdb9RslRWdRuB5VM3O/j0j5y+2lriHQ6/FpiQMKzoWoMkSO/W9aVRNKulZ4l5B06BYJhrGtIU5A4jzKqZgT4XK+FffoQiBcV5HRnx7kyu//WkXf+Lx6cuBlaATcpGTzKPYXoF4ArDaJKLaFOWImQbunsv3ZDPY1KK/six7kxg==;31:ngKoVJdDiGaOTgY7zeex2mQ/vveohcAq42TCpHD9xOBPD+iXctpnnsAmabw3t6woWv9+zAhAJeUVwdSV0wSWxbTTmK2C8+7OkHuCahZkRBGvgAkqzIqJGjahDZZUBokel19VGh5PQ6kIAMyN3tbi1IT7S0bbKIz3PJJcCbjPhJZR7CMC5/lMKL8AC9GtjUSfKsKDnrJbyKfQzOkU2oD+kTCKB5qHq18wegsZ1bBqkwE= X-MS-TrafficTypeDiagnostic: AM5PR0801MB1332: Authentication-Results: spf=none (sender IP is ) smtp.mailfrom=ktkhai@virtuozzo.com; X-Microsoft-Exchange-Diagnostics: 1;AM5PR0801MB1332;20:UUJq+NOdsxRPZjpnmjIRntMgkl/7AfpwSSSaqAKjsYXttnlID+UijKfokTEIx2+8bmoikTvYxlNtrbzynf2nLWozfA0xRTGlDfDndJZ4AxXl8eM8d1KYLlt4RSdRkY3EDFqW9Ez3m+H/IBrtnBXxpLlqLJ5OkRxcAAVOWIGm5SsdxXW44hxqsNvDxpsv9FUZFYWALShIpzwdylEa5MJw5RT+sFhfN5Nxu0xaXPnWNJR+3KousHmpJcKpNxhwTsYd4WiJfP6LUioFyRbYSw/BYv7tMpM9ofcNNV6HGImZTwWUWdO1GJ0prUzyFrg56b9i7oUooE2Tkyy0SwdCNJhpZpuWT3Pu6WrVNS5sZEQDeQI16s4a6Ba9jN/dulEaS+KbAjQwDTLM3PzAxV6u8pC+SjqXn7QeU4RTgu9QKg5gpMg=;4:rJlLMIfLsBuaABqriysGRTJ+3AxYG2izaHOjGd1xsZpshdYoZfxrWKXiM7cpMRzljw2x7RtNt0dIeFQjSww2GEMK6G4pqLDRRQKxzri6/setCOdKAyxwBukKckOWwu796fvtkoqGqVJKvQnmxJXKePwpRVf/PosNXiKYDCDFuxcoWYWhhFkSm0kkcHiQZ9OaFEfT8vmwYnCthBFidFjEKT6ERJUjyJRD70ZKJ67Tu5fIuxmZzWRQsN0QYpeBOdtw/DWGdp9LypxCFQCGcfVUFQ== X-Microsoft-Antispam-PRVS: X-Exchange-Antispam-Report-Test: UriScan:; X-Exchange-Antispam-Report-CFA-Test: BCL:0;PCL:0;RULEID:(100000700101)(100105000095)(100000701101)(100105300095)(100000702101)(100105100095)(6040450)(2401047)(5005006)(8121501046)(100000703101)(100105400095)(3002001)(93006095)(93001095)(10201501046)(3231022)(6041248)(20161123555025)(20161123562025)(201703131423075)(201702281528075)(201703061421075)(201703061406153)(20161123560025)(20161123558100)(20161123564025)(6072148)(201708071742011)(100000704101)(100105200095)(100000705101)(100105500095);SRVR:AM5PR0801MB1332;BCL:0;PCL:0;RULEID:(100000800101)(100110000095)(100000801101)(100110300095)(100000802101)(100110100095)(100000803101)(100110400095)(100000804101)(100110200095)(100000805101)(100110500095);SRVR:AM5PR0801MB1332; X-Forefront-PRVS: 049486C505 X-Forefront-Antispam-Report: SFV:NSPM;SFS:(10019020)(6049001)(6009001)(376002)(346002)(51914003)(24454002)(189002)(199003)(51444003)(6246003)(305945005)(53936002)(2950100002)(23676003)(6916009)(6116002)(107886003)(8676002)(36756003)(31696002)(229853002)(16576012)(25786009)(93886005)(3846002)(86362001)(58126008)(47776003)(316002)(65806001)(66066001)(65956001)(6666003)(81156014)(76176999)(50986999)(2906002)(478600001)(50466002)(65826007)(33646002)(54356999)(81166006)(189998001)(53546010)(5660300001)(230700001)(39060400002)(7416002)(101416001)(31686004)(83506002)(16526018)(77096006)(106356001)(4326008)(97736004)(8936002)(68736007)(105586002)(64126003)(6486002)(7736002)(55236003);DIR:OUT;SFP:1102;SCL:1;SRVR:AM5PR0801MB1332;H:[172.16.25.196];FPR:;SPF:None;PTR:InfoNoRecords;MX:1;A:1;LANG:en; X-Microsoft-Exchange-Diagnostics: =?utf-8?B?MTtBTTVQUjA4MDFNQjEzMzI7MjM6RmJWNVN5Q3ZqUFFyekVISkp2NzFRaDZ6?= =?utf-8?B?SDkwYTVJRG5YVHRKdHVkYklSRmplRWlTRTdmTU9IMUFhdEdQWDdRaWhpbzdS?= =?utf-8?B?TDNOVnlrRFQ4VllQakxvUUlaM3kyNWlLMlpXK1NvWlJtcytWTFg0MHNXS1JT?= =?utf-8?B?UitOT2FNazJ6VTZIQWI1aEhlQzdQcGVRK2hybldsSHpVUWlPclRUeGYyQkNk?= =?utf-8?B?ai81c1hkeVN4Q1hqWTB1Mld5eFFGWVArVkppM1czUkMzelAxNHQxbXdHU1Fr?= =?utf-8?B?Nk5PMEF5bXI3Rm5pRnZUQ3VkQnZxOWJ4UTMrQ0RQekJXd2Zidk9EZFpjNCtw?= =?utf-8?B?TlpDaGY5ZDVLM3pTbXN4SVFRaUtvMzlPdWUrUzh6M1FWanQyaTlrbktwV1Ir?= =?utf-8?B?bXVySGRSVHBudENmYkViRGJBZEcxYjM4aXBhWVAwSWhtcW9IQWg1UElwNGlX?= =?utf-8?B?UDB0WlMzdUZwZHMzZ2lLMWRqc3BEZXRCUVJjaGJHdGFob1dKOHJWcVk4Y3Z0?= =?utf-8?B?TjdwSFAyb0NuL2hhTTg0YUZQK0tNbm1KY1J5TlRYclNEZXZ2eEx1cDdIWTJW?= =?utf-8?B?dFJUTG4wYmRTSVllZ2lZWWRPMHhPbFM1NXZ5MDNqNk93dkVwYUNpUTdBdWwy?= =?utf-8?B?OHE3REpCU1lVYmdlU1hQTk80bU0rNVQ4VEZvSDZFRmxqSWNlVlVrMFk1OXY4?= =?utf-8?B?K1ZzaUFtUG5iWjc4ZlBCd2Fla1FXSWNlUTIzdVdWQTJ1cEM0czMvYVJWZFl2?= =?utf-8?B?OXU5bnR3aVkrUWtCSmpJbENOamJaeDBmeFNmWHFWUm8rVEhOUHVua3VKcnNW?= =?utf-8?B?K0tkZFhyclUvVEhUMnpVL2lMclZLdDdnTFhLUDBFY0x3emdnMUw0UkVVZTY0?= =?utf-8?B?S05mcjNpRk1GOGR4ek8wdzVWVU5MVmpLeUZWRmNOcjF1S3BLLytYM0hWQlhq?= =?utf-8?B?cjJQQjhjdkFmdkFnM0M3VnV1UVd1bURDYnA4VUdKRDRMaWR3SkNjZHdRYWg1?= =?utf-8?B?dVVvQlpGeTVGdU9uYUt3dlVYR0FQMkEvODEzNTk1UVpQN3NkNUxsR3NFdlkv?= =?utf-8?B?WXlEdDZvWk5peWVodGZLUk1jYndjWVJESGFzV25vQVpZcktOTHhRMXhXRCtn?= =?utf-8?B?YXcyQ1hmcis4T0hmYWxZNVJaV3o0Nm9ZRVdoMGhFNW9NemcwVE1Sb3JnbmY4?= =?utf-8?B?cXErRjFFUWkvTXQ0eTU4QmVkSUlkVlpmZ1plR0RUNXE0eXpnT3RRdkFoUlV4?= =?utf-8?B?OEpwS3NweGkySW94Tmd0YjFBUDVHekdhRzZJbEV6NDMzS005YzE4VGtBeU1i?= =?utf-8?B?bW5uSzBKb0VNNXd5SzNrTjE2YUVvWW5yYWc0cFhNZzVndmZTdjN2ZWtCQ0NT?= =?utf-8?B?K0pBOTd1WnhrY01ya2VOcllCZHp1U0xwc3JqQzNrVTUrWUFNRThYek5yZ2ht?= =?utf-8?B?V0s3MDJ0QmpaY3ZvZzBHSzUrbkF6SHloNnNNdzZmcGtLd3RoWlk3VTh4MGpH?= =?utf-8?B?MTBWc0M4SUt4eE02TVBEbFZEWFZzNzVBSVBKRFlQeW9BM0psTWNpQ1MxNi95?= =?utf-8?B?cWFJU3c3OFdUVzZZbWxQbFRKbTV5QzFpbjFGaHRJYTlVZnFzYW1PeCs1Wm56?= =?utf-8?B?THdSR0NiUFJvL1lYVHMrM0pxYUtLWVV1M2tiL3FZbCszMTBNUG1FcHJKSWZW?= =?utf-8?B?SWdjWXVpZ0N0dzNUelNiaDhGM3Q4NE5yUzhDVWpZT3EwOEpsRWNyaENHTlUw?= =?utf-8?B?SXh4VmpyR1c1NXUxMkswTW5MdWk0M3QzZVhFZnpBR1RUejJrU0ZHTzRXVkV1?= =?utf-8?B?akpBOWhxaUh3TVRIZjMwYWhNNXJTMWRPZ05YN3dDTTBldXgxb0p0T3pqQWJw?= =?utf-8?B?T1NOVWIvamI1YlgrR3lyNGxjZmZYSUVzaWF2T0V0TGNQdk1ZVVdwQTRMa3NT?= =?utf-8?B?NGNuejV2d0IvTllKaUQ1MnFKcm40emdhM1ByZGNXWWo4YWVSNUZiMUNLTlJ0?= =?utf-8?Q?mfeBW0DY?= X-Microsoft-Exchange-Diagnostics: 1;AM5PR0801MB1332;6:jlRLtAUln+1hOKoH6YyJUm7XPDNtJkKPy12mrw//m1BB7ZEQbGzEJg3qlz9RwUPeyXp/jSqE5lhEj/oL0tVhavYwVezY/FX3rT3zW+gkSkJdx0iEOCiAxNwO5DNe65ObmzlJdBMVvbgqTqd1qyRMnT9UZZEjLeDA3ZMUYIs1e0f1pl4vkyizOuYnqx0Z6yzQOer4ZiSTuoNnq2OsuqhyZbzwMoorQNHLcRqklHSt0KfR01tk2SymWRRDmSvyWME+a4kr50mImjE05q8f1dguUxoBRhQqNowuSw9N3E9MiIoDCbpntbWtVXVl1nlPQza0n7JBTxRz6cg/VKTxfZ1bmlvo8y0uWMTiV+bDncEFFQA=;5:8foT18y32V5w8BsqLfF4ppI3z6wZQGzEcU2t24ZAHNw3HwEECr0+tit6HN/PWl3PoQI4DAEYZWhDFmc52mpi0pRxAfFpg0PlxyjwYF3LzEm9nfECQXed9ghYReugHKkizrf+gEclLCwCzcpgoWrIXePQCfKE6AOhk1BMJIDhUv8=;24:yfMpWUYiUfMwBEO8aNk5FoDqapjtzSpZoU5cnxLfUyxOwtYAyVJkLLKVIrzTO1cdj+Z8T92LWNzjUU0tTrVwMlBtQniXhQjGCooj/vTm9/c=;7:vF0HWBrtJ1oO2KL74/0wJucDgnIx2ng4xtiQmx5hnugCTMVB3ML2PCMYe1gtPyJZsbEZf8R5wheaSC9EeHJrspRL7XJU5q5C2XCctkOzGBe3sX0SUvB7cmTeApzMilkNx9R3Q77/KOQsIqUsbUYSGy50MUStGQF2j+9WRFgDvAEVtDxQPM2lQfsGj9wvBxgY8nCk61Vilzo71/mWmFVGdBnAaHUs0C6pyeUTndqJ55CK1dK1tPd3/iwmGpI49QjL SpamDiagnosticOutput: 1:99 SpamDiagnosticMetadata: NSPM X-Microsoft-Exchange-Diagnostics: 1;AM5PR0801MB1332;20:BbAscygjnSMxdwhK+5vwGAfyIH2rWXwY/k06pF9nmTN8xYCTwW6KtjDx1veFJMH6ODgemI6LX6LqJ8WPtKv/pSwXII2y2Ln/mqECS0kcNbiuDTsLW98CgdcODnoOManeflCsL3I6qYZtUD6QdRNJ9p/WQsv16GXDnVeLdvn43Hc= X-OriginatorOrg: virtuozzo.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 17 Nov 2017 16:46:58.0391 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: aa861412-6242-4822-9076-08d52ddad236 X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: 0bc7f26d-0264-416e-a6fc-8352af79c58f X-MS-Exchange-Transport-CrossTenantHeadersStamped: AM5PR0801MB1332 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 15.11.2017 19:29, Eric W. Biederman wrote: > Kirill Tkhai writes: > >> On 15.11.2017 09:25, Eric W. Biederman wrote: >>> Kirill Tkhai writes: >>> >>>> Curently mutex is used to protect pernet operations list. It makes >>>> cleanup_net() to execute ->exit methods of the same operations set, >>>> which was used on the time of ->init, even after net namespace is >>>> unlinked from net_namespace_list. >>>> >>>> But the problem is it's need to synchronize_rcu() after net is removed >>>> from net_namespace_list(): >>>> >>>> Destroy net_ns: >>>> cleanup_net() >>>> mutex_lock(&net_mutex) >>>> list_del_rcu(&net->list) >>>> synchronize_rcu() <--- Sleep there for ages >>>> list_for_each_entry_reverse(ops, &pernet_list, list) >>>> ops_exit_list(ops, &net_exit_list) >>>> list_for_each_entry_reverse(ops, &pernet_list, list) >>>> ops_free_list(ops, &net_exit_list) >>>> mutex_unlock(&net_mutex) >>>> >>>> This primitive is not fast, especially on the systems with many processors >>>> and/or when preemptible RCU is enabled in config. So, all the time, while >>>> cleanup_net() is waiting for RCU grace period, creation of new net namespaces >>>> is not possible, the tasks, who makes it, are sleeping on the same mutex: >>>> >>>> Create net_ns: >>>> copy_net_ns() >>>> mutex_lock_killable(&net_mutex) <--- Sleep there for ages >>>> >>>> The solution is to convert net_mutex to the rw_semaphore. Then, >>>> pernet_operations::init/::exit methods, modifying the net-related data, >>>> will require down_read() locking only, while down_write() will be used >>>> for changing pernet_list. >>>> >>>> This gives signify performance increase, like you may see below. There >>>> is measured sequential net namespace creation in a cycle, in single >>>> thread, without other tasks (single user mode): >>>> >>>> 1)int main(int argc, char *argv[]) >>>> { >>>> unsigned nr; >>>> if (argc < 2) { >>>> fprintf(stderr, "Provide nr iterations arg\n"); >>>> return 1; >>>> } >>>> nr = atoi(argv[1]); >>>> while (nr-- > 0) { >>>> if (unshare(CLONE_NEWNET)) { >>>> perror("Can't unshare"); >>>> return 1; >>>> } >>>> } >>>> return 0; >>>> } >>>> >>>> Origin, 100000 unshare(): >>>> 0.03user 23.14system 1:39.85elapsed 23%CPU >>>> >>>> Patched, 100000 unshare(): >>>> 0.03user 67.49system 1:08.34elapsed 98%CPU >>>> >>>> 2)for i in {1..10000}; do unshare -n bash -c exit; done >>>> >>>> Origin: >>>> real 1m24,190s >>>> user 0m6,225s >>>> sys 0m15,132s >>>> >>>> Patched: >>>> real 0m18,235s (4.6 times faster) >>>> user 0m4,544s >>>> sys 0m13,796s >>>> >>>> This patch requires commit 76f8507f7a64 "locking/rwsem: Add down_read_killable()" >>>> from Linus tree (not in net-next yet). >>> >>> Using a rwsem to protect the list of operations makes sense. >>> >>> That should allow removing the sing >>> >>> I am not wild about taking a the rwsem down_write in >>> rtnl_link_unregister, and net_ns_barrier. I think that works but it >>> goes from being a mild hack to being a pretty bad hack and something >>> else that can kill the parallelism you are seeking it add. >>> >>> There are about 204 instances of struct pernet_operations. That is a >>> lot of code to have carefully audited to ensure it can in parallel all >>> at once. The existence of the exit_batch method, net_ns_barrier, >>> for_each_net and taking of net_mutex in rtnl_link_unregister all testify >>> to the fact that there are data structures accessed by multiple network >>> namespaces. >>> >>> My preference would be to: >>> >>> - Add the net_sem in addition to net_mutex with down_write only held in >>> register and unregister, and maybe net_ns_barrier and >>> rtnl_link_unregister. >>> >>> - Factor out struct pernet_ops out of struct pernet_operations. With >>> struct pernet_ops not having the exit_batch method. With pernet_ops >>> being embedded an anonymous member of the old struct pernet_operations. >>> >>> - Add [un]register_pernet_{sys,dev} functions that take a struct >>> pernet_ops, that don't take net_mutex. Have them order the >>> pernet_list as: >>> >>> pernet_sys >>> pernet_subsys >>> pernet_device >>> pernet_dev >>> >>> With the chunk in the middle taking the net_mutex. >> >> I think this approach will work. Thanks for the suggestion. Some more >> thoughts to the plan below. >> >> The only difficult thing there will be to choose the right order >> to move ops from pernet_subsys to pernet_sys and from pernet_device >> to pernet_dev one by one. >> >> This is rather easy in case of tristate drivers, as modules may be loaded >> at any time, and the only important order is dependences between them. >> So, it's possible to start from a module, who has no dependences, >> and move it to pernet_sys, and then continue with modules, >> who have no more dependences in pernet_subsys. For pernet_device >> it's vise versa. >> >> In case of bool drivers, the situation may be worse, because >> the order is not so clear there. The same priority initcalls >> (for example, initcalls, registered via core_initcall()) may require >> the certain order, driven by linking order. I know one example from >> device mapper code, which lives here: drivers/md/Makefile. >> This problem is also solvable, even if such places do not contain >> comments about linking order. It's just need to respect Makefile >> order, when choosing a new candidate to move. >> >>> I think I would enforce the ordering with a failure to register >>> if a subsystem or a device tried to register out of order. >>> >>> - Disable use of the single threaded workqueue if nothing needs the >>> net_mutex. >> >> We may use per-cpu worqueues in the future. The idea to refuse using >> worqueue doesn't seem good for me, because asynchronous net destroying >> looks very useful. > > per-cpu workqueues are fine, and definitely what I am expecting. If we > are doing this I want to get us off the single threaded workqueue that > serializes all of the cleanup. That has a huge potential for > simplifying things and reducing maintenance if running everything in > parallel is actually safe. > > I forget how the modern per-cpu workqueues work with respect to sleeps > and locking (I don't remember if when piece of work sleeps for a long > time we spin up another thread per-cpu workqueue thread, and thus avoid > priority inversion problems). > > If locks between workqueues are not a problem we could start the > transition off of the single-threaded serializing workqueue sooner > rather than later. > >>> - Add a test mode that deliberartely spawns threads on multiple >>> processors and deliberately creates multiple network namespaces >>> at the same time. >>> >>> - Add a test mode that deliberately spawns threads on multiple >>> processors and delibearate destrosy multiple network namespaces >>> at the same time.> >>> - Convert the code to unlocked operation one pernet_operations to at a >>> time. Being careful with the loopback device it's order in the list >>> strongly matters. >>> >>> - Finally remove the unnecessary code. >>> >>> >>> At the end of the day because all of the operations for one network >>> namespace will run in parallel with all of the operations for another >>> network namespace all of the sophistication that goes into batching the >>> cleanup of multiple network namespaces can be removed. As different >>> tasks (not sharing a lock) can wait in syncrhonize_rcu at the same time >>> without slowing each other down. >>> >>> I think we can remove the batching but I am afraid that will lead us into >>> rtnl_lock contention. >> >> I've looked into this lock. It used in many places for many reasons. >> It's a little strange, nobody tried to break it up in several small locks.. > > It gets tricky. At this point getting net_mutex is enough to start > with. Mostly the rtnl_lock covers the slow path which tends to keep it > from rising to the top of the priority list. > >>> I am definitely not comfortable with changing the locking on so much >>> code without an explanation at all why it is safe in the commit comments >>> in all 204 instances. Which combined equal most of the well maintained >>> and regularly used part of the networking stack. >> >> David, >> >> could you please check the plan above and say whether 1)it's OK for you, >> and 2)if so, will you expect all the changes are made in one big 200+ patch set >> or we may go sequentially(50 patches a time, for example)? > > I think I would start with the something covering the core networking > pieces so the improvements can be seen. Since the main problem is mutex held during synchronize_rcu(), the performance improvements may become seen only after it's completely removed.