From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 39D18215055 for ; Wed, 4 Feb 2026 04:49:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770180542; cv=none; b=qOvoIHILX/CABwGkGUIIzuyrP/dGeQ9Z9zbEgpm/foQkfE7JpGAWniqWPcv5VPQK0QlDSAPRm9XEs+mu4fctOy9HotEYZaQaQgaqsDuYHuVzvvr71KsTFhp6/zhrIL2SRQwFjA5wNwONLDmG2CM3w0W7pVeTEfOgVDMbwee8A4Y= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770180542; c=relaxed/simple; bh=+czzITXfo7mastxUUGo/aH9kzFCTW8TcppL4zk1k7Fc=; h=From:Message-ID:Date:MIME-Version:Subject:To:Cc:References: In-Reply-To:Content-Type; b=MIIgN6RT2q8yQea8DIx1cQ1SrTV2yCKVx12IJe5wDl9miXrSXq7XkcO1xnSHe9X7NOLWr9oJ2nIIRwMxA7sqGHRCMPNODdyfutRuI/5vih3dT05CYxqzgFfTFoooK7AjKLaWsZbMnv7h5Mq5O/aObsCsEn3/Kkz8LYkSKDtx9xQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=cCQo1LBZ; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=iVMoyMtD; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="cCQo1LBZ"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="iVMoyMtD" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1770180541; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=q/2+xTfIjtGMBVQrmAcGgdqv6JTdvdwaafvX4/TR3+8=; b=cCQo1LBZqlfjBkg4gysCS+BbhYlcka2gDJoi2JuvkfbXxdd/Bpd6wjuYKxmiAgsRiEGPXt tK37x7QkPkqEibximlxDIrePo/KVWkJ9u1T7iiI3fBd2groycbG3xoiil4McpbR6gJL8k+ wAXzbHToa49myNxKFS+Fe/xTKJO1fa0= Received: from mail-qt1-f197.google.com (mail-qt1-f197.google.com [209.85.160.197]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-694-PUeTZIkkP-yBuj9iIJzhGQ-1; Tue, 03 Feb 2026 23:49:00 -0500 X-MC-Unique: PUeTZIkkP-yBuj9iIJzhGQ-1 X-Mimecast-MFC-AGG-ID: PUeTZIkkP-yBuj9iIJzhGQ_1770180540 Received: by mail-qt1-f197.google.com with SMTP id d75a77b69052e-5014ad65e3eso23642391cf.1 for ; Tue, 03 Feb 2026 20:49:00 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1770180539; x=1770785339; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:content-language:references :cc:to:subject:user-agent:mime-version:date:message-id:from:from:to :cc:subject:date:message-id:reply-to; bh=q/2+xTfIjtGMBVQrmAcGgdqv6JTdvdwaafvX4/TR3+8=; b=iVMoyMtDP2bIkLDMvS9EqhP2HTWpRqs2UX3PlcEJsIXMKrJFQDKv0t0X1inkwI1EnJ qNPbGc5p5brkXgaH3xK8OSLaDCohQ716rlQAto9AEf7SG7sGy6f0AAM5bsvyaop2TuTn Steg2sRjFL95gVqNiRO9BR1R6De5ts5FCFwWjPMNzNx5104RRYpFppckOtz0syAtnpe2 OtS4a0ifC3JRDY1Hf7IhNWT0VzmwXEOKGBhPhLCdZkytYokZ/ansm0E3lOKgW1WoSEgd n6g2GnHvptwbVqMUlnam57BSLuQtSZKTpNohmtfIUN2jIZL6OCZO0O33Ucy41PvmRKRr Ofbg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1770180539; x=1770785339; h=content-transfer-encoding:in-reply-to:content-language:references :cc:to:subject:user-agent:mime-version:date:message-id:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=q/2+xTfIjtGMBVQrmAcGgdqv6JTdvdwaafvX4/TR3+8=; b=oSsF9KxWi3XLKCucCwJJy8ysGOhZc+RuMSYvKzIX/y39CPPNSw/F1tgu+mLO7xfTfv 1n2p8Qk8Cd/WndK/7tD520C8y+TPSNUSrc7/Kpc/jTh53PUWezDuqnRDNnHllK41gzOS o8MeFIjzzVXl65fp54iXcI8EFsyjY1z9pzmupWp8TeyC5fWgYMMhH2yK3lV51mpJcLJv UCBOqjMcQGskVszizBsPvSNgLdD0P8JzztVvtw+/pbTtsxLnmaz8rAxxJfvqQTTnZQE+ w2+EhFTDRFDyNpGhdzhw+MPIUK6P9T9kWhdeS4ycj4LTSO7n9ws1Mwr8IuaGXPGDiZEK hzJw== X-Forwarded-Encrypted: i=1; AJvYcCUr2Z8c6CoeKl3qR+Lf5pwMXVVMK2FVnpb5eaCUsqQVF4FlrClIxD9a+adWZeGWtPIau+RAN78Ays3Xjpk=@vger.kernel.org X-Gm-Message-State: AOJu0YzW3N0T4zUF4UUYqa6KnoiMVXKoPfur/sbdZvvhoTI4gpVZDFcb kgdZLX7IoW6LlIdOJqYuO28sNREXix0pPX6LdCHNVmwMcgJ/jSHCMXPuzmFa1NLFuVohif8J9zh LYJOS61ovhNQMajWvfSKTgRIzKSbzRvcK8xjVnvWgiFKArDFr9Ng7aAk8JMm9DyKVpg== X-Gm-Gg: AZuq6aIsiUHeffCTR1R3Ml+OiS4vQuoAr7AOqvhpPjWTl/QWBoNSVEdEq4IfU97P419 19aqGgNl9OrbJf56Cx7uLR6uVIfyUfltGY03gwxfVj6ZDGSlFeNtANs8+uISonjtk/GN6DPDb9p 0NRbByb4qJXrm07a3cFzxva/Axbk5NRvHDUP1kU4z/4/Vj/Zx7tLQfHXnsnV70QzJXJAi5EzsVe y6GakjEQS1/tw1VXNJvP5uV/ssb9JBiDSbEbMKtyCfnyRyXCQ0r+3eigUV1xQQxbhac8EVDIBEI Voa2XSHksb9kVUV6qb36qQKlNzPBK74RyQ0kICZsXdSwW/qoPuQJ2K1I+DlpAYWtTYT6DRAArx4 qTI/9+pp8dRLzu6TxMdqHBuujOjEW+VjqaHIW7cKpUkxGZ2U5O/aD6kl5 X-Received: by 2002:a05:622a:11c9:b0:505:e529:11e9 with SMTP id d75a77b69052e-5060930a73cmr78411491cf.36.1770180539525; Tue, 03 Feb 2026 20:48:59 -0800 (PST) X-Received: by 2002:a05:622a:11c9:b0:505:e529:11e9 with SMTP id d75a77b69052e-5060930a73cmr78411211cf.36.1770180539100; Tue, 03 Feb 2026 20:48:59 -0800 (PST) Received: from ?IPV6:2601:188:c102:b180:1f8b:71d0:77b1:1f6e? ([2601:188:c102:b180:1f8b:71d0:77b1:1f6e]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-89521d3cbb1sm11553216d6.55.2026.02.03.20.48.57 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 03 Feb 2026 20:48:58 -0800 (PST) From: Waiman Long X-Google-Original-From: Waiman Long Message-ID: Date: Tue, 3 Feb 2026 23:48:57 -0500 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH/for-next v3 3/3] cgroup/cpuset: Call housekeeping_update() without holding cpus_read_lock To: Chen Ridong , Tejun Heo , Johannes Weiner , =?UTF-8?Q?Michal_Koutn=C3=BD?= , Ingo Molnar , Peter Zijlstra , Juri Lelli , Vincent Guittot , Steven Rostedt , Ben Segall , Mel Gorman , Valentin Schneider , Anna-Maria Behnsen , Frederic Weisbecker , Thomas Gleixner , Shuah Khan Cc: cgroups@vger.kernel.org, linux-kernel@vger.kernel.org, linux-kselftest@vger.kernel.org References: <20260202201144.1669260-1-longman@redhat.com> <20260202201144.1669260-4-longman@redhat.com> Content-Language: en-US In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 2/3/26 9:44 PM, Chen Ridong wrote: > > On 2026/2/3 4:11, Waiman Long wrote: >> The current cpuset partition code is able to dynamically update >> the sched domains of a running system and the corresponding >> HK_TYPE_DOMAIN housekeeping cpumask to perform what is essentally the >> "isolcpus=domain,..." boot command line feature at run time. >> >> The housekeeping cpumask update requires flushing a number of different >> workqueues which may not be safe with cpus_read_lock() held as the >> workqueue flushing code may acquire cpus_read_lock() or acquiring locks >> which have locking dependency with cpus_read_lock() down the chain. Below >> is an example of such circular locking problem. >> >> ====================================================== >> WARNING: possible circular locking dependency detected >> 6.18.0-test+ #2 Tainted: G S >> ------------------------------------------------------ >> test_cpuset_prs/10971 is trying to acquire lock: >> ffff888112ba4958 ((wq_completion)sync_wq){+.+.}-{0:0}, at: touch_wq_lockdep_map+0x7a/0x180 >> >> but task is already holding lock: >> ffffffffae47f450 (cpuset_mutex){+.+.}-{4:4}, at: cpuset_partition_write+0x85/0x130 >> >> which lock already depends on the new lock. >> >> the existing dependency chain (in reverse order) is: >> -> #4 (cpuset_mutex){+.+.}-{4:4}: >> -> #3 (cpu_hotplug_lock){++++}-{0:0}: >> -> #2 (rtnl_mutex){+.+.}-{4:4}: >> -> #1 ((work_completion)(&arg.work)){+.+.}-{0:0}: >> -> #0 ((wq_completion)sync_wq){+.+.}-{0:0}: >> >> Chain exists of: >> (wq_completion)sync_wq --> cpu_hotplug_lock --> cpuset_mutex >> >> 5 locks held by test_cpuset_prs/10971: >> #0: ffff88816810e440 (sb_writers#7){.+.+}-{0:0}, at: ksys_write+0xf9/0x1d0 >> #1: ffff8891ab620890 (&of->mutex#2){+.+.}-{4:4}, at: kernfs_fop_write_iter+0x260/0x5f0 >> #2: ffff8890a78b83e8 (kn->active#187){.+.+}-{0:0}, at: kernfs_fop_write_iter+0x2b6/0x5f0 >> #3: ffffffffadf32900 (cpu_hotplug_lock){++++}-{0:0}, at: cpuset_partition_write+0x77/0x130 >> #4: ffffffffae47f450 (cpuset_mutex){+.+.}-{4:4}, at: cpuset_partition_write+0x85/0x130 >> >> Call Trace: >> >> : >> touch_wq_lockdep_map+0x93/0x180 >> __flush_workqueue+0x111/0x10b0 >> housekeeping_update+0x12d/0x2d0 >> update_parent_effective_cpumask+0x595/0x2440 >> update_prstate+0x89d/0xce0 >> cpuset_partition_write+0xc5/0x130 >> cgroup_file_write+0x1a5/0x680 >> kernfs_fop_write_iter+0x3df/0x5f0 >> vfs_write+0x525/0xfd0 >> ksys_write+0xf9/0x1d0 >> do_syscall_64+0x95/0x520 >> entry_SYSCALL_64_after_hwframe+0x76/0x7e >> >> To avoid such a circular locking dependency problem, we have to >> call housekeeping_update() without holding the cpus_read_lock() and >> cpuset_mutex. The current set of wq's flushed by housekeeping_update() >> may not have work functions that call cpus_read_lock() directly, >> but we are likely to extend the list of wq's that are flushed in the >> future. Moreover, the current set of work functions may hold locks that >> may have cpu_hotplug_lock down the dependency chain. >> >> One way to do that is to defer the housekeeping_update() call after >> the current cpuset critical section has finished without holding >> cpus_read_lock. For cpuset control file write, this can be done by >> deferring it using task_work right before returning to userspace. >> >> To enable mutual exclusion between the housekeeping_update() call and >> other cpuset control file write actions, a new top level cpuset_top_mutex >> is introduced. This new mutex will be acquired first to allow sharing >> variables used by both code paths. However, cpuset update from CPU >> hotplug can still happen in parallel with the housekeeping_update() >> call, though that should be rare in production environment. >> >> As cpus_read_lock() is now no longer held when >> tmigr_isolated_exclude_cpumask() is called, it needs to acquire it >> directly. >> >> The lockdep_is_cpuset_held() is also updated to check the new >> cpuset_top_mutex. >> >> Signed-off-by: Waiman Long >> --- >> kernel/cgroup/cpuset.c | 103 +++++++++++++++++++++++++++------- >> kernel/sched/isolation.c | 4 +- >> kernel/time/timer_migration.c | 3 +- >> 3 files changed, 86 insertions(+), 24 deletions(-) >> >> diff --git a/kernel/cgroup/cpuset.c b/kernel/cgroup/cpuset.c >> index e98a2e953392..d2f51f40f87e 100644 >> --- a/kernel/cgroup/cpuset.c >> +++ b/kernel/cgroup/cpuset.c >> @@ -65,14 +65,28 @@ static const char * const perr_strings[] = { >> * CPUSET Locking Convention >> * ------------------------- >> * >> - * Below are the three global locks guarding cpuset structures in lock >> + * Below are the four global/local locks guarding cpuset structures in lock >> * acquisition order: >> + * - cpuset_top_mutex >> * - cpu_hotplug_lock (cpus_read_lock/cpus_write_lock) >> * - cpuset_mutex >> * - callback_lock (raw spinlock) >> * >> - * A task must hold all the three locks to modify externally visible or >> - * used fields of cpusets, though some of the internally used cpuset fields >> + * As cpuset will now indirectly flush a number of different workqueues in >> + * housekeeping_update() to update housekeeping cpumasks when the set of >> + * isolated CPUs is going to be changed, it may be vulnerable to deadlock >> + * if we hold cpus_read_lock while calling into housekeeping_update(). >> + * >> + * The first cpuset_top_mutex will be held except when calling into >> + * cpuset_handle_hotplug() from the CPU hotplug code where cpus_write_lock >> + * and cpuset_mutex will be held instead. The main purpose of this mutex >> + * is to prevent regular cpuset control file write actions from interfering >> + * with the call to housekeeping_update(), though CPU hotplug operation can >> + * still happen in parallel. This mutex also provides protection for some >> + * internal variables. >> + * >> + * A task must hold all the remaining three locks to modify externally visible >> + * or used fields of cpusets, though some of the internally used cpuset fields >> * and internal variables can be modified without holding callback_lock. If only >> * reliable read access of the externally used fields are needed, a task can >> * hold either cpuset_mutex or callback_lock which are exposed to other >> @@ -100,6 +114,7 @@ static const char * const perr_strings[] = { >> * cpumasks and nodemasks. >> */ >> >> +static DEFINE_MUTEX(cpuset_top_mutex); >> static DEFINE_MUTEX(cpuset_mutex); >> >> /* >> @@ -111,6 +126,8 @@ static DEFINE_MUTEX(cpuset_mutex); >> * >> * CSCB: Readable by holding either cpuset_mutex or callback_lock. Writable >> * by holding both cpuset_mutex and callback_lock. >> + * >> + * T: Read/write-able by holding the cpuset_top_mutex. >> */ >> >> /* >> @@ -135,6 +152,13 @@ static cpumask_var_t isolated_cpus; /* CSCB */ >> */ >> static bool isolated_cpus_updating; /* RWCS */ >> >> +/* >> + * Copy of isolated_cpus to be processed by housekeeping_update() >> + */ >> +static cpumask_var_t isolated_hk_cpus; /* T */ >> +static bool isolcpus_twork_queued; /* T */ >> + >> + >> /* >> * A flag to force sched domain rebuild at the end of an operation. >> * It can be set in >> @@ -298,6 +322,7 @@ void lockdep_assert_cpuset_lock_held(void) >> */ >> void cpuset_full_lock(void) >> { >> + mutex_lock(&cpuset_top_mutex); >> cpus_read_lock(); >> mutex_lock(&cpuset_mutex); >> } >> @@ -306,12 +331,13 @@ void cpuset_full_unlock(void) >> { >> mutex_unlock(&cpuset_mutex); >> cpus_read_unlock(); >> + mutex_unlock(&cpuset_top_mutex); >> } >> >> #ifdef CONFIG_LOCKDEP >> bool lockdep_is_cpuset_held(void) >> { >> - return lockdep_is_held(&cpuset_mutex); >> + return lockdep_is_held(&cpuset_top_mutex); >> } >> #endif >> > void cpuset_lock(void) > { > mutex_lock(&cpuset_mutex); > } > > void cpuset_unlock(void) > { > mutex_unlock(&cpuset_mutex); > } > > void lockdep_assert_cpuset_lock_held(void) > { > lockdep_assert_held(&cpuset_mutex); > } > > A potential issue is that lockdep_is_cpuset_held() only checks cpuset_top_mutex. > In the call chain below, only cpuset_mutex is acquired: > > rebuild_sched_domains_cpuslocked ---only cpuset_mutex is acquired > rebuild_sched_domains_locked > partition_sched_domains > dl_rebuild_rd_accounting > dl_rebuild_rd_accounting > dl_update_tasks_root_domain > dl_add_task_root_domain > dl_get_task_effective_cpus > housekeeping_cpumask > housekeeping_dereference_check > if (IS_ENABLED(CONFIG_CPUSETS) && lockdep_is_cpuset_held()) > > Since lockdep_is_cpuset_held() validates cpuset_top_mutex rather than > cpuset_mutex, could this lead to false lockdep warnings? Right, it should check for either cpuset_mutex or cpuset_top_mutex. Thanks, Longman