mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v2 0/2] x86/resctrl: Implement rename to help move containers' tasks
@ 2022-11-29 12:01 Peter Newman
  2022-11-29 12:01 ` [PATCH v2 1/2] x86/resctrl: Factor rdtgroup lock for multi-file ops Peter Newman
  2022-11-29 12:01 ` [PATCH v2 2/2] x86/resctrl: Implement rename op for mon groups Peter Newman
  0 siblings, 2 replies; 4+ messages in thread
From: Peter Newman @ 2022-11-29 12:01 UTC (permalink / raw)
  To: fenghua.yu, reinette.chatre
  Cc: Babu.Moger, bp, dave.hansen, eranian, gupasani, hpa, james.morse,
	linux-kernel, mingo, tglx, x86, Peter Newman

Hi Reinette, Fenghua,

This patch series implements the solution Reinette suggested in the
earlier RFD thread[1] for the problem of moving a container's tasks to a
different control group on systems that don't provide enough CLOSIDs to
give every container its own control group.

This change originally depended on the CLOSID update race fix[2] to
provide a race-free mechanism for notifying the CPUs where moved tasks
were residing. However, now that the current guidance for group task
movement is to broadcast IPIs to all CPUs, this patch can be applied
independently.

This patch series assumes that a MON group's CLOSID can simply be
changed to that of a new parent CTRL_MON group. This is allowed on Intel
and AMD, but not MPAM implementations. While we (Google) only foresee
needing this functionality on Intel and AMD systems, this series should
hopefully be a good starting point for supporting MPAM.

Thanks!
-Peter

Updates:

v2: reworded change logs based on what I've learned from review comments
    in another patch series[3]

[v1] https://lore.kernel.org/lkml/20221115154515.952783-1-peternewman@google.com/

[1] https://lore.kernel.org/lkml/7b09fb62-e61a-65b9-a71e-ab725f527ded@intel.com/
[2] https://lore.kernel.org/lkml/20221103141641.3055981-2-peternewman@google.com/
[3] https://lore.kernel.org/lkml/54e50a9b-268f-2020-f54c-d38312489e2f@intel.com/

Peter Newman (2):
  x86/resctrl: Factor rdtgroup lock for multi-file ops
  x86/resctrl: Implement rename op for mon groups

 arch/x86/kernel/cpu/resctrl/rdtgroup.c | 101 +++++++++++++++++++++----
 1 file changed, 88 insertions(+), 13 deletions(-)


base-commit: b7b275e60bcd5f89771e865a8239325f86d9927d
-- 
2.38.1.584.g0f3c55d4c2-goog


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH v2 1/2] x86/resctrl: Factor rdtgroup lock for multi-file ops
  2022-11-29 12:01 [PATCH v2 0/2] x86/resctrl: Implement rename to help move containers' tasks Peter Newman
@ 2022-11-29 12:01 ` Peter Newman
  2022-11-29 12:01 ` [PATCH v2 2/2] x86/resctrl: Implement rename op for mon groups Peter Newman
  1 sibling, 0 replies; 4+ messages in thread
From: Peter Newman @ 2022-11-29 12:01 UTC (permalink / raw)
  To: fenghua.yu, reinette.chatre
  Cc: Babu.Moger, bp, dave.hansen, eranian, gupasani, hpa, james.morse,
	linux-kernel, mingo, tglx, x86, Peter Newman

rdtgroup_kn_lock_live() can only release a kernfs lock for a single file
before waiting on the rdtgroup_mutex, limiting its usefulness for
operations on multiple files, such as rename.

Factor the work needed to respectively break and unbreak active
protection on an individual file into rdtgroup_kn_{get,put}().

This refactoring should not result in any functional change.

Signed-off-by: Peter Newman <peternewman@google.com>
---
 arch/x86/kernel/cpu/resctrl/rdtgroup.c | 35 ++++++++++++++++----------
 1 file changed, 22 insertions(+), 13 deletions(-)

diff --git a/arch/x86/kernel/cpu/resctrl/rdtgroup.c b/arch/x86/kernel/cpu/resctrl/rdtgroup.c
index e5a48f05e787..03b51543c26d 100644
--- a/arch/x86/kernel/cpu/resctrl/rdtgroup.c
+++ b/arch/x86/kernel/cpu/resctrl/rdtgroup.c
@@ -2026,6 +2026,26 @@ static struct rdtgroup *kernfs_to_rdtgroup(struct kernfs_node *kn)
 	}
 }
 
+static void rdtgroup_kn_get(struct rdtgroup *rdtgrp, struct kernfs_node *kn)
+{
+	atomic_inc(&rdtgrp->waitcount);
+	kernfs_break_active_protection(kn);
+}
+
+static void rdtgroup_kn_put(struct rdtgroup *rdtgrp, struct kernfs_node *kn)
+{
+	if (atomic_dec_and_test(&rdtgrp->waitcount) &&
+	    (rdtgrp->flags & RDT_DELETED)) {
+		if (rdtgrp->mode == RDT_MODE_PSEUDO_LOCKSETUP ||
+		    rdtgrp->mode == RDT_MODE_PSEUDO_LOCKED)
+			rdtgroup_pseudo_lock_remove(rdtgrp);
+		kernfs_unbreak_active_protection(kn);
+		rdtgroup_remove(rdtgrp);
+	} else {
+		kernfs_unbreak_active_protection(kn);
+	}
+}
+
 struct rdtgroup *rdtgroup_kn_lock_live(struct kernfs_node *kn)
 {
 	struct rdtgroup *rdtgrp = kernfs_to_rdtgroup(kn);
@@ -2033,8 +2053,7 @@ struct rdtgroup *rdtgroup_kn_lock_live(struct kernfs_node *kn)
 	if (!rdtgrp)
 		return NULL;
 
-	atomic_inc(&rdtgrp->waitcount);
-	kernfs_break_active_protection(kn);
+	rdtgroup_kn_get(rdtgrp, kn);
 
 	mutex_lock(&rdtgroup_mutex);
 
@@ -2053,17 +2072,7 @@ void rdtgroup_kn_unlock(struct kernfs_node *kn)
 		return;
 
 	mutex_unlock(&rdtgroup_mutex);
-
-	if (atomic_dec_and_test(&rdtgrp->waitcount) &&
-	    (rdtgrp->flags & RDT_DELETED)) {
-		if (rdtgrp->mode == RDT_MODE_PSEUDO_LOCKSETUP ||
-		    rdtgrp->mode == RDT_MODE_PSEUDO_LOCKED)
-			rdtgroup_pseudo_lock_remove(rdtgrp);
-		kernfs_unbreak_active_protection(kn);
-		rdtgroup_remove(rdtgrp);
-	} else {
-		kernfs_unbreak_active_protection(kn);
-	}
+	rdtgroup_kn_put(rdtgrp, kn);
 }
 
 static int mkdir_mondata_all(struct kernfs_node *parent_kn,
-- 
2.38.1.584.g0f3c55d4c2-goog


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH v2 2/2] x86/resctrl: Implement rename op for mon groups
  2022-11-29 12:01 [PATCH v2 0/2] x86/resctrl: Implement rename to help move containers' tasks Peter Newman
  2022-11-29 12:01 ` [PATCH v2 1/2] x86/resctrl: Factor rdtgroup lock for multi-file ops Peter Newman
@ 2022-11-29 12:01 ` Peter Newman
  2022-12-16 15:12   ` Peter Newman
  1 sibling, 1 reply; 4+ messages in thread
From: Peter Newman @ 2022-11-29 12:01 UTC (permalink / raw)
  To: fenghua.yu, reinette.chatre
  Cc: Babu.Moger, bp, dave.hansen, eranian, gupasani, hpa, james.morse,
	linux-kernel, mingo, tglx, x86, Peter Newman

To change the class of service for a large group of tasks, such as an
application container, a container manager must write all of the tasks'
IDs into the tasks file interface of the new control group.

If a container manager is tracking containers' bandwidth usage by
placing tasks from each into their own monitoring group, it must first
move the tasks to the default monitoring group of the new control group
before it can move the tasks into their new monitoring groups. This is
undesirable because it makes bandwidth usage during the move
unattributable to the correct tasks and resets monitoring event counters
and cache usage information for the group.

To address this, implement the rename operation for resctrlfs mon groups
to effect a change in CLOSID for a MON group while otherwise leaving the
monitoring group intact.

It's important to note that this solution relies on the fact that Intel
and AMD hardware allow the RMID to be assigned independently of the
CLOSID. Without this, the operation may not be as useful.

Signed-off-by: Peter Newman <peternewman@google.com>
---
 arch/x86/kernel/cpu/resctrl/rdtgroup.c | 66 ++++++++++++++++++++++++++
 1 file changed, 66 insertions(+)

diff --git a/arch/x86/kernel/cpu/resctrl/rdtgroup.c b/arch/x86/kernel/cpu/resctrl/rdtgroup.c
index 03b51543c26d..d6562d98b816 100644
--- a/arch/x86/kernel/cpu/resctrl/rdtgroup.c
+++ b/arch/x86/kernel/cpu/resctrl/rdtgroup.c
@@ -3230,6 +3230,71 @@ static int rdtgroup_rmdir(struct kernfs_node *kn)
 	return ret;
 }
 
+static void mongrp_move(struct rdtgroup *rdtgrp, struct rdtgroup *new_prdtgrp)
+{
+	struct rdtgroup *prdtgrp = rdtgrp->mon.parent;
+	struct task_struct *p, *t;
+
+	WARN_ON(list_empty(&prdtgrp->mon.crdtgrp_list));
+	list_del(&rdtgrp->mon.crdtgrp_list);
+
+	list_add_tail(&rdtgrp->mon.crdtgrp_list,
+		      &new_prdtgrp->mon.crdtgrp_list);
+	rdtgrp->mon.parent = new_prdtgrp;
+
+	read_lock(&tasklist_lock);
+	for_each_process_thread(p, t) {
+		if (is_closid_match(t, prdtgrp) && is_rmid_match(t, rdtgrp))
+			WRITE_ONCE(t->closid, new_prdtgrp->closid);
+	}
+	read_unlock(&tasklist_lock);
+
+	update_closid_rmid(cpu_online_mask, NULL);
+}
+
+static int rdtgroup_rename(struct kernfs_node *kn,
+			   struct kernfs_node *new_parent, const char *new_name)
+{
+	struct rdtgroup *new_prdtgrp;
+	struct rdtgroup *rdtgrp;
+	int ret;
+
+	rdtgrp = kernfs_to_rdtgroup(kn);
+	new_prdtgrp = kernfs_to_rdtgroup(new_parent);
+	if (!rdtgrp || !new_prdtgrp)
+		return -EPERM;
+
+	/* Release both kernfs active_refs before obtaining rdtgroup mutex. */
+	rdtgroup_kn_get(rdtgrp, kn);
+	rdtgroup_kn_get(new_prdtgrp, new_parent);
+
+	mutex_lock(&rdtgroup_mutex);
+
+	if ((rdtgrp->flags & RDT_DELETED) || (new_prdtgrp->flags & RDT_DELETED)) {
+		ret = -ESRCH;
+		goto out;
+	}
+
+	/* Only a mon group can be moved to a new mon_groups directory. */
+	if (rdtgrp->type != RDTMON_GROUP ||
+	    !is_mon_groups(new_parent, kn->name)) {
+		ret = -EPERM;
+		goto out;
+	}
+
+	ret = kernfs_rename(kn, new_parent, new_name);
+	if (ret)
+		goto out;
+
+	mongrp_move(rdtgrp, new_prdtgrp);
+
+out:
+	mutex_unlock(&rdtgroup_mutex);
+	rdtgroup_kn_put(rdtgrp, kn);
+	rdtgroup_kn_put(new_prdtgrp, new_parent);
+	return ret;
+}
+
 static int rdtgroup_show_options(struct seq_file *seq, struct kernfs_root *kf)
 {
 	if (resctrl_arch_get_cdp_enabled(RDT_RESOURCE_L3))
@@ -3247,6 +3312,7 @@ static int rdtgroup_show_options(struct seq_file *seq, struct kernfs_root *kf)
 static struct kernfs_syscall_ops rdtgroup_kf_syscall_ops = {
 	.mkdir		= rdtgroup_mkdir,
 	.rmdir		= rdtgroup_rmdir,
+	.rename		= rdtgroup_rename,
 	.show_options	= rdtgroup_show_options,
 };
 
-- 
2.38.1.584.g0f3c55d4c2-goog


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH v2 2/2] x86/resctrl: Implement rename op for mon groups
  2022-11-29 12:01 ` [PATCH v2 2/2] x86/resctrl: Implement rename op for mon groups Peter Newman
@ 2022-12-16 15:12   ` Peter Newman
  0 siblings, 0 replies; 4+ messages in thread
From: Peter Newman @ 2022-12-16 15:12 UTC (permalink / raw)
  To: fenghua.yu, reinette.chatre
  Cc: Babu.Moger, bp, dave.hansen, eranian, gupasani, hpa, james.morse,
	linux-kernel, mingo, tglx, x86

On Tue, Nov 29, 2022 at 1:02 PM Peter Newman <peternewman@google.com> wrote:
> diff --git a/arch/x86/kernel/cpu/resctrl/rdtgroup.c b/arch/x86/kernel/cpu/resctrl/rdtgroup.c
> index 03b51543c26d..d6562d98b816 100644
> --- a/arch/x86/kernel/cpu/resctrl/rdtgroup.c
> +++ b/arch/x86/kernel/cpu/resctrl/rdtgroup.c
> @@ -3230,6 +3230,71 @@ static int rdtgroup_rmdir(struct kernfs_node *kn)
>         return ret;
>  }
>
> +static void mongrp_move(struct rdtgroup *rdtgrp, struct rdtgroup *new_prdtgrp)
> +{
> +       struct rdtgroup *prdtgrp = rdtgrp->mon.parent;
> +       struct task_struct *p, *t;
> +
> +       WARN_ON(list_empty(&prdtgrp->mon.crdtgrp_list));
> +       list_del(&rdtgrp->mon.crdtgrp_list);
> +
> +       list_add_tail(&rdtgrp->mon.crdtgrp_list,
> +                     &new_prdtgrp->mon.crdtgrp_list);
> +       rdtgrp->mon.parent = new_prdtgrp;
> +
> +       read_lock(&tasklist_lock);
> +       for_each_process_thread(p, t) {
> +               if (is_closid_match(t, prdtgrp) && is_rmid_match(t, rdtgrp))
> +                       WRITE_ONCE(t->closid, new_prdtgrp->closid);
> +       }
> +       read_unlock(&tasklist_lock);
> +
> +       update_closid_rmid(cpu_online_mask, NULL);

I will need to refresh this patch now that we're back to building an
update mask.

This will once again depend on
https://lore.kernel.org/lkml/20221216133125.3159406-1-peternewman@google.com/

^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2022-12-16 15:13 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2022-11-29 12:01 [PATCH v2 0/2] x86/resctrl: Implement rename to help move containers' tasks Peter Newman
2022-11-29 12:01 ` [PATCH v2 1/2] x86/resctrl: Factor rdtgroup lock for multi-file ops Peter Newman
2022-11-29 12:01 ` [PATCH v2 2/2] x86/resctrl: Implement rename op for mon groups Peter Newman
2022-12-16 15:12   ` Peter Newman

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®