From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from BL0PR03CU003.outbound.protection.outlook.com (mail-eastusazon11012059.outbound.protection.outlook.com [52.101.53.59]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 943622D5A19 for ; Thu, 23 Jul 2026 22:08:47 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.53.59 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784844530; cv=fail; b=mfWQmPPWN+Fzt0FVpRxqXnzuQ3u8mby6x0UaRaxNjyhfj2BTAd1ZZGeE61pOTag3whR+vnYEIT1miSNcdHVr6TVYLuOirkL6fco0QsI1B6TiHYFgmiramkyrSGphHMi+ReB1TdUoiPAnXKYT4p82Hm/mx2hvvMepJQwP7ZflfvA= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784844530; c=relaxed/simple; bh=a44tyn0QvsHWN+TJO2piyRZKawj7ARJxNoQ3uxmE70E=; h=Message-ID:Date:Subject:To:Cc:References:From:In-Reply-To: Content-Type:MIME-Version; b=P4/QpI9FMj/yJ6Vi0rSkVhpXGoFGWNQj4jxKPItaN1QXtZ/Y4IgtcmSTGtykWuZdvc+c02sJyC+2s7I0qfCEHyeDxK9cvNM0Evww4tqyvvfQXK8Rd8BuE3cdlu2ow4fMlkK1eK28tk4jbzEFKMl0YLGuED6oqch9Xq9c1GAeA5I= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com; spf=fail smtp.mailfrom=nvidia.com; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b=l7jy8Mt+; arc=fail smtp.client-ip=52.101.53.59 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=nvidia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b="l7jy8Mt+" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=YoNmrFdBFHielPSQ5PcTnDiIgQQKliW+k1iQrm0on3SSYbJSF9OUdxU0GqIz9afg3cqL6fW/s4sgGyFxyoEVNNErv9LS4ASXB17EtnynclBBiCI53Yp/dYTJ9aEOfGxcFUXOHy/ovXuelx3a+W6UkpV2aRGC+awCalNLoAlXlJ1R6M01TEqObG5Wrq3E+bVx9arnsHCVxahALkHvxyOXjHgi+4YN3pS9djtnMeLL7E0GvbRmw9n5ywChE5U5yUBuQuYewQEGfbMHDD6LxrQfyzb2M8nTMPqKdAx29m6Yce51YdlsLB0tcc3norzOPB/+2X95Qxy03cdTeCCTCnZb6A== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=bFFx+ZO0f4/FrF+jkBynI+3DUT7inNCqDYTU6oE2jQY=; b=gyqK/FWMjT47wHgvFM40QaahLsiC3g6rtPCZSvX4WNUe9nM6DiYKh3vIdkTuG8HWn+SXn6blvdGThCF/BLLeFljRGS05z+jYra4z9BOmdI5EPatJZ2qQ1Q8ko5Xl3FlZEr240f3ABPCduJMwCQzDCro1uasLMs7ZU9fc/fjwutPhaxOIeSjtK8yuEzGJbzvhoXhZEz7nP29r6TzbheoKXVitwoMn3eqq4FwAENhnw6CvxdyQD4wqM25VwG/UgK5vOsLAiQ1yusYiIa1fo+iyEPso48IO9ZjRX1HiENNmSjP/tI7E415eS7q55lrDhcOXfG9j1uEAAV68UKPaxExVsw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=nvidia.com; dmarc=pass action=none header.from=nvidia.com; dkim=pass header.d=nvidia.com; arc=none DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=bFFx+ZO0f4/FrF+jkBynI+3DUT7inNCqDYTU6oE2jQY=; b=l7jy8Mt+QpXCbMn2guqrzIuIJRaCialTUDGpEYum/y/CHkdXqgtVibRzl/jXKG26OxRk9ajSHVkFnKp3j4VVGMEjcC7G/i6FElSLD20HNax5Qmjx3tthQQ60hJLVsp+NdYqVpZBVNVK5ehHdznt+Vs28wQygeNc/7LloXpJnx91t8XhQTux379jjx4y5iZSG3bxe5VPxysan2ylIT//xvQnfP4WrBGLZj3I6ttdhollah4G1gYVnxRu9c5fl3yIGmAe7fDj6zBVsMad8B41kM9Fe+BCn+nAV5gf6NVaf6AdtFrSfvTkoxhXPRaT4WKmm4liiCOhAL/6rk3D41QZlEQ== Authentication-Results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=nvidia.com; Received: from DM4PR12MB5230.namprd12.prod.outlook.com (2603:10b6:5:399::11) by DS2PR12MB9663.namprd12.prod.outlook.com (2603:10b6:8:27a::13) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.223.18; Thu, 23 Jul 2026 22:08:38 +0000 Received: from DM4PR12MB5230.namprd12.prod.outlook.com ([fe80::6e87:1bde:1853:3b73]) by DM4PR12MB5230.namprd12.prod.outlook.com ([fe80::6e87:1bde:1853:3b73%5]) with mapi id 15.21.0245.010; Thu, 23 Jul 2026 22:08:32 +0000 Message-ID: Date: Thu, 23 Jul 2026 15:08:30 -0700 User-Agent: Mozilla Thunderbird Subject: Re: [RFC] mpam,x86,fs/resctrl: Generic schema description Proof of Concept To: Ben Horgan , Reinette Chatre , Tony Luck , James Morse , Dave Martin , Babu Moger , Drew Fustini , Chen Yu Cc: Borislav Petkov , Thomas Gleixner , Dave Hansen , Peter Newman , "x86@kernel.org" , "linux-kernel@vger.kernel.org" References: <5ee87762-1898-4b62-94da-85b3e9917ecc@intel.com> <62701203-c4a3-4ec2-a9af-602e1fc15863@nvidia.com> <8f9f78dd-e3f5-4b35-bc72-0eb5dafdcedf@nvidia.com> <36163a81-9737-49e3-93ef-6c392f7272f0@intel.com> <0fc6df54-26c7-43fa-948a-528cd94937f1@arm.com> <9049378c-699a-4155-b1e4-737a1d7265d5@intel.com> <57740b97-80ee-4632-bca3-dc43cd7776c2@arm.com> <44f26cd4-be79-476e-b002-7ccfb7705179@intel.com> <749bd904-523d-4e9d-8493-0e8cfd79949e@arm.com> <9db33feb-cf04-420c-a99a-e31e4b8e4954@arm.com> <8fd6caed-820f-457a-a1ef-a0a006fa52aa@intel.com> <4ef15dde-2fbb-4763-93b6-4333b02d6859@arm.com> <7b751c28-2f04-42b7-b957-af6447e7f824@intel.com> <34b95afb-8b60-4680-9ad1-90c5b24e8fb7@arm.com> Content-Language: en-US From: Fenghua Yu In-Reply-To: <34b95afb-8b60-4680-9ad1-90c5b24e8fb7@arm.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit X-ClientProxiedBy: SJ0PR13CA0170.namprd13.prod.outlook.com (2603:10b6:a03:2c7::25) To DM4PR12MB5230.namprd12.prod.outlook.com (2603:10b6:5:399::11) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: DM4PR12MB5230:EE_|DS2PR12MB9663:EE_ X-MS-Office365-Filtering-Correlation-Id: 5c227ac6-5504-4aeb-0d59-08dee906eeac X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|7416014|23010399003|366016|376014|1800799024|6133799003|18002099003|22082099003|3023799007|5023799004|11063799006|56012099006|4143699003|10067099003; X-Microsoft-Antispam-Message-Info: /VH7L4gi0Gc58mSSag7HYaHAddrACd7TZKZkT0l9/8LzbK3yvZXJXT28MhRe9+3t5Gegg+GVlahrWP9GFPNeOGlt+LHa4ae4PnoVOEZ7tQcj9kQwCgQEjLWT73+JUlHyMVv2pPJDO6h2EZwp5blwZH/y2mAamZFgPFo9Cyp+kVE9H3jj+U6qTL9Pkruj/igs7CyeyIAGsb1JtpLJi6ZsmhON0zb8lgV5SsDGaPEHOyOfm5sy4amnwTwmXjJ7OXzFaTSXs3mzhubEa1WMEE7cNv2sQf8CsSRUaCRFTXkwqc+/VKi+yLxOAl8TDbAUFQXqY4rUmAZIJoCQqT029jBdbDJ5slFrU3J9f5x6oropDWfji3ijXZjQNPNO2Vp8DAMlTaXJrhWvbm47rHqMVM5ECdZTjYQaig6eEE/skmJJ4rn9rt0iCw3G76Brxt4kkL6X2LBKHdYZuftuxzeNg6UVwQA8W7+P+adSJhJsZRJ8JSX0Y9Q/O/LofPc2evvA4U5keu53ye2yAo2ORfKPYA08J5K7ygaiai1XCBimz14/sudNE7EM0s32j9JAQuT2oPaCJB3cINierT7U0znDg+7Vin14FWYnRhLnfu2HuyVkMiPMKbNyj9Nca72Sa6HtGgjFmE/FSPPTX75SYSMBZJ3UoemTqBbfQYrBFW6ht2AM9ms= X-Forefront-Antispam-Report: CIP:255.255.255.255;CTRY:;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:DM4PR12MB5230.namprd12.prod.outlook.com;PTR:;CAT:NONE;SFS:(13230040)(7416014)(23010399003)(366016)(376014)(1800799024)(6133799003)(18002099003)(22082099003)(3023799007)(5023799004)(11063799006)(56012099006)(4143699003)(10067099003);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: =?utf-8?B?c0ZzaTA3UVVrcVJZRUFjMTM4MTdNMG1pd0ZnMDMwTVJNbFFXN29iTXM2OFoy?= =?utf-8?B?dHg2UC8yZVJqY0I5a2ZKNzRKSmJiMUdMcWJVd0NiWXAxaXpzemRKczZUbXNu?= =?utf-8?B?ZmZJQzZXUXE3TDk5MTN2VkRlZVBwT0pzZzVNaXlxazU3T21qa0NEOGdaK29E?= =?utf-8?B?KzZIbDUwQ0xGblY1dlBsRi8xd0F0bGhOeXl1cmRSdi9wU21GZzByb29QUE1y?= =?utf-8?B?aC9KV0ErZTRQYjIzTTMxOXZwd3lheFp5QWU0MjlIaitRWUlhRXlWQ3o0cU81?= =?utf-8?B?RTBWSnVrVWRqMTRkRkdLTkZ2Sy92WnNwdTVIUFVzZ0Y2VG52a2tQVmNTK1pN?= =?utf-8?B?UVpuS3dDMUNja3U5d25jbTdZWkZpc2FlYnkyQkxrLzBuUU5EYmVoc2V0TVpi?= =?utf-8?B?ZFllSkxHNzFwS0h3R2l6NGhtMlhKWWl2R0RuSklpQU1IZ01FM0pDckFOZlVO?= =?utf-8?B?MEV0eG1wcHIrUnJTSFZzMGpYV2Fjd3ZSSGxzSkM5a002eXFsOHhFVnVWemNM?= =?utf-8?B?b05tcVJJMW5jUVkvbkIxL0JDNFNHUjRVeVFuNllMRDh1U1gxZElEUG5KNWJH?= =?utf-8?B?dzJSRHFiQVdCai84c1d3SmdNeW5BcHRQem04c1B2ZHNhazlxbTV0UHhjelAr?= =?utf-8?B?aExLQnJhUjk4V0ViUUNOM2tRdkZDOThVNmIxUG5JVFpQVTBaa2ppeXRyRytV?= =?utf-8?B?MWdXSzhWRHgzWWR1a0lJbmpmTDBHaTdHenJaN1hIT2NwcHpoUEdmRHNNWnhY?= =?utf-8?B?M0ErKzdJaGJVVGRjMktJZG9hQXFsbXlyTFZ4emp6UW55SmJ2UlhHem40c2Zr?= =?utf-8?B?NWNja1k1Y3RjaTJ4Q0FMR05RTmt5bFZJbUh4RUZieUF1d09aWFJBNTRYbFFD?= =?utf-8?B?dkUxZFphU2xkZDFoM3JtcTVyUzNkN2dFUHVQczc0dVEvQzlFSktIVTgzenBi?= =?utf-8?B?UUJOS25EWDdDNG44ZGptUjJlSXhkNURLM0hoVnYvOU1SSG16UjhVSDlZaFNx?= =?utf-8?B?L0M5alJoMmdNTDJnalFjNW5NeTZjYkFlMTQ2TWFsV1Z5MkZXdVJzRXZuMjJa?= =?utf-8?B?SStMSUhxVzBDVGpZcFZuSksxSjZ5OWVVS2puSE5CblZFY0FhMVBVYVl1MHVu?= =?utf-8?B?Sm9MdTdjTGx3NmdoZFZDSkVvTEhnakV1dmtRVzlRMXJZTWRCSWhLVlFoOFRl?= =?utf-8?B?c1QvUVNCYk53bGNDTGtWMnF3dFlFKzU4c1RBKzIxVXMwbTNEREVyQklubUNN?= =?utf-8?B?Tnl3dWNCbTdWcEdTdE5YVWN5RzZHYXJ2ODU1NzEwWnVSRERUYkJ1YXBRam1S?= =?utf-8?B?SUtQUHBnUmhWcXh6b3ZSY1FvOFl1azlZSDNGUFlCbkpwM0l6aHFsSmxPZUFp?= =?utf-8?B?dHo2ek9uYzI3aEhqZm9mQnZqWjNiRGJ5d0w3R1VKcSsyNHhVeVYwTnpOREZp?= =?utf-8?B?NWRsN3JaWTcxMGZYMG9KMUpyelVxN0NUNFVrdFRSb3VXMUlTeEZYY1JSdEJa?= =?utf-8?B?Rm5TWDh1QlRZYU1seWRKVUpTbXBmVFY2YjcvZkozY2ZNdmk0ZjlqQUdwdjFW?= =?utf-8?B?VjBpQ3d6TFIrZlVIUEhHcDAxQmZRZ3hUMTVXMldNVHdaQzVPR3BidDUzNjRn?= =?utf-8?B?SFZLL3JnT2dpek9nN1gvOHhsKy9OeG1KNnZKUkVFRkZNeXlpendXLzVmVE9P?= =?utf-8?B?T1UrbE9CcHEwWGVIRW9OTXZwU3BoTnFVUFY1ckp5L3FrTXMwZ1Bvak5GZU1Q?= =?utf-8?B?U21YeUdWVlp1NmEyamxwejFVamZYVkk3UXNEVUFJTVdINi9CRUtyVmt4dUVl?= =?utf-8?B?MFoxdjAvY1BIakxnYzhYV3hUQ2xwRDZGdDRGWHN1eXM5OHRhR3NzVi81eWo1?= =?utf-8?B?S0ZWY1pvUkFjc3JBOENCOWhCTUQ2ajY2bmU3eE1hK1c2QW03c0l4UnhwY0VT?= =?utf-8?B?WUxTMVlJVVFTVWNESlF0bUV0eEVqYll4U0ZZQU9qMkVyYmdRaHJJdWw1RXh2?= =?utf-8?B?VCtYZXZzbk9tanZPWW1WUUR1ZmhNQXNwR3RGOUtDdWxIeUdPLy9pVm9xTnZq?= =?utf-8?B?b1BXSTdmM1NRVjVISFAwZDhpbFVEcTZEOG1kUDM2dXhwVFZLc3I1U2lpN1k1?= =?utf-8?B?anhIQmNKUER0ODlGdDM5ZVdMSE5RN21TN2pTTnhwT3d5NDR4V21IMGsvUVNj?= =?utf-8?B?Q05SUFVCYjVoRjNxa1Q5ZWtUbTJBUzQwS25ZTVVwMXNsbmZYVEV0NHhWbG5n?= =?utf-8?B?eGV2ZUxCN2RkQTczaVpVWmQwOVF1azhSM0pONCswMFJZb1F6dWQ0cmtPb3I5?= =?utf-8?B?R1Z3Q0lleHpybG42TGovMlZLcmpsdXBKSWxsTnJqdnB1VXBZNEp6dz09?= X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-Network-Message-Id: 5c227ac6-5504-4aeb-0d59-08dee906eeac X-MS-Exchange-CrossTenant-AuthSource: DM4PR12MB5230.namprd12.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Internal X-MS-Exchange-CrossTenant-OriginalArrivalTime: 23 Jul 2026 22:08:31.8614 (UTC) X-MS-Exchange-CrossTenant-FromEntityHeader: Hosted X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-MailboxType: HOSTED X-MS-Exchange-CrossTenant-UserPrincipalName: +4xZdZLG6aH9YHCxpCrK+EJMZ9yBXC4SrHydzFFCRii+XDQigld7yNn99PUNoSvsX8S7BfCkXkJShtInPlqnIg== X-MS-Exchange-Transport-CrossTenantHeadersStamped: DS2PR12MB9663 Hi, Ben, On 7/21/26 06:23, Ben Horgan wrote: > Hi Reinette, > > On 7/20/26 23:54, Reinette Chatre wrote: >> Hi Ben, >> >> On 7/20/26 6:30 AM, Ben Horgan wrote: >>> Hi Reinette, >>> >>> On 7/17/26 17:00, Reinette Chatre wrote: >>>> Hi Ben, >>>> >>>> On 7/17/26 5:20 AM, Ben Horgan wrote: >>>>> On 7/16/26 18:07, Reinette Chatre wrote: >>>>>> On 7/16/26 9:44 AM, Ben Horgan wrote: >>>>>>> On 7/16/26 17:04, Reinette Chatre wrote: >>>>>>>> Hi Ben, Chenyu, and Tony, >>>>>>>> >>>>>>>> On 7/16/26 7:59 AM, Ben Horgan wrote: >>>>>>>>> On 7/15/26 16:41, Reinette Chatre wrote: >>>>>>>>>> On 7/15/26 1:34 AM, Ben Horgan wrote: >>>>>>>>>>> On 7/14/26 23:06, Reinette Chatre wrote: >>>> >>>> ... >>>> >>>>>>>>>>> Alternatively, could the "mode" file be used to switch between "MB" and "MB_MAX" and the two never >>>>>>>>>>> need be shown at the same time. The user opts in to using the new interface, "MB_MAX" by setting >>>>>>>>>>> "mode" and can just toggle back if they want to use "MB" directly again. >>>>>>>>>> Interesting. So far the "mode" options have been "legacy" and "native" where "legacy" would show the >>>>>>>>>> legacy as well as emulated controls in the schemata file and "native" will only show the new controls. >>>>>>>>>> resctrl could make the default of "legacy" mean that *only* the legacy control is shown without insight >>>>>>>>>> into the controls it is being emulated with. The only insight to this would continue to be via the >>>>>>>>>> hierarchy in the info directory, which would show relationship but not the actual control values. >>>>>>>>>> Do you think there could be a need for users (excluding validation?) that may want to see the underlying >>>>>>>>>> control values used to emulate a legacy control? >>>>>>>>> >>>>>>>>> They may want to see the underlying values to be able to migrate there legacy configuration to the >>>>>>>>> native configuration but they can just write their legacy configuration and then toggle the mode to >>>>>>>>> see the values in the native mode. >>>>>>>>> >>>>>>>>>> >>>>>>>>>> I assume for backward compatibility that "legacy" would remain the default so a possible inconvenience >>>>>>>>>> here would be that users familiar with the new controls would forever need to switch the mode before being >>>>>>>>>> able to use them. >>>>>>>>> >>>>>>>>> This does also make the new controls slightly harder to discover. >>>>>>>>> >>>>>>>> What if resctrl combines the two suggestions? Specifically, for backward compatibility the mode will be >>>>>>>> "legacy" and if there are emulated controls then resctrl displays them in schemata file with "#" prefix >>>>>>>> but not(*) support any changes from user space to the underlying hardware controls. This will make new controls >>>>>>>> easier to discover and let user space see the underlying values, but not break a user space that >>>>>>>> may read schemata file, change a few values, and write entire file back. >>>>>>> >>>>>>> My initial impression is that this would work but is probably unnecessary. >>>>>> >>>>>> ok. My goal was to present ideas to address the issues raised so far. Ideally this would result in discussion of >>>>>> pros/cons. If you find this unnecessary, could you please expand with some insight into which parts you find >>>>>> unnecessary or how you would prefer this solved? I am finding it difficult to interpret above response. >>>>> >>>>> Sorry, yes. I was a bit lazy in my reply yesterday. If the mode is easy and non-destructive to >>>>> change then doesn't the user get the same information with just the burden of toggling the mode and >>>>> then reading the schemata a second time. This does rely on the user being aware of the new behaviour >>>>> though. >>>>> >>>>> If the emulation isn't a, one to one, bijection then is changing mode expected to be destructive? >>>>> More on this below. >>>> >>>> There may be some corner cases (region aware may have some as you highlight below), but in general I do not expect >>>> changing mode to be destructive. The idea behind giving insight into underlying control values while legacy interface >>>> is in use is to address your earlier point that not doing so will make new controls slightly harder to discover. >>>> (more below) >>>> >>>>>>>> How the "enabled" vs "disabled" state of a control works with this needs some confirmation since region-aware >>>>>>>> MBA support has this extra caveat of supporting MSR and ACPI interfaces which results in the relationship >>>>>>>> between "mode" and "status" of underlying controls not being consistent. This may be ok but please consider >>>>>>>> example below. >>>>>>>> >>>>>>>> Thinking through this with examples as I understand MPAM and RDT region aware so far. Could this work for >>>>>>>> MPAM and region aware MBA? >>>>>>>> >>>>>>>> (*) Should resctrl allow user space to change underlying control value of an emulated control in "legacy" mode? >>>>>>> >>>>>>> I don't think this causes problems for MPAM but for RDT region aware controls couldn't you end up >>>>>>> with a control state that isn't reachable just by configuring the legacy schema. >>>>>> >>>>>> Could you please highlight the scenario you refer to? >>>>> >>>>> My expectation here was that for any value of MB, X, the value for each of the regions would be the >>>>> same, Y. (Not sure if this is actually the case.) >>>>> >>>>> MB: X >>>>> # MB_REGION0_MAX:Y >>>>> # MB_REGION1_MAX:Y >>>>> # MB_REGION2_MAX:Y >>>>> # MB_REGION3_MAX:Y >>>>> >>>>> So, if one of the regions was to be changed individually then it would be in a state not reachable >>>>> by just changing the legacy control, MB. Potentially this complicated switching mode as well as >>>>> uncommenting and writing a schemata. >>>> >>>> Indeed. This scenario was highlighted in slide 8 of >>>> https://lpc.events/event/19/contributions/2093/attachments/1958/4172/resctrl%20Microconference%20LPC%202025%20Tokyo.pdf >>>> and discussed between Tony and Dave Martin in the thread starting at >>>> https://lore.kernel.org/lkml/aPf0OKwDZ4XbmVRB@agluck-desk3/ >>>> >>>> Tony and Dave discussed a few scenarios and how resctrl could behave with different user interactions. >>>> On a high level I understood that the underlying controls will always, as you state below, "show the >>>> full story" and if user space interacts with them (instead of using the legacy control) then the >>>> control values associated with the legacy control cannot be relied on to represent accurate state to >>>> the point that it may even be better to not be visible. >>>> You are right, this is a good motivation to not allow user to modify underlying controls when in >>>> "legacy" mode. >>>> >>>> The remaining open is whether resctrl should display the (read-only) underlying control >>>> values in schemata file when in "legacy" mode. By default this does not work for region-aware since >>>> it will at least initially use different hardware interfaces for the different controls, essentially >>>> this means that there is no actual emulation and the legacy and region aware control are both >>>> legitimate hardware controls. Specifically, considering your example, there is no mapping from "X" to "Y". >>>> >>>> The way region-aware is planning to address this is to use the "enabled" vs "disabled" status of a >>>> control to start with underlying controls disabled so that they are not displayed in schemata file >>>> when legacy mode is enabled. >>>> >>>> If resctrl instead only displays legacy controls when in "legacy" mode then the "status" may no longer >>>> be needed for region-aware. RISC-V also considered using this "status" that I think may also be >>>> solved with this approach, I am not sure though. >>>> >>>>> I was thinking that in legacy mode that the region values would always be kept the same but perhaps >>>>> the legacy (MB) value could just be the 'best estimate' and the native (MB_REGION0) values show the >>>>> full story. 'best estimate' could be difficult to choose but allows switching mode to be >>>>> non-destructive. >>>>> >>>>> For MPAM, at least until Fenghua sent his series today on MBA control emulation [1], I was expecting >>>>> that MPAM would only use emulation for exposing finer grade controls to the user and not allowing >>>>> different underlying hardware to be used for the emulation. To me, it seems reasonable that a new >>>>> control would involve a new schema but I do appreciate the benefit of being able to continue to use >>>>> the old interface without stopping a new interface being used. The danger is that there are >>>>> unexpected user visible side effects, e.g. domain lifetime, and that control behaviour is different >>>>> from the user expectations. >>>> >>>> I have not looked at Fenghua's series in detail but from earlier discussion I understood the high level >>>> problem to be that user space may have expectation that for every "resource" represented by a directory in >>>> info/ there is a matching entry in the schemata file. Whether this is an actual expectation from user >>>> space tools is not obvious to me though: there is already an exception since there is the >>>> "L3_MON" "resource" that does not have a schemata file entry. >>>> >>>> To ensure backward compatibility the safest would be for resctrl to always provide a "MB" control via >>>> an entry in the schemata file when the "MB" resource is exposed via info/. If on the other hand resctrl >>>> is not expected to provide a "MB" entry in schemata file when there is a "MB" resource described in info/MB >>>> then resctrl should not be forced to always provide such emulation. >>>> >>>> I would appreciate your thoughts here. >>> >>> Ok, thanks for pointing this out, I hadn't previously properly understood the motivation for >>> resource emulation. >>> >>> As emulation adds complication and has the potential to break other user expectations I hope we can >>> manage to keep it to only simple cases. For instance as I mentioned before, the lifetime of a >>> resctrl domain lifetime is necessary different (when there are multiple NUMA nodes) for domains >>> backed by MSC at the memory to MSC at the cache. At the memory they would be inaccessible when the >>> NUMA node is powered off. Also, any partitioning is happening at a different point in the topology. >>> >>> Can we get around this by by choosing a different organization of the info/ directory? We already >>> have separate resource for L2 and L3. Perhaps, the same for MB, MB_NODE, (not sure about MB_REGION) >>> and move the scope to be a property at info//scope rather than >>> info//resource_schemata/ctrl/scope. >> >> Good point that L2 and L3, which is the same resource at different scope, are treated as different >> resources. With that view it seems reasonable to also consider MB and MB_NODE as different resources. >> The point here is that it is at a different scope so more specifically what a "domain ID" in the >> schemata file represents. With this view, MB_REGION is L3 scope that matches to the MB resource. > > Ok. This makes sense to me. I had confused myself about what MB_REGION is. > >> >> You will notice I get a bit lost in the discussion below so getting back here I would like to >> clarify if you propose that info/ contains a directory for each allocation scope of each resource with that >> directory containing descriptions of all the controls for that resource at the indicated scope or do >> you propose that info/ contains a directory for each control? > > The former, info/ contains a directory for each allocation scope of each resource. > > > > info > ├── L2 > │   ├── resource_schemata > │   │   ├── L2 > │   │   ├── L2_CMAX > │   │   └── L2_CMIN > │   └── scope : L2 > ├── L3 > │   ├── resource_schemata > │   │   ├── L3 > │   │   ├── L3_CMAX > │   │   └── L3_CMIN > │   └── scope : L3 > ├── MB > │   ├── resource_schemata > │   │   ├── MB > │   │   │   └── MB_MAX > │   │   ├── MB_MIN > │   │   ├── MB_PBM > │   │   └── MB_PROP > │   └── scope : L3 > └── MB_NODE > ├── resource_schemata > │   ├── MB_NODE_MAX > │   ├── MB_NODE_MIN > │   ├── MB_NODE_PBM > │   └── MB_NODE_PROP > └── scope : NUMA NODE > >> >>> >>> Can MB_REGION be considered an orthogonal new control or does using it require that the traditional >>> intel MB (delay) not be configured? >> >> The Intel systems that support both MSR ("traditional") and ACPI (region aware MBA) cannot use both >> concurrently. Not sure if this answers your question. > > Thanks, this answers my question. > >> >> Just to clarify, there is no single "MB_REGION" control. When considering the "region aware MBA" >> feature I am currently aware of the following possible controls: >> MB_REGION0_MIN >> MB_REGION0_MAX >> MB_REGION0_OPT >> MB_REGION1_MIN >> MB_REGION1_MAX >> MB_REGION1_OPT >> MB_REGION2_MIN >> MB_REGION2_MAX >> MB_REGION2_OPT >> MB_REGION3_MIN >> MB_REGION3_MAX >> MB_REGION3_OPT >> >> >>> On an MPAM system: >>> >>> info >>> ├── MB >>> │   ├── resource_schemata >>> │   │   ├── MB >>> │   │   └── MB_MAX >>> │   └── scope >>> └── MB_NODE >>> ├── resource_schemata >>> │   └── MB_NODE >>> └── scope >> >> Above MB and MB_MAX are represented on the same level. My understanding is that MB_MAX is >> used by driver as the underlying hardware control for the percentage based MB exposed to >> user space. Thus, when user space changes the "MB" control via schemata file it is expected >> to also impact the underlying "MB_MAX" control. By representing them as above this relationship >> is lost. So far I understood that the "MB_MAX" control will be shown as a child of the "MB" control >> to show this relationship. Looks like you plan to change this but from what I understand this >> emulation is still relevant? > > Please just consider this a mistake. MB_MAX should be a child of MB. > > info/ > ├── MB > │   ├── resource_schemata > │   │   └── MB > │   │   └── MB_MAX > │   └── scope > └── MB_NODE > ├── resource_schemata > │   └── MB_NODE_MAX (Changed from MB_NODE) > └── scope > >> >> When thinking about MPAM, what would the underlying hardware control of "MB_NODE" be? It looks >> from above that it would either start out by itself having the properties of the underlying >> "MAX" control or is the plan to have it be a percentage based control backed by the >> underlying "MAX" hardware control? > > The underlying hardware of MB_NODE would be essentially the same hardware as that backing MB_MAX, > but at a different location in the SoC, at the memory controller rather than in the L3. > > To correct myself slightly, I don't think we should have a control called MB_NODE, rather, it should > be MB_NODE_MAX. > > My understanding of previous discussions is that __ it the > pattern for control names and the pattern for resource names being _. > (allowing for or being missing to match existing naming.) > > I don't think we should introduce more percent based controls and for new controls we can introduce > a new format to describe them. Perhaps just the positive integer with a resolution supplied in info/ > as discussed previously. Although, I have been pondering on whether we can do a bit better. > > We could use hexadecimal point based format for controls which are a proportion of a resource and > have a resolution which is a power of 2. The advantage of this is that the meaning of the value is > independent of the granularity of the control (number of parts). > > 0 is represented as 0x0 > 1 as 0x1 > 1/2 as 0x0.8 > 7/256 as 0x0.07 > 1/2**28 0x0.00000001 > etc > > This maps well to the MPAM fixed-point fraction point format without having the weirdness of having > values forced to 1 or 0 not being really 0. These MPAM h/w oddities can be hidden just by using the > mbw_min mbw_max of a control. In MPAM this could be used in CMIN, CMAX, MB_MAX, MB_MIN and I would > hope this would be useful for other architectures too. I am preparing some RFC patches on top of > your PoC for consideration of this idea and to explore some of the proposals discussed relating to > generic schemata and how they land in practice from the MPAM side. > >> >> I also understand MPAM to support more memory bandwidth controls ("MIN", "HARDMAX"/"HARDLIM", etc.). >> Do you envision them to exist within info/MB/resource_schemata/ as well as within >> info/MB_NODE/resource_schemata/? > > Yes, at least for MIN, see the info/ tree above. For HARDLIM, perhaps, but HARDLIM has the added > complications that it is a property of the MBW_MAX control and that it may be configurable for each > PARTID or a fixed property of the h/w. When HARDLIM is configurable the control name could be of the > form ___ where is HARDLIM and the full > name for the HARDLIM configuration on the MB_NODE resource is MB_NODE_MAX_HARDLIM. There can also be > an info//resource_schemata//lim file which has values, soft, hard, configurable. > > For CMAX, maximum cache capacity, there is an equivalent control SOFTLIM, which behaves as HARDLIM > except the meaning of the bit is reversed. We can just use a consistent name in s/w though. > Is it possible to view MBW_MAX hard limit, CMAX soft limit, or future similar things as a "configuration" for a control, instead of a "control" itself? A "configurations" is different from a "control" in that: 1. The configuration configures the control, e.g. toggle hard limit on MB control. 2. The configuration doesn't have the control's properties (e.g. gran). 3. The configuration is similar to "event_configs" in MB_MON, that can also configure MB_MON's event_filter. MBW_MAX hard limit configures MB/MB_MAX control's hard limit. CMAX soft limit configures L3 control's soft limit. Is it OK to implement configurations in a control? Then in struct resctrl_ctrl, add "list_head configs" to track all configurations in this control. $ cat schemata MB_NODE:0=100;1=100;2=100;10=100;18=100;26=100;34=100;35=100 # MB_MAXHLIM_NODE is a configuration of MB_NODE control MB_MAXHLIM_NODE:0=1;1=0;2=0;10=0;18=0;26=0;34=0;35=0 MB:0=100;1=100;2=100;10=100;18=100;26=100;34=100;35=100 # MB_MAXHLIM is a configuration of MB control MB_MAXHLIM:0=1;1=0;2=0;10=0;18=0;26=0;34=0;35=0 L3:1=ffff;2=ffff info/ │   └── resource_schemata │   ├── MB (simulated by MB_NODE) │   │   ├── configs # all configs in MB control │   │   │   └── MB_MAXHLIM # MB_MAXHLIM config in MB control │   │   │   └── type # boolean │   │   ├── max │   │   ├── MB_NODE # simulated MB │   │   │   ├── configs # all configs in MB_NODE control │   │   │   │   └── MB_MAXHLIM_NODE # MB_MAXHLIM config in MB_NODE │   │   │   │   └── type │   │   │   ├── max │   │   │   ├── min │   │   │   ├── resolution │   │   │   ├── scale │   │   │   ├── scope │   │   │   ├── status │   │   │   ├── tolerance │   │   │   ├── type │   │   │   └── unit │   │   ├── min │   │   ├── resolution │   │   ├── scale │   │   ├── scope │   │   ├── status │   │   ├── tolerance │   │   ├── type │   │   └── unit │   └── mode >> >>> >>> On an x86 system: >>> >>> info >>> ├── MB >>> │   ├── resource_schemata >>> │   │   ├── MB >>> │   │   └── MB_MAX (Finer grained MB, more below) >>> │   └── scope >>> └── MB_REGION >>> ├── resource_schemata >>> │   └── MB_REGION >>> └── scope >>> >>> >>> Do you think this helps? >> >> "REGION" is not a new scope but instead region-aware MBA is controlled and manages bandwidth at L3 scope. >> Combine that with up to (currently) four regions each with three controls I find an interface like above >> potentially confusing to document in an intuitive way. Unless you are perhaps saying that we should introduce >> a new separate "MB_REGION" L3 scope resource (so let resctrl support multiple "MB" resources at the same >> scope?) and then *it* contains the twelve new controls within its resource_schemata directory? >> >> Since the region-aware controls are orthogonal to the MSR based legacy control resctrl would still need a >> way for user space to switch from one to the other which implies a dependency between "MB" and "MB_REGION" >> that is not presented in above hierarchy. > > OK. The /sys/fs/resctrl/info/MB/schemata/mode and the MB_REGION controls a child of MB you described > previously seem s better fit than what I suggested. > >> >> I think I am missing quite a bit here as I try to navigate an interface so different from what we have >> discussed so far. >> I would like to explore with more detail how this interface can handle the different scenarios we have >> discussed so far. > > Certainly, I don't think we have got to the bottom of this yet. > >> >> The other x86 feature to consider is AMD's upcoming "Global" MBA/SMBA that exposes memory bandwidth allocation >> in "groups of L3" that I understand could usually be mapped to NODE scope (but it remains controlled at L3 scope), >> except for one configuration where it is "SYSTEM"(?) scope. >> Ref.: https://lore.kernel.org/lkml/8f77f498b1c77fa8fd8f5d5687f03ae598068544.1776980182.git.babu.moger@amd.com/ > > Hmmm, I'm not sure that the scope can be considered to be NODE scope for GMBA. To me it seems to be > accidental that it maps to the NUMA node but really the scope is just a grouping of L3 instances. > For a control to NUMA scope I would expect the resctrl domains to go offline and online in sync with > the NUMA nodes. For GMBA it looks like it would just going offline/online based on whether any of > the CPUs and so L3 instances in the group are online. Am I correct here? > > Assuming the domains are on L3 groups rather than NUMA also changes which end of the link the > traffic is regulated and so how cross-NUMA traffic behaves differently. If the domain is an L3 group > then a task running on a CPU affine to that L3 group won't be throttled unless that particular > domain is throttled but with NUMA node domains it may be throttled if it has traffic going to that > domain. >> >>> >>> This also brings another question. On MPAM systems the 'MB_MAX' is backed by the same MSC h/w as MB >>> but it exposed a different interface to the user. If I understand correctly intel have an option to >>> have finer grained control of MB (delay) as well and so it would make sense to use a common name >>> rather than just going for the MPAM centric name of MB_MAX. >> >> Apologies but I was not able to parse above. > > Ok, let me try to explain again (although it's probably not what we want to do). My intent here was > to try and explore whether we can reuse naming and controls across architectures in the same way we > already have for the L2/L3 cache portion bitmap and the existing MB control. > > To quote from a previous mail of yours: > https://lore.kernel.org/lkml/a84af037-6439-4362-be07-d45143e06309@intel.com/ > """ > For example, on an MPAM system (if I understand correctly) the user may see: > info/ > └── MB/ > └── resource_schemata/ > ├── MB/ > │ └── MB_MAX/ > └── MB_MIN/ > > Compared with a possible implementation on Intel that looks like: > info/ > └── MB/ > └── resource_schemata/ > ├── MB/ > │ └── MB_OPT/ > ├── MB_MAX/ > └── MB_MIN/ > """ > > In the two setups MPAM MB_MAX and intel MB_OPT play the same role, a finer grained control of the > legacy MB control. I was thinking these could share a name (MB-PRECISE), but it probably doesn't > make sense as MB_OPT and MB_MAX have different relationships to MB_MIN. > > Thanks, > > Ben > >> >> Reinette >> > > Thanks. -Fenghua