mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Reinette Chatre <reinette.chatre@intel.com>
To: "Luck, Tony" <tony.luck@intel.com>
Cc: Fenghua Yu <fenghuay@nvidia.com>,
	Maciej Wieczor-Retman <maciej.wieczor-retman@intel.com>,
	Peter Newman <peternewman@google.com>,
	James Morse <james.morse@arm.com>,
	Babu Moger <babu.moger@amd.com>,
	"Drew Fustini" <dfustini@baylibre.com>,
	Dave Martin <Dave.Martin@arm.com>, Chen Yu <yu.c.chen@intel.com>,
	David E Box <david.e.box@intel.com>, <x86@kernel.org>,
	Christoph Hellwig <hch@infradead.org>,
	<linux-kernel@vger.kernel.org>, <patches@lists.linux.dev>
Subject: Re: [PATCH v9 06/12] arm,x86,fs/resctrl: Handle change in number of RMIDs on each mount
Date: Tue, 14 Jul 2026 14:07:17 -0700	[thread overview]
Message-ID: <1dac7a0e-8b69-4b4f-81e1-c010d186e69d@intel.com> (raw)
In-Reply-To: <alacCy4Ja4gF97PU@agluck-desk3>

Hi Tony,

On 7/14/26 1:28 PM, Luck, Tony wrote:
> Hi Reinette,
> 
> On Tue, Jul 14, 2026 at 11:16:53AM -0700, Reinette Chatre wrote:
>> Hi Tony,
>>
>> On 7/14/26 10:50 AM, Luck, Tony wrote:
>>> On Tue, Jul 14, 2026 at 09:45:04AM -0700, Reinette Chatre wrote:
>>>> Hi Tony,
>>>
>>> ... trimming to open issue ...
>>>
>>>>>> How much to rely on CPUID is not clear to me. The direction seems to
>>>>>> be to move away from CPUID, which makes adding new CPUID dependencies
>>>>>> less ideal? 
>>>
>>> Where is this direction to move away from CPUID coming from?
>>
>> This is the impression I got from the discussions during series that adds support for
>> LLC occupancy monitoring via ERDT (https://lore.kernel.org/lkml/cover.1782866200.git.yu.c.chen@intel.com/).
>> If ERDT is the future direction then I do not think resctrl should assume that the
>> same data will always be available via CPUID also. With ERDT there is a separate
>> per-domain maximum RMID that should be discovered via ACPI.
> 
> My view of system design today is that architects have a box full of Lego(TM)
> bricks for features that they clip together. The bricks often have different
> parameters - sized based on the maximum values across all systems where they
> might be used.
> 
> In this case you end up with different components supporting different
> numbers of RMIDs.

ack.

> 
> For RMID the absolute maximum usable value is defined by CPUID(0xF).EBX.

Would this be the case even if CPUID(0xF).EDX == 0?

> A WRMSR to IA32_PQR_ASSOC.RMID of any value greater than that will #GP
> fault. So, while you might have an ERDT table saying that it supports 500
> RMIDs, if that Lego brick is plugged into an SoC that has CPUID(0xF).EBX =
> 399, then you can only use 400 RMIDs, and the extra 100 counters on the ERDT
> device will sit unused as there is no way for them to be accessed.
> 
>>>>> Here's my current work-in-progress version:
>>>>>
>>>>> u32 resctrl_arch_system_max_rmid_idx(void)
>>>>> {
>>>>> 	struct rdt_resource *r = &rdt_resources_all[RDT_RESOURCE_L3].r_resctrl;
>>>>> 	u32 ret;
>>>>>
>>>>> 	/* CPUID provides maximum possible RMID value */
>>>>> 	ret = cpuid_ebx(0xf) + 1;
>>>>
>>>> There is also boot_cpu_data.x86_cache_max_rmid initialized in resctrl_cpu_detect() that
>>>> could be used directly? Looks like resctrl_cpu_detect() already scales the number of
>>>> RMID down if needed when L3 monitoring is supported, but it does not take SNC into account.
>>>
>>> I looked at that, but rejected it because it isn't adjusted for SNC.
>>
>> Right. Neither is cpuid_ebx(0xf). Is it necessary for resctrl_arch_system_max_rmid_idx()
>> to call CPUID every time or could it use the cached value in boot_cpu_data.x86_cache_max_rmid?
> 
> Yes. Necessary. See how boot_cpu_data.x86_cache_max_rmid is derived:
> 
> /* Runs once on the BSP during boot. */
> void resctrl_cpu_detect(struct cpuinfo_x86 *c)
> {
> 	if (!cpu_has(c, X86_FEATURE_CQM_LLC) && !cpu_has(c, X86_FEATURE_ABMC)) {
> 		c->x86_cache_max_rmid  = -1;
> 		c->x86_cache_occ_scale = -1;
> 		c->x86_cache_mbm_width_offset = -1;
> 		return;
> 	}
> 
> On a theoretical platform that doesn't support X86_FEATURE_CQM_LLC or X86_FEATURE_ABMC
> x86_cache_max_rmid will be -1.
> 
> In fact you can get this with boot argument of "clearcpuid=cqm_occup_llc"
> 
> I wonder if that test should really be for X86_FEATURE_CQM and thus
> check CPUID(0x7).EBX[12]?

It seems that doing so would match the definition of IA32_PQR_ASSOC found below that
you highlighted earlier and should result in a valid x86_cache_max_rmid as intended?
I am not familiar with the history of the original enabling
(commit cbc82b172638 ("x86: Add support for Intel Cache QoS Monitoring (CQM) detection"))
that resulted in X86_FEATURE_CQM not used.


> 
>>>>> 	/*
>>>>> 	 * if system is capable of L3 monitoring the maximum RMID value may
>>>>> 	 * be lower that system maximum. Either because the L3 monitoring
>>>>> 	 * feature supports fewer RMIDs (CPUID(0xF, 0x1).ECX), or because SNC
>>>>> 	 * (Sub-NUMA Cluster) is enabled and divides RMIDs per cluster.
>>>>> 	 */
>>>>> 	if (r->mon_capable)
>>>>> 		ret = r->mon.num_rmid;
>>>>>
>>>>> 	return ret;
>>>>> }
>>>>>
>>>>> CPUID seems unavoidable in the case that the platform doesn't support
>>>>> (or has disabled the various L3 monitoring events). In that case using
>>>>
>>>> Since L3 monitoring is the only CPUID supported monitoring resource I do not
>>>> think cpuid_ebx(0xf) can be used if the platform does not support L3 monitoring.
>>>> I expect that would mean that leaf 0x7 would indicate that monitoring is
>>>> not supported which would make leaf 0xf invalid? Interestingly resctrl does
>>>> not seem to consider X86_FEATURE_CQM from leaf 0x7 at all and just goes straight
>>>> to leaf 0xf.
>>>
>>> I disagree. The definition of IA32_PQR_ASSOC says:
>>>
>>> 1) The MSR exists if either of CPUID.07H.00H:EBX[12] or CPUID.07H.00H:EBX[15] is set
>>>    (in this case bit 12 for monitoring, It does appear to be a bug that resctrl is
>>>    not checking bit 12 before looking at leaf 0xF).
>>>
>>> 2) The supported width of the RMID field is Ceil(Log2 (CPUID.0FH.00H:EBX[31:0] +1))
>>>
>>> So I believe that a theoretical system that supported AET (or some other
>>> monitoring) but didn't support L3 monitoring, would have to set CPUID.07H.00H:EBX[12]
>>> and provide the max RMID value in CPUID.0FH.00H:EBX[31:0].
>>
>> I see. That would mean that CPUID would report a max RMID value while also reporting
>> that there are no resources being monitored. I assume ERDT is an example of "some other
>> monitoring" and it has its own per-domain max RMID that is separate from the different
>> monitoring features. How do these different max RMID values relate? Should this
>> architectural "get the max RMID" be expanded to also consider the per-RMDD maximum when
>> ERDT support lands?
> 
> Yes, ERDT is an example of a monitoring resource enumerated by ACPI.
> 
> AET is enumerated via a combination of PCIe configuration space VSEC
> entries and an XML file.
> 
> See above "Lego" explanation for why they don't factor into the "max RMID"
> calculation.
> 
>>>
>>> For resctrl we have the practical case of booting with "rdt=!cmt,!mbmtotal,!mbmlocal"
>>> to software disable all L3 monitoring.
>> ack.
>>
>> Reinette
>>
> 
> -Tony

Reinette

  reply	other threads:[~2026-07-14 21:07 UTC|newest]

Thread overview: 43+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-07-01 21:35 [PATCH v9 00/12] Allow AET to use PMT as loadable module Tony Luck
2026-07-01 21:35 ` [PATCH v9 01/12] platform/x86/intel/{pmt,vsec}: Prevent unbind via sysfs Tony Luck
2026-07-08 22:45   ` Reinette Chatre
2026-07-09 17:41     ` Luck, Tony
2026-07-09 20:48       ` Reinette Chatre
2026-07-09 21:12         ` Luck, Tony
2026-07-10 17:01           ` Luck, Tony
2026-07-10 20:24             ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 02/12] fs/resctrl: Remove redundant calls to resctrl_arch_mon_capable() Tony Luck
2026-07-01 21:35 ` [PATCH v9 03/12] x86/resctrl: Honor rdt=perf option to force enable AET perf events Tony Luck
2026-07-08 22:46   ` Reinette Chatre
2026-07-10 20:29     ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 04/12] fs/resctrl: Add interface to disable a monitor event Tony Luck
2026-07-01 21:35 ` [PATCH v9 05/12] x86/resctrl: Drop global 'rdt_mon_capable' flag Tony Luck
2026-07-08 22:47   ` Reinette Chatre
2026-07-10 20:31     ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 06/12] arm,x86,fs/resctrl: Handle change in number of RMIDs on each mount Tony Luck
2026-07-08 22:50   ` Reinette Chatre
2026-07-10 20:51     ` Luck, Tony
2026-07-14 15:18       ` Reinette Chatre
2026-07-14 16:01         ` Luck, Tony
2026-07-14 16:45           ` Reinette Chatre
2026-07-14 17:50             ` Luck, Tony
2026-07-14 18:16               ` Reinette Chatre
2026-07-14 20:28                 ` Luck, Tony
2026-07-14 21:07                   ` Reinette Chatre [this message]
2026-07-14 22:14                     ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 07/12] x86/resctrl: Add PMT registration API for AET enumeration callbacks Tony Luck
2026-07-08 22:51   ` Reinette Chatre
2026-07-10 20:54     ` Luck, Tony
2026-07-14 15:18       ` Reinette Chatre
2026-07-14 15:41         ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 08/12] platform/x86/intel/pmt: Register enumeration functions with resctrl Tony Luck
2026-07-01 21:35 ` [PATCH v9 09/12] arm,x86/resctrl: Resolve INTEL_PMT_TELEMETRY symbols at runtime Tony Luck
2026-07-08 22:52   ` Reinette Chatre
2026-07-10 20:59     ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 10/12] fs/resctrl: Call architecture hooks for every mount/unmount Tony Luck
2026-07-08 22:53   ` Reinette Chatre
2026-07-10 21:01     ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 11/12] x86/resctrl: Simplify Kconfig options for resctrl Tony Luck
2026-07-08 23:01   ` Reinette Chatre
2026-07-10 21:08     ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 12/12] Documentation/filesystems/resctrl: Document telemetry mount timing caveat Tony Luck

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=1dac7a0e-8b69-4b4f-81e1-c010d186e69d@intel.com \
    --to=reinette.chatre@intel.com \
    --cc=Dave.Martin@arm.com \
    --cc=babu.moger@amd.com \
    --cc=david.e.box@intel.com \
    --cc=dfustini@baylibre.com \
    --cc=fenghuay@nvidia.com \
    --cc=hch@infradead.org \
    --cc=james.morse@arm.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=maciej.wieczor-retman@intel.com \
    --cc=patches@lists.linux.dev \
    --cc=peternewman@google.com \
    --cc=tony.luck@intel.com \
    --cc=x86@kernel.org \
    --cc=yu.c.chen@intel.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome