From: Reinette Chatre <reinette.chatre@intel.com>
To: "Luck, Tony" <tony.luck@intel.com>
Cc: Fenghua Yu <fenghuay@nvidia.com>,
Maciej Wieczor-Retman <maciej.wieczor-retman@intel.com>,
Peter Newman <peternewman@google.com>,
James Morse <james.morse@arm.com>,
Babu Moger <babu.moger@amd.com>,
"Drew Fustini" <dfustini@baylibre.com>,
Dave Martin <Dave.Martin@arm.com>, Chen Yu <yu.c.chen@intel.com>,
David E Box <david.e.box@intel.com>, <x86@kernel.org>,
Christoph Hellwig <hch@infradead.org>,
<linux-kernel@vger.kernel.org>, <patches@lists.linux.dev>
Subject: Re: [PATCH v9 06/12] arm,x86,fs/resctrl: Handle change in number of RMIDs on each mount
Date: Tue, 14 Jul 2026 14:07:17 -0700 [thread overview]
Message-ID: <1dac7a0e-8b69-4b4f-81e1-c010d186e69d@intel.com> (raw)
In-Reply-To: <alacCy4Ja4gF97PU@agluck-desk3>
Hi Tony,
On 7/14/26 1:28 PM, Luck, Tony wrote:
> Hi Reinette,
>
> On Tue, Jul 14, 2026 at 11:16:53AM -0700, Reinette Chatre wrote:
>> Hi Tony,
>>
>> On 7/14/26 10:50 AM, Luck, Tony wrote:
>>> On Tue, Jul 14, 2026 at 09:45:04AM -0700, Reinette Chatre wrote:
>>>> Hi Tony,
>>>
>>> ... trimming to open issue ...
>>>
>>>>>> How much to rely on CPUID is not clear to me. The direction seems to
>>>>>> be to move away from CPUID, which makes adding new CPUID dependencies
>>>>>> less ideal?
>>>
>>> Where is this direction to move away from CPUID coming from?
>>
>> This is the impression I got from the discussions during series that adds support for
>> LLC occupancy monitoring via ERDT (https://lore.kernel.org/lkml/cover.1782866200.git.yu.c.chen@intel.com/).
>> If ERDT is the future direction then I do not think resctrl should assume that the
>> same data will always be available via CPUID also. With ERDT there is a separate
>> per-domain maximum RMID that should be discovered via ACPI.
>
> My view of system design today is that architects have a box full of Lego(TM)
> bricks for features that they clip together. The bricks often have different
> parameters - sized based on the maximum values across all systems where they
> might be used.
>
> In this case you end up with different components supporting different
> numbers of RMIDs.
ack.
>
> For RMID the absolute maximum usable value is defined by CPUID(0xF).EBX.
Would this be the case even if CPUID(0xF).EDX == 0?
> A WRMSR to IA32_PQR_ASSOC.RMID of any value greater than that will #GP
> fault. So, while you might have an ERDT table saying that it supports 500
> RMIDs, if that Lego brick is plugged into an SoC that has CPUID(0xF).EBX =
> 399, then you can only use 400 RMIDs, and the extra 100 counters on the ERDT
> device will sit unused as there is no way for them to be accessed.
>
>>>>> Here's my current work-in-progress version:
>>>>>
>>>>> u32 resctrl_arch_system_max_rmid_idx(void)
>>>>> {
>>>>> struct rdt_resource *r = &rdt_resources_all[RDT_RESOURCE_L3].r_resctrl;
>>>>> u32 ret;
>>>>>
>>>>> /* CPUID provides maximum possible RMID value */
>>>>> ret = cpuid_ebx(0xf) + 1;
>>>>
>>>> There is also boot_cpu_data.x86_cache_max_rmid initialized in resctrl_cpu_detect() that
>>>> could be used directly? Looks like resctrl_cpu_detect() already scales the number of
>>>> RMID down if needed when L3 monitoring is supported, but it does not take SNC into account.
>>>
>>> I looked at that, but rejected it because it isn't adjusted for SNC.
>>
>> Right. Neither is cpuid_ebx(0xf). Is it necessary for resctrl_arch_system_max_rmid_idx()
>> to call CPUID every time or could it use the cached value in boot_cpu_data.x86_cache_max_rmid?
>
> Yes. Necessary. See how boot_cpu_data.x86_cache_max_rmid is derived:
>
> /* Runs once on the BSP during boot. */
> void resctrl_cpu_detect(struct cpuinfo_x86 *c)
> {
> if (!cpu_has(c, X86_FEATURE_CQM_LLC) && !cpu_has(c, X86_FEATURE_ABMC)) {
> c->x86_cache_max_rmid = -1;
> c->x86_cache_occ_scale = -1;
> c->x86_cache_mbm_width_offset = -1;
> return;
> }
>
> On a theoretical platform that doesn't support X86_FEATURE_CQM_LLC or X86_FEATURE_ABMC
> x86_cache_max_rmid will be -1.
>
> In fact you can get this with boot argument of "clearcpuid=cqm_occup_llc"
>
> I wonder if that test should really be for X86_FEATURE_CQM and thus
> check CPUID(0x7).EBX[12]?
It seems that doing so would match the definition of IA32_PQR_ASSOC found below that
you highlighted earlier and should result in a valid x86_cache_max_rmid as intended?
I am not familiar with the history of the original enabling
(commit cbc82b172638 ("x86: Add support for Intel Cache QoS Monitoring (CQM) detection"))
that resulted in X86_FEATURE_CQM not used.
>
>>>>> /*
>>>>> * if system is capable of L3 monitoring the maximum RMID value may
>>>>> * be lower that system maximum. Either because the L3 monitoring
>>>>> * feature supports fewer RMIDs (CPUID(0xF, 0x1).ECX), or because SNC
>>>>> * (Sub-NUMA Cluster) is enabled and divides RMIDs per cluster.
>>>>> */
>>>>> if (r->mon_capable)
>>>>> ret = r->mon.num_rmid;
>>>>>
>>>>> return ret;
>>>>> }
>>>>>
>>>>> CPUID seems unavoidable in the case that the platform doesn't support
>>>>> (or has disabled the various L3 monitoring events). In that case using
>>>>
>>>> Since L3 monitoring is the only CPUID supported monitoring resource I do not
>>>> think cpuid_ebx(0xf) can be used if the platform does not support L3 monitoring.
>>>> I expect that would mean that leaf 0x7 would indicate that monitoring is
>>>> not supported which would make leaf 0xf invalid? Interestingly resctrl does
>>>> not seem to consider X86_FEATURE_CQM from leaf 0x7 at all and just goes straight
>>>> to leaf 0xf.
>>>
>>> I disagree. The definition of IA32_PQR_ASSOC says:
>>>
>>> 1) The MSR exists if either of CPUID.07H.00H:EBX[12] or CPUID.07H.00H:EBX[15] is set
>>> (in this case bit 12 for monitoring, It does appear to be a bug that resctrl is
>>> not checking bit 12 before looking at leaf 0xF).
>>>
>>> 2) The supported width of the RMID field is Ceil(Log2 (CPUID.0FH.00H:EBX[31:0] +1))
>>>
>>> So I believe that a theoretical system that supported AET (or some other
>>> monitoring) but didn't support L3 monitoring, would have to set CPUID.07H.00H:EBX[12]
>>> and provide the max RMID value in CPUID.0FH.00H:EBX[31:0].
>>
>> I see. That would mean that CPUID would report a max RMID value while also reporting
>> that there are no resources being monitored. I assume ERDT is an example of "some other
>> monitoring" and it has its own per-domain max RMID that is separate from the different
>> monitoring features. How do these different max RMID values relate? Should this
>> architectural "get the max RMID" be expanded to also consider the per-RMDD maximum when
>> ERDT support lands?
>
> Yes, ERDT is an example of a monitoring resource enumerated by ACPI.
>
> AET is enumerated via a combination of PCIe configuration space VSEC
> entries and an XML file.
>
> See above "Lego" explanation for why they don't factor into the "max RMID"
> calculation.
>
>>>
>>> For resctrl we have the practical case of booting with "rdt=!cmt,!mbmtotal,!mbmlocal"
>>> to software disable all L3 monitoring.
>> ack.
>>
>> Reinette
>>
>
> -Tony
Reinette
next prev parent reply other threads:[~2026-07-14 21:07 UTC|newest]
Thread overview: 43+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-07-01 21:35 [PATCH v9 00/12] Allow AET to use PMT as loadable module Tony Luck
2026-07-01 21:35 ` [PATCH v9 01/12] platform/x86/intel/{pmt,vsec}: Prevent unbind via sysfs Tony Luck
2026-07-08 22:45 ` Reinette Chatre
2026-07-09 17:41 ` Luck, Tony
2026-07-09 20:48 ` Reinette Chatre
2026-07-09 21:12 ` Luck, Tony
2026-07-10 17:01 ` Luck, Tony
2026-07-10 20:24 ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 02/12] fs/resctrl: Remove redundant calls to resctrl_arch_mon_capable() Tony Luck
2026-07-01 21:35 ` [PATCH v9 03/12] x86/resctrl: Honor rdt=perf option to force enable AET perf events Tony Luck
2026-07-08 22:46 ` Reinette Chatre
2026-07-10 20:29 ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 04/12] fs/resctrl: Add interface to disable a monitor event Tony Luck
2026-07-01 21:35 ` [PATCH v9 05/12] x86/resctrl: Drop global 'rdt_mon_capable' flag Tony Luck
2026-07-08 22:47 ` Reinette Chatre
2026-07-10 20:31 ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 06/12] arm,x86,fs/resctrl: Handle change in number of RMIDs on each mount Tony Luck
2026-07-08 22:50 ` Reinette Chatre
2026-07-10 20:51 ` Luck, Tony
2026-07-14 15:18 ` Reinette Chatre
2026-07-14 16:01 ` Luck, Tony
2026-07-14 16:45 ` Reinette Chatre
2026-07-14 17:50 ` Luck, Tony
2026-07-14 18:16 ` Reinette Chatre
2026-07-14 20:28 ` Luck, Tony
2026-07-14 21:07 ` Reinette Chatre [this message]
2026-07-14 22:14 ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 07/12] x86/resctrl: Add PMT registration API for AET enumeration callbacks Tony Luck
2026-07-08 22:51 ` Reinette Chatre
2026-07-10 20:54 ` Luck, Tony
2026-07-14 15:18 ` Reinette Chatre
2026-07-14 15:41 ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 08/12] platform/x86/intel/pmt: Register enumeration functions with resctrl Tony Luck
2026-07-01 21:35 ` [PATCH v9 09/12] arm,x86/resctrl: Resolve INTEL_PMT_TELEMETRY symbols at runtime Tony Luck
2026-07-08 22:52 ` Reinette Chatre
2026-07-10 20:59 ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 10/12] fs/resctrl: Call architecture hooks for every mount/unmount Tony Luck
2026-07-08 22:53 ` Reinette Chatre
2026-07-10 21:01 ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 11/12] x86/resctrl: Simplify Kconfig options for resctrl Tony Luck
2026-07-08 23:01 ` Reinette Chatre
2026-07-10 21:08 ` Luck, Tony
2026-07-01 21:35 ` [PATCH v9 12/12] Documentation/filesystems/resctrl: Document telemetry mount timing caveat Tony Luck
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1dac7a0e-8b69-4b4f-81e1-c010d186e69d@intel.com \
--to=reinette.chatre@intel.com \
--cc=Dave.Martin@arm.com \
--cc=babu.moger@amd.com \
--cc=david.e.box@intel.com \
--cc=dfustini@baylibre.com \
--cc=fenghuay@nvidia.com \
--cc=hch@infradead.org \
--cc=james.morse@arm.com \
--cc=linux-kernel@vger.kernel.org \
--cc=maciej.wieczor-retman@intel.com \
--cc=patches@lists.linux.dev \
--cc=peternewman@google.com \
--cc=tony.luck@intel.com \
--cc=x86@kernel.org \
--cc=yu.c.chen@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome