From: Ben Horgan <ben.horgan@arm.com>
To: Andre Przywara <andre.przywara@arm.com>,
Lorenzo Pieralisi <lpieralisi@kernel.org>,
Hanjun Guo <guohanjun@huawei.com>,
Sudeep Holla <sudeep.holla@kernel.org>,
Catalin Marinas <catalin.marinas@arm.com>,
Will Deacon <will@kernel.org>,
"Rafael J . Wysocki" <rafael@kernel.org>,
Len Brown <lenb@kernel.org>, James Morse <james.morse@arm.com>,
Reinette Chatre <reinette.chatre@intel.com>,
Fenghua Yu <fenghuay@nvidia.com>
Cc: Jonathan Cameron <jic23@kernel.org>,
Srivathsa L Rao <srivathsa.rao@oss.qualcomm.com>,
Ganapatrao Kulkarni <ganapatrao.kulkarni@oss.qualcomm.com>,
Trilok Soni <tsoni@quicinc.com>,
Srinivas Ramana <sramana@qti.qualcomm.com>,
Niyas Sait <niyas.sait@arm.com>, Lee Trager <lee@trager.us>,
Ritwick Sharma <ritwick.sharma@arm.com>,
Gavin Shan <gshan@redhat.com>,
linux-acpi@vger.kernel.org, linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH v11 02/13] arm_mpam: propagate MSC access errors for hw_probe functions
Date: Wed, 30 Sep 2026 15:01:46 +0100 [thread overview]
Message-ID: <4cc35ee7-c4d3-40c0-b1f9-24ef5a2f5844@arm.com> (raw)
In-Reply-To: <20260924152739.2865510-3-andre.przywara@arm.com>
Hi Andre,
On 24/09/2026 16:27, Andre Przywara wrote:
> Allow the functions probing for MSC hardware and features to return an
> error, and propagate read and write errors from the lower level up.
> This uses some "scoped cleanup" functions like scoped_guard() and
> ACQUIRE() to avoid the complexity of error handling when a lock has been
> taken. Since the mon_sel_lock is a bit special (even more so in an
> upcoming patch), we define a new GUARD type for it.
>
> Signed-off-by: Andre Przywara <andre.przywara@arm.com>
> Reviewed-by: Jonathan Cameron <jonathan.cameron@oss.qualcomm.com>
> Reviewed-by: Ben Horgan <ben.horgan@arm.com>
> Reviewed-by: Gavin Shan <gshan@redhat.com>
> Tested-by: Gavin Shan <gshan@redhat.com> # on NVIDIA Grace Hopper
> ---
> drivers/resctrl/mpam_devices.c | 126 +++++++++++++++++++++-----------
> drivers/resctrl/mpam_internal.h | 4 +
> 2 files changed, 87 insertions(+), 43 deletions(-)
>
> diff --git a/drivers/resctrl/mpam_devices.c b/drivers/resctrl/mpam_devices.c
> index c7686d3799653..22069e7f7bb05 100644
> --- a/drivers/resctrl/mpam_devices.c
> +++ b/drivers/resctrl/mpam_devices.c
> @@ -798,38 +798,46 @@ static bool mpam_ris_hw_probe_csu_nrdy(struct mpam_msc_ris *ris)
> bool can_set, can_clear;
> struct mpam_msc *msc = ris->vmsc->msc;
>
> - if (WARN_ON_ONCE(!mpam_mon_sel_lock(msc)))
> + ACQUIRE(mon_sel_lock, guard)(msc);
> + if (ACQUIRE_ERR(mon_sel_lock, &guard))
> return false;
>
> mon_sel = FIELD_PREP(MSMON_CFG_MON_SEL_MON_SEL, 0) |
> FIELD_PREP(MSMON_CFG_MON_SEL_RIS, ris->ris_idx);
> - mpam_write_monsel_reg(msc, CFG_MON_SEL, mon_sel);
> + if (mpam_write_monsel_reg(msc, CFG_MON_SEL, mon_sel))
> + return false;
>
> /* Hardware might ignore nrdy if it's not enabled */
> ctl_val = MSMON_CFG_CSU_CTL_TYPE_CSU;
> ctl_val |= MSMON_CFG_x_CTL_MATCH_PARTID;
> ctl_val |= MSMON_CFG_x_CTL_MATCH_PMG;
> ctl_val |= MSMON_CFG_x_CTL_EN;
> - mpam_write_monsel_reg(msc, CFG_CSU_FLT, 0);
> - mpam_write_monsel_reg(msc, CFG_CSU_CTL, ctl_val);
> + if (mpam_write_monsel_reg(msc, CFG_CSU_FLT, 0))
> + return false;
> + if (mpam_write_monsel_reg(msc, CFG_CSU_CTL, ctl_val))
> + return false;
>
> - _mpam_write_monsel_reg(msc, MSMON_CSU, MSMON___NRDY);
> - _mpam_read_monsel_reg(msc, MSMON_CSU, &now);
> + if (_mpam_write_monsel_reg(msc, MSMON_CSU, MSMON___NRDY))
> + return false;
> + if (_mpam_read_monsel_reg(msc, MSMON_CSU, &now))
> + return false;
> can_set = now & MSMON___NRDY;
>
> - _mpam_write_monsel_reg(msc, MSMON_CSU, 0);
> + if (_mpam_write_monsel_reg(msc, MSMON_CSU, 0))
> + return false;
> /* Configuration change to try and coax hardware into setting nrdy */
> - mpam_write_monsel_reg(msc, CFG_CSU_FLT, 0x1);
> - _mpam_read_monsel_reg(msc, MSMON_CSU, &now);
> + if (mpam_write_monsel_reg(msc, CFG_CSU_FLT, 0x1))
> + return false;
> + if (_mpam_read_monsel_reg(msc, MSMON_CSU, &now))
> + return false;
> can_clear = !(now & MSMON___NRDY);
> - mpam_mon_sel_unlock(msc);
>
> return (!can_set || !can_clear);
> }
I've had a look at the Sashiko reports [1] for these patches. I'll call out the things that I feel
need answering.
As sashiko says, mpam_ris_hw_probe_csu_nrdy(), this swallows the error rather than propagating. It
would be better if this returned success/error and the answer to whether or not the h/w supports
nrdy separately. As we make other accesses to the msc later it is unlikely this will ever cause a
failing probe to report success though.
Thanks,
Ben
[1] https://sashiko.dev/#/patchset/20260924152739.2865510-1-andre.przywara%40arm.com
>
> -static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> +static int mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> {
> - int err;
> + int fw_has_nrdy_err, err;
> struct mpam_msc *msc = ris->vmsc->msc;
> struct device *dev = &msc->pdev->dev;
> struct mpam_props *props = &ris->props;
> @@ -842,7 +850,9 @@ static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> if (FIELD_GET(MPAMF_IDR_HAS_CCAP_PART, ris->idr)) {
> u32 ccap_features;
>
> - mpam_read_partsel_reg(msc, CCAP_IDR, &ccap_features);
> + err = mpam_read_partsel_reg(msc, CCAP_IDR, &ccap_features);
> + if (err)
> + return err;
>
> props->cmax_wd = FIELD_GET(MPAMF_CCAP_IDR_CMAX_WD, ccap_features);
> if (props->cmax_wd &&
> @@ -867,7 +877,9 @@ static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> if (FIELD_GET(MPAMF_IDR_HAS_CPOR_PART, ris->idr)) {
> u32 cpor_features;
>
> - mpam_read_partsel_reg(msc, CPOR_IDR, &cpor_features);
> + err = mpam_read_partsel_reg(msc, CPOR_IDR, &cpor_features);
> + if (err)
> + return err;
>
> props->cpbm_wd = FIELD_GET(MPAMF_CPOR_IDR_CPBM_WD, cpor_features);
> if (props->cpbm_wd)
> @@ -877,8 +889,9 @@ static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> /* Memory bandwidth partitioning */
> if (FIELD_GET(MPAMF_IDR_HAS_MBW_PART, ris->idr)) {
> u32 mbw_features;
> -
> - mpam_read_partsel_reg(msc, MBW_IDR, &mbw_features);
> + err = mpam_read_partsel_reg(msc, MBW_IDR, &mbw_features);
> + if (err)
> + return err;
>
> /* portion bitmap resolution */
> props->mbw_pbm_bits = FIELD_GET(MPAMF_MBW_IDR_BWPBM_WD, mbw_features);
> @@ -907,8 +920,9 @@ static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> /* Priority partitioning */
> if (FIELD_GET(MPAMF_IDR_HAS_PRI_PART, ris->idr)) {
> u32 pri_features;
> -
> - mpam_read_partsel_reg(msc, PRI_IDR, &pri_features);
> + err = mpam_read_partsel_reg(msc, PRI_IDR, &pri_features);
> + if (err)
> + return err;
>
> props->intpri_wd = FIELD_GET(MPAMF_PRI_IDR_INTPRI_WD, pri_features);
> if (props->intpri_wd && FIELD_GET(MPAMF_PRI_IDR_HAS_INTPRI, pri_features)) {
> @@ -929,20 +943,24 @@ static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> if (FIELD_GET(MPAMF_IDR_HAS_MSMON, ris->idr)) {
> u32 msmon_features;
>
> - mpam_read_partsel_reg(msc, MSMON_IDR, &msmon_features);
> + err = mpam_read_partsel_reg(msc, MSMON_IDR, &msmon_features);
> + if (err)
> + return err;
>
> /*
> * If the firmware max-nrdy-us property is missing, the
> * CSU counters can't be used. Should we wait forever?
> */
> - err = device_property_read_u32(&msc->pdev->dev,
> - "arm,not-ready-us",
> - &msc->nrdy_usec);
> + fw_has_nrdy_err = device_property_read_u32(&msc->pdev->dev,
> + "arm,not-ready-us",
> + &msc->nrdy_usec);
>
> if (FIELD_GET(MPAMF_MSMON_IDR_MSMON_CSU, msmon_features)) {
> u32 csumonidr;
>
> - mpam_read_partsel_reg(msc, CSUMON_IDR, &csumonidr);
> + err = mpam_read_partsel_reg(msc, CSUMON_IDR, &csumonidr);
> + if (err)
> + return err;
>
> props->num_csu_mon = FIELD_GET(MPAMF_CSUMON_IDR_NUM_MON, csumonidr);
> if (props->num_csu_mon) {
> @@ -960,7 +978,7 @@ static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> * Accept the missing firmware property if NRDY appears
> * un-implemented.
> */
> - if (err && hw_managed)
> + if (fw_has_nrdy_err && hw_managed)
> dev_err_once(dev, "Counters are not usable because not-ready timeout was not provided by firmware.");
> }
> }
> @@ -968,7 +986,9 @@ static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> bool has_long;
> u32 mbwumon_idr;
>
> - mpam_read_partsel_reg(msc, MBWUMON_IDR, &mbwumon_idr);
> + err = mpam_read_partsel_reg(msc, MBWUMON_IDR, &mbwumon_idr);
> + if (err)
> + return err;
>
> props->num_mbwu_mon = FIELD_GET(MPAMF_MBWUMON_IDR_NUM_MON, mbwumon_idr);
> if (props->num_mbwu_mon) {
> @@ -1001,16 +1021,22 @@ static void mpam_ris_hw_probe(struct mpam_msc_ris *ris)
> u16 partid_max;
> u32 nrwidr;
>
> - mpam_read_partsel_reg(msc, PARTID_NRW_IDR, &nrwidr);
> + err = mpam_read_partsel_reg(msc, PARTID_NRW_IDR, &nrwidr);
> + if (err)
> + return err;
> +
> partid_max = FIELD_GET(MPAMF_PARTID_NRW_IDR_INTPARTID_MAX, nrwidr);
>
> mpam_set_feature(mpam_feat_partid_nrw, props);
> msc->partid_max = min(msc->partid_max, partid_max);
> }
> +
> + return 0;
> }
>
> static int mpam_msc_hw_probe(struct mpam_msc *msc)
> {
> + int ret;
> u64 idr;
> u16 partid_max;
> u8 ris_idx, pmg_max;
> @@ -1025,11 +1051,15 @@ static int mpam_msc_hw_probe(struct mpam_msc *msc)
> }
>
> /* Grab an IDR value to find out how many RIS there are */
> - mutex_lock(&msc->part_sel_lock);
> - mpam_msc_read_idr(msc, &idr);
> - mpam_read_partsel_reg(msc, IIDR, &msc->iidr);
> + scoped_guard(mutex, &msc->part_sel_lock) {
> + ret = mpam_msc_read_idr(msc, &idr);
> + if (ret)
> + return ret;
>
> - mutex_unlock(&msc->part_sel_lock);
> + ret = mpam_read_partsel_reg(msc, IIDR, &msc->iidr);
> + if (ret)
> + return ret;
> + }
>
> mpam_enable_quirks(msc);
>
> @@ -1040,10 +1070,15 @@ static int mpam_msc_hw_probe(struct mpam_msc *msc)
> msc->pmg_max = FIELD_GET(MPAMF_IDR_PMG_MAX, idr);
>
> for (ris_idx = 0; ris_idx <= msc->ris_max; ris_idx++) {
> - mutex_lock(&msc->part_sel_lock);
> - __mpam_part_sel(ris_idx, 0, msc);
> - mpam_msc_read_idr(msc, &idr);
> - mutex_unlock(&msc->part_sel_lock);
> + scoped_guard(mutex, &msc->part_sel_lock) {
> + ret = __mpam_part_sel(ris_idx, 0, msc);
> + if (ret)
> + return ret;
> +
> + ret = mpam_msc_read_idr(msc, &idr);
> + if (ret)
> + return ret;
> + }
>
> partid_max = FIELD_GET(MPAMF_IDR_PARTID_MAX, idr);
> pmg_max = FIELD_GET(MPAMF_IDR_PMG_MAX, idr);
> @@ -1051,17 +1086,22 @@ static int mpam_msc_hw_probe(struct mpam_msc *msc)
> msc->pmg_max = min(msc->pmg_max, pmg_max);
> msc->has_extd_esr = FIELD_GET(MPAMF_IDR_HAS_EXTD_ESR, idr);
>
> - mutex_lock(&mpam_list_lock);
> - ris = mpam_get_or_create_ris(msc, ris_idx);
> - mutex_unlock(&mpam_list_lock);
> - if (IS_ERR(ris))
> - return PTR_ERR(ris);
> + scoped_guard(mutex, &mpam_list_lock) {
> + ris = mpam_get_or_create_ris(msc, ris_idx);
> + if (IS_ERR(ris))
> + return PTR_ERR(ris);
> + }
> ris->idr = idr;
>
> - mutex_lock(&msc->part_sel_lock);
> - __mpam_part_sel(ris_idx, 0, msc);
> - mpam_ris_hw_probe(ris);
> - mutex_unlock(&msc->part_sel_lock);
> + scoped_guard(mutex, &msc->part_sel_lock) {
> + ret = __mpam_part_sel(ris_idx, 0, msc);
> + if (ret)
> + return ret;
> +
> + ret = mpam_ris_hw_probe(ris);
> + if (ret)
> + return ret;
> + }
> }
>
> /* Clear any stale errors */
> diff --git a/drivers/resctrl/mpam_internal.h b/drivers/resctrl/mpam_internal.h
> index def0e3a65c230..68a6cf2b9cc73 100644
> --- a/drivers/resctrl/mpam_internal.h
> +++ b/drivers/resctrl/mpam_internal.h
> @@ -162,6 +162,10 @@ static inline void mpam_mon_sel_lock_init(struct mpam_msc *msc)
> raw_spin_lock_init(&msc->_mon_sel_lock);
> }
>
> +DEFINE_GUARD(mon_sel, struct mpam_msc *,
> + mpam_mon_sel_lock(_T), mpam_mon_sel_unlock(_T));
> +DEFINE_GUARD_COND(mon_sel, _lock, mpam_mon_sel_lock(_T), _RET);
> +
> /* Bits for mpam features bitmaps */
> enum mpam_device_features {
> mpam_feat_cpor_part,
next prev parent reply other threads:[~2026-09-30 14:01 UTC|newest]
Thread overview: 26+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-24 15:27 [PATCH v11 00/13] arm_mpam: Add MPAM-Fb firmware support Andre Przywara
2026-09-24 15:27 ` [PATCH v11 01/13] arm_mpam: let low level MSC accessors return an error Andre Przywara
2026-09-24 15:27 ` [PATCH v11 02/13] arm_mpam: propagate MSC access errors for hw_probe functions Andre Przywara
2026-09-30 14:01 ` Ben Horgan [this message]
2026-09-30 15:57 ` Andre Przywara
2026-09-30 16:16 ` Ben Horgan
2026-09-24 15:27 ` [PATCH v11 03/13] arm_mpam: propagate MSC access errors for MBWU counters Andre Przywara
2026-09-24 15:27 ` [PATCH v11 04/13] arm_mpam: propagate MSC access errors for msmon helpers Andre Przywara
2026-09-24 15:27 ` [PATCH v11 05/13] arm_mpam: propagate MSC access errors for __ris_msmon_read() Andre Przywara
2026-09-24 15:27 ` [PATCH v11 06/13] arm_mpam: propagate MSC access errors for state saving function Andre Przywara
2026-09-24 15:27 ` [PATCH v11 07/13] arm_mpam: propagate MSC access errors for mpam_reprogram_ris_partid() Andre Przywara
2026-09-24 15:27 ` [PATCH v11 08/13] arm_mpam: propagate MSC access errors for interrupt control Andre Przywara
2026-09-24 22:20 ` Jonathan Cameron
2026-09-24 15:27 ` [PATCH v11 09/13] arm_mpam: propagate MSC access errors in mpam_reset_class_locked() Andre Przywara
2026-09-24 22:21 ` Jonathan Cameron
2026-09-24 15:27 ` [PATCH v11 10/13] arm_mpam: prepare mon_sel locking for MPAM-Fb Andre Przywara
2026-09-24 15:27 ` [PATCH v11 11/13] arm_mpam: add MPAM-Fb MSC firmware access support Andre Przywara
2026-09-30 14:07 ` Ben Horgan
2026-09-30 15:09 ` Andre Przywara
2026-09-30 16:14 ` Ben Horgan
2026-09-24 15:27 ` [PATCH v11 12/13] arm_mpam: change MPAM-Fb error IRQ to use a threaded IRQ handler Andre Przywara
2026-09-24 22:36 ` Jonathan Cameron
2026-09-24 15:27 ` [PATCH v11 13/13] arm_mpam: detect and enable MPAM-Fb PCC support Andre Przywara
2026-09-30 14:15 ` Ben Horgan
2026-09-30 16:41 ` Andre Przywara
2026-09-30 16:46 ` Ben Horgan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=4cc35ee7-c4d3-40c0-b1f9-24ef5a2f5844@arm.com \
--to=ben.horgan@arm.com \
--cc=andre.przywara@arm.com \
--cc=catalin.marinas@arm.com \
--cc=fenghuay@nvidia.com \
--cc=ganapatrao.kulkarni@oss.qualcomm.com \
--cc=gshan@redhat.com \
--cc=guohanjun@huawei.com \
--cc=james.morse@arm.com \
--cc=jic23@kernel.org \
--cc=lee@trager.us \
--cc=lenb@kernel.org \
--cc=linux-acpi@vger.kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lpieralisi@kernel.org \
--cc=niyas.sait@arm.com \
--cc=rafael@kernel.org \
--cc=reinette.chatre@intel.com \
--cc=ritwick.sharma@arm.com \
--cc=sramana@qti.qualcomm.com \
--cc=srivathsa.rao@oss.qualcomm.com \
--cc=sudeep.holla@kernel.org \
--cc=tsoni@quicinc.com \
--cc=will@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®