From: Xianwei Zhao <xianwei.zhao@amlogic.com>
To: sashiko-reviews@lists.linux.dev
Cc: robh@kernel.org, neil.armstrong@linaro.org,
linux-amlogic@lists.infradead.org, devicetree@vger.kernel.org,
conor+dt@kernel.org
Subject: Re: [PATCH v5 2/2] thermal: amlogic: Add support for A9 thermal controller
Date: Thu, 10 Sep 2026 16:03:15 +0800 [thread overview]
Message-ID: <6671af4d-1229-4277-a355-71247dd696f8@amlogic.com> (raw)
In-Reply-To: <20260910070634.60B5F1F000FF@smtp.kernel.org>
On 2026/9/10 15:06, sashiko-bot@kernel.org wrote:
> Thank you for your contribution! Sashiko AI review found 4 potential issue(s) to consider:
>
> New issues:
> - [Critical] The A9-specific hardware over-temperature reset mechanism is enabled unconditionally on all SoCs, causing unclocked MMIO accesses and undefined register writes that will crash existing platforms (e.g., G12A) on resume.
>
> Pre-existing issues:
> - [High] This problem wasn't introduced by this patch, but the thermal zone and hwmon sysfs interfaces are registered before the hardware and clocks are fully initialized, creating a race condition that leads to unclocked MMIO reads.
> - [High] This problem wasn't introduced by this patch, but incorrect sign extension of a bitwise NOT operation on a 32-bit unsigned mask causes massively inflated negative trim values on 32-bit architectures, breaking thermal calibration.
> - [High] This problem wasn't introduced by this patch, but disabling hardware clocks in `remove()` before the devres-managed thermal zone is unregistered introduces a use-after-disable race condition that can crash the system.
> --
>
> commit c409169ce8a3938d784d6682dd656715a5d6d606
> Author: Xianwei Zhao<xianwei.zhao@amlogic.com>
>
> thermal: amlogic: Add support for A9 thermal controller
>
> Add support for the Amlogic A9 thermal controller.
>
> The A9 thermal controller uses different calibration parameters, requires
> separate core and bus clocks, and supports a hardware over-temperature
> reset mechanism. It also restores the hardware reset configuration after
> resume.
>
>> diff --git a/drivers/thermal/amlogic_thermal.c b/drivers/thermal/amlogic_thermal.c
>> index a0b530624b60c..e7cedb9443151 100644
>> --- a/drivers/thermal/amlogic_thermal.c
>> +++ b/drivers/thermal/amlogic_thermal.c
> [ ... ]
>> @@ -138,6 +158,51 @@ static int amlogic_thermal_code_to_millicelsius(struct amlogic_thermal *pdata,
>> return temp;
>> }
>>
>> +/*
>> + * Calculate a temperature code from a temperature value .
>> + * The unit of the temperature is degree milliCelsius.
>> + */
>> +static u32 amlogic_thermal_millicelsius_to_code(struct amlogic_thermal *pdata, int millicelsius)
>> +{
>> + const struct amlogic_thermal_soc_calib_data *param =
>> + pdata->data->calibration_parameters;
>> + s64 factor, uptat, uefuse;
>> + u32 temp_code;
>> +
>> + uefuse = pdata->trim_info & TSENSOR_TRIM_SIGN_MASK ?
>> + ~(pdata->trim_info & TSENSOR_TRIM_TEMP_MASK) + 1 :
>> + (pdata->trim_info & TSENSOR_TRIM_TEMP_MASK);
> [Severity: High]
> This is a pre-existing issue, but by copying this logic from
> amlogic_thermal_code_to_millicelsius(), the new function inherits a sign
> extension bug.
>
> When the hardware provides a negative trim value (bit 15 set in trim_info),
> the bitwise NOT operation on the 32-bit unsigned TSENSOR_TRIM_TEMP_MASK mask
> evaluates to a large positive value on 32-bit systems (e.g., 0xFFFFFFFF).
>
> When this is assigned to the s64 uefuse variable, C promotion rules
> zero-extend it to 4294967295 rather than sign-extending to -1.
>
> This breaks thermal calibration for hardware instances with negative factory
> calibration offsets, rendering those devices unusable.
>
> Could we fix this integer conversion so that the negative value is properly
> sign-extended into the s64 variable?
>
> [ ... ]
>> @@ -346,6 +440,17 @@ static int amlogic_thermal_probe(struct platform_device *pdev)
>>
>> devm_thermal_add_hwmon_sysfs(&pdev->dev, pdata->tzd);
>>
> [Severity: High]
> This isn't a bug introduced by this patch, but registering the thermal zone
> device here before the hardware clocks are fully initialized creates a race
> condition.
>
> The thermal core invokes the .get_temp callback immediately upon registration
> to establish the initial temperature state.
>
> If the driver is built as a loadable module and loaded after boot, when
> bootloader-enabled clocks have already been disabled by the clock framework's
> late init, the .get_temp call will attempt to read from the hardware before
> its clock is prepared and enabled.
>
> Will this result in an unclocked MMIO read and subsequent bus fault?
>
>> + if (pdata->data->reset_temp) {
>> + reset_temp = pdata->data->reset_temp;
>> + } else {
>> + dev_dbg(dev, "using default hardware reset temperature\n");
>> + reset_temp = TSENSOR_HW_RESET_DEFAULT_TEMP;
>> + }
>> +
>> + pdata->temp_code = amlogic_thermal_millicelsius_to_code(pdata, reset_temp);
>> +
>> + amlogic_tsensor_setup_hw_reset(pdata);
>> +
>> ret = amlogic_thermal_enable(pdata);
> [Severity: Critical]
> Does unconditionally applying the A9-specific hardware reset setup here
> cause issues for older non-A9 platforms?
>
Non-A9 HW is support too, just that this function wasn't added before.
Here, the default value is used.
> For existing SoCs where has_sysclk is false, the peripheral clock (data->clk)
> isn't enabled until amlogic_thermal_enable() is called on the next line.
>
The existing SoCs sys clock is not control by this driver, it is enabled
by default.
> Since the clock is physically off during probe, does this regmap_update_bits()
> write inside amlogic_tsensor_setup_hw_reset() trigger a bus fault?
>
> Also, could writing to TSENSOR_CFG_REG2 cause undefined register modification
> on hardware that doesn't support this over-temperature reset mechanism?
>
The Non-SoC has TSENSOR_CFG_REG2 regiter, the register was hidden before
because the related functions were not supported.
> [ ... ]
>> @@ -370,7 +477,15 @@ static int amlogic_thermal_suspend(struct device *dev)
>> static int amlogic_thermal_resume(struct device *dev)
>> {
>> struct amlogic_thermal *data = dev_get_drvdata(dev);
>> + int ret;
>> +
>> + if (data->data->has_sysclk) {
>> + ret = clk_prepare_enable(data->sysclk);
>> + if (ret)
>> + return ret;
>> + }
>>
>> + amlogic_tsensor_setup_hw_reset(data);
>> return amlogic_thermal_enable(data);
> [Severity: Critical]
> Similarly, does this execute an MMIO write before the peripheral clock is
> enabled during the resume path for non-A9 platforms?
>
> On non-A9 platforms, data->clk is only enabled inside
> amlogic_thermal_enable().
>
The sysclock is keeping openning all time for other SoC.
> If this write occurs while the clock is still physically disabled during
> system resume, could this cause a synchronous external abort and break power
> management completely?
>
Non A9 clock is aways_on ,so for A9 if enable sysclock failed, it will
be return.
> [Severity: High]
> This is a pre-existing issue, but is there a use-after-disable race condition
> in the driver's remove path?
>
> drivers/thermal/amlogic_thermal.c:amlogic_thermal_remove() {
> struct amlogic_thermal *data = platform_get_drvdata(pdev);
>
> amlogic_thermal_disable(data);
> }
>
> Disabling the hardware clocks in remove() before the devres-managed thermal
> zone is cleanly unregistered could crash the system.
>
> Because pdata->tzd was registered with devm_thermal_of_zone_register(), it
> remains fully active and exposed to the thermal core until the devres cleanup
> phase, which executes after remove() returns.
>
> If the thermal core's polling timer fires, or if a user reads the temperature
> sysfs attribute during this window, amlogic_thermal_get_temp() will attempt
> to read from the unclocked hardware.
_______________________________________________
linux-amlogic mailing list
linux-amlogic@lists.infradead.org
http://lists.infradead.org/mailman/listinfo/linux-amlogic
prev parent reply other threads:[~2026-09-10 8:03 UTC|newest]
Thread overview: 5+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-10 6:38 [PATCH v5 0/2] thermal: Add support A9 Xianwei Zhao via B4 Relay
2026-09-10 6:38 ` [PATCH v5 1/2] dt-bindings: thermal: amlogic: Add A9 thermal bindings Xianwei Zhao via B4 Relay
2026-09-10 6:38 ` [PATCH v5 2/2] thermal: amlogic: Add support for A9 thermal controller Xianwei Zhao via B4 Relay
2026-09-10 7:06 ` sashiko-bot
2026-09-10 8:03 ` Xianwei Zhao [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=6671af4d-1229-4277-a355-71247dd696f8@amlogic.com \
--to=xianwei.zhao@amlogic.com \
--cc=conor+dt@kernel.org \
--cc=devicetree@vger.kernel.org \
--cc=linux-amlogic@lists.infradead.org \
--cc=neil.armstrong@linaro.org \
--cc=robh@kernel.org \
--cc=sashiko-reviews@lists.linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®