* [PATCH] clk: qcom: dispcc-x1e80100: don't disable display PLLs in the unused-clock sweep
@ 2026-09-30 9:28 Jianfeng Liu
2026-10-01 8:18 ` Konrad Dybcio
0 siblings, 1 reply; 3+ messages in thread
From: Jianfeng Liu @ 2026-09-30 9:28 UTC (permalink / raw)
To: linux-clk
Cc: linux-arm-msm, Bjorn Andersson, linux-kernel, Dmitry Baryshkov,
Abel Vesa, Stephen Boyd, Rajendra Nayak, Brian Masney,
Jerome Brunet, Jianfeng Liu
UEFI leaves disp_cc_pll0 locked and running for the handover
framebuffer. clk_disable_unused() runs before any
display driver probes, finds it unused (no consumers, no
CLK_IGNORE_UNUSED) and powers it down through
alpha_pll_reset_lucid_evo_disable(). When the DPU driver later
enables mdp_clk, alpha_pll_reset_lucid_evo_prepare() performs a
full PLL reset and relock.
On X1E80100 that relock can end up marginal, and whether it does is
a per-boot analog lottery. A marginal PLL degrades the DPU core
timing just enough that the DisplayPort controller drops every
other MTP (micro transfer packet) of the pixel stream fed by the
DPU. The result is static black vertical bands on the eDP panel:
at 2560x1600, 24bpp, 4 lanes, one MTP is 4 * 64 symbols / 3 bytes
per pixel = 85.33 pixels, so the screen splits into exactly 30
bands of which every second one is black (20 bands at 16bpp).
This reproduces on roughly half of all boots and persists for the
whole boot.
The corruption is invisible to every register or clock-state dump:
the PLL locks either way, only the lock quality differs, and all
programmed values (DPU, DP controller MSA/TU, PHY, dispcc) are
byte-identical between good and bad boots. The controller's
internal BIST pattern stays clean on a corrupted boot, pinning the
failure to the pixel ingest timing fed by the DPU.
Verified on an Acer Swift SFA14-11 (X1E78100):
- clk_ignore_unused on the cmdline: 7/7 boots clean
- keeping only disp_cc_pll0 from the sweep: 10/10 boots clean
- keeping any other clock from the sweep: still corrupts
- with this patch: 10/10 boots clean
Keep the firmware-provided PLL state by marking disp_cc_pll0
CLK_IGNORE_UNUSED, the same way gcc drivers protect CPU hfplls
(cf. gcc-msm8960). Other drivers using
clk_alpha_pll_reset_lucid_ole_ops for their display PLLs (sm8550,
sm8650) may carry the same latent issue.
Fixes: ee3f0739035f5 ("clk: qcom: Add dispcc clock driver for x1e80100")
Signed-off-by: Jianfeng Liu <liujianfeng1994@gmail.com>
Assisted-by: LLM
---
drivers/clk/qcom/dispcc-x1e80100.c | 15 +++++++++++++++
1 file changed, 15 insertions(+)
diff --git a/drivers/clk/qcom/dispcc-x1e80100.c b/drivers/clk/qcom/dispcc-x1e80100.c
index aed06203a886a..5613f1fdb2e4a 100644
--- a/drivers/clk/qcom/dispcc-x1e80100.c
+++ b/drivers/clk/qcom/dispcc-x1e80100.c
@@ -94,6 +94,21 @@ static struct clk_alpha_pll disp_cc_pll0 = {
.clkr = {
.hw.init = &(const struct clk_init_data) {
.name = "disp_cc_pll0",
+ /*
+ * The boot firmware (UEFI) leaves disp_cc_pll0 locked
+ * and running for the handover framebuffer.
+ * clk_disable_unused() runs before any display driver
+ * probes, finds the PLL unused (no consumers, no
+ * CLK_IGNORE_UNUSED) and powers it down. The DPU
+ * driver then has to relock it via the reset-prepare
+ * sequence. On X1E80100 that relock can end up
+ * marginal (per-boot analog lottery), degrading the
+ * DPU core timing enough to corrupt eDP output at
+ * MTP granularity (every other micro transfer packet
+ * delivered blank, on ~50% of boots). Keeping the
+ * firmware PLL state avoids the relock entirely.
+ */
+ .flags = CLK_IGNORE_UNUSED,
.parent_data = &(const struct clk_parent_data) {
.index = DT_BI_TCXO,
},
---
base-commit: 6474fa070f2b8013b4b87350b775b8c3be6e8aac
branch: dispcc-x1e80100-keep-disp-plls
--
2.47.3
^ permalink raw reply [flat|nested] 3+ messages in thread* Re: [PATCH] clk: qcom: dispcc-x1e80100: don't disable display PLLs in the unused-clock sweep
2026-09-30 9:28 [PATCH] clk: qcom: dispcc-x1e80100: don't disable display PLLs in the unused-clock sweep Jianfeng Liu
@ 2026-10-01 8:18 ` Konrad Dybcio
2026-10-03 14:02 ` Jianfeng Liu
0 siblings, 1 reply; 3+ messages in thread
From: Konrad Dybcio @ 2026-10-01 8:18 UTC (permalink / raw)
To: Jianfeng Liu, linux-clk
Cc: linux-arm-msm, Bjorn Andersson, linux-kernel, Dmitry Baryshkov,
Abel Vesa, Stephen Boyd, Rajendra Nayak, Brian Masney,
Jerome Brunet
On 9/30/26 11:28 AM, Jianfeng Liu wrote:
> UEFI leaves disp_cc_pll0 locked and running for the handover
> framebuffer. clk_disable_unused() runs before any
> display driver probes, finds it unused (no consumers, no
> CLK_IGNORE_UNUSED) and powers it down through
> alpha_pll_reset_lucid_evo_disable(). When the DPU driver later
> enables mdp_clk, alpha_pll_reset_lucid_evo_prepare() performs a
> full PLL reset and relock.
Can you give
https://lore.kernel.org/linux-arm-msm/20260626-clk-sync-state-v1-0-4156d8196dc8@redhat.com/
a shot?
Konrad
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH] clk: qcom: dispcc-x1e80100: don't disable display PLLs in the unused-clock sweep
2026-10-01 8:18 ` Konrad Dybcio
@ 2026-10-03 14:02 ` Jianfeng Liu
0 siblings, 0 replies; 3+ messages in thread
From: Jianfeng Liu @ 2026-10-03 14:02 UTC (permalink / raw)
To: konrad.dybcio
Cc: abelvesa, andersson, bmasney+clk, jbrunet+clk, linux-arm-msm,
linux-clk, linux-kernel, liujianfeng1994, lumag, quic_rjendra,
sboyd
Hi Konrad,
Thanks for the pointer! I gave Brian's series [1] a shot on a
7.3-rc4 based tree with my pll0 patch reverted (the trivial
conflict in the pmdomain hunk resolved by hand), and on this
platform it does not solve the issue - it actually makes things
worse.
With the series applied, the unused-clock sweeps for the display
providers fire at 1.4-1.5s, which lands *inside* the display
bring-up sequence instead of before it:
[ 1.381597] tcsrcc-x1e80100 1fc0000.clock-controller: clk: Disabling unused clocks
[ 1.468038] qcom-edp-phy aec5a00.phy: clk: Disabling unused clocks
[ 1.468194] dispcc-x1e80100 af00000.clock-controller: clk: Disabling unused clocks
[ 1.807670] msm_dpu ae01000.display-controller: bound ae90000.displayport-controller (ops msm_dp_display_comp_ops [msm])
[ 1.808271] msm_dpu ae01000.display-controller: bound aea0000.displayport-controller (ops msm_dp_display_comp_ops [msm])
[ 1.810010] msm_dpu ae01000.display-controller: bound ae9a000.displayport-controller (ops msm_dp_display_comp_ops [msm])
[ 1.860460] [drm:dpu_kms_hw_init:1168] dpu hardware revision:0x90020000
[ 2.002334] msm_dpu ae01000.display-controller: [drm] fb0: msmdrmfb frame buffer device
The reason is that the mdss device's probe returns early; the
actual display bring-up happens later through the component
framework (controllers binding at 1.8s, first modeset after
that). The driver core only sees "probed", so sync_state fires
before any of the display clocks have been claimed:
- tcsrcc's sweep at 1.381s disables tcsr_edp_clkref_en while
the eDP PHY is still mid-probe (the PHY is about to claim it
as its "ref" clock)
- the eDP PHY provider's own sweep at 1.468s disables its
link/pixel output clocks
- dispcc's sweep at 1.468s disables all of
disp_cc_mdss_dptx3_{aux,link,pixel0}*, whose consumer (the
DP controller) only binds at 1.8s
- disp_cc_pll0 is still swept at 1.468s, so the reset-relock
lottery from my patch's commit message remains as well
The outcome is again a ~50% per-boot lottery, just with a
different and harsher failure mode. In 2 out of 4 boots the eDP
panel came up broken: the top third of the screen shows fine
green/black striping and the lower two thirds stay black. On
the broken boots the final clk_summary shows the whole eDP link
clock chain disabled and unclaimed (disp_cc_mdss_dptx3_{aux,
link,pixel0}*, tcsr_edp_clkref_en, and even the eDP PHY's own
link/vco_div clocks), i.e. the bring-up failed right where the
sweeps had just fired. On the good boots the same clocks are all
claimed and enabled. The per-boot variable is simply whether the
sweeps at 1.38-1.47s happen to land before or after the DP
controller and the PHY claim their clocks.
So for this class of problems - firmware handover state, where
the hardware consumes clocks that the CCF does not know about
yet, and whose drivers enable them later than probe() - sync_state
trades the old race for a new one: the trigger ("all consumers
probed") still precedes the moment the clocks actually get
enabled, and the decision (enable_count == 0) still cannot see
the firmware-provided state. The original relock lottery remains
as well, since disp_cc_pll0 is still swept at 1.468s, before
mdp_clk is first enabled at ~1.86s.
I do think the series is valuable for the module / late-probe
cases it targets, and I'm happy to help testing it further. But
the pll0 fix still seems needed, not least because it can be
backported to stable branches while the sync_state work lands.
Full dmesg and clk_summary captures from both the good and broken
boots are available on request.
[1] https://lore.kernel.org/linux-arm-msm/20260626-clk-sync-state-v1-0-4156d8196dc8@redhat.com/
Jianfeng
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-03 14:03 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 9:28 [PATCH] clk: qcom: dispcc-x1e80100: don't disable display PLLs in the unused-clock sweep Jianfeng Liu
2026-10-01 8:18 ` Konrad Dybcio
2026-10-03 14:02 ` Jianfeng Liu
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®