From: Mauro Carvalho Chehab <mauro.chehab@linux.intel.com>
To: Tvrtko Ursulin <tvrtko.ursulin@linux.intel.com>
Cc: Mauro Carvalho Chehab <mchehab@kernel.org>,
Andi Shyti <andi.shyti@linux.intel.com>,
Daniel Vetter <daniel@ffwll.ch>,
Daniele Ceraolo Spurio <daniele.ceraolospurio@intel.com>,
David Airlie <airlied@linux.ie>,
Jani Nikula <jani.nikula@linux.intel.com>,
John Harrison <John.C.Harrison@Intel.com>,
Joonas Lahtinen <joonas.lahtinen@linux.intel.com>,
Lucas De Marchi <lucas.demarchi@intel.com>,
Matt Roper <matthew.d.roper@intel.com>,
Matthew Auld <matthew.auld@intel.com>,
Rodrigo Vivi <rodrigo.vivi@intel.com>,
dri-devel@lists.freedesktop.org, intel-gfx@lists.freedesktop.org,
linux-kernel@vger.kernel.org,
Tvrtko Ursulin <tvrtko.ursulin@intel.com>,
Sushma Venkatesh Reddy <sushma.venkatesh.reddy@intel.com>,
Daniel Vetter <daniel.vetter@ffwll.ch>,
Dave Airlie <airlied@redhat.com>,
Jon Bloomfield <jon.bloomfield@intel.com>,
Jani Nikula <jani.nikula@intel.com>,
stable@vger.kernel.org
Subject: Re: [PATCH] drm/i915: don't flush TLB on GEN8
Date: Fri, 27 May 2022 13:50:22 +0200 [thread overview]
Message-ID: <20220527135022.0dd0891d@maurocar-mobl2> (raw)
In-Reply-To: <d981f429-d01f-4576-2e5c-0ae153d24df1@linux.intel.com>
On Fri, 27 May 2022 11:55:42 +0100
Tvrtko Ursulin <tvrtko.ursulin@linux.intel.com> wrote:
> On 27/05/2022 10:09, Mauro Carvalho Chehab wrote:
> > i915 selftest hangcheck is causing the i915 driver timeouts, as
> > reported by Intel CI:
> >
> > http://gfx-ci.fi.intel.com/cibuglog-ng/issuefilterassoc/24297?query_key=42a999f48fa6ecce068bc8126c069be7c31153b4
> >
> > When such test runs, the only output is:
> >
> > [ 68.811639] i915: Performing live selftests with st_random_seed=0xe138eac7 st_timeout=500
> > [ 68.811792] i915: Running hangcheck
> > [ 68.811859] i915: Running intel_hangcheck_live_selftests/igt_hang_sanitycheck
> > [ 68.816910] i915 0000:00:02.0: [drm] Cannot find any crtc or sizes
> > [ 68.841597] i915: Running intel_hangcheck_live_selftests/igt_reset_nop
> > [ 69.346347] igt_reset_nop: 80 resets
> > [ 69.362695] i915: Running intel_hangcheck_live_selftests/igt_reset_nop_engine
> > [ 69.863559] igt_reset_nop_engine(rcs0): 709 resets
> > [ 70.364924] igt_reset_nop_engine(bcs0): 903 resets
> > [ 70.866005] igt_reset_nop_engine(vcs0): 659 resets
> > [ 71.367934] igt_reset_nop_engine(vcs1): 549 resets
> > [ 71.869259] igt_reset_nop_engine(vecs0): 553 resets
> > [ 71.882592] i915: Running intel_hangcheck_live_selftests/igt_reset_idle_engine
> > [ 72.383554] rcs0: Completed 16605 idle resets
> > [ 72.884599] bcs0: Completed 18641 idle resets
> > [ 73.385592] vcs0: Completed 17517 idle resets
> > [ 73.886658] vcs1: Completed 15474 idle resets
> > [ 74.387600] vecs0: Completed 17983 idle resets
> > [ 74.387667] i915: Running intel_hangcheck_live_selftests/igt_reset_active_engine
> > [ 74.889017] rcs0: Completed 747 active resets
> > [ 75.174240] intel_engine_reset(bcs0) failed, err:-110
> > [ 75.174301] bcs0: Completed 525 active resets
> >
> > After that, the machine just silently hangs.
> >
> > The root cause is that the flush TLB logic is not working as
> > expected on GEN8.
> >
> > Tested on an Intel NUC5i7RYB with an i7-5557U Broadwell CPU.
> >
> > This patch partially reverts the logic by skipping GEN8 from
> > the TLB cache flush.
>
> Since I am pretty sure no such failures were spotted when merging the
> feature I assume the failure is sporadic and/or limited to some
> configurations? Do you have any details there? Because it is an
> important security issue we should not revert it lightly.
It occurs every time here:
https://intel-gfx-ci.01.org/tree/drm-tip/fi-bdw-5557u.html
It also happens on my own NUC5i7RYB every time when the TLB patch is
applied. Reverting it (or applying this fix) is enough for hangcheck
to pass.
I suspect that TLB flush never happens there, causing ETIMEOUT at
hangcheck.
It could indeed be limited to some specific setups. I dunno.
The only Gen8 machine I have access is my own NUC. So, I can't
test it elsewhere.
Regards,
Mauro
next prev parent reply other threads:[~2022-05-27 12:07 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2022-05-27 9:09 Mauro Carvalho Chehab
2022-05-27 10:55 ` Tvrtko Ursulin
2022-05-27 11:50 ` Mauro Carvalho Chehab [this message]
-- strict thread matches above, loose matches on Subject: below --
2022-05-27 8:58 Mauro Carvalho Chehab
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20220527135022.0dd0891d@maurocar-mobl2 \
--to=mauro.chehab@linux.intel.com \
--cc=John.C.Harrison@Intel.com \
--cc=airlied@linux.ie \
--cc=airlied@redhat.com \
--cc=andi.shyti@linux.intel.com \
--cc=daniel.vetter@ffwll.ch \
--cc=daniel@ffwll.ch \
--cc=daniele.ceraolospurio@intel.com \
--cc=dri-devel@lists.freedesktop.org \
--cc=intel-gfx@lists.freedesktop.org \
--cc=jani.nikula@intel.com \
--cc=jani.nikula@linux.intel.com \
--cc=jon.bloomfield@intel.com \
--cc=joonas.lahtinen@linux.intel.com \
--cc=linux-kernel@vger.kernel.org \
--cc=lucas.demarchi@intel.com \
--cc=matthew.auld@intel.com \
--cc=matthew.d.roper@intel.com \
--cc=mchehab@kernel.org \
--cc=rodrigo.vivi@intel.com \
--cc=stable@vger.kernel.org \
--cc=sushma.venkatesh.reddy@intel.com \
--cc=tvrtko.ursulin@intel.com \
--cc=tvrtko.ursulin@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®