From: Thomas Gleixner <tglx@linutronix.de>
To: Linus Torvalds <torvalds@linux-foundation.org>
Cc: LKML <linux-kernel@vger.kernel.org>,
Andrew Morton <akpm@linux-foundation.org>,
Ingo Molnar <mingo@kernel.org>, "H. Peter Anvin" <hpa@zytor.com>
Subject: Re: [GIT pull] timer updates for 4.9
Date: Mon, 24 Oct 2016 11:39:53 +0200 (CEST) [thread overview]
Message-ID: <alpine.DEB.2.20.1610240947160.4872@nanos> (raw)
In-Reply-To: <CA+55aFyDAWBNfyYM_BG2D5-g1skj=iCkhfW9ZNA_h2zZDdfy_g@mail.gmail.com>
On Sun, 23 Oct 2016, Linus Torvalds wrote:
> So I found what looks like a bug in lock_timer_base() wrt migration.
>
> This code:
>
> for (;;) {
> struct timer_base *base;
> u32 tf = timer->flags;
>
> if (!(tf & TIMER_MIGRATING)) {
> base = get_timer_base(tf);
> spin_lock_irqsave(&base->lock, *flags);
> if (timer->flags == tf)
> return base;
> spin_unlock_irqrestore(&base->lock, *flags);
> }
> cpu_relax();
> }
>
> looks subtly buggy. I think that load of "tf" needs a READ_ONCE() to
> make sure that gcc doesn't simply reload the valid of "timer->flags"
> at random points.
You are right, that needs a READ_ONCE(). Stupid me.
> Yes, the spin_lock_irqsave() is a barrier, but that's the only one.
> Afaik, gcc could decide that "I need to spill tf, so I'll just reload
> it" after looking up get_timer_base().
>
> And no, I don't think this is the cause of my problem, but I suspect
> that something _like_ fragility in lock_timer_base() could cause this.
It might explain it, when this really ends up with the wrong base.
> I dunno. That whole thing looks very fragile to begin with: is it
> really ok to change the expiry time of a timer without holding any
> locks what-so-ever? The timer may just be firing on another CPU, and
> you may be setting the expiry time on a timer that isn't ever going to
> fire again.
A timer firing on the other CPU is not an issue, but yes, we should not do
that unlocked. This certainly is a naive over optimization.
I'll go through the locking once more with a fine comb and send out fixes
later today along with another NOHZ issue which I decoded over the weekend.
> And there may well be valid reasons why I'm full of crap, and there's
> some reason why this is all safe. Maybe the GP fault I saw was my
> fault after all, in some way that I can't for the life of me figure
> out right now..
Is your pr_cont thing serialized or does it rely on the timer locking (or
the lack of it)?
Thanks,
tglx
next prev parent reply other threads:[~2016-10-24 9:42 UTC|newest]
Thread overview: 12+ messages / expand[flat|nested] mbox.gz Atom feed top
2016-10-22 12:02 Thomas Gleixner
2016-10-23 22:39 ` Linus Torvalds
2016-10-23 23:20 ` Linus Torvalds
2016-10-24 9:39 ` Thomas Gleixner [this message]
2016-10-24 14:51 ` Thomas Gleixner
2016-10-24 15:13 ` Thomas Gleixner
2016-10-25 14:57 ` [tip:timers/urgent] timers: Plug locking race vs. timer migration tip-bot for Thomas Gleixner
2016-10-25 14:58 ` [tip:timers/urgent] timers: Lock base for same bucket optimization tip-bot for Thomas Gleixner
2016-10-24 17:16 ` [GIT pull] timer updates for 4.9 Linus Torvalds
2016-10-24 19:09 ` Thomas Gleixner
2016-10-24 19:30 ` Linus Torvalds
2016-10-24 21:36 ` Thomas Gleixner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=alpine.DEB.2.20.1610240947160.4872@nanos \
--to=tglx@linutronix.de \
--cc=akpm@linux-foundation.org \
--cc=hpa@zytor.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@kernel.org \
--cc=torvalds@linux-foundation.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®