From: Peter Zijlstra <peterz@infradead.org>
To: Don Zickus <dzickus@redhat.com>
Cc: x86@kernel.org, Andi Kleen <andi@firstfloor.org>,
gong.chen@linux.intel.com, LKML <linux-kernel@vger.kernel.org>,
Elliott@hp.com, fweisbec@gmail.com
Subject: Re: [PATCH 1/6] x86, nmi: Implement delayed irq_work mechanism to handle lost NMIs
Date: Wed, 21 May 2014 12:29:34 +0200 [thread overview]
Message-ID: <20140521102934.GZ2485@laptop.programming.kicks-ass.net> (raw)
In-Reply-To: <1400181949-157170-2-git-send-email-dzickus@redhat.com>
On Thu, May 15, 2014 at 03:25:44PM -0400, Don Zickus wrote:
> +DEFINE_PER_CPU(bool, nmi_delayed_work_pending);
> +
> +static void nmi_delayed_work_func(struct irq_work *irq_work)
> +{
> + DECLARE_BITMAP(nmi_mask, NR_CPUS);
That's _far_ too big for on-stack, 4k cpus would make that 512 bytes.
> + cpumask_t *mask;
> +
> + preempt_disable();
That's superfluous, irq_work's are guaranteed to be called with IRQs
disabled.
> +
> + /*
> + * Can't use send_IPI_self here because it will
> + * send an NMI in IRQ context which is not what
> + * we want. Create a cpumask for local cpu and
> + * force an IPI the normal way (not the shortcut).
> + */
> + bitmap_zero(nmi_mask, NR_CPUS);
> + mask = to_cpumask(nmi_mask);
> + cpu_set(smp_processor_id(), *mask);
> +
> + __this_cpu_xchg(nmi_delayed_work_pending, true);
Why is this xchg and not __this_cpu_write() ?
> + apic->send_IPI_mask(to_cpumask(nmi_mask), NMI_VECTOR);
What's wrong with apic->send_IPI_self(NMI_VECTOR); ?
> +
> + preempt_enable();
> +}
> +
> +struct irq_work nmi_delayed_work =
> +{
> + .func = nmi_delayed_work_func,
> + .flags = IRQ_WORK_LAZY,
> +};
OK, so I don't particularly like the LAZY stuff and was hoping to remove
it before more users could show up... apparently I'm too late :-(
Frederic, I suppose this means dual lists.
> +static bool nmi_queue_work_clear(void)
> +{
> + bool set = __this_cpu_read(nmi_delayed_work_pending);
> +
> + __this_cpu_write(nmi_delayed_work_pending, false);
> +
> + return set;
> +}
That's a test-and-clear, the name doesn't reflect this. And here you do
_not_ use xchg where you actually could have.
That said, try and avoid using xchg() its unconditionally serialized.
> +
> +static int nmi_queue_work(void)
> +{
> + bool queued = irq_work_queue(&nmi_delayed_work);
> +
> + if (queued) {
> + /*
> + * If the delayed NMI actually finds a 'dropped' NMI, the
> + * work pending bit will never be cleared. A new delayed
> + * work NMI is supposed to be sent in that case. But there
> + * is no guarantee that the same cpu will be used. So
> + * pro-actively clear the flag here (the new self-IPI will
> + * re-set it.
> + *
> + * However, there is a small chance that a real NMI and the
> + * simulated one occur at the same time. What happens is the
> + * simulated IPI NMI sets the work_pending flag and then sends
> + * the IPI. At this point the irq_work allows a new work
> + * event. So when the simulated IPI is handled by a real NMI
> + * handler it comes in here to queue more work. Because
> + * irq_work returns success, the work_pending bit is cleared.
> + * The second part of the back-to-back NMI is kicked off, the
> + * work_pending bit is not set and an unknown NMI is generated.
> + * Therefore check the BUSY bit before clearing. The theory is
> + * if the BUSY bit is set, then there should be an NMI for this
> + * cpu latched somewhere and will be cleared when it runs.
> + */
> + if (!(nmi_delayed_work.flags & IRQ_WORK_BUSY))
> + nmi_queue_work_clear();
So I'm utterly and completely failing to parse that. It just doesn't
make sense.
> + }
> +
> + return 0;
> +}
Why does this function have a return value if all it can return is 0 and
everybody ignores it?
> +
> static int __kprobes nmi_handle(unsigned int type, struct pt_regs *regs, bool b2b)
> {
> struct nmi_desc *desc = nmi_to_desc(type);
> @@ -341,6 +441,9 @@ static __kprobes void default_do_nmi(struct pt_regs *regs)
> */
> if (handled > 1)
> __this_cpu_write(swallow_nmi, true);
> +
> + /* kick off delayed work in case we swallowed external NMI */
That's inaccurate, there's no guarantee we actually swallowed one
afaict, this is where we have to assume we lost one because there's
really no other place.
> + nmi_queue_work();
> return;
> }
>
> @@ -362,10 +465,16 @@ static __kprobes void default_do_nmi(struct pt_regs *regs)
> #endif
> __this_cpu_add(nmi_stats.external, 1);
> raw_spin_unlock(&nmi_reason_lock);
> + /* kick off delayed work in case we swallowed external NMI */
> + nmi_queue_work();
Again, inaccurate, there's no guarantee we did swallow an external NMI,
but the thing is, there's no guarantee we didn't either, which is why we
need to do this.
> return;
> }
> raw_spin_unlock(&nmi_reason_lock);
>
> + /* expected delayed queued NMI? Don't flag as unknown */
> + if (nmi_queue_work_clear())
> + return;
> +
Right, so here we effectively swallow the extra nmi and avoid the
unknown_nmi_error() bits.
next prev parent reply other threads:[~2014-05-21 10:29 UTC|newest]
Thread overview: 17+ messages / expand[flat|nested] mbox.gz Atom feed top
2014-05-15 19:25 [PATCH 0/6 V2] x86, nmi: Various fixes and cleanups Don Zickus
2014-05-15 19:25 ` [PATCH 1/6] x86, nmi: Implement delayed irq_work mechanism to handle lost NMIs Don Zickus
2014-05-21 10:29 ` Peter Zijlstra [this message]
2014-05-21 16:45 ` Don Zickus
2014-05-21 17:51 ` Peter Zijlstra
2014-05-21 19:02 ` Don Zickus
2014-05-21 19:38 ` Peter Zijlstra
2014-05-15 19:25 ` [PATCH 2/6] x86, nmi: Add new nmi type 'external' Don Zickus
2014-05-15 19:25 ` [PATCH 3/6] x86, nmi: Add boot line option 'panic_on_unrecovered_nmi' and 'panic_on_io_nmi' Don Zickus
2014-05-15 19:25 ` [PATCH 4/6] x86, nmi: Remove 'reason' value from unknown nmi output Don Zickus
2014-05-15 19:25 ` [PATCH 5/6] x86, nmi: Move default external NMI handler to its own routine Don Zickus
2014-05-21 10:38 ` Peter Zijlstra
2014-05-21 16:48 ` Don Zickus
2014-05-21 18:17 ` Peter Zijlstra
2014-05-21 19:13 ` Don Zickus
2014-05-15 19:25 ` [PATCH 6/6 V2] x86, nmi: Add better NMI stats to /proc/interrupts and show handlers Don Zickus
2014-05-15 20:28 ` [PATCH 0/6 V2] x86, nmi: Various fixes and cleanups Don Zickus
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20140521102934.GZ2485@laptop.programming.kicks-ass.net \
--to=peterz@infradead.org \
--cc=Elliott@hp.com \
--cc=andi@firstfloor.org \
--cc=dzickus@redhat.com \
--cc=fweisbec@gmail.com \
--cc=gong.chen@linux.intel.com \
--cc=linux-kernel@vger.kernel.org \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome