From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752065AbdKZElV (ORCPT ); Sat, 25 Nov 2017 23:41:21 -0500 Received: from mx1.redhat.com ([209.132.183.28]:43114 "EHLO mx1.redhat.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751902AbdKZElT (ORCPT ); Sat, 25 Nov 2017 23:41:19 -0500 Date: Sat, 25 Nov 2017 22:41:15 -0600 From: Josh Poimboeuf To: Andy Lutomirski Cc: Thomas Gleixner , X86 ML , Borislav Petkov , "linux-kernel@vger.kernel.org" , Brian Gerst , Dave Hansen , Linus Torvalds Subject: Re: [PATCH] x86/orc: Don't bail on stack overflow Message-ID: <20171126044115.tper4nvt47tsxr2j@treble> References: <7b2c3ad96fb62afbc095bd7ba8a93022480153cb.1511630836.git.luto@kernel.org> <20171126024031.uxi4numpbjm5rlbr@treble> MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline In-Reply-To: User-Agent: Mutt/1.6.0.1 (2016-04-01) X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16 (mx1.redhat.com [10.5.110.28]); Sun, 26 Nov 2017 04:41:19 +0000 (UTC) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Sat, Nov 25, 2017 at 08:25:12PM -0800, Andy Lutomirski wrote: > On Sat, Nov 25, 2017 at 6:40 PM, Josh Poimboeuf wrote: > > On Sat, Nov 25, 2017 at 04:16:23PM -0800, Andy Lutomirski wrote: > >> Can you send me whatever config and exact commit hash generated this? > >> I can try to figure out why it failed. > > > > Sorry, I've been traveling. I just got some time to take a look at > > this. I think there are at least two unwinder issues here: > > > > - It doesn't deal gracefully with the case where the stack overflows and > > the stack pointer itself isn't on a valid stack but the > > to-be-dereferenced data *is*. > > > > - The oops dump code doesn't know how to print partial pt_regs, for the > > case where if we get an interrupt/exception in *early* entry code > > before the full pt_regs have been saved. > > > > (Andy, I'm not quite sure about your patch, and whether it's still > > needed after these patches. I'll need to look at it later when I have > > more time.) > > I haven't tested yet, but I think my patch is probably still needed. > The issue I fixed is that unwind_start() would bail out early if sp > was below the stack. Also: Makes sense, maybe both are needed. Your patch deals with a bad SP at the beginning and mine deals with a bad SP in the middle. > > -static bool stack_access_ok(struct unwind_state *state, unsigned long addr, > > +static bool stack_access_ok(struct unwind_state *state, unsigned long _addr, > > size_t len) > > { > > struct stack_info *info = &state->stack_info; > > + void *addr = (void *)_addr; > > > > - /* > > - * If the address isn't on the current stack, switch to the next one. > > - * > > - * We may have to traverse multiple stacks to deal with the possibility > > - * that info->next_sp could point to an empty stack and the address > > - * could be on a subsequent stack. > > - */ > > - while (!on_stack(info, (void *)addr, len)) > > - if (get_stack_info(info->next_sp, state->task, info, > > - &state->stack_mask)) > > - return false; > > + if (!on_stack(info, addr, len) && > > + (get_stack_info(addr, state->task, info, &state->stack_mask))) > > + return false; > > > > return true; > > } > > This looks odd to me before and after. Shouldn't this be side-effect > free? That is, shouldn't it create a copy of info and stack_mask and > point that to get_stack_info() rather than allowing get_stack_info() > to modify the unwind state? I think the side effects are ok, but maybe stack_access_ok() should be renamed to make it clearer that it has side effects. -- Josh