From mboxrd@z Thu Jan 1 00:00:00 1970 Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S932177AbeAMNxn (ORCPT + 1 other); Sat, 13 Jan 2018 08:53:43 -0500 Received: from mail.linuxfoundation.org ([140.211.169.12]:43624 "EHLO mail.linuxfoundation.org" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1755056AbeAMNxm (ORCPT ); Sat, 13 Jan 2018 08:53:42 -0500 Date: Sat, 13 Jan 2018 14:53:42 +0100 From: Greg KH To: Pavel Tatashin Cc: steven.sistare@oracle.com, linux-kernel@vger.kernel.org, tglx@linutronix.de, mingo@redhat.com, hpa@zytor.com, x86@kernel.org, jkosina@suse.cz, hughd@google.com, dave.hansen@linux.intel.com, luto@kernel.org, torvalds@linux-foundation.org Subject: Re: [PATCH 4.4 v2] x86/pti/efi: broken conversion from efi to kernel page table Message-ID: <20180113135342.GB28839@kroah.com> References: <20180112200002.25907-1-pasha.tatashin@oracle.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20180112200002.25907-1-pasha.tatashin@oracle.com> User-Agent: Mutt/1.9.2 (2017-12-15) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Return-Path: On Fri, Jan 12, 2018 at 03:00:02PM -0500, Pavel Tatashin wrote: > In entry_64.S we have code like this: > > /* Unconditionally use kernel CR3 for do_nmi() */ > /* %rax is saved above, so OK to clobber here */ > ALTERNATIVE "jmp 2f", "movq %cr3, %rax", X86_FEATURE_KAISER > /* If PCID enabled, NOFLUSH now and NOFLUSH on return */ > ALTERNATIVE "", "bts $63, %rax", X86_FEATURE_PCID > pushq %rax > /* mask off "user" bit of pgd address and 12 PCID bits: */ > andq $(~(X86_CR3_PCID_ASID_MASK | KAISER_SHADOW_PGD_OFFSET)), %rax > movq %rax, %cr3 > 2: > > /* paranoidentry do_nmi, 0; without TRACE_IRQS_OFF */ > call do_nmi > > With this instruction: > andq $(~(X86_CR3_PCID_ASID_MASK | KAISER_SHADOW_PGD_OFFSET)), %rax > > We unconditionally switch from whatever our CR3 was to kernel page table. > But, in arch/x86/platform/efi/efi_64.c We temporarily set a different page > table, that does not have the kernel page table with 0x1000 offset from it. > > Look in efi_thunk() and efi_thunk_set_virtual_address_map(). > > So, while CR3 points to the other page table, we get an NMI interrupt, > and clear 0x1000 from CR3, resulting in a bogus CR3 if the 0x1000 bit was > set. > > The efi page table comes from realmode/rm/trampoline_64.S: > > arch/x86/realmode/rm/trampoline_64.S > > 141 .bss > 142 .balign PAGE_SIZE > 143 GLOBAL(trampoline_pgd) .space PAGE_SIZE > > Notice: alignment is PAGE_SIZE, so after applying KAISER_SHADOW_PGD_OFFSET > which equal to PAGE_SIZE, we can get a different page table. > > But, even if we fix alignment, here the trampoline binary is later copied > into dynamically allocated memory in reserve_real_mode(), so we need to > fix that place as well. > > Fixes: 8a43ddfb93a0 ("KAISER: Kernel Address Isolation") > > Signed-off-by: Pavel Tatashin > Reviewed-by: Steven Sistare > --- > arch/x86/include/asm/kaiser.h | 10 ++++++++++ > arch/x86/realmode/init.c | 4 +++- > arch/x86/realmode/rm/trampoline_64.S | 3 ++- > 3 files changed, 15 insertions(+), 2 deletions(-) > > Changelog: > v1 - v2: Fixed compiling issue when PTI config is disabled. This one is now queued up, thanks. greg k-h