From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f54.google.com (mail-wm1-f54.google.com [209.85.128.54]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4AB8E3346BE for ; Wed, 12 Aug 2026 15:50:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.54 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786549824; cv=none; b=Rc2dkJ8Rz2Cr54xhOwDrB0g1PRyWJTiLG5O5ydt7gZY3aUiRD8HCHlwfHbyQuVxQ8cLSCAK1yGIXinlxVlns3B12ll33u+BuCXhny1Msc234PHSLkMdb0Xsg5pQkkhI4IvDzEMEyh+a9HHxRMzZQa87RRB4qOmL/L44FtfQQlwM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786549824; c=relaxed/simple; bh=GaLKOlHjERL1TzDNZS2lBkaCSUMxuyr863vvPUuAS0M=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=WxU9JTAtSg+j5BGRUBdU+qkgY2WVFWlKDOEvi6ODXIGpMlBAX5PVQvBVYiIXMOFH8VLgpiTIBZ3mONB/iLX8f5jBtrCajztKM7RkV15y7KGX87anjdS2/HaS8b+nbPpyV2va3wF9dlu83XZwsWqNdElqyrq50sXSRinpGy0dhlM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=ITweSItJ; arc=none smtp.client-ip=209.85.128.54 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="ITweSItJ" Received: by mail-wm1-f54.google.com with SMTP id 5b1f17b1804b1-496b7622a83so7941125e9.2 for ; Wed, 12 Aug 2026 08:50:22 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786549820; x=1787154620; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:from:to:cc:subject:date:message-id:reply-to:content-type; bh=2JouJ2E5oqqDA7AVxZedljtEksx3HgSAB/3D+SB+buY=; b=ITweSItJ3hXLBGntz43cm4Qy6xeD1f3K9Ah1fJlQAZSHaOXpe6ck68gm+iZRwHf1G+ 1xxi+ioUmSuW61pUAfpuZjeEDEuNh71mRye6DhTtEnZrXBz/5qT49Km9Ra3H3vF+1wFs 5EMzGl7KKB/QaXuCk/MBz0t0mCvYfCam2JNg7StCOXCrNGoqZHZkqzHqGTpvCgvAaJOS KdbW6XgfOfaWTQDp/3v3v/UP35MUPRmcHyDyyXt59d6gbWEWPZ+Wsz+KVEy4LCm+3K8s saUKiHV5O2suVNb0Tfj/wKl5plWImaZpMKGSnZx5zNq/9QBFOkzI0NSDd/V9/ZInB1cu sDPQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786549820; x=1787154620; h=in-reply-to:content-transfer-encoding:content-disposition :content-type:mime-version:references:message-id:subject:cc:to:from :date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to:content-type; bh=2JouJ2E5oqqDA7AVxZedljtEksx3HgSAB/3D+SB+buY=; b=LB9gpvK0JtOSB1Q1gF8qU6OiS+wNaF+ZEzVc9VV1yNF6imCx2Pl6WHukrDk8fC59Rb OwJHtE+EbsuckSWsnL/sUb0+psFY4zKmuoW0vhPa2NhTdoeCTd9q1c2wlEnLY/QP8U9B eld+Jck1WGMqAkK/G+Oigvzr+UtJEtLlc5YlWSzS+rt1ssS4tMEg3ooHK0wcjPWIGBuI n4xmOXGT6RLnfj3by398qb4lnbuV2tXavBLEIW1E6C8cuJl2//Ll26dWZxh/qR2P4PaV 9JIBjXBTbFTyYVmhQSPDkSSXv2Y9g2UUyylEL8xX+sSFRiRbY9buzHql8d5ZGvb1Flyu xu1g== X-Forwarded-Encrypted: i=1; AHgh+RrZQkoFz2MbOS+6IcPvXN/fWnepzyjFTkscODuRLqds0OY+BTn0bKGLuBOjlLONctv2231y822oieyz3a8=@vger.kernel.org X-Gm-Message-State: AOJu0Yw5u2tnIY2Ef1KfsQ6a03PU5dSG5fR5EQJLnGwEm8xnpzq+ZFvv MsEP8mQOKisLYGZM/WPt+t7MB9VUgzV8FL0f9ueotvVrG+H06o2fdl6pYgxo/HJZGA== X-Gm-Gg: AR+sD13oZM9tt9IYW3qc1TvkgUrBPswiSLYDttRX1V5AMjxlgvvVYZ6vghPawpiZfH7 s8divBZjcSUCiUdPFv60ABNb7MzxgMYnmf3Ei605MrPS7e6lnpUsGk9PtPjUc4KN7egz5tDm3AJ TByMIcNnzm+l3aTqivIZmGOmv0GLnpOoNkJLJHQQb2csalheiBSZozuA0Ykuy2wibUiLXpvtVfA g1B4ahemFJgz0EQGVGpXi2/ZcgeEgXh8Cp/ksaJ6pTvZ1K1285JXU/twxvXrK50mvVY1BMd6haX adGsuEz2n8fTCo5egraU++9X5aDehsncVGaKmz/aOw1rEiZmMKWR03gPMVansYVc+47W9iyHcyt G5IjFr45Ff+Wu1QpfYs3HXTGbBphwM5HWirWwzFso4aCP/2DfnPLchxpczOa+rXfT3e+a5sp2SI DHaL7uw6QOjV9p01I+5J84DLZcEADR4aT7sIvUMZfGipdJn0UL/4XIW0KVczEFKtGkwsRw3IQjg f0WDp5p6ma9XgVL8gc5j8cPqSlM2hM= X-Received: by 2002:a05:600c:4ecc:b0:495:5e3d:15e1 with SMTP id 5b1f17b1804b1-4997c114d7emr68437175e9.16.1786549819656; Wed, 12 Aug 2026 08:50:19 -0700 (PDT) Received: from google.com (110.121.148.146.bc.googleusercontent.com. [146.148.121.110]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-4997c993d80sm60760175e9.12.2026.08.12.08.50.17 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 12 Aug 2026 08:50:18 -0700 (PDT) Date: Wed, 12 Aug 2026 15:50:11 +0000 From: Dmytro Maluka To: Sergey Senozhatsky Cc: Xin Li , Chuanxiao Dong , "H. Peter Anvin" , Thomas Gleixner , Peter Zijlstra , x86@kernel.org, linux-kernel@vger.kernel.org, Grzegorz Jaszczyk , Vineeth Pillai Subject: Re: x86: missing FRED #PF event data? Message-ID: References: <29FD2DAB-C771-4E91-95C4-435B5DF90802@zytor.com> <507D6E38-4328-4DFB-BB9B-1CE220A22075@zytor.com> <74290FF8-ED7D-473D-9A0B-4288B7EB7F4A@zytor.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: On Tue, Aug 11, 2026 at 04:05:36PM +0900, Sergey Senozhatsky wrote: > On (26/08/10 22:37), Xin Li wrote: > > > On August 10, 2026 6:47:13 PM PDT, Sergey Senozhatsky wrote: > > >> On (26/08/10 08:40), H. Peter Anvin wrote: > > >>> On August 10, 2026 1:58:18 AM PDT, Sergey Senozhatsky wrote: > > >>>> On (26/08/10 16:38), Sergey Senozhatsky wrote: > > >>>>> [..] > > >>>>>> All the crashes are reported as NULL ptr derefs, however, I believe this > > >>>>>> is not exactly the case. In all crashes CR2 is 0x1000 aligned (we always > > >>>>>> crash accessing first byte of a page). It seems that csum_partial() calls > > >>>>>> load_unaligned_zeropad() and we hit what load_unaligned_zeropad() comment > > >>>>>> describes as very unlikely) case: "word being a page-crosser and the > > >>>>>> next page not being mapped"). So instead of reading 4 remaining bytes > > >>>>>> of the page and zeroes for trailing 4 bytes, we panic(). It appears that > > >>>>>> FRED #PF is set to 0 while CR2 points to a correct page address. I added > > >>>>>> a simple printk to exc_page_fault: > > >>>>>> > > >>>>>> address = cpu_feature_enabled(X86_FEATURE_FRED) ? fred_event_data(regs) : read_cr2(); > > >>>>>> /* Fall back to CR2 if FRED event data was empty */ > > >>>>>> if (unlikely(!address)) { > > >>>>>> address = read_cr2(); > > >>>>>> pr_err(":: fixed up address to %lx [[fred: %lx cr2: %lx]]\n", address, fred_event_data(regs), read_cr2()); > > >>>>>> } > > >>>>>> > > >>>>>> and got the following while running my tests (and well, we don't crash > > >>>>>> anymore): > > >>>>>> > > >>>>>> [ 254.040223] :: fixed up address to ffff9c4d64af4000 [[fred: 0 cr2: ffff9c4d64af4000]] > > >>>>>> ... > > >>>>>> [ 1821.904563] :: fixed up address to ffff9c4e9dd0a000 [[fred: 0 cr2: ffff9c4e9dd0a000]] > > >>>>>> > > >>>>>> Does any of this make sense to you? > > >>>>> > > >>>>> I think the explanation is some pKVM shenanigans. Sorry for the noise. > > >>>> > > >>>> No, I think we are back at square one. I thought that maybe pKVM > > >>>> was disabling FRED and that was causing issues. But I actually see > > >>>> that both cpu_feature_enabled(X86_FEATURE_FRED) and (cr4 & X86_CR4_FRED) > > >>>> claim FRED is enabled, yet fred #PF data is 0 while CR2 holds the correct > > >>>> address. > > >>> > > >>> What is pKVM? Paravirtualized KVM? > > >> > > >> Protected KVM. > > >> > > >>> In that case, it is most likely pKVM not filling in the relevant fields > > >>> in the FRED stack frame, which would be a very serious bug. > > >>> > > >>> I cannot think of any other way that that could possibly happen otherwise; > > >>> on bare metal those fields are set by hardware and Linux only consumes them. > > >> > > >> I agree. I'll look at it from the pKVM side. I was not aware of pKVM > > >> when I started this discussion, I found out about it later. > > > > > > If that code calls the FRED entry from KVM routine, that routine doesn't have support for setting event_data in upstream. This would be fixed if necessary. > > > > Per Sean, it’s “host” running in a VM, so it’s kind of like a filter > > hypervisor you ever mentioned; part of the “host" running in non-root mode. > > > > So where is this page fault from? If it’s from non-root mode, does this > > page fault cause a VM exit? If yes and pKVM forwards it to FRED entry, I > > would guess it is exactly the case. > > Added Chuanxiao and Dmytro, folks please correct me. > > What I see: the page fault is happening in the non-root mode (native > MMU?). I don't see a VM exit - I tried injecting FRED #PF data but > exc_page_fault() still reads 0x00 FRED #PF data. What I also see is > that... it seems to be a hybrid configurations. From what I can tell, > guests have FRED enabled in CR4, while hypervisor has FRED disabled in > CR4. So maybe this mix of FRED modes is what pushes empty FRED #PF frame? > > Sorry if I babbled complete nonsense. I'll happily hand it over to > Chuanxiao and Dmytro at this point. We already figured the problem out offline, let me describe it here for posterity. There is actually a VM exit. What is happening is: with our out-of-tree pKVM-x86 patches, load_unaligned_zeropad() legitimately crosses a page that is protected from the host by pKVM (i.e. unmapped in the host VM's EPT page tables) yet still mapped in the host's own stage-1 page tables (e.g. as a part of the kernel direct map). So this doesn't trigger a native #PF within the host, it triggers an EPT violation, and then pKVM synthesizes a #PF and injects it into the host VM to let load_unaligned_zeropad() work seamlessly. And basically hpa's guess is spot on: the problem is that for injecting this #PF, pKVM is reusing KVM's vmx_inject_exception(), which doesn't support the case when FRED is enabled in the guest and thus doesn't set the event_data (not until Xin's patches [1] are merged). I've quickly patched that up in [2]. As for the FRED setup in hypervisor vs host VM (which is rather orthogonal to the above problem): indeed, FRED is disabled in the pKVM hypervisor [3] while in the host VM it is kept enabled if it was enabled before deprivileging [4], and the host VM "owns" FRED, i.e. all FRED-related MSRs are passed-through to the host VM, and VM_EXIT_SAVE_IA32_FRED and VM_EXIT_LOAD_IA32_FRED are *not* enabled for the host VM. Please anyone let me know if such a setup is problematic in any way. [1] https://lore.kernel.org/kvm/20251026201911.505204-1-xin@zytor.com/ [2] https://android-review.googlesource.com/4225278 [3] https://android.googlesource.com/kernel/common/+/d13d0c68ee9106a26a20cbab4a653f0d8dd4691f/arch/x86/kvm/vmx/pkvm_init.c#909 [4] https://android.googlesource.com/kernel/common/+/d13d0c68ee9106a26a20cbab4a653f0d8dd4691f/arch/x86/kvm/vmx/pkvm_init.c#794