From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0189FE571 for ; Thu, 27 Feb 2025 01:54:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1740621264; cv=none; b=jpI20e+m0aig+odMu3Jd6zAQExAxdPsMrioh8J0ypOBWDyE7ZtjPoZhlNz3DwEk+QjD2glronXxdGeNManNhQc+dNT5CV/PSGfwzVgrmgsO4gGYXdSNUbcJAi3qKjfSH0NWsb/w7Dx+wCIldS9dJKboztOTReeBNaew4OXrtVe0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1740621264; c=relaxed/simple; bh=BUX4aPOvtEa1+6oCrTSXTpfOjw/MuX+wsTkMmn6zUAg=; h=Message-ID:Subject:From:To:Cc:Date:In-Reply-To:References: Content-Type:MIME-Version; b=T2ab9cMDM5Fg5zOn6Ky10LmbmvgSgK21uEmx/Kswt/hR2Atwxfx0X3rPd0LmXOlZ0QLoeIZblIq6NR9+CBj0o53K/59q1OjNhb3cAl09vdzdT5f4iuaOGeRoYjcUjWsbmZ/6V9s4J1lQPJZCW8RygS4ZKMogNUJNRdkmIojIsU8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=WcP/g6h/; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="WcP/g6h/" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1740621260; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Ou+95ZKQM5HeiQtErwcnpxUHmSwBFhNR4ZItLSh1obw=; b=WcP/g6h/9y2S2Zy/T0o9q4cXKbfGtAVVaUwig/1Ja6OZjMxw3KBG+J7VX+OX/UIqBIehQS qYmGBL8pjhp6yaYlZSxe2eKHYxGrA9nzcYxO8eWtUQL9G38DgzWTZ7A6f3dq3rdvx+bM6u ZH/TyT0VmOpRkMeUWCVy5Kh/IYTwynI= Received: from mail-qv1-f69.google.com (mail-qv1-f69.google.com [209.85.219.69]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-502-kM4yLzBuN4GTBHjD5NcWMA-1; Wed, 26 Feb 2025 20:54:18 -0500 X-MC-Unique: kM4yLzBuN4GTBHjD5NcWMA-1 X-Mimecast-MFC-AGG-ID: kM4yLzBuN4GTBHjD5NcWMA_1740621258 Received: by mail-qv1-f69.google.com with SMTP id 6a1803df08f44-6e637c24051so8905016d6.0 for ; Wed, 26 Feb 2025 17:54:18 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1740621258; x=1741226058; h=content-transfer-encoding:mime-version:user-agent:references :in-reply-to:date:cc:to:from:subject:message-id:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=Ou+95ZKQM5HeiQtErwcnpxUHmSwBFhNR4ZItLSh1obw=; b=HtKTk2OLd/ZRadoBxfXO/H2T0qZNtwzmZLGA7RrqPnixjoP+1YpQkDMVIqLO6iLDxj LnJHd5hWE4kUViKf0F0UYtQRhFDPsWQpgREqgLtBpiCdpl5Jg6F6VtwUOqfAzqZ1g5Yu gmFKOV9ewZZz1rALssmssgXvD9U7mT5fJy7C+hOznT35R7AixJ7xTopnTsap64QDtqqg uOnv5/2k6nO9oWNcD9uXoymdChrSMRVVqAQ7ac3Wb/Oc5D665CbfLK2sBe/y8jD373y+ M0O0jwGd3npC7DyuGsk0I9TzO/BHNPWsTL11fiBqgjkDIAzL2Nnq60G+rq+A4dAGFrmJ 4kKQ== X-Forwarded-Encrypted: i=1; AJvYcCXYcBipxXXQTUtc9TaMzjvHv7WhGUJyx9CBXudWNxfn1yniofMDMybXXKEgzjetayWWGckJgzBkjYgLnQE=@vger.kernel.org X-Gm-Message-State: AOJu0Yy1pnzuM8eMT1osvv4Ow9Cka0JPfSoBoICbqIUmsMmNeRyhDJlt UztKpCmbmkty/P73iS2PsTeEsp7JJowEuZXc7U19d6pV0qs6PuIFU+K72MHefbXO9aTPUan2Fy+ vprJe4PUVxxMWlTFoXXsjjS7ZnuYc7IctyxV6AKwB0P4fJnl9nwDdwGkV/eli7g== X-Gm-Gg: ASbGnctqn8YHHqEeyYKo2nyW/TGi3RyLzdcpZRd4+Apar6ACF3oEdeLSEd8ohvjQ859 m69yfS4DOuyajhHy7OCZSU4QZu2lyEKYaxSIhP3aVWKHKftvZ6XI3hWYtCeq5DJpuYvxHBzFXUR qI+OcpRaery8/WglvpT5Ils7K51hCbDyY4nBD+Qt+NR0Hk4daOsPAqemjAO68cX0ciGnCVN9DZ6 hJ1LwMstIuBNjkDEzYmSoU+Sxz/9mp5bdBvatA/1A0S417Ex1SEKmtL4bh/PVN8wh7mfGM5Iwv5 +59ezwBzzN9YGjg= X-Received: by 2002:ad4:5bcf:0:b0:6cb:d1ae:27a6 with SMTP id 6a1803df08f44-6e6b010e18bmr279715466d6.24.1740621258375; Wed, 26 Feb 2025 17:54:18 -0800 (PST) X-Google-Smtp-Source: AGHT+IFBqZ/D95K1s9QpCXseRrPcX1BdgmuGX/rQVnBiRdXhBp6GsljdcZd9Jzp4AScOPLQ7SlfGGQ== X-Received: by 2002:ad4:5bcf:0:b0:6cb:d1ae:27a6 with SMTP id 6a1803df08f44-6e6b010e18bmr279715256d6.24.1740621258073; Wed, 26 Feb 2025 17:54:18 -0800 (PST) Received: from starship ([2607:fea8:fc01:8d8d:6adb:55ff:feaa:b156]) by smtp.gmail.com with ESMTPSA id d75a77b69052e-47472409b32sm4405971cf.61.2025.02.26.17.54.17 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 26 Feb 2025 17:54:17 -0800 (PST) Message-ID: <8d4ef8fdf192bba60c7c31f22925270d62c87c54.camel@redhat.com> Subject: Re: [PATCH v9 00/11] KVM: x86/mmu: Age sptes locklessly From: Maxim Levitsky To: Sean Christopherson Cc: James Houghton , Paolo Bonzini , David Matlack , David Rientjes , Marc Zyngier , Oliver Upton , Wei Xu , Yu Zhao , Axel Rasmussen , kvm@vger.kernel.org, linux-kernel@vger.kernel.org Date: Wed, 26 Feb 2025 20:54:16 -0500 In-Reply-To: References: <20250204004038.1680123-1-jthoughton@google.com> <025b409c5ca44055a5f90d2c67e76af86617e222.camel@redhat.com> <07788b85473e24627131ffe1a8d1d01856dd9cb5.camel@redhat.com> <4c605b4e395a3538d9a2790918b78f4834912d72.camel@redhat.com> Content-Type: text/plain; charset="UTF-8" User-Agent: Evolution 3.36.5 (3.36.5-2.fc32) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 7bit On Wed, 2025-02-26 at 16:51 -0800, Sean Christopherson wrote: > On Wed, Feb 26, 2025, Maxim Levitsky wrote: > > On Tue, 2025-02-25 at 16:50 -0800, Sean Christopherson wrote: > > > On Tue, Feb 25, 2025, Maxim Levitsky wrote: > > > What if we make the assertion user controllable? I.e. let the user opt-out (or > > > off-by-default and opt-in) via command line? We did something similar for the > > > rseq test, because the test would run far fewer iterations than expected if the > > > vCPU task was migrated to CPU(s) in deep sleep states. > > > > > > TEST_ASSERT(skip_sanity_check || i > (NR_TASK_MIGRATIONS / 2), > > > "Only performed %d KVM_RUNs, task stalled too much?\n\n" > > > " Try disabling deep sleep states to reduce CPU wakeup latency,\n" > > > " e.g. via cpuidle.off=1 or setting /dev/cpu_dma_latency to '0',\n" > > > " or run with -u to disable this sanity check.", i); > > > > > > This is quite similar, because as you say, it's impractical for the test to account > > > for every possible environmental quirk. > > > > No objections in principle, especially if sanity check is skipped by default, > > although this does sort of defeats the purpose of the check. > > I guess that the check might still be used for developers. > > A middle ground would be to enable the check by default if NUMA balancing is off. > We can always revisit the default setting if it turns out there are other problematic > "features". That works for me. I can send a patch for this then. > > > > > > Aha! I wonder if in the failing case, the vCPU gets migrated to a pCPU on a > > > > > different node, and that causes NUMA balancing to go crazy and zap pretty much > > > > > all of guest memory. If that's what's happening, then a better solution for the > > > > > NUMA balancing issue would be to affine the vCPU to a single NUMA node (or hard > > > > > pin it to a single pCPU?). > > > > > > > > Nope. I pinned main thread to CPU 0 and VM thread to CPU 1 and the problem > > > > persists. On 6.13, the only way to make the test consistently work is to > > > > disable NUMA balancing. > > > > > > Well that's odd. While I'm quite curious as to what's happening, > > Gah, chatting about this offline jogged my memory. NUMA balancing doesn't zap > (mark PROT_NONE/PROT_NUMA) PTEs for paging the kernel thinks are being accessed > remotely, it zaps PTEs to see if they're are being accessed remotely. So yeah, > whenever NUMA balancing kicks in, the guest will see a large amount of its memory > get re-faulted. > > Which is why it's such a terribly feature to pair with KVM, at least as-is. NUMA > balancing is predicated on inducing and resolving the #PF being relatively cheap, > but that doesn't hold true for secondary MMUs due to the coarse nature of mmu_notifiers. > I also think so, not to mention that VM exits aren't that cheap either, and the general direction is to avoid them as much as possible, and here this feature pretty much yanks the memory out of the guest every two seconds or so, causing lots of VM exits. Best regards, Maxim Levitsky