From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9D73C1773A for ; Sat, 13 Jul 2024 01:39:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1720834750; cv=none; b=WoLbww0+xKw0Cie1IxUC1Bo0/DUEvdSQ1N1DHP8afPCyoKGshO/ndbc0Roc95NkNg8ysG/3NhAt26lPITx8XGS+6AxhxVrqf43vayRk2KXveo6XXLP63w8ykridRVo4DYN+WrKJe1y52HbjCLoCZBL5EdO0LAxtYuY6xXidOPpc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1720834750; c=relaxed/simple; bh=ZLnJzFlCPFz+aFSkdpAsEfC2J8bneJIy6qVFc/o0Kg0=; h=From:To:Cc:Subject:Date:Message-Id:Content-Type:MIME-Version; b=orbtR7TiDJapDIBcAeaPEmkhyID5qmK8sTXigqjI6GuP2FIvC57UlOsHAS1uujoWKfbTqlQzqnZ8Bp1XN+swDkVinJRoHUOFJK/JQlxLfXCHsxBTQ1pMJ+HtwkaBnay0xEZmJ9Gphozm4hmYYs+KWk3wmhj+0lKpkayjibf6WJw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=WP8Pv2B6; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="WP8Pv2B6" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1720834747; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding; bh=MUkwksWZM4bpo+VerhOiEyiXHghLuvfzYDHRObdTPyU=; b=WP8Pv2B6rSEy2Gz6H0TUlkIXuj7QDnPYvFTwqK9XOlE74w/QQsvqYjkgr6V+Vku+r70lgT iEg0JuGfRGZfpb7mW47vvTHPdVmXLHwM3tPr3Zqf7gEGRdMendVNxAJsjLmmKaZ4jJ1A0x iP0iPE31PzaxjMnFOJsy60VWJedSspU= Received: from mx-prod-mc-04.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-509-TjfmI7CGMJqB2UM6TOXzqA-1; Fri, 12 Jul 2024 21:39:01 -0400 X-MC-Unique: TjfmI7CGMJqB2UM6TOXzqA-1 Received: from mx-prod-int-01.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-01.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.4]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-04.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 874DD19560AD; Sat, 13 Jul 2024 01:38:59 +0000 (UTC) Received: from starship.lan (unknown [10.22.18.76]) by mx-prod-int-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id D5DC13000181; Sat, 13 Jul 2024 01:38:56 +0000 (UTC) From: Maxim Levitsky To: kvm@vger.kernel.org Cc: Dave Hansen , Thomas Gleixner , Paolo Bonzini , Borislav Petkov , x86@kernel.org, linux-kernel@vger.kernel.org, Sean Christopherson , Ingo Molnar , "H. Peter Anvin" , Maxim Levitsky Subject: [PATCH 0/2] Fix for a very old KVM bug in the segment cache Date: Fri, 12 Jul 2024 21:38:54 -0400 Message-Id: <20240713013856.1568501-1-mlevitsk@redhat.com> Content-Type: text/plain; charset="utf-8" Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: quoted-printable X-Scanned-By: MIMEDefang 3.4.1 on 10.30.177.4 Hi,=0D =0D Recently, while trying to understand why the pmu_counters_test=0D selftest sometimes fails when run nested I stumbled=0D upon a very interesting and old bug:=0D =0D It turns out that KVM caches guest segment state,=0D but this cache doesn't have any protection against concurrent use.=0D =0D This usually works because the cache is per vcpu, and should=0D only be accessed by vCPU thread, however there is an exception:=0D =0D If the full preemption is enabled in the host kernel,=0D it is possible that vCPU thread will be preempted, for=0D example during the vmx_vcpu_reset.=0D =0D vmx_vcpu_reset resets the segment cache bitmask and then initializes=0D the segments in the vmcs, however if the vcpus is preempted in the=0D middle of this code, the kvm_arch_vcpu_put is called which=0D reads SS's AR bytes to determine if the vCPU is in the kernel mode,=0D which caches the old value.=0D =0D Later vmx_vcpu_reset will set the SS's AR field to the correct value=0D in vmcs but the cache still contains an invalid value which=0D can later for example leak via KVM_GET_SREGS and such.=0D =0D In particular, kvm selftests will do KVM_GET_SREGS,=0D and then KVM_SET_SREGS, with a broken SS's AR field passed as is,=0D which will lead to vm entry failure.=0D =0D This issue is not a nested issue, and actually I was able=0D to reproduce it on bare metal, but due to timing it happens=0D much more often nested. The only requirement for this to happen=0D is to have full preemption enabled in the kernel which runs the selftest.=0D =0D pmu_counters_test reproduces this issue well, because it creates=0D lots of short lived VMs, but the issue as was noted=0D about is not related to pmu.=0D =0D To fix this issue, I wrapped the places which write the segment=0D fields with preempt_disable/enable. It's not an ideal fix, other options ar= e=0D possible. Please tell me if you prefer these:=0D =0D 1. Getting rid of the segment cache. I am not sure how much it helps=0D these days - this code is very old.=0D =0D 2. Using a read/write lock - IMHO the cleanest solution but might=0D also affect performance.=0D =0D 3. Making the kvm_arch_vcpu_in_kernel not touch the cache=0D and instead do a vmread directly.=0D This is a shorter solution but probably less future proof.=0D =0D Best regards,=0D Maxim Levitsky=0D =0D Maxim Levitsky (2):=0D KVM: nVMX: use vmx_segment_cache_clear=0D KVM: VMX: disable preemption when writing guest segment state=0D =0D arch/x86/kvm/vmx/nested.c | 7 ++++++-=0D arch/x86/kvm/vmx/vmx.c | 22 ++++++++++++++++++----=0D arch/x86/kvm/vmx/vmx.h | 5 +++++=0D 3 files changed, 29 insertions(+), 5 deletions(-)=0D =0D -- =0D 2.26.3=0D =0D