From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f197.google.com (mail-pf1-f197.google.com [209.85.210.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CE6FB47044F for ; Mon, 17 Aug 2026 20:12:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.197 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786997528; cv=none; b=Wy61ynCE9/AQNyvNyKVea2+LPNr5zCZbYMPXyMnsLLxVM5uhDHLUe74UX5w4K/O4JlPZvB2sRmg3I34CrvwFCfroI7WVflbqBSlOKJU0vIFHxyYLE/zLWx+UPoyf2xdlUgp9s2cS3fa6nOtGFEXUSTrilPY6dzjlJ9+DR9HgyII= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786997528; c=relaxed/simple; bh=21kzg6NN40Ce/8AdJ6lIoY6WV7VpheLmKQXAIMBvZso=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=hsGsiPj3Lk4uT3+uxuDGRfE87RVJKTqaKRisqbLLXavReDyrvHTCaBvCuLsaBl6GcRIw6GAg7RmqI8BB+uPGcB/TPMSWetf9v6YzksdgDwzEFM8PMZICNVyhpb6BneIE0+C9fJ5oUD/7peqTN1OJ1QsODOK4aV94OyIGU4mFGHw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=iJTodV0a; arc=none smtp.client-ip=209.85.210.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="iJTodV0a" Received: by mail-pf1-f197.google.com with SMTP id d2e1a72fcca58-8487eb67173so6566121b3a.2 for ; Mon, 17 Aug 2026 13:12:06 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786997526; x=1787602326; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=N0+O1Vxdvyhj7pmjvbt8iDMZ8zrd75NX4nQ7WNMkzm4=; b=iJTodV0aKmCVsLMnddLyz/1Zki4oM726DwxCPiRebwkqR3bRk4Lf97qwgRfOmmBtUm 89fv2Wy2Y+O8Ih7dA2cbg+uPsLVXTeIOIsljApkWJ8IUWLvjOCzh+J2q8cIc4sXYPxcl 7EeRr0HyYjvPOrLfPvTOyigo9kHGTPFs6L8nKfrxwmOowgmd3lra7KdGkfcivsDe4p8e USYbY3VXHsXhMGOo6KsQ60WU4blkAf6iW5CrHy6couujo3s8Gf/A4glUpkXcnCSAy6z+ UywIsUzWcRVKSNG9DYbFBoloOW0QikRNOJ/sTZYyDd2ckL25Y2LeDyXcnlw7xfBKni07 2wDg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786997526; x=1787602326; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=N0+O1Vxdvyhj7pmjvbt8iDMZ8zrd75NX4nQ7WNMkzm4=; b=pAIB+N5lHmW/KaALv94pc0b0eZPV3WyxjJzfHDE9pwvRPJahFhjrr9XraTJPHZQ7Px MQv2/Tqs6mmxL9MlWRjfqptCznd6iicIaO2iLad03AVdQB7EJHAgfke/bS0FaYEA7ibs Zo2P1XhXs35BP9gy0gEZtQSExuTSw+icJ9k8oRPuBM5d3mNPUniGZgjG/YAiK8DSnbUk 0xqtcmT02TowNj9DNWdbX0meGJLOoHSExnGSzpTbyR3croNgPB0zFkRjaRf3dgFu+ukh uuUmoN4sa1nY2raySNv7UvLtSzWLgz/COWl3ZwbYxp6fKHdM4GqmXslij+K40BhFfgc2 UHTQ== X-Forwarded-Encrypted: i=1; AHgh+Ro3UUu7ckUL6kdvguhy238WqeivE9FJFf2DwmtCbBllqWg2Y9kSYd6wXUqYO1CVQLsG8PkkGJeMR4Ivwm4=@vger.kernel.org X-Gm-Message-State: AOJu0Yy9qGwKs9FkwKE/9vM77Srw8RdmADNE44qRTKrwLuPOGz7Iqb8S VQx9cpJ70CNYXBip4c+qH9ZNNHsdA6h2/z1L7QuNQh0cF5VUCXkO0vmVI7JAfEtX5Ui6M6jMd9L I9My9tQ== X-Received: from pgmp38.prod.google.com ([2002:a63:1e66:0:b0:cbe:e160:452d]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:23c1:b0:847:968d:b101 with SMTP id d2e1a72fcca58-851b88ce701mr2522658b3a.21.1786997525814; Mon, 17 Aug 2026 13:12:05 -0700 (PDT) Date: Mon, 17 Aug 2026 13:12:05 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <0c80b9b0e13e3ab2cbe4f9eaf4a02ebba25a7001.camel@intel.com> Message-ID: Subject: Re: [PATCH v10 11/41] KVM: guest_memfd: Ensure pages are not in use before conversion From: Sean Christopherson To: Yan Zhao Cc: Rick P Edgecombe , "ackerleytng@google.com" , "david@kernel.org" , "kvm@vger.kernel.org" , "steven.price@arm.com" , "peterx@redhat.com" , "forkloop@google.com" , "tabba@google.com" , "linux-trace-kernel@vger.kernel.org" , "dave.hansen@linux.intel.com" , "x86@kernel.org" , Vishal Annapurve , "willy@infradead.org" , "tglx@kernel.org" , "wyihan@google.com" , "pratyush@kernel.org" , "aik@amd.com" , "jmattson@google.com" , "aneesh.kumar@kernel.org" , "linux-kernel@vger.kernel.org" , "akpm@linux-foundation.org" , "binbin.wu@linux.intel.com" , "rientjes@google.com" , "andrew.jones@linux.dev" , "linux-kselftest@vger.kernel.org" , "chrisl@kernel.org" , "shakeel.butt@linux.dev" , "mathieu.desnoyers@efficios.com" , "oupton@kernel.org" , "mhiramat@kernel.org" , "baohua@kernel.org" , "tarunsahu@google.com" , "linux-coco@lists.linux.dev" , "jhubbard@nvidia.com" , "jgg@ziepe.ca" , "jthoughton@google.com" , "yuanchu@google.com" , "hpa@zytor.com" , "shikemeng@huaweicloud.com" , "nphamcs@gmail.com" , "linux-doc@vger.kernel.org" , "shivankg@amd.com" , "shuah@kernel.org" , "youngjun.park@lge.com" , "kasong@tencent.com" , "pankaj.gupta@amd.com" , "suzuki.poulose@arm.com" , "chao.p.peng@linux.intel.com" , "pbonzini@redhat.com" , "vbabka@kernel.org" , "weixugc@google.com" , "michael.roth@amd.com" , "rostedt@goodmis.org" , "mingo@redhat.com" , "qperret@google.com" , "brauner@kernel.org" , "bp@alien8.de" , "baoquan.he@linux.dev" , "corbet@lwn.net" , "skhan@linuxfoundation.org" , "liam@infradead.org" , "axelrasmussen@google.com" , "kas@kernel.org" , "qi.zheng@linux.dev" , "linux-mm@kvack.org" Content-Type: text/plain; charset="us-ascii" On Mon, Aug 17, 2026, Yan Zhao wrote: > On Thu, Aug 13, 2026 at 04:20:05PM -0700, Sean Christopherson wrote: > > On Thu, Aug 13, 2026, Rick P Edgecombe wrote: > > > > > nicely so we actually just run an old branch's TDX selftests against newer > > > > > kernels. So the branch is a bit of a pile, and not really suitable for sharing. > > > > > We plan to clean it and upstream it when the path clears. So it would really > > > > > help to get those basic ones upstream. We remain happy to help, so please let us > > > > > know. > > > > > > > > I guess at this point I'm hoping y'all and Sean are okay that this > > > > conversions series merges, and we let this stress test failure be > > > > handled later. I'll be around to fix things :) > > > > > > > > I'd say the line of sight to fixing this would be when the KVM MMU only > > > > gets PFNs (and no pages at all) from guest_memfd. > > > > > > Hmm, I think we shouldn't upstream a uABI that we don't have line of sight to > > > making robust. So it would be good to settle this thread at least. > > > > This isn't uABI. You're talking about hitting a race condition between one task > Hmm. Perhaps it is not a uABI issue, since users are allowed to retry. However, > it is hard to convince me that it makes sense to require users to retry a > private-to-shared conversion before a GFN has ever been mapped, given that a > retry is not required when the GFN is currently in use by the guest. > > > converting a page and another faulting in the same page. An NMI, SMI, or IRQ at > > just the right/wrong time, especially on a preemptible kernel, could lead to the > > same test failures, even if KVM drops the refcount "immediately". > Could you elaborate on how an NMI, SMI, or IRQ at just the right/wrong time > could lead to the same test failures? Ah, sorry, my bad. I was speed reading and missed that the key to your suggested "*page = NULL" change was that the reference was put _before_ filemap_invalidate_unlock_shared(), i.e. before dropping the invalidate lock and thus before __kvm_gmem_set_attributes() will walk the folios to look for outstanding references. I was thinking that putting the reference right away was just shrinking the timing window, but putting the reference while still holding the invalidate lock closes the window entirely. So, I take back what I said about this not being ABI, and about this not blocking in-place conversion. It most definitely affects ABI, and so needs to be addressed before merging in-place conversion. The only question is if we want to commit to guaranteeing that conversion will succeed in this scenario, or if we want to take the easy way out and formally document that conversion can fail with EAGAIN at any time, even if userspace has never mmap()'d the memory in question. I'm leaning pretty strongly towards guaranteeing conversion will succeed. We'll still need to document the EAGAIN behavior, but IMO there's a massive difference between conversion failing if there's a lingering reference acquired via a VMA, conversion failing because a vCPU page fault raced with conversion. E.g. being able to assert success in a very curated test, as the stress test presumably does, would be extremely valuable for helping detect/prevent edge case bugs. The argument against guaranteeing success is that we might make our future lives harder, e.g. if it turns out there are legitimate, hard-to-solve edge cases. But I'm ok with that risk, as it seems highly unlikely to be problematic in practice, and there is real benefit to guaranteeing success. > > As for in-place conversion, this is not a blocker. > Sorry. I didn't intend to block in-place conversion. LOL, what we intend and what happens aren't always the same. :-) > I encountered this issue during testing, so reported it. Thanks for doing so! I'd *much* rather sort these issues out *before* merging code, even if it means delaying the merge by a bit.