From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f197.google.com (mail-pl1-f197.google.com [209.85.214.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id AEB21438019 for ; Tue, 11 Aug 2026 17:28:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.197 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786469337; cv=none; b=DzHCKrS3YLz7Yn4ZYQ8oSMPy8L6IoiX9OFYS/Qzi5VpOzb8DvbZpiFodO0eTHvPi6Vh5bJABmb1vbMu+tV4FqQ83uul5VJtJSFrcOcZ3zh+6PCHthNw0xoYKRpUZ1swdVBFOtlivzgPxOJVSNbGVEZmjrXxJ8dEoeir/00J7uHo= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786469337; c=relaxed/simple; bh=gEPHLTrISbAwWv2wzSbXTZKZJ71uRtEhHNVm+uhlAPw=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=ShA3IOnhBflBNPEgNsgCc1C0hacGOb/rE6KE+pYx3uve4M8Nh+q/QZjHxEEX56DLXbuBTq3rUEz0vL1MivioyUqt3toAUQX6B1kMAPZzGzi0j7IGAviLgVauuDd5mBZ/AGJNjJtb1nGJ6upowwlGB19da7daKl+lD8j1GgCsaxA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=cRH3hL0u; arc=none smtp.client-ip=209.85.214.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="cRH3hL0u" Received: by mail-pl1-f197.google.com with SMTP id d9443c01a7336-2d004f13426so2001085ad.3 for ; Tue, 11 Aug 2026 10:28:53 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786469332; x=1787074132; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=kMDt/Sj6YjXvFiWQr62Lu+RNKL1NV+pa3PqCNwc1t5I=; b=cRH3hL0u7KwkHEw+Ah/fr3udLdHkRuuhpmHOOLbBJgKQWvUtMBtLW0AbkOoyuX+IEC G+7Fdbl5YuqDXzmS7+FTZRa3GmQna9q8ttfkJNncc0w7hxacywXKle4qZIbG0KCidu2p g6034ZQgJKBzJlqA8LXqU2vEl9XCB0/z1ldb6BOFo6EC4bByA6HUjILzyD7kU/S8MEnH REWhFY3IfVYRiyCqzfEHLxzgDkGKj00Sl0Jf4V+6fmJsQQze+QeK5/zxNWYJF4sy/AIp ElpU0wSDW6W38QZRpt8ivbFoZEsZ6NXlJW7QsaBfzestiq6pxaqqxzH44lvqHblIbafa 6qJQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786469332; x=1787074132; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=kMDt/Sj6YjXvFiWQr62Lu+RNKL1NV+pa3PqCNwc1t5I=; b=hS1eebrjFl9DgxLtigq60XtyejxOEm+OgguRUTleqJq/PSdu79/jIoLQUhosQGwmwP p/BdgV35TNq5Mr5U7feg31Ebf1inqZ1ppUbvHfPxqRa/WlwsnL2I+qs3FUJlqYPVT+4M C9hTOO2r7xCMj/N4TLCIFKonXAocKlXnfqGaOgFYtbIuSteZjk3FRvaFa6ndqHz7Q/IS 7ICvGvORyBLS2H8lkY6UgIEM+h3PbuqO7BmX95ALvRjKoAWAmDAOZCUM9wD9KIFh4Qi0 IAMzsUZHyYn2tzfy6GuZf2ypGIgYc4wowpRi+zLMM+WuxNjCpoZwpCbqM4ZiE4AHT8HG eU9g== X-Forwarded-Encrypted: i=1; AHgh+RqzMPp/mKfW1KTDMg7bLPGurJ1S/zOMZHBG3C2PKFKgvKzrHg4ro9ktnu6fSerxLWNwDXSW0vvY2Tv8e38=@vger.kernel.org X-Gm-Message-State: AOJu0YwzGGMFlCffN3g+XPycgifG74l5hApQsKOujO+Kh5JWaTbUzSTz kLv/Qi95sMsIyU71m4MFSiqlcn3fEIo/R/pvJ+ne/6I9IS+HnJFNt7jBv5nhRLHFhLhIysSpSGF sJRX5vA== X-Received: from pleu12.prod.google.com ([2002:a17:903:41cc:b0:2cb:97d4:2e28]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:350c:b0:2c9:bf82:dd11 with SMTP id d9443c01a7336-2d3177f6e83mr61151075ad.7.1786469331912; Tue, 11 Aug 2026 10:28:51 -0700 (PDT) Date: Tue, 11 Aug 2026 10:28:51 -0700 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20260728144954.355376-18-dwmw2@infradead.org> Message-ID: Subject: Re: [PATCH v7 17/36] KVM: x86: Allow KVM master clock mode when TSCs are offset from each other From: Sean Christopherson To: David Woodhouse Cc: Paolo Bonzini , Jonathan Corbet , Shuah Khan , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Vitaly Kuznetsov , Juergen Gross , Boris Ostrovsky , Paul Durrant , Jonathan Cameron , Sascha Bischoff , Marc Zyngier , Joey Gouly , Jack Allister , Dongli Zhang , joe.jin@oracle.com, kvm@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, xen-devel@lists.xenproject.org, linux-kselftest@vger.kernel.org Content-Type: text/plain; charset="us-ascii" On Tue, Aug 11, 2026, David Woodhouse wrote: > On Tue, 2026-08-11 at 09:40 -0700, Sean Christopherson wrote: > > > > > > > > FWIW in my local tree I've just extended the pvclock_migration_test to > > > test precisely the thing you were concerned about: three vCPUs with > > > divergent TSC offsets, migrated by setting each vCPU's TSC and then > > > invoking KVM_SET_CLOCK_GUEST once, through vCPU0. > > > > I wasn't actually concerned about migration, I was concerned about time going > > backwards from the guest's perspective. > > But KVM_[SG]ET_CLOCK_GUEST is *purely* for migration. Huh? I raised my concern in the context of "Allow KVM master clock mode when TSCs are offset from each other", and AFAICT, nothing ensures that won't cause problems. Aaah, it clears PVCLOCK_TSC_STABLE_BIT and relies on the guest to clean up the mess. So the guest won't see time go backwards, but it could see time stop for an extended duration, or jump forward. E.g. if tsc_timestamp is set to a per-VM value, and vCPU0's offset matches the effectively offset, but vCPU1's offset does not, then vCPU1 could see time that is either far in the past or far in the future when doing __pvclock_clocksource_read(). If vCPU1 computes time that's in the future, the guest would see a potentially massive jump forward in time. And after that, or if vCPU1 computes time that's in the past, if the "future" vCPU stopped running, the "past" vCPU would see time stop as last_value would be ahead of the "past" vCPU's current time for a very long while. if ((valid_flags & PVCLOCK_TSC_STABLE_BIT) && (flags & PVCLOCK_TSC_STABLE_BIT)) return ret; /* * Assumption here is that last_value, a global accumulator, always goes * forward. If we are less than that, we should not be much smaller. * We assume there is an error margin we're inside, and then the correction * does not sacrifice accuracy. * * For reads: global may have changed between test and return, * but this means someone else updated poked the clock at a later time. * We just need to make sure we are not seeing a backwards event. * * For updates: last_value = ret is not enough, since two vcpus could be * updating at the same time, and one of them could be slightly behind, * making the assumption that last_value always go forward fail to hold. */ last = raw_atomic64_read(&last_value); do { if (ret <= last) return last; } while (!raw_atomic64_try_cmpxchg(&last_value, &last, ret)); > And your variant just added a dependency on wallclock time back into it Can you elaborate? I'm guessing I don't entirely understand what you mean by wallclock time. > again, where wallclock should *only* be used for setting the TSC, and even > then *only* for a live *migration* to a different host, not a live *update* > via kexec/KHO on the same host, where the TSC should be restored as an offset > from the host TSC.