From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pf1-f197.google.com (mail-pf1-f197.google.com [209.85.210.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1F7C03793CB for ; Tue, 11 Aug 2026 18:41:40 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.210.197 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786473701; cv=none; b=XEKNDOm10NQYtG73vwTO9H8Rba8jicd7umJ5exwF0tU8A989VDNQG7eX0fq/PEBwGPWTKnLg5cLT3eFvqrpkcrcWgiMdwK2MmthqVbpaFaL5dq4Ck+GSl2aNJIq1PosQb3m62+czr+ZERGgbcm4ZbRBeYx2PupQ/vRBhH/E9HRk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786473701; c=relaxed/simple; bh=94Lt8TmaqiNs5FtCiFHbzvvHftTmibgvQkuT4qx1sEQ=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=QKrb4dBYGlU9QgO52cvCMPoV+boPnBLAZcWhc9OaPtaKJA909KpIczofBKFQMaV+yy6XlN5Zpa6W9QSyeqHnDJaCX+eApalUaf244ak2uYZ56tVH0yMeQ8qyFlbKrxddcjeXMv1v6xRJHImwUsP405Th+nukuqXneh1ZjSN3rPM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=jdib0Te2; arc=none smtp.client-ip=209.85.210.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--seanjc.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="jdib0Te2" Received: by mail-pf1-f197.google.com with SMTP id d2e1a72fcca58-84859a64079so271591b3a.3 for ; Tue, 11 Aug 2026 11:41:40 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1786473699; x=1787078499; darn=vger.kernel.org; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date:from:to:cc :subject:date:message-id:reply-to:content-type; bh=fJOzqFK/vfJSw1ado9uoaagkeuADN2fLa/8s2FkuGPk=; b=jdib0Te2Cm3baCHT5yBcxvk0rx7m+lSp0Gemjd24HzyFGYX/y2vtTteCsbsT6yqrZ/ gM5hZgaDgWBgXpvT+LIOOLxPza63VZa7YBt5RC7chDqhyvmjYh7ixTJtSWGg285tBNHC pqj5MDGvbnpnwJZ+qDNodKm+P9WGUB8yWU2cK2jY5o80B3bmFc1ZndE4ZwKBpoPp1Dv1 eItYviMa07/cvxn9kn3pC5nJaksLQEo+aoe98Z8j7jmouTVP3Uuf29C+2gVpQg6azIXn PFx0gkGRmAvTbAzh9QioDf3wUJKLLsSSGb/QqcfAgcfo1qJgmPqonMgE6onMAVNDjWes 2Upg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1786473699; x=1787078499; h=content-transfer-encoding:content-type:cc:to:from:subject :message-id:references:mime-version:in-reply-to:date :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=fJOzqFK/vfJSw1ado9uoaagkeuADN2fLa/8s2FkuGPk=; b=eRDENyX9+VoDxMIBTi3wSi4EC3e79hsTuKNxXoI4aJwpDJCT6/LBevh0emQj8v+q4G onOELkBNJC5xIinDkKFzS73AdENA59pF2lbBHTqdlIjlvHA8dBVuy+h7Qc43IVFMEUK5 i57TrMCqLD4xj4QXTrxoJ/4CuFOItjCbUEUL8cgumDSNMx9hF7X7F/zX+XcymvKMxemW Rs1UHiGr2eY6mijSxK/+hC//Z2sxptXkPVS3Hyh6hxZhvuLKvRqvO5FSPdBWhOIuH7Bq kMaKtR99Vrw8/U8iI0CDoB3n+27HN8CjnstdRGU8wfvY1lRySuXf6uRCLm3fx/wmvoSq zXRg== X-Forwarded-Encrypted: i=1; AHgh+Rpc+mtK9wU4Fr9MMGmBVOb/M6xfp+Wq3u0QX8oblovW3BEVbMh+yaJ8qrF0rUVIevkb2koScrH4cORvbhs=@vger.kernel.org X-Gm-Message-State: AOJu0YxSLAgLiG4kYchtMTBXRpYpA0E8RrCPnYBxdQrN/WFlPnmPgHCf gghShAk23i9uSg6/xGEf3LlMN44Q1HypbtM322OxwDjGuRKtHnzAcShnGC1hfRwTxvZM9cJqU6K Kg00few== X-Received: from pfbih2.prod.google.com ([2002:a05:6a00:8c02:b0:84e:530a:923a]) (user=seanjc job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a00:2355:b0:848:2c6c:dfe3 with SMTP id d2e1a72fcca58-84fa86c28d9mr5596667b3a.17.1786473699196; Tue, 11 Aug 2026 11:41:39 -0700 (PDT) Date: Tue, 11 Aug 2026 11:41:37 -0700 In-Reply-To: <6ab49538675d97f1f4bf01574b1b066aae0bc05c.camel@infradead.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <6ab49538675d97f1f4bf01574b1b066aae0bc05c.camel@infradead.org> Message-ID: Subject: Re: [PATCH v7 17/36] KVM: x86: Allow KVM master clock mode when TSCs are offset from each other From: Sean Christopherson To: David Woodhouse Cc: Paolo Bonzini , Jonathan Corbet , Shuah Khan , Thomas Gleixner , Ingo Molnar , Borislav Petkov , Dave Hansen , x86@kernel.org, "H. Peter Anvin" , Vitaly Kuznetsov , Juergen Gross , Boris Ostrovsky , Paul Durrant , Jonathan Cameron , Sascha Bischoff , Marc Zyngier , Joey Gouly , Jack Allister , Dongli Zhang , joe.jin@oracle.com, kvm@vger.kernel.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, xen-devel@lists.xenproject.org, linux-kselftest@vger.kernel.org Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: quoted-printable On Tue, Aug 11, 2026, David Woodhouse wrote: > On Tue, 2026-08-11 at 10:28 -0700, Sean Christopherson wrote: > > On Tue, Aug 11, 2026, David Woodhouse wrote: > > > On Tue, 2026-08-11 at 09:40 -0700, Sean Christopherson wrote: > > > >=20 > > > > >=20 > > > > > FWIW in my local tree I've just extended the pvclock_migration_te= st to > > > > > test precisely the thing you were concerned about: three vCPUs wi= th > > > > > divergent TSC offsets, migrated by setting each vCPU's TSC and th= en > > > > > invoking KVM_SET_CLOCK_GUEST once, through vCPU0.=20 > > > >=20 > > > > I wasn't actually concerned about migration, I was concerned about = time going > > > > backwards from the guest's perspective. > > >=20 > > > But KVM_[SG]ET_CLOCK_GUEST is *purely* for migration.=20 > >=20 > > Huh?=C2=A0 I raised my concern in the context of "Allow KVM master cloc= k mode when > > TSCs are offset from each other", and AFAICT, nothing ensures that won'= t cause > > problems. > >=20 > > Aaah, it clears PVCLOCK_TSC_STABLE_BIT and relies on the guest to clean= up the > > mess.=C2=A0 So the guest won't see time go backwards, but it could see = time stop for > > an extended duration, or jump forward. >=20 > It shouldn't. Each vCPU gets its *own* pvclock structure, tailored to > *its* offset. They should all see *identical* results. OMG, I hate this code. After literally hours of staring at this, and even = typing up a lengthy example of why guest time would go off the rails, I finally sp= otted that l1_tsc_offset is accounted for by the call to kvm_read_l1_tsc(). FML. Thanks for being patient and not flaming me too much :-) > But yeah, if the guest does that then we can't set the > PVCLOCK_TSC_STABLE_BIT. > >=20 > > > And your variant just added a dependency on wallclock time back into = it > >=20 > > Can you elaborate?=C2=A0 I'm guessing I don't entirely understand what = you mean by > > wallclock time. >=20 > The system_time field? The unspecified might-be-UTC-might-have-leap-secon= ds one :) Ok, I think I finally understand the goal. I got turned around by the comb= ination of the name SET_CLOCK_GUEST and the full pvclock structure being passed to = the guest. I was expecting SET_CLOCK_GUEST to literally set the entire clock, = e.g. mul+shift, timestamp, etc. But all of that metadata is just a means to an end: the one and only goal i= s to calculate the per-VM kvmclock_offset for the "new" host's TSC+time snapshot= , by computing the nanoseconds delta for the new snapshot as if it the guest obs= erved the TSC while running on the old host. And that is done in the kernel instead of in userspace to minimize the amou= nt of slop introduced due to delay between taking the snapshot and computing the = offset. And for similar reasons, it's undesirable for userspace to redo the GET on = the target to compute the explicit offset, because there would be a massive TOC= TOU issue, e.g. if KVM took a new masterclock snapshot between GET and SET. After working through that, I feel quite strongly that we should have asymm= etric names for the GET vs. SET flows, because the actually functionality is also= very asymmetric. And setting a field, and only that field, that's doesn't have = a direct association in the userspace payload is very surprising when the nam= e of the ioctl suggests a restoration of the entire payload. E.g. KVM_GET_REFERENCE_PVCLOCK and then KVM_SET_REFERENCE_PVCLOCK_OFFSET? = Or maybe UPDATE, CALCULATE, REFRESH, or COMPUTE instead of SET? I think I'd v= ote for SET even though it's convoluted, because pretty much everyone associate= s SET with the restore side of save/restore.