* Re: [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC [not found] <20260826214952.129BB1F000E9@smtp.kernel.org> @ 2026-08-28 22:46 ` David Woodhouse 2026-09-01 19:33 ` Sean Christopherson 0 siblings, 1 reply; 4+ messages in thread From: David Woodhouse @ 2026-08-28 22:46 UTC (permalink / raw) To: seanjc, pbonzini, Thomas Gleixner Cc: kvm, linux-kernel, paul, dongli.zhang, sashiko-reviews [-- Attachment #1: Type: text/plain, Size: 691 bytes --] On Wed, 2026-08-26 at 21:49 +0000, sashiko-bot@kernel.org wrote: > [Severity: High] > Does removing KVM's local TSC clamping expose guests to mismatched time pairs? > > When establishing the guest's PV clock, KVM provides a master clock > snapshot pairing a host TSC with a system time. The previous read_tsc() > logic explicitly clamped the TSC to cycle_last to maintain mathematical > consistency. I think this is basically the same class of issue I brought up in https://lore.kernel.org/all/87v7beb7s3.ffs@fw13/ and Thomas said the right thing is just to ignore it? If the discrepancy was more than a few cycles, we shouldn't be in masterclock mode anyway, should we? [-- Attachment #2: smime.p7s --] [-- Type: application/pkcs7-signature, Size: 6179 bytes --] ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC 2026-08-28 22:46 ` [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC David Woodhouse @ 2026-09-01 19:33 ` Sean Christopherson 0 siblings, 0 replies; 4+ messages in thread From: Sean Christopherson @ 2026-09-01 19:33 UTC (permalink / raw) To: David Woodhouse Cc: pbonzini, Thomas Gleixner, kvm, linux-kernel, paul, dongli.zhang, sashiko-reviews On Fri, Aug 28, 2026, David Woodhouse wrote: > On Wed, 2026-08-26 at 21:49 +0000, sashiko-bot@kernel.org wrote: > > [Severity: High] > > Does removing KVM's local TSC clamping expose guests to mismatched time pairs? > > > > When establishing the guest's PV clock, KVM provides a master clock > > snapshot pairing a host TSC with a system time. The previous read_tsc() > > logic explicitly clamped the TSC to cycle_last to maintain mathematical > > consistency. > > I think this is basically the same class of issue I brought up in > https://lore.kernel.org/all/87v7beb7s3.ffs@fw13/ and Thomas said the > right thing is just to ignore it? > > If the discrepancy was more than a few cycles, we shouldn't be in > masterclock mode anyway, should we? Agreed. If a host CPU really has a TSC that is observably behind others, KVM will mark the TSC unstable when loading a vCPU on that pCPU. s64 tsc_delta = !vcpu->arch.last_host_tsc ? 0 : rdtsc() - vcpu->arch.last_host_tsc; if (tsc_delta < 0) mark_tsc_unstable("KVM discovered backwards TSC"); And as Thomas pointed out, the discrepancy will show up at some point. E.g. even if the kernel provided a perfect pair, RDTSC from the guest would read a too-low value since KVM has historically required identical offsets to use masterclock, i.e. the guest would still be able to observe time going backwards if a vCPU did __pvclock_clocksource_read() with a TSC that is behind the masterclock reference. ^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH v10 00/21] KVM: x86: Cleaning up the KVM clock mess, part 1
@ 2026-08-26 21:32 Sean Christopherson
2026-08-26 21:33 ` [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC Sean Christopherson
0 siblings, 1 reply; 4+ messages in thread
From: Sean Christopherson @ 2026-08-26 21:32 UTC (permalink / raw)
To: Sean Christopherson, Paolo Bonzini
Cc: kvm, linux-kernel, Paul Durrant, David Woodhouse, Dongli Zhang
Part 1 of David's series to clean up the host-side mess of kvmclock. The cut
point in this version is still quite arbitrary, but what's left for "part 2"
fits nicely into 2-3 buckets:
1: Masterclock overhaul
2. New uAPI and tests
v10:
- Drop the stale reference to using div_u64() for the now-removed CPUID leaf.
[Sashiko]
- Don't bother with a enabling the pvclock_gtod notifier on 32-bit, and go
straight to ktime_mono_to_any(). [Sashiko]
v9:
- https://lore.kernel.org/all/20260810225500.869288-1-seanjc@google.com
- Pull in "Compute kvmclock base without pvclock_gtod_data" before "Avoid NTP
frequency skew for KVM clock on 32-bit host" to avoid having to provide what
is effectively an open-coded version of ktime_mono_to_any() to avoid data
races. [Sashiko].
- Pull in use of kernel timekeeping snapshots, split in three patches purely
to aid review and bisection.
- Check and clear KVM_REQ_MASTERCLOCK_UPDATE in kvm_arch_vcpu_postcreate().
[me, Sashiko]
- Add the missing return value in "Move "no master clock" fallback from...".
[Sashiko]
v8:
- Obviously cut partway through the full series.
- Update commit references for 633d7652f80f ("KVM: x86/xen: Do not corrupt
KVM clock in kvm_xen_shared_info_init()").
- Chunk the "Restructure ..." patches into individual logical changes
("Restructure kvm_guest_time_update() for TSC upscaling" in particular was
quite gnarly).
- Ensure the TSC+freq pair during guest time updates happens on a single
CPU. [Sashiko]
- Keep __get_kvmclock() as a masterclock-specific helper to avoid deep
indentation and gotos.
- Do waaaaay too much spelunking to try piece together the history of some
of the many warts.
- Ensure preemption is disabled in get_kvmclock() when getting the TSC freq.
v7: https://lore.kernel.org/all/20260728144954.355376-1-dwmw2@infradead.org
David Woodhouse (14):
KVM: x86: Improve accuracy of KVM clock when TSC scaling is in force
KVM: x86: Explicitly disable TSC scaling without CONSTANT_TSC
KVM: x86: Activate master clock immediately on vCPU creation
KVM: x86: Compute kvmclock base without pvclock_gtod_data
KVM: x86: Avoid NTP frequency skew for KVM clock on 32-bit host
KVM: x86: Wrap all of __get_kvmclock_master_clock() with
CONFIG_X86_64=y
KVM: x86: Fall back to non-master-clock if clockread fails in
get_kvmclock()
KVM: x86: Fix KVM clock precision in get_kvmclock() with TSC scaling
KVM: x86: Use get_kvmclock() in kvm_get_wall_clock_epoch()
KVM: x86: Fix compute_guest_tsc() to handle negative time deltas
KVM: x86: Upscale TSC to "now", not master clock when updating PV
clocks
KVM: x86: Simplify and comment kvm_get_time_scale()
KVM: x86: Remove implicit rdtsc() from kvm_compute_l1_tsc_offset()
KVM: x86: Use kernel timekeeping snapshots for getting kvmclock time
since boot
Sean Christopherson (7):
KVM: x86: Update "last guest TSC" snapshot prior to enabling
IRQs/preemption
KVM: x86: Drop unnecessary CPU pinning when computing/getting kvmclock
KVM: x86: Move "no master clock" fallback from __get_kvmclock() to
get_kvmclock()
KVM: x86: Disable preemption, not IRQs, when getting TSC+freq pair
KVM: x86: Make master clock logic in guest PV clock updates 64-bit
only
KVM: x86: Use kernel timekeeping snapshot for monotonic clock
KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC
arch/x86/include/asm/kvm_host.h | 6 +-
arch/x86/kvm/cpuid.c | 1 +
arch/x86/kvm/msrs.c | 3 +-
arch/x86/kvm/svm/svm.c | 3 +-
arch/x86/kvm/vmx/vmx.c | 10 +
arch/x86/kvm/x86.c | 480 +++++++++++++++-----------------
arch/x86/kvm/x86.h | 3 +-
7 files changed, 240 insertions(+), 266 deletions(-)
base-commit: 76671054f9a1ff6abb976583cd8da37650acdc97
--
2.55.0.887.g758fc8c411-goog
^ permalink raw reply [flat|nested] 4+ messages in thread* [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC 2026-08-26 21:32 [PATCH v10 00/21] KVM: x86: Cleaning up the KVM clock mess, part 1 Sean Christopherson @ 2026-08-26 21:33 ` Sean Christopherson 2026-09-01 19:17 ` Sean Christopherson 0 siblings, 1 reply; 4+ messages in thread From: Sean Christopherson @ 2026-08-26 21:33 UTC (permalink / raw) To: Sean Christopherson, Paolo Bonzini Cc: kvm, linux-kernel, Paul Durrant, David Woodhouse, Dongli Zhang Replace the KVM-private vgettsc()+/do_do_realtime() timekeeping reimplementation with calls to the recently crafted, generic ktime_get_snapshot_id() interface. As noted previously, the snapshot provides both the system time and the raw_cycles (TSC), atomically paired using a sequence counter. With great pleasure, delete the now unused read_tsc() and vgettsc() Signed-off-by: David Woodhouse <dwmw@amazon.co.uk> [sean: separate from other conversions, express joy at vgettsc()'s demise] Signed-off-by: Sean Christopherson <seanjc@google.com> --- arch/x86/kvm/x86.c | 85 ++++------------------------------------------ 1 file changed, 6 insertions(+), 79 deletions(-) diff --git a/arch/x86/kvm/x86.c b/arch/x86/kvm/x86.c index 0352bd147c40..34eadc75fee4 100644 --- a/arch/x86/kvm/x86.c +++ b/arch/x86/kvm/x86.c @@ -1383,82 +1383,6 @@ void kvm_synchronize_tsc(struct kvm_vcpu *vcpu, u64 *user_value) #ifdef CONFIG_X86_64 -static u64 read_tsc(void) -{ - u64 ret = (u64)rdtsc_ordered(); - u64 last = pvclock_gtod_data.clock.cycle_last; - - if (likely(ret >= last)) - return ret; - - /* - * GCC likes to generate cmov here, but this branch is extremely - * predictable (it's just a function of time and the likely is - * very likely) and there's a data dependence, so force GCC - * to generate a branch instead. I don't barrier() because - * we don't actually need a barrier, and if this function - * ever gets inlined it will generate worse code. - */ - asm volatile (""); - return last; -} - -static inline u64 vgettsc(struct pvclock_clock *clock, u64 *tsc_timestamp, - int *mode) -{ - u64 tsc_pg_val; - long v; - - switch (clock->vclock_mode) { - case VDSO_CLOCKMODE_HVCLOCK: - if (hv_read_tsc_page_tsc(hv_get_tsc_page(), - tsc_timestamp, &tsc_pg_val)) { - /* TSC page valid */ - *mode = VDSO_CLOCKMODE_HVCLOCK; - v = (tsc_pg_val - clock->cycle_last) & - clock->mask; - } else { - /* TSC page invalid */ - *mode = VDSO_CLOCKMODE_NONE; - } - break; - case VDSO_CLOCKMODE_TSC: - *mode = VDSO_CLOCKMODE_TSC; - *tsc_timestamp = read_tsc(); - v = (*tsc_timestamp - clock->cycle_last) & - clock->mask; - break; - default: - *mode = VDSO_CLOCKMODE_NONE; - } - - if (*mode == VDSO_CLOCKMODE_NONE) - *tsc_timestamp = v = 0; - - return v * clock->mult; -} - -static int do_realtime(struct timespec64 *ts, u64 *tsc_timestamp) -{ - struct pvclock_gtod_data *gtod = &pvclock_gtod_data; - unsigned long seq; - int mode; - u64 ns; - - do { - seq = read_seqcount_begin(>od->seq); - ts->tv_sec = gtod->wall_time_sec; - ns = gtod->clock.base_cycles; - ns += vgettsc(>od->clock, tsc_timestamp, &mode); - ns >>= gtod->clock.shift; - } while (unlikely(read_seqcount_retry(>od->seq, seq))); - - ts->tv_sec += __iter_div_u64_rem(ns, NSEC_PER_SEC, &ns); - ts->tv_nsec = ns; - - return mode; -} - static bool kvm_snapshot_has_tsc(struct system_time_snapshot *snap, u64 *tsc_timestamp) { @@ -1525,11 +1449,14 @@ bool kvm_get_monotonic_and_clockread(s64 *kernel_ns, u64 *tsc_timestamp) static bool kvm_get_walltime_and_clockread(struct timespec64 *ts, u64 *tsc_timestamp) { - /* checked again under seqlock below */ - if (!gtod_is_based_on_tsc(pvclock_gtod_data.clock.vclock_mode)) + struct system_time_snapshot snap = {}; + + ktime_get_snapshot_id(CLOCK_REALTIME, &snap); + if (!kvm_snapshot_has_tsc(&snap, tsc_timestamp)) return false; - return gtod_is_based_on_tsc(do_realtime(ts, tsc_timestamp)); + *ts = ktime_to_timespec64(snap.systime); + return true; } #endif -- 2.55.0.887.g758fc8c411-goog ^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC 2026-08-26 21:33 ` [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC Sean Christopherson @ 2026-09-01 19:17 ` Sean Christopherson 0 siblings, 0 replies; 4+ messages in thread From: Sean Christopherson @ 2026-09-01 19:17 UTC (permalink / raw) To: Paolo Bonzini, kvm, linux-kernel, Paul Durrant, David Woodhouse, Dongli Zhang On Wed, Aug 26, 2026, Sean Christopherson wrote: > Replace the KVM-private vgettsc()+/do_do_realtime() timekeeping > reimplementation with calls to the recently crafted, generic > ktime_get_snapshot_id() interface. As noted previously, the snapshot > provides both the system time and the raw_cycles (TSC), atomically paired > using a sequence counter. > > With great pleasure, delete the now unused read_tsc() and vgettsc() > > Signed-off-by: David Woodhouse <dwmw@amazon.co.uk> > [sean: separate from other conversions, express joy at vgettsc()'s demise] > Signed-off-by: Sean Christopherson <seanjc@google.com> Note to self, this needs to be fixed up so that David is the Author. I'll check the other patches too, I suspect I forgot to specify --author when splitting the original patch into 3. ^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-01 19:33 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
[not found] <20260826214952.129BB1F000E9@smtp.kernel.org>
2026-08-28 22:46 ` [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC David Woodhouse
2026-09-01 19:33 ` Sean Christopherson
2026-08-26 21:32 [PATCH v10 00/21] KVM: x86: Cleaning up the KVM clock mess, part 1 Sean Christopherson
2026-08-26 21:33 ` [PATCH v10 21/21] KVM: x86: Use kernel timekeeping snapshot to get walltime+TSC Sean Christopherson
2026-09-01 19:17 ` Sean Christopherson
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®