mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision
@ 2026-09-16  2:32 Zhan Xusheng
  2026-09-16  2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
                   ` (4 more replies)
  0 siblings, 5 replies; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16  2:32 UTC (permalink / raw)
  To: tglx, thomas.weissschuh
  Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng

Patch 3/4 is the fix.  vdso_time_update_aux() rounds the CLOCK_AUX base
down to whole nanoseconds, so the vDSO reports 0 or 1 ns below
ktime_get_aux() for the same clock.  1/4 and 2/4 prepare the division
helper it uses.  4/4 is independent of the fix and can be dropped without
affecting the rest.

Changes since v4:

 - Every patch carries a single From: now.  format.from was set to a
   different address than user.email, so format-patch emitted a header
   From plus an in-body From, and send-email then added a second in-body
   one.  That setting is gone; git am records zhanxusheng@xiaomi.com for
   all four patches.

 - 1/4: rewrite the changelog, which referred to a loop it had not
   introduced.  Also put __iter_div_u64_rem()'s signature on one line, so
   that the two helpers in the file do not end up in different styles.

 - 2/4: put the __iter_div64_u64_rem() signature on one line, 90
   characters.  Rewrite the changelog and correct a size claim that did
   not survive re-measurement: the 16 bytes come from folding the two
   open-coded loops into one inlined helper, and the u32 return type makes
   no difference to vsyscall.o on x86-64.

 - 4/4: rewrite the subject and changelog using the wording the code uses
   for this.  "clock id dispatch" and "dispatch mask" were mine, not the
   kernel's.

 - 3/4 is unchanged.

 - The tag block follows the maintainer-tip.rst order now, which puts the
   author's Signed-off-by ahead of Reviewed-by.

Testing on x86-64: checkpatch --strict, gcc W=1 and make LLVM=1 are all
clean.  Built and booted with kernel/configs/x86_debug.config plus
panic_on_warn=1; no warning fired and the log holds no lockdep or
debug-object complaint.  A guest with an aux clock enabled saw no vDSO
reading ahead of a later ktime_get_aux() in 100000 samples at each of
offset 0, +5.123456789 and a negative offs_aux.

The generated code in vsyscall.o, vdso64 and vdso32 is byte-identical to
v4, so the reflowed lines changed nothing.

v4: https://lore.kernel.org/all/20260902033801.2912699-1-zhanxusheng@xiaomi.com

Zhan Xusheng (4):
  vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
  vdso/math64: Add and use __iter_div64_u64_rem()
  vdso/vsyscall: Keep the CLOCK_AUX base scaled
  vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask

 include/vdso/math64.h   | 31 ++++++++++++++++++++++++++-----
 kernel/time/vsyscall.c  | 26 ++++++++++----------------
 lib/vdso/gettimeofday.c |  2 ++
 3 files changed, 38 insertions(+), 21 deletions(-)


base-commit: 9b87fdc9af2fbfcdb5c24a64139685ef80f6573f
-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
  2026-09-16  2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
@ 2026-09-16  2:32 ` Zhan Xusheng
  2026-09-25  7:44   ` Thomas Weißschuh
  2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
  2026-09-16  2:32 ` [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem() Zhan Xusheng
                   ` (3 subsequent siblings)
  4 siblings, 2 replies; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16  2:32 UTC (permalink / raw)
  To: tglx, thomas.weissschuh
  Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng

__iter_div_u64_rem() divides by repeated subtraction, because its callers
only ever produce a small quotient.  A barrier inside the subtraction loop
keeps the compiler from replacing the loop with a division:

	asm("" : "+rm"(dividend));

The "rm" constraint permits a memory operand.  clang picks it and spills
the dividend inside the loop, at 32-bit and 64-bit alike.  The helper sits
on the clock_gettime() fast path through vdso_set_timespec(), so the spill
is not free: with clang 18 the x86 vDSO is 64 bytes larger in vdso64 text
and 112 bytes larger in vdso32 text than with a register-only barrier.

Use OPTIMIZER_HIDE_VAR(), which is the register-only form of the same
barrier and the form the rest of the kernel uses.  gcc 13 generates
identical code either way, and neither compiler turns the loop into a
division at either width.

Put the function signature on one line while touching it, and reformat the
comment, which described the asm() that is now gone.

Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
---
 include/vdso/math64.h | 11 ++++++-----
 1 file changed, 6 insertions(+), 5 deletions(-)

diff --git a/include/vdso/math64.h b/include/vdso/math64.h
index 22ae212f8b28..c628d6cf447c 100644
--- a/include/vdso/math64.h
+++ b/include/vdso/math64.h
@@ -2,15 +2,16 @@
 #ifndef __VDSO_MATH64_H
 #define __VDSO_MATH64_H
 
-static __always_inline u32
-__iter_div_u64_rem(u64 dividend, u32 divisor, u64 *remainder)
+static __always_inline u32 __iter_div_u64_rem(u64 dividend, u32 divisor, u64 *remainder)
 {
 	u32 ret = 0;
 
 	while (dividend >= divisor) {
-		/* The following asm() prevents the compiler from
-		   optimising this loop into a modulo operation.  */
-		asm("" : "+rm"(dividend));
+		/*
+		 * Prevent the compiler from optimising this loop into a
+		 * modulo operation.
+		 */
+		OPTIMIZER_HIDE_VAR(dividend);
 
 		dividend -= divisor;
 		ret++;
-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem()
  2026-09-16  2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
  2026-09-16  2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
@ 2026-09-16  2:32 ` Zhan Xusheng
  2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
  2026-09-16  2:32 ` [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled Zhan Xusheng
                   ` (2 subsequent siblings)
  4 siblings, 1 reply; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16  2:32 UTC (permalink / raw)
  To: tglx, thomas.weissschuh
  Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng

The vDSO basetimes for CLOCK_MONOTONIC and CLOCK_BOOTTIME are stored in
the scaled nanoseconds of tkr_mono, so normalising them requires a
division by NSEC_PER_SEC << shift.  That divisor does not fit the u32
parameter of __iter_div_u64_rem(), so update_vdso_time_data() open-codes
the same iterative division twice.

Repeated subtraction is the appropriate form at these two sites because
the quotient never exceeds one.  accumulate_nsecs_to_secs() keeps
xtime_nsec below one scaled second, and the offset added to it is a
normalised timespec64 fraction, so the dividend stays below twice the
divisor.  That bound holds on every architecture, which matters more here
than the cost of a division on any particular one.

Add __iter_div64_u64_rem(), the u64-divisor counterpart of
__iter_div_u64_rem(), and use it at both sites.  Return the quotient as a
u32 for consistency with the u32-divisor variant; the bound above leaves
no use for a wider type.

Store the remainder directly into the basetime, as the coarse clocks
already do.  The CLOCK_BOOTTIME copy of the CLOCK_MONOTONIC values then
takes both fields from the same place.

Folding the two open-coded loops into one inlined helper removes 16 bytes
of vsyscall.o text on x86-64 with gcc 13.

No functional change.

Suggested-by: David Laight <david.laight.linux@gmail.com>
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
---
 include/vdso/math64.h  | 20 ++++++++++++++++++++
 kernel/time/vsyscall.c | 16 +++++-----------
 2 files changed, 25 insertions(+), 11 deletions(-)

diff --git a/include/vdso/math64.h b/include/vdso/math64.h
index c628d6cf447c..55b45f5cf615 100644
--- a/include/vdso/math64.h
+++ b/include/vdso/math64.h
@@ -22,6 +22,26 @@ static __always_inline u32 __iter_div_u64_rem(u64 dividend, u32 divisor, u64 *re
 	return ret;
 }
 
+static __always_inline u32 __iter_div64_u64_rem(u64 dividend, u64 divisor, u64 *remainder)
+{
+	u32 ret = 0;
+
+	while (dividend >= divisor) {
+		/*
+		 * Prevent the compiler from optimising this loop into a
+		 * modulo operation.
+		 */
+		OPTIMIZER_HIDE_VAR(dividend);
+
+		dividend -= divisor;
+		ret++;
+	}
+
+	*remainder = dividend;
+
+	return ret;
+}
+
 #if defined(CONFIG_ARCH_SUPPORTS_INT128) && defined(__SIZEOF_INT128__)
 
 #ifndef mul_u64_u32_add_u64_shr
diff --git a/kernel/time/vsyscall.c b/kernel/time/vsyscall.c
index aa59919b8f2c..f43dd3f4744b 100644
--- a/kernel/time/vsyscall.c
+++ b/kernel/time/vsyscall.c
@@ -41,14 +41,12 @@ static inline void update_vdso_time_data(struct vdso_time_data *vdata, struct ti
 
 	nsec = tk->tkr_mono.xtime_nsec;
 	nsec += ((u64)tk->wall_to_monotonic.tv_nsec << tk->tkr_mono.shift);
-	while (nsec >= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift)) {
-		nsec -= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift);
-		vdso_ts->sec++;
-	}
-	vdso_ts->nsec	= nsec;
+	vdso_ts->sec	+= __iter_div64_u64_rem(nsec, (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+						&vdso_ts->nsec);
 
 	/* Copy MONOTONIC time for BOOTTIME */
 	sec	= vdso_ts->sec;
+	nsec	= vdso_ts->nsec;
 	/* Add the boot offset */
 	sec	+= tk->monotonic_to_boot.tv_sec;
 	nsec	+= (u64)tk->monotonic_to_boot.tv_nsec << tk->tkr_mono.shift;
@@ -56,12 +54,8 @@ static inline void update_vdso_time_data(struct vdso_time_data *vdata, struct ti
 	/* CLOCK_BOOTTIME */
 	vdso_ts		= &vc[CS_HRES_COARSE].basetime[CLOCK_BOOTTIME];
 	vdso_ts->sec	= sec;
-
-	while (nsec >= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift)) {
-		nsec -= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift);
-		vdso_ts->sec++;
-	}
-	vdso_ts->nsec	= nsec;
+	vdso_ts->sec	+= __iter_div64_u64_rem(nsec, (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+						&vdso_ts->nsec);
 
 	/* CLOCK_MONOTONIC_RAW */
 	vdso_ts		= &vc[CS_RAW].basetime[CLOCK_MONOTONIC_RAW];
-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled
  2026-09-16  2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
  2026-09-16  2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
  2026-09-16  2:32 ` [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem() Zhan Xusheng
@ 2026-09-16  2:32 ` Zhan Xusheng
  2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
  2026-09-16  2:32 ` [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask Zhan Xusheng
  2026-09-30 12:41 ` [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Thomas Gleixner
  4 siblings, 1 reply; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16  2:32 UTC (permalink / raw)
  To: tglx, thomas.weissschuh
  Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng

The vDSO basetime of a clock is stored in the scaled nanoseconds of
tkr_mono, so that the reader can floor the base and the cycle delta
together in vdso_calc_ns().

vdso_time_update_aux() instead shifts the base down to nanoseconds, adds
the offset, and shifts it back up, which zeroes the fractional nanoseconds
of xtime_nsec.  The reader then floors the base and the delta separately:

  ktime_get_aux():  base + ((delta * mult + xtime_nsec) >> shift)
  vdso:             base + (xtime_nsec >> shift)
                         + ((delta * mult) >> shift)

Since floor(a) + floor(b) <= floor(a + b), the vDSO reports 0 or 1 ns
below the syscall for the same clock.  It is not a monotonicity problem:
across an update the step is floor(a + d) - floor(a) - floor(d), which is
0 or 1, never negative.

Add the offset in scaled nanoseconds as the other high resolution clocks
do, and normalise with __iter_div64_u64_rem() so that the stored base
stays below one second and the userspace fast-path does not iterate more
in __iter_div_u64_rem().

Only the sub-second field changes.  (a + (b << shift)) >> shift is exactly
(a >> shift) + b, so the seconds carried out of the normalisation are the
same as before; what the old form dropped was the low shift bits of the
remainder.

monotonic_to_aux.tv_nsec is a normalised timespec64 fraction, so it stays
below NSEC_PER_SEC even for a negative offset, and the sum stays below
2 * (NSEC_PER_SEC << shift).  The largest shift clocks_calc_mult_shift()
can pick is 32, which makes that 8.6e18 against a u64 limit of 1.8e19.

Fixes: 380b84e168e5 ("vdso/vsyscall: Update auxiliary clock data in the datapage")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
---
 kernel/time/vsyscall.c | 10 +++++-----
 1 file changed, 5 insertions(+), 5 deletions(-)

diff --git a/kernel/time/vsyscall.c b/kernel/time/vsyscall.c
index f43dd3f4744b..0e4b499328c0 100644
--- a/kernel/time/vsyscall.c
+++ b/kernel/time/vsyscall.c
@@ -155,11 +155,11 @@ void vdso_time_update_aux(struct timekeeper *tk)
 
 		vdso_ts->sec = tk->xtime_sec + tk->monotonic_to_aux.tv_sec;
 
-		nsec = tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift;
-		nsec += tk->monotonic_to_aux.tv_nsec;
-		vdso_ts->sec += __iter_div_u64_rem(nsec, NSEC_PER_SEC, &nsec);
-		nsec = nsec << tk->tkr_mono.shift;
-		vdso_ts->nsec = nsec;
+		nsec = tk->tkr_mono.xtime_nsec;
+		nsec += (u64)tk->monotonic_to_aux.tv_nsec << tk->tkr_mono.shift;
+		vdso_ts->sec += __iter_div64_u64_rem(nsec,
+						     (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+						     &vdso_ts->nsec);
 	}
 
 	__arch_update_vdso_clock(vc);
-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask
  2026-09-16  2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
                   ` (2 preceding siblings ...)
  2026-09-16  2:32 ` [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled Zhan Xusheng
@ 2026-09-16  2:32 ` Zhan Xusheng
  2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
  2026-09-30 12:41 ` [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Thomas Gleixner
  4 siblings, 1 reply; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16  2:32 UTC (permalink / raw)
  To: tglx, thomas.weissschuh
  Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng

__cvdso_clock_gettime_common() and __cvdso_clock_getres_common() convert
the clockid into a bitmask and match it against VDSO_HRES, VDSO_COARSE,
VDSO_RAW and VDSO_AUX:

	msk = 1U << clock;

vdso_clockid_valid() rejects anything above CLOCK_AUX_LAST beforehand,
and CLOCK_AUX_LAST is 23, so the shift count is in range.  Nothing
records that dependency though.  Raising MAX_AUX_CLOCKS beyond 16 moves
CLOCK_AUX_LAST to 32 and makes the shift undefined.

Add a BUILD_BUG_ON() at both conversion sites.  The condition is on a
function parameter rather than a constant, so it relies on the compiler
deriving the range from the vdso_clockid_valid() bail-out above it.  gcc
13 and clang 18 both do: the x86 vdso64 and vdso32 builds stay clean, and
raising MAX_AUX_CLOCKS to 17 trips the assert.

Suggested-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
---
 lib/vdso/gettimeofday.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/lib/vdso/gettimeofday.c b/lib/vdso/gettimeofday.c
index f7a591aba59f..ef4dcc614489 100644
--- a/lib/vdso/gettimeofday.c
+++ b/lib/vdso/gettimeofday.c
@@ -285,6 +285,7 @@ __cvdso_clock_gettime_common(const struct vdso_time_data *vd, clockid_t clock,
 	 * Convert the clockid to a bitmask and use it to check which
 	 * clocks are handled in the VDSO directly.
 	 */
+	BUILD_BUG_ON(clock >= BITS_PER_TYPE(msk));
 	msk = 1U << clock;
 	if (likely(msk & VDSO_HRES))
 		vc = &vc[CS_HRES_COARSE];
@@ -438,6 +439,7 @@ bool __cvdso_clock_getres_common(const struct vdso_time_data *vd, clockid_t cloc
 	 * Convert the clockid to a bitmask and use it to check which
 	 * clocks are handled in the VDSO directly.
 	 */
+	BUILD_BUG_ON(clock >= BITS_PER_TYPE(msk));
 	msk = 1U << clock;
 	if (msk & (VDSO_HRES | VDSO_RAW)) {
 		/*
-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
  2026-09-16  2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
@ 2026-09-25  7:44   ` Thomas Weißschuh
  2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
  1 sibling, 0 replies; 11+ messages in thread
From: Thomas Weißschuh @ 2026-09-25  7:44 UTC (permalink / raw)
  To: Zhan Xusheng
  Cc: tglx, luto, vincenzo.frascino, david.laight.linux, linux-kernel,
	zhanxusheng

On Wed, Sep 16, 2026 at 10:32:49AM +0800, Zhan Xusheng wrote:
> __iter_div_u64_rem() divides by repeated subtraction, because its callers
> only ever produce a small quotient.  A barrier inside the subtraction loop
> keeps the compiler from replacing the loop with a division:
> 
> 	asm("" : "+rm"(dividend));
> 
> The "rm" constraint permits a memory operand.  clang picks it and spills
> the dividend inside the loop, at 32-bit and 64-bit alike.  The helper sits
> on the clock_gettime() fast path through vdso_set_timespec(), so the spill
> is not free: with clang 18 the x86 vDSO is 64 bytes larger in vdso64 text
> and 112 bytes larger in vdso32 text than with a register-only barrier.
> 
> Use OPTIMIZER_HIDE_VAR(), which is the register-only form of the same
> barrier and the form the rest of the kernel uses.  gcc 13 generates
> identical code either way, and neither compiler turns the loop into a
> division at either width.
> 
> Put the function signature on one line while touching it, and reformat the
> comment, which described the asm() that is now gone.
> 
> Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>

Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>

> ---
>  include/vdso/math64.h | 11 ++++++-----
>  1 file changed, 6 insertions(+), 5 deletions(-)

(...)

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [tip: timers/vdso] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask
  2026-09-16  2:32 ` [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask Zhan Xusheng
@ 2026-09-29 19:27   ` tip-bot2 for Zhan Xusheng
  0 siblings, 0 replies; 11+ messages in thread
From: tip-bot2 for Zhan Xusheng @ 2026-09-29 19:27 UTC (permalink / raw)
  To: linux-tip-commits
  Cc: thomas.weissschuh, Zhan Xusheng, Thomas Gleixner, x86, linux-kernel

The following commit has been merged into the timers/vdso branch of tip:

Commit-ID:     ca46a07c4b68246c4e400ce3649ab7228a878d8f
Gitweb:        https://git.kernel.org/tip/ca46a07c4b68246c4e400ce3649ab7228a878d8f
Author:        Zhan Xusheng <zhanxusheng1024@gmail.com>
AuthorDate:    Wed, 16 Sep 2026 10:32:52 +08:00
Committer:     Thomas Gleixner <tglx@kernel.org>
CommitterDate: Tue, 29 Sep 2026 21:23:30 +02:00

vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask

__cvdso_clock_gettime_common() and __cvdso_clock_getres_common() convert
the clockid into a bitmask and match it against VDSO_HRES, VDSO_COARSE,
VDSO_RAW and VDSO_AUX:

	msk = 1U << clock;

vdso_clockid_valid() rejects anything above CLOCK_AUX_LAST beforehand,
and CLOCK_AUX_LAST is 23, so the shift count is in range.  Nothing
records that dependency though.  Raising MAX_AUX_CLOCKS beyond 16 moves
CLOCK_AUX_LAST to 32 and makes the shift undefined.

Add a BUILD_BUG_ON() at both conversion sites.  The condition is on a
function parameter rather than a constant, so it relies on the compiler
deriving the range from the vdso_clockid_valid() bail-out above it.  gcc
13 and clang 18 both do: the x86 vdso64 and vdso32 builds stay clean, and
raising MAX_AUX_CLOCKS to 17 trips the assert.

Suggested-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260916023252.418473-5-zhanxusheng@xiaomi.com
---
 lib/vdso/gettimeofday.c | 2 ++
 1 file changed, 2 insertions(+)

diff --git a/lib/vdso/gettimeofday.c b/lib/vdso/gettimeofday.c
index f7a591a..ef4dcc6 100644
--- a/lib/vdso/gettimeofday.c
+++ b/lib/vdso/gettimeofday.c
@@ -285,6 +285,7 @@ __cvdso_clock_gettime_common(const struct vdso_time_data *vd, clockid_t clock,
 	 * Convert the clockid to a bitmask and use it to check which
 	 * clocks are handled in the VDSO directly.
 	 */
+	BUILD_BUG_ON(clock >= BITS_PER_TYPE(msk));
 	msk = 1U << clock;
 	if (likely(msk & VDSO_HRES))
 		vc = &vc[CS_HRES_COARSE];
@@ -438,6 +439,7 @@ bool __cvdso_clock_getres_common(const struct vdso_time_data *vd, clockid_t cloc
 	 * Convert the clockid to a bitmask and use it to check which
 	 * clocks are handled in the VDSO directly.
 	 */
+	BUILD_BUG_ON(clock >= BITS_PER_TYPE(msk));
 	msk = 1U << clock;
 	if (msk & (VDSO_HRES | VDSO_RAW)) {
 		/*

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [tip: timers/vdso] vdso/vsyscall: Keep the CLOCK_AUX base scaled
  2026-09-16  2:32 ` [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled Zhan Xusheng
@ 2026-09-29 19:27   ` tip-bot2 for Zhan Xusheng
  0 siblings, 0 replies; 11+ messages in thread
From: tip-bot2 for Zhan Xusheng @ 2026-09-29 19:27 UTC (permalink / raw)
  To: linux-tip-commits
  Cc: Zhan Xusheng, Thomas Gleixner, thomas.weissschuh, x86, linux-kernel

The following commit has been merged into the timers/vdso branch of tip:

Commit-ID:     4e6aa9d672875aaa20cfb64ffb6fcd79e5165522
Gitweb:        https://git.kernel.org/tip/4e6aa9d672875aaa20cfb64ffb6fcd79e5165522
Author:        Zhan Xusheng <zhanxusheng1024@gmail.com>
AuthorDate:    Wed, 16 Sep 2026 10:32:51 +08:00
Committer:     Thomas Gleixner <tglx@kernel.org>
CommitterDate: Tue, 29 Sep 2026 21:23:30 +02:00

vdso/vsyscall: Keep the CLOCK_AUX base scaled

The vDSO basetime of a clock is stored in the scaled nanoseconds of
tkr_mono, so that the reader can floor the base and the cycle delta
together in vdso_calc_ns().

vdso_time_update_aux() instead shifts the base down to nanoseconds, adds
the offset, and shifts it back up, which zeroes the fractional nanoseconds
of xtime_nsec.  The reader then floors the base and the delta separately:

  ktime_get_aux():  base + ((delta * mult + xtime_nsec) >> shift)
  vdso:             base + (xtime_nsec >> shift)
                         + ((delta * mult) >> shift)

Since floor(a) + floor(b) <= floor(a + b), the vDSO reports 0 or 1 ns
below the syscall for the same clock.  It is not a monotonicity problem:
across an update the step is floor(a + d) - floor(a) - floor(d), which is
0 or 1, never negative.

Add the offset in scaled nanoseconds as the other high resolution clocks
do, and normalise with __iter_div64_u64_rem() so that the stored base
stays below one second and the userspace fast-path does not iterate more
in __iter_div_u64_rem().

Only the sub-second field changes.  (a + (b << shift)) >> shift is exactly
(a >> shift) + b, so the seconds carried out of the normalisation are the
same as before; what the old form dropped was the low shift bits of the
remainder.

monotonic_to_aux.tv_nsec is a normalised timespec64 fraction, so it stays
below NSEC_PER_SEC even for a negative offset, and the sum stays below
2 * (NSEC_PER_SEC << shift).  The largest shift clocks_calc_mult_shift()
can pick is 32, which makes that 8.6e18 against a u64 limit of 1.8e19.

Fixes: 380b84e168e5 ("vdso/vsyscall: Update auxiliary clock data in the datapage")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260916023252.418473-4-zhanxusheng@xiaomi.com
---
 kernel/time/vsyscall.c | 10 +++++-----
 1 file changed, 5 insertions(+), 5 deletions(-)

diff --git a/kernel/time/vsyscall.c b/kernel/time/vsyscall.c
index f43dd3f..0e4b499 100644
--- a/kernel/time/vsyscall.c
+++ b/kernel/time/vsyscall.c
@@ -155,11 +155,11 @@ void vdso_time_update_aux(struct timekeeper *tk)
 
 		vdso_ts->sec = tk->xtime_sec + tk->monotonic_to_aux.tv_sec;
 
-		nsec = tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift;
-		nsec += tk->monotonic_to_aux.tv_nsec;
-		vdso_ts->sec += __iter_div_u64_rem(nsec, NSEC_PER_SEC, &nsec);
-		nsec = nsec << tk->tkr_mono.shift;
-		vdso_ts->nsec = nsec;
+		nsec = tk->tkr_mono.xtime_nsec;
+		nsec += (u64)tk->monotonic_to_aux.tv_nsec << tk->tkr_mono.shift;
+		vdso_ts->sec += __iter_div64_u64_rem(nsec,
+						     (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+						     &vdso_ts->nsec);
 	}
 
 	__arch_update_vdso_clock(vc);

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [tip: timers/vdso] vdso/math64: Add and use __iter_div64_u64_rem()
  2026-09-16  2:32 ` [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem() Zhan Xusheng
@ 2026-09-29 19:27   ` tip-bot2 for Zhan Xusheng
  0 siblings, 0 replies; 11+ messages in thread
From: tip-bot2 for Zhan Xusheng @ 2026-09-29 19:27 UTC (permalink / raw)
  To: linux-tip-commits
  Cc: David Laight, Zhan Xusheng, Thomas Gleixner, thomas.weissschuh,
	x86, linux-kernel

The following commit has been merged into the timers/vdso branch of tip:

Commit-ID:     5c82c995b63748cfa5026f4ca813eee754263f62
Gitweb:        https://git.kernel.org/tip/5c82c995b63748cfa5026f4ca813eee754263f62
Author:        Zhan Xusheng <zhanxusheng1024@gmail.com>
AuthorDate:    Wed, 16 Sep 2026 10:32:50 +08:00
Committer:     Thomas Gleixner <tglx@kernel.org>
CommitterDate: Tue, 29 Sep 2026 21:23:30 +02:00

vdso/math64: Add and use __iter_div64_u64_rem()

The vDSO basetimes for CLOCK_MONOTONIC and CLOCK_BOOTTIME are stored in
the scaled nanoseconds of tkr_mono, so normalising them requires a
division by NSEC_PER_SEC << shift.  That divisor does not fit the u32
parameter of __iter_div_u64_rem(), so update_vdso_time_data() open-codes
the same iterative division twice.

Repeated subtraction is the appropriate form at these two sites because
the quotient never exceeds one.  accumulate_nsecs_to_secs() keeps
xtime_nsec below one scaled second, and the offset added to it is a
normalised timespec64 fraction, so the dividend stays below twice the
divisor.  That bound holds on every architecture, which matters more here
than the cost of a division on any particular one.

Add __iter_div64_u64_rem(), the u64-divisor counterpart of
__iter_div_u64_rem(), and use it at both sites.  Return the quotient as a
u32 for consistency with the u32-divisor variant; the bound above leaves
no use for a wider type.

Store the remainder directly into the basetime, as the coarse clocks
already do.  The CLOCK_BOOTTIME copy of the CLOCK_MONOTONIC values then
takes both fields from the same place.

Folding the two open-coded loops into one inlined helper removes 16 bytes
of vsyscall.o text on x86-64 with gcc 13.

No functional change.

Suggested-by: David Laight <david.laight.linux@gmail.com>
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260916023252.418473-3-zhanxusheng@xiaomi.com
---
 include/vdso/math64.h  | 20 ++++++++++++++++++++
 kernel/time/vsyscall.c | 16 +++++-----------
 2 files changed, 25 insertions(+), 11 deletions(-)

diff --git a/include/vdso/math64.h b/include/vdso/math64.h
index c628d6c..55b45f5 100644
--- a/include/vdso/math64.h
+++ b/include/vdso/math64.h
@@ -22,6 +22,26 @@ static __always_inline u32 __iter_div_u64_rem(u64 dividend, u32 divisor, u64 *re
 	return ret;
 }
 
+static __always_inline u32 __iter_div64_u64_rem(u64 dividend, u64 divisor, u64 *remainder)
+{
+	u32 ret = 0;
+
+	while (dividend >= divisor) {
+		/*
+		 * Prevent the compiler from optimising this loop into a
+		 * modulo operation.
+		 */
+		OPTIMIZER_HIDE_VAR(dividend);
+
+		dividend -= divisor;
+		ret++;
+	}
+
+	*remainder = dividend;
+
+	return ret;
+}
+
 #if defined(CONFIG_ARCH_SUPPORTS_INT128) && defined(__SIZEOF_INT128__)
 
 #ifndef mul_u64_u32_add_u64_shr
diff --git a/kernel/time/vsyscall.c b/kernel/time/vsyscall.c
index aa59919..f43dd3f 100644
--- a/kernel/time/vsyscall.c
+++ b/kernel/time/vsyscall.c
@@ -41,14 +41,12 @@ static inline void update_vdso_time_data(struct vdso_time_data *vdata, struct ti
 
 	nsec = tk->tkr_mono.xtime_nsec;
 	nsec += ((u64)tk->wall_to_monotonic.tv_nsec << tk->tkr_mono.shift);
-	while (nsec >= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift)) {
-		nsec -= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift);
-		vdso_ts->sec++;
-	}
-	vdso_ts->nsec	= nsec;
+	vdso_ts->sec	+= __iter_div64_u64_rem(nsec, (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+						&vdso_ts->nsec);
 
 	/* Copy MONOTONIC time for BOOTTIME */
 	sec	= vdso_ts->sec;
+	nsec	= vdso_ts->nsec;
 	/* Add the boot offset */
 	sec	+= tk->monotonic_to_boot.tv_sec;
 	nsec	+= (u64)tk->monotonic_to_boot.tv_nsec << tk->tkr_mono.shift;
@@ -56,12 +54,8 @@ static inline void update_vdso_time_data(struct vdso_time_data *vdata, struct ti
 	/* CLOCK_BOOTTIME */
 	vdso_ts		= &vc[CS_HRES_COARSE].basetime[CLOCK_BOOTTIME];
 	vdso_ts->sec	= sec;
-
-	while (nsec >= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift)) {
-		nsec -= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift);
-		vdso_ts->sec++;
-	}
-	vdso_ts->nsec	= nsec;
+	vdso_ts->sec	+= __iter_div64_u64_rem(nsec, (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+						&vdso_ts->nsec);
 
 	/* CLOCK_MONOTONIC_RAW */
 	vdso_ts		= &vc[CS_RAW].basetime[CLOCK_MONOTONIC_RAW];

^ permalink raw reply	[flat|nested] 11+ messages in thread

* [tip: timers/vdso] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
  2026-09-16  2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
  2026-09-25  7:44   ` Thomas Weißschuh
@ 2026-09-29 19:27   ` tip-bot2 for Zhan Xusheng
  1 sibling, 0 replies; 11+ messages in thread
From: tip-bot2 for Zhan Xusheng @ 2026-09-29 19:27 UTC (permalink / raw)
  To: linux-tip-commits
  Cc: Zhan Xusheng, Thomas Gleixner, thomas.weissschuh, x86, linux-kernel

The following commit has been merged into the timers/vdso branch of tip:

Commit-ID:     37db375ad90467d7e3ce4e450b3d1c4f297d4ed3
Gitweb:        https://git.kernel.org/tip/37db375ad90467d7e3ce4e450b3d1c4f297d4ed3
Author:        Zhan Xusheng <zhanxusheng1024@gmail.com>
AuthorDate:    Wed, 16 Sep 2026 10:32:49 +08:00
Committer:     Thomas Gleixner <tglx@kernel.org>
CommitterDate: Tue, 29 Sep 2026 21:23:30 +02:00

vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()

__iter_div_u64_rem() divides by repeated subtraction, because its callers
only ever produce a small quotient.  A barrier inside the subtraction loop
keeps the compiler from replacing the loop with a division:

	asm("" : "+rm"(dividend));

The "rm" constraint permits a memory operand.  clang picks it and spills
the dividend inside the loop, at 32-bit and 64-bit alike.  The helper sits
on the clock_gettime() fast path through vdso_set_timespec(), so the spill
is not free: with clang 18 the x86 vDSO is 64 bytes larger in vdso64 text
and 112 bytes larger in vdso32 text than with a register-only barrier.

Use OPTIMIZER_HIDE_VAR(), which is the register-only form of the same
barrier and the form the rest of the kernel uses.  gcc 13 generates
identical code either way, and neither compiler turns the loop into a
division at either width.

Put the function signature on one line while touching it, and reformat the
comment, which described the asm() that is now gone.

Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260916023252.418473-2-zhanxusheng@xiaomi.com
---
 include/vdso/math64.h | 11 ++++++-----
 1 file changed, 6 insertions(+), 5 deletions(-)

diff --git a/include/vdso/math64.h b/include/vdso/math64.h
index 22ae212..c628d6c 100644
--- a/include/vdso/math64.h
+++ b/include/vdso/math64.h
@@ -2,15 +2,16 @@
 #ifndef __VDSO_MATH64_H
 #define __VDSO_MATH64_H
 
-static __always_inline u32
-__iter_div_u64_rem(u64 dividend, u32 divisor, u64 *remainder)
+static __always_inline u32 __iter_div_u64_rem(u64 dividend, u32 divisor, u64 *remainder)
 {
 	u32 ret = 0;
 
 	while (dividend >= divisor) {
-		/* The following asm() prevents the compiler from
-		   optimising this loop into a modulo operation.  */
-		asm("" : "+rm"(dividend));
+		/*
+		 * Prevent the compiler from optimising this loop into a
+		 * modulo operation.
+		 */
+		OPTIMIZER_HIDE_VAR(dividend);
 
 		dividend -= divisor;
 		ret++;

^ permalink raw reply	[flat|nested] 11+ messages in thread

* Re: [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision
  2026-09-16  2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
                   ` (3 preceding siblings ...)
  2026-09-16  2:32 ` [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask Zhan Xusheng
@ 2026-09-30 12:41 ` Thomas Gleixner
  4 siblings, 0 replies; 11+ messages in thread
From: Thomas Gleixner @ 2026-09-30 12:41 UTC (permalink / raw)
  To: Zhan Xusheng, thomas.weissschuh
  Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng

On Wed, Sep 16 2026 at 10:32, Zhan Xusheng wrote:
> Testing on x86-64: checkpatch --strict, gcc W=1 and make LLVM=1 are all
> clean.  Built and booted with kernel/configs/x86_debug.config plus
> panic_on_warn=1; no warning fired and the log holds no lockdep or
> debug-object complaint.  A guest with an aux clock enabled saw no vDSO
> reading ahead of a later ktime_get_aux() in 100000 samples at each of
> offset 0, +5.123456789 and a negative offs_aux.
>
> The generated code in vsyscall.o, vdso64 and vdso32 is byte-identical to
> v4, so the reflowed lines changed nothing.

I've removed the series from tip timers/vdso due to the build fails
which were reported by 0-day.

Please address them and repost an updated series.

Thanks,

        tgl

^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-09-30 12:41 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-16  2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
2026-09-16  2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
2026-09-25  7:44   ` Thomas Weißschuh
2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16  2:32 ` [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem() Zhan Xusheng
2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16  2:32 ` [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled Zhan Xusheng
2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16  2:32 ` [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask Zhan Xusheng
2026-09-29 19:27   ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-30 12:41 ` [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Thomas Gleixner

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®