* [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision
@ 2026-09-16 2:32 Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
` (4 more replies)
0 siblings, 5 replies; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16 2:32 UTC (permalink / raw)
To: tglx, thomas.weissschuh
Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng
Patch 3/4 is the fix. vdso_time_update_aux() rounds the CLOCK_AUX base
down to whole nanoseconds, so the vDSO reports 0 or 1 ns below
ktime_get_aux() for the same clock. 1/4 and 2/4 prepare the division
helper it uses. 4/4 is independent of the fix and can be dropped without
affecting the rest.
Changes since v4:
- Every patch carries a single From: now. format.from was set to a
different address than user.email, so format-patch emitted a header
From plus an in-body From, and send-email then added a second in-body
one. That setting is gone; git am records zhanxusheng@xiaomi.com for
all four patches.
- 1/4: rewrite the changelog, which referred to a loop it had not
introduced. Also put __iter_div_u64_rem()'s signature on one line, so
that the two helpers in the file do not end up in different styles.
- 2/4: put the __iter_div64_u64_rem() signature on one line, 90
characters. Rewrite the changelog and correct a size claim that did
not survive re-measurement: the 16 bytes come from folding the two
open-coded loops into one inlined helper, and the u32 return type makes
no difference to vsyscall.o on x86-64.
- 4/4: rewrite the subject and changelog using the wording the code uses
for this. "clock id dispatch" and "dispatch mask" were mine, not the
kernel's.
- 3/4 is unchanged.
- The tag block follows the maintainer-tip.rst order now, which puts the
author's Signed-off-by ahead of Reviewed-by.
Testing on x86-64: checkpatch --strict, gcc W=1 and make LLVM=1 are all
clean. Built and booted with kernel/configs/x86_debug.config plus
panic_on_warn=1; no warning fired and the log holds no lockdep or
debug-object complaint. A guest with an aux clock enabled saw no vDSO
reading ahead of a later ktime_get_aux() in 100000 samples at each of
offset 0, +5.123456789 and a negative offs_aux.
The generated code in vsyscall.o, vdso64 and vdso32 is byte-identical to
v4, so the reflowed lines changed nothing.
v4: https://lore.kernel.org/all/20260902033801.2912699-1-zhanxusheng@xiaomi.com
Zhan Xusheng (4):
vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
vdso/math64: Add and use __iter_div64_u64_rem()
vdso/vsyscall: Keep the CLOCK_AUX base scaled
vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask
include/vdso/math64.h | 31 ++++++++++++++++++++++++++-----
kernel/time/vsyscall.c | 26 ++++++++++----------------
lib/vdso/gettimeofday.c | 2 ++
3 files changed, 38 insertions(+), 21 deletions(-)
base-commit: 9b87fdc9af2fbfcdb5c24a64139685ef80f6573f
--
2.43.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
2026-09-16 2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
@ 2026-09-16 2:32 ` Zhan Xusheng
2026-09-25 7:44 ` Thomas Weißschuh
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem() Zhan Xusheng
` (3 subsequent siblings)
4 siblings, 2 replies; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16 2:32 UTC (permalink / raw)
To: tglx, thomas.weissschuh
Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng
__iter_div_u64_rem() divides by repeated subtraction, because its callers
only ever produce a small quotient. A barrier inside the subtraction loop
keeps the compiler from replacing the loop with a division:
asm("" : "+rm"(dividend));
The "rm" constraint permits a memory operand. clang picks it and spills
the dividend inside the loop, at 32-bit and 64-bit alike. The helper sits
on the clock_gettime() fast path through vdso_set_timespec(), so the spill
is not free: with clang 18 the x86 vDSO is 64 bytes larger in vdso64 text
and 112 bytes larger in vdso32 text than with a register-only barrier.
Use OPTIMIZER_HIDE_VAR(), which is the register-only form of the same
barrier and the form the rest of the kernel uses. gcc 13 generates
identical code either way, and neither compiler turns the loop into a
division at either width.
Put the function signature on one line while touching it, and reformat the
comment, which described the asm() that is now gone.
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
---
include/vdso/math64.h | 11 ++++++-----
1 file changed, 6 insertions(+), 5 deletions(-)
diff --git a/include/vdso/math64.h b/include/vdso/math64.h
index 22ae212f8b28..c628d6cf447c 100644
--- a/include/vdso/math64.h
+++ b/include/vdso/math64.h
@@ -2,15 +2,16 @@
#ifndef __VDSO_MATH64_H
#define __VDSO_MATH64_H
-static __always_inline u32
-__iter_div_u64_rem(u64 dividend, u32 divisor, u64 *remainder)
+static __always_inline u32 __iter_div_u64_rem(u64 dividend, u32 divisor, u64 *remainder)
{
u32 ret = 0;
while (dividend >= divisor) {
- /* The following asm() prevents the compiler from
- optimising this loop into a modulo operation. */
- asm("" : "+rm"(dividend));
+ /*
+ * Prevent the compiler from optimising this loop into a
+ * modulo operation.
+ */
+ OPTIMIZER_HIDE_VAR(dividend);
dividend -= divisor;
ret++;
--
2.43.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem()
2026-09-16 2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
@ 2026-09-16 2:32 ` Zhan Xusheng
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled Zhan Xusheng
` (2 subsequent siblings)
4 siblings, 1 reply; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16 2:32 UTC (permalink / raw)
To: tglx, thomas.weissschuh
Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng
The vDSO basetimes for CLOCK_MONOTONIC and CLOCK_BOOTTIME are stored in
the scaled nanoseconds of tkr_mono, so normalising them requires a
division by NSEC_PER_SEC << shift. That divisor does not fit the u32
parameter of __iter_div_u64_rem(), so update_vdso_time_data() open-codes
the same iterative division twice.
Repeated subtraction is the appropriate form at these two sites because
the quotient never exceeds one. accumulate_nsecs_to_secs() keeps
xtime_nsec below one scaled second, and the offset added to it is a
normalised timespec64 fraction, so the dividend stays below twice the
divisor. That bound holds on every architecture, which matters more here
than the cost of a division on any particular one.
Add __iter_div64_u64_rem(), the u64-divisor counterpart of
__iter_div_u64_rem(), and use it at both sites. Return the quotient as a
u32 for consistency with the u32-divisor variant; the bound above leaves
no use for a wider type.
Store the remainder directly into the basetime, as the coarse clocks
already do. The CLOCK_BOOTTIME copy of the CLOCK_MONOTONIC values then
takes both fields from the same place.
Folding the two open-coded loops into one inlined helper removes 16 bytes
of vsyscall.o text on x86-64 with gcc 13.
No functional change.
Suggested-by: David Laight <david.laight.linux@gmail.com>
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
---
include/vdso/math64.h | 20 ++++++++++++++++++++
kernel/time/vsyscall.c | 16 +++++-----------
2 files changed, 25 insertions(+), 11 deletions(-)
diff --git a/include/vdso/math64.h b/include/vdso/math64.h
index c628d6cf447c..55b45f5cf615 100644
--- a/include/vdso/math64.h
+++ b/include/vdso/math64.h
@@ -22,6 +22,26 @@ static __always_inline u32 __iter_div_u64_rem(u64 dividend, u32 divisor, u64 *re
return ret;
}
+static __always_inline u32 __iter_div64_u64_rem(u64 dividend, u64 divisor, u64 *remainder)
+{
+ u32 ret = 0;
+
+ while (dividend >= divisor) {
+ /*
+ * Prevent the compiler from optimising this loop into a
+ * modulo operation.
+ */
+ OPTIMIZER_HIDE_VAR(dividend);
+
+ dividend -= divisor;
+ ret++;
+ }
+
+ *remainder = dividend;
+
+ return ret;
+}
+
#if defined(CONFIG_ARCH_SUPPORTS_INT128) && defined(__SIZEOF_INT128__)
#ifndef mul_u64_u32_add_u64_shr
diff --git a/kernel/time/vsyscall.c b/kernel/time/vsyscall.c
index aa59919b8f2c..f43dd3f4744b 100644
--- a/kernel/time/vsyscall.c
+++ b/kernel/time/vsyscall.c
@@ -41,14 +41,12 @@ static inline void update_vdso_time_data(struct vdso_time_data *vdata, struct ti
nsec = tk->tkr_mono.xtime_nsec;
nsec += ((u64)tk->wall_to_monotonic.tv_nsec << tk->tkr_mono.shift);
- while (nsec >= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift)) {
- nsec -= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift);
- vdso_ts->sec++;
- }
- vdso_ts->nsec = nsec;
+ vdso_ts->sec += __iter_div64_u64_rem(nsec, (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+ &vdso_ts->nsec);
/* Copy MONOTONIC time for BOOTTIME */
sec = vdso_ts->sec;
+ nsec = vdso_ts->nsec;
/* Add the boot offset */
sec += tk->monotonic_to_boot.tv_sec;
nsec += (u64)tk->monotonic_to_boot.tv_nsec << tk->tkr_mono.shift;
@@ -56,12 +54,8 @@ static inline void update_vdso_time_data(struct vdso_time_data *vdata, struct ti
/* CLOCK_BOOTTIME */
vdso_ts = &vc[CS_HRES_COARSE].basetime[CLOCK_BOOTTIME];
vdso_ts->sec = sec;
-
- while (nsec >= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift)) {
- nsec -= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift);
- vdso_ts->sec++;
- }
- vdso_ts->nsec = nsec;
+ vdso_ts->sec += __iter_div64_u64_rem(nsec, (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+ &vdso_ts->nsec);
/* CLOCK_MONOTONIC_RAW */
vdso_ts = &vc[CS_RAW].basetime[CLOCK_MONOTONIC_RAW];
--
2.43.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled
2026-09-16 2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem() Zhan Xusheng
@ 2026-09-16 2:32 ` Zhan Xusheng
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask Zhan Xusheng
2026-09-30 12:41 ` [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Thomas Gleixner
4 siblings, 1 reply; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16 2:32 UTC (permalink / raw)
To: tglx, thomas.weissschuh
Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng
The vDSO basetime of a clock is stored in the scaled nanoseconds of
tkr_mono, so that the reader can floor the base and the cycle delta
together in vdso_calc_ns().
vdso_time_update_aux() instead shifts the base down to nanoseconds, adds
the offset, and shifts it back up, which zeroes the fractional nanoseconds
of xtime_nsec. The reader then floors the base and the delta separately:
ktime_get_aux(): base + ((delta * mult + xtime_nsec) >> shift)
vdso: base + (xtime_nsec >> shift)
+ ((delta * mult) >> shift)
Since floor(a) + floor(b) <= floor(a + b), the vDSO reports 0 or 1 ns
below the syscall for the same clock. It is not a monotonicity problem:
across an update the step is floor(a + d) - floor(a) - floor(d), which is
0 or 1, never negative.
Add the offset in scaled nanoseconds as the other high resolution clocks
do, and normalise with __iter_div64_u64_rem() so that the stored base
stays below one second and the userspace fast-path does not iterate more
in __iter_div_u64_rem().
Only the sub-second field changes. (a + (b << shift)) >> shift is exactly
(a >> shift) + b, so the seconds carried out of the normalisation are the
same as before; what the old form dropped was the low shift bits of the
remainder.
monotonic_to_aux.tv_nsec is a normalised timespec64 fraction, so it stays
below NSEC_PER_SEC even for a negative offset, and the sum stays below
2 * (NSEC_PER_SEC << shift). The largest shift clocks_calc_mult_shift()
can pick is 32, which makes that 8.6e18 against a u64 limit of 1.8e19.
Fixes: 380b84e168e5 ("vdso/vsyscall: Update auxiliary clock data in the datapage")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
---
kernel/time/vsyscall.c | 10 +++++-----
1 file changed, 5 insertions(+), 5 deletions(-)
diff --git a/kernel/time/vsyscall.c b/kernel/time/vsyscall.c
index f43dd3f4744b..0e4b499328c0 100644
--- a/kernel/time/vsyscall.c
+++ b/kernel/time/vsyscall.c
@@ -155,11 +155,11 @@ void vdso_time_update_aux(struct timekeeper *tk)
vdso_ts->sec = tk->xtime_sec + tk->monotonic_to_aux.tv_sec;
- nsec = tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift;
- nsec += tk->monotonic_to_aux.tv_nsec;
- vdso_ts->sec += __iter_div_u64_rem(nsec, NSEC_PER_SEC, &nsec);
- nsec = nsec << tk->tkr_mono.shift;
- vdso_ts->nsec = nsec;
+ nsec = tk->tkr_mono.xtime_nsec;
+ nsec += (u64)tk->monotonic_to_aux.tv_nsec << tk->tkr_mono.shift;
+ vdso_ts->sec += __iter_div64_u64_rem(nsec,
+ (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+ &vdso_ts->nsec);
}
__arch_update_vdso_clock(vc);
--
2.43.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask
2026-09-16 2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
` (2 preceding siblings ...)
2026-09-16 2:32 ` [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled Zhan Xusheng
@ 2026-09-16 2:32 ` Zhan Xusheng
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-30 12:41 ` [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Thomas Gleixner
4 siblings, 1 reply; 11+ messages in thread
From: Zhan Xusheng @ 2026-09-16 2:32 UTC (permalink / raw)
To: tglx, thomas.weissschuh
Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng
__cvdso_clock_gettime_common() and __cvdso_clock_getres_common() convert
the clockid into a bitmask and match it against VDSO_HRES, VDSO_COARSE,
VDSO_RAW and VDSO_AUX:
msk = 1U << clock;
vdso_clockid_valid() rejects anything above CLOCK_AUX_LAST beforehand,
and CLOCK_AUX_LAST is 23, so the shift count is in range. Nothing
records that dependency though. Raising MAX_AUX_CLOCKS beyond 16 moves
CLOCK_AUX_LAST to 32 and makes the shift undefined.
Add a BUILD_BUG_ON() at both conversion sites. The condition is on a
function parameter rather than a constant, so it relies on the compiler
deriving the range from the vdso_clockid_valid() bail-out above it. gcc
13 and clang 18 both do: the x86 vdso64 and vdso32 builds stay clean, and
raising MAX_AUX_CLOCKS to 17 trips the assert.
Suggested-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
---
lib/vdso/gettimeofday.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/lib/vdso/gettimeofday.c b/lib/vdso/gettimeofday.c
index f7a591aba59f..ef4dcc614489 100644
--- a/lib/vdso/gettimeofday.c
+++ b/lib/vdso/gettimeofday.c
@@ -285,6 +285,7 @@ __cvdso_clock_gettime_common(const struct vdso_time_data *vd, clockid_t clock,
* Convert the clockid to a bitmask and use it to check which
* clocks are handled in the VDSO directly.
*/
+ BUILD_BUG_ON(clock >= BITS_PER_TYPE(msk));
msk = 1U << clock;
if (likely(msk & VDSO_HRES))
vc = &vc[CS_HRES_COARSE];
@@ -438,6 +439,7 @@ bool __cvdso_clock_getres_common(const struct vdso_time_data *vd, clockid_t cloc
* Convert the clockid to a bitmask and use it to check which
* clocks are handled in the VDSO directly.
*/
+ BUILD_BUG_ON(clock >= BITS_PER_TYPE(msk));
msk = 1U << clock;
if (msk & (VDSO_HRES | VDSO_RAW)) {
/*
--
2.43.0
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
2026-09-16 2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
@ 2026-09-25 7:44 ` Thomas Weißschuh
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
1 sibling, 0 replies; 11+ messages in thread
From: Thomas Weißschuh @ 2026-09-25 7:44 UTC (permalink / raw)
To: Zhan Xusheng
Cc: tglx, luto, vincenzo.frascino, david.laight.linux, linux-kernel,
zhanxusheng
On Wed, Sep 16, 2026 at 10:32:49AM +0800, Zhan Xusheng wrote:
> __iter_div_u64_rem() divides by repeated subtraction, because its callers
> only ever produce a small quotient. A barrier inside the subtraction loop
> keeps the compiler from replacing the loop with a division:
>
> asm("" : "+rm"(dividend));
>
> The "rm" constraint permits a memory operand. clang picks it and spills
> the dividend inside the loop, at 32-bit and 64-bit alike. The helper sits
> on the clock_gettime() fast path through vdso_set_timespec(), so the spill
> is not free: with clang 18 the x86 vDSO is 64 bytes larger in vdso64 text
> and 112 bytes larger in vdso32 text than with a register-only barrier.
>
> Use OPTIMIZER_HIDE_VAR(), which is the register-only form of the same
> barrier and the form the rest of the kernel uses. gcc 13 generates
> identical code either way, and neither compiler turns the loop into a
> division at either width.
>
> Put the function signature on one line while touching it, and reformat the
> comment, which described the asm() that is now gone.
>
> Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
> ---
> include/vdso/math64.h | 11 ++++++-----
> 1 file changed, 6 insertions(+), 5 deletions(-)
(...)
^ permalink raw reply [flat|nested] 11+ messages in thread
* [tip: timers/vdso] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask
2026-09-16 2:32 ` [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask Zhan Xusheng
@ 2026-09-29 19:27 ` tip-bot2 for Zhan Xusheng
0 siblings, 0 replies; 11+ messages in thread
From: tip-bot2 for Zhan Xusheng @ 2026-09-29 19:27 UTC (permalink / raw)
To: linux-tip-commits
Cc: thomas.weissschuh, Zhan Xusheng, Thomas Gleixner, x86, linux-kernel
The following commit has been merged into the timers/vdso branch of tip:
Commit-ID: ca46a07c4b68246c4e400ce3649ab7228a878d8f
Gitweb: https://git.kernel.org/tip/ca46a07c4b68246c4e400ce3649ab7228a878d8f
Author: Zhan Xusheng <zhanxusheng1024@gmail.com>
AuthorDate: Wed, 16 Sep 2026 10:32:52 +08:00
Committer: Thomas Gleixner <tglx@kernel.org>
CommitterDate: Tue, 29 Sep 2026 21:23:30 +02:00
vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask
__cvdso_clock_gettime_common() and __cvdso_clock_getres_common() convert
the clockid into a bitmask and match it against VDSO_HRES, VDSO_COARSE,
VDSO_RAW and VDSO_AUX:
msk = 1U << clock;
vdso_clockid_valid() rejects anything above CLOCK_AUX_LAST beforehand,
and CLOCK_AUX_LAST is 23, so the shift count is in range. Nothing
records that dependency though. Raising MAX_AUX_CLOCKS beyond 16 moves
CLOCK_AUX_LAST to 32 and makes the shift undefined.
Add a BUILD_BUG_ON() at both conversion sites. The condition is on a
function parameter rather than a constant, so it relies on the compiler
deriving the range from the vdso_clockid_valid() bail-out above it. gcc
13 and clang 18 both do: the x86 vdso64 and vdso32 builds stay clean, and
raising MAX_AUX_CLOCKS to 17 trips the assert.
Suggested-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260916023252.418473-5-zhanxusheng@xiaomi.com
---
lib/vdso/gettimeofday.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/lib/vdso/gettimeofday.c b/lib/vdso/gettimeofday.c
index f7a591a..ef4dcc6 100644
--- a/lib/vdso/gettimeofday.c
+++ b/lib/vdso/gettimeofday.c
@@ -285,6 +285,7 @@ __cvdso_clock_gettime_common(const struct vdso_time_data *vd, clockid_t clock,
* Convert the clockid to a bitmask and use it to check which
* clocks are handled in the VDSO directly.
*/
+ BUILD_BUG_ON(clock >= BITS_PER_TYPE(msk));
msk = 1U << clock;
if (likely(msk & VDSO_HRES))
vc = &vc[CS_HRES_COARSE];
@@ -438,6 +439,7 @@ bool __cvdso_clock_getres_common(const struct vdso_time_data *vd, clockid_t cloc
* Convert the clockid to a bitmask and use it to check which
* clocks are handled in the VDSO directly.
*/
+ BUILD_BUG_ON(clock >= BITS_PER_TYPE(msk));
msk = 1U << clock;
if (msk & (VDSO_HRES | VDSO_RAW)) {
/*
^ permalink raw reply [flat|nested] 11+ messages in thread
* [tip: timers/vdso] vdso/vsyscall: Keep the CLOCK_AUX base scaled
2026-09-16 2:32 ` [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled Zhan Xusheng
@ 2026-09-29 19:27 ` tip-bot2 for Zhan Xusheng
0 siblings, 0 replies; 11+ messages in thread
From: tip-bot2 for Zhan Xusheng @ 2026-09-29 19:27 UTC (permalink / raw)
To: linux-tip-commits
Cc: Zhan Xusheng, Thomas Gleixner, thomas.weissschuh, x86, linux-kernel
The following commit has been merged into the timers/vdso branch of tip:
Commit-ID: 4e6aa9d672875aaa20cfb64ffb6fcd79e5165522
Gitweb: https://git.kernel.org/tip/4e6aa9d672875aaa20cfb64ffb6fcd79e5165522
Author: Zhan Xusheng <zhanxusheng1024@gmail.com>
AuthorDate: Wed, 16 Sep 2026 10:32:51 +08:00
Committer: Thomas Gleixner <tglx@kernel.org>
CommitterDate: Tue, 29 Sep 2026 21:23:30 +02:00
vdso/vsyscall: Keep the CLOCK_AUX base scaled
The vDSO basetime of a clock is stored in the scaled nanoseconds of
tkr_mono, so that the reader can floor the base and the cycle delta
together in vdso_calc_ns().
vdso_time_update_aux() instead shifts the base down to nanoseconds, adds
the offset, and shifts it back up, which zeroes the fractional nanoseconds
of xtime_nsec. The reader then floors the base and the delta separately:
ktime_get_aux(): base + ((delta * mult + xtime_nsec) >> shift)
vdso: base + (xtime_nsec >> shift)
+ ((delta * mult) >> shift)
Since floor(a) + floor(b) <= floor(a + b), the vDSO reports 0 or 1 ns
below the syscall for the same clock. It is not a monotonicity problem:
across an update the step is floor(a + d) - floor(a) - floor(d), which is
0 or 1, never negative.
Add the offset in scaled nanoseconds as the other high resolution clocks
do, and normalise with __iter_div64_u64_rem() so that the stored base
stays below one second and the userspace fast-path does not iterate more
in __iter_div_u64_rem().
Only the sub-second field changes. (a + (b << shift)) >> shift is exactly
(a >> shift) + b, so the seconds carried out of the normalisation are the
same as before; what the old form dropped was the low shift bits of the
remainder.
monotonic_to_aux.tv_nsec is a normalised timespec64 fraction, so it stays
below NSEC_PER_SEC even for a negative offset, and the sum stays below
2 * (NSEC_PER_SEC << shift). The largest shift clocks_calc_mult_shift()
can pick is 32, which makes that 8.6e18 against a u64 limit of 1.8e19.
Fixes: 380b84e168e5 ("vdso/vsyscall: Update auxiliary clock data in the datapage")
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260916023252.418473-4-zhanxusheng@xiaomi.com
---
kernel/time/vsyscall.c | 10 +++++-----
1 file changed, 5 insertions(+), 5 deletions(-)
diff --git a/kernel/time/vsyscall.c b/kernel/time/vsyscall.c
index f43dd3f..0e4b499 100644
--- a/kernel/time/vsyscall.c
+++ b/kernel/time/vsyscall.c
@@ -155,11 +155,11 @@ void vdso_time_update_aux(struct timekeeper *tk)
vdso_ts->sec = tk->xtime_sec + tk->monotonic_to_aux.tv_sec;
- nsec = tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift;
- nsec += tk->monotonic_to_aux.tv_nsec;
- vdso_ts->sec += __iter_div_u64_rem(nsec, NSEC_PER_SEC, &nsec);
- nsec = nsec << tk->tkr_mono.shift;
- vdso_ts->nsec = nsec;
+ nsec = tk->tkr_mono.xtime_nsec;
+ nsec += (u64)tk->monotonic_to_aux.tv_nsec << tk->tkr_mono.shift;
+ vdso_ts->sec += __iter_div64_u64_rem(nsec,
+ (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+ &vdso_ts->nsec);
}
__arch_update_vdso_clock(vc);
^ permalink raw reply [flat|nested] 11+ messages in thread
* [tip: timers/vdso] vdso/math64: Add and use __iter_div64_u64_rem()
2026-09-16 2:32 ` [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem() Zhan Xusheng
@ 2026-09-29 19:27 ` tip-bot2 for Zhan Xusheng
0 siblings, 0 replies; 11+ messages in thread
From: tip-bot2 for Zhan Xusheng @ 2026-09-29 19:27 UTC (permalink / raw)
To: linux-tip-commits
Cc: David Laight, Zhan Xusheng, Thomas Gleixner, thomas.weissschuh,
x86, linux-kernel
The following commit has been merged into the timers/vdso branch of tip:
Commit-ID: 5c82c995b63748cfa5026f4ca813eee754263f62
Gitweb: https://git.kernel.org/tip/5c82c995b63748cfa5026f4ca813eee754263f62
Author: Zhan Xusheng <zhanxusheng1024@gmail.com>
AuthorDate: Wed, 16 Sep 2026 10:32:50 +08:00
Committer: Thomas Gleixner <tglx@kernel.org>
CommitterDate: Tue, 29 Sep 2026 21:23:30 +02:00
vdso/math64: Add and use __iter_div64_u64_rem()
The vDSO basetimes for CLOCK_MONOTONIC and CLOCK_BOOTTIME are stored in
the scaled nanoseconds of tkr_mono, so normalising them requires a
division by NSEC_PER_SEC << shift. That divisor does not fit the u32
parameter of __iter_div_u64_rem(), so update_vdso_time_data() open-codes
the same iterative division twice.
Repeated subtraction is the appropriate form at these two sites because
the quotient never exceeds one. accumulate_nsecs_to_secs() keeps
xtime_nsec below one scaled second, and the offset added to it is a
normalised timespec64 fraction, so the dividend stays below twice the
divisor. That bound holds on every architecture, which matters more here
than the cost of a division on any particular one.
Add __iter_div64_u64_rem(), the u64-divisor counterpart of
__iter_div_u64_rem(), and use it at both sites. Return the quotient as a
u32 for consistency with the u32-divisor variant; the bound above leaves
no use for a wider type.
Store the remainder directly into the basetime, as the coarse clocks
already do. The CLOCK_BOOTTIME copy of the CLOCK_MONOTONIC values then
takes both fields from the same place.
Folding the two open-coded loops into one inlined helper removes 16 bytes
of vsyscall.o text on x86-64 with gcc 13.
No functional change.
Suggested-by: David Laight <david.laight.linux@gmail.com>
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260916023252.418473-3-zhanxusheng@xiaomi.com
---
include/vdso/math64.h | 20 ++++++++++++++++++++
kernel/time/vsyscall.c | 16 +++++-----------
2 files changed, 25 insertions(+), 11 deletions(-)
diff --git a/include/vdso/math64.h b/include/vdso/math64.h
index c628d6c..55b45f5 100644
--- a/include/vdso/math64.h
+++ b/include/vdso/math64.h
@@ -22,6 +22,26 @@ static __always_inline u32 __iter_div_u64_rem(u64 dividend, u32 divisor, u64 *re
return ret;
}
+static __always_inline u32 __iter_div64_u64_rem(u64 dividend, u64 divisor, u64 *remainder)
+{
+ u32 ret = 0;
+
+ while (dividend >= divisor) {
+ /*
+ * Prevent the compiler from optimising this loop into a
+ * modulo operation.
+ */
+ OPTIMIZER_HIDE_VAR(dividend);
+
+ dividend -= divisor;
+ ret++;
+ }
+
+ *remainder = dividend;
+
+ return ret;
+}
+
#if defined(CONFIG_ARCH_SUPPORTS_INT128) && defined(__SIZEOF_INT128__)
#ifndef mul_u64_u32_add_u64_shr
diff --git a/kernel/time/vsyscall.c b/kernel/time/vsyscall.c
index aa59919..f43dd3f 100644
--- a/kernel/time/vsyscall.c
+++ b/kernel/time/vsyscall.c
@@ -41,14 +41,12 @@ static inline void update_vdso_time_data(struct vdso_time_data *vdata, struct ti
nsec = tk->tkr_mono.xtime_nsec;
nsec += ((u64)tk->wall_to_monotonic.tv_nsec << tk->tkr_mono.shift);
- while (nsec >= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift)) {
- nsec -= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift);
- vdso_ts->sec++;
- }
- vdso_ts->nsec = nsec;
+ vdso_ts->sec += __iter_div64_u64_rem(nsec, (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+ &vdso_ts->nsec);
/* Copy MONOTONIC time for BOOTTIME */
sec = vdso_ts->sec;
+ nsec = vdso_ts->nsec;
/* Add the boot offset */
sec += tk->monotonic_to_boot.tv_sec;
nsec += (u64)tk->monotonic_to_boot.tv_nsec << tk->tkr_mono.shift;
@@ -56,12 +54,8 @@ static inline void update_vdso_time_data(struct vdso_time_data *vdata, struct ti
/* CLOCK_BOOTTIME */
vdso_ts = &vc[CS_HRES_COARSE].basetime[CLOCK_BOOTTIME];
vdso_ts->sec = sec;
-
- while (nsec >= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift)) {
- nsec -= (((u64)NSEC_PER_SEC) << tk->tkr_mono.shift);
- vdso_ts->sec++;
- }
- vdso_ts->nsec = nsec;
+ vdso_ts->sec += __iter_div64_u64_rem(nsec, (u64)NSEC_PER_SEC << tk->tkr_mono.shift,
+ &vdso_ts->nsec);
/* CLOCK_MONOTONIC_RAW */
vdso_ts = &vc[CS_RAW].basetime[CLOCK_MONOTONIC_RAW];
^ permalink raw reply [flat|nested] 11+ messages in thread
* [tip: timers/vdso] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
2026-09-16 2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
2026-09-25 7:44 ` Thomas Weißschuh
@ 2026-09-29 19:27 ` tip-bot2 for Zhan Xusheng
1 sibling, 0 replies; 11+ messages in thread
From: tip-bot2 for Zhan Xusheng @ 2026-09-29 19:27 UTC (permalink / raw)
To: linux-tip-commits
Cc: Zhan Xusheng, Thomas Gleixner, thomas.weissschuh, x86, linux-kernel
The following commit has been merged into the timers/vdso branch of tip:
Commit-ID: 37db375ad90467d7e3ce4e450b3d1c4f297d4ed3
Gitweb: https://git.kernel.org/tip/37db375ad90467d7e3ce4e450b3d1c4f297d4ed3
Author: Zhan Xusheng <zhanxusheng1024@gmail.com>
AuthorDate: Wed, 16 Sep 2026 10:32:49 +08:00
Committer: Thomas Gleixner <tglx@kernel.org>
CommitterDate: Tue, 29 Sep 2026 21:23:30 +02:00
vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem()
__iter_div_u64_rem() divides by repeated subtraction, because its callers
only ever produce a small quotient. A barrier inside the subtraction loop
keeps the compiler from replacing the loop with a division:
asm("" : "+rm"(dividend));
The "rm" constraint permits a memory operand. clang picks it and spills
the dividend inside the loop, at 32-bit and 64-bit alike. The helper sits
on the clock_gettime() fast path through vdso_set_timespec(), so the spill
is not free: with clang 18 the x86 vDSO is 64 bytes larger in vdso64 text
and 112 bytes larger in vdso32 text than with a register-only barrier.
Use OPTIMIZER_HIDE_VAR(), which is the register-only form of the same
barrier and the form the rest of the kernel uses. gcc 13 generates
identical code either way, and neither compiler turns the loop into a
division at either width.
Put the function signature on one line while touching it, and reformat the
comment, which described the asm() that is now gone.
Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Signed-off-by: Thomas Gleixner <tglx@kernel.org>
Reviewed-by: Thomas Weißschuh <thomas.weissschuh@linutronix.de>
Link: https://patch.msgid.link/20260916023252.418473-2-zhanxusheng@xiaomi.com
---
include/vdso/math64.h | 11 ++++++-----
1 file changed, 6 insertions(+), 5 deletions(-)
diff --git a/include/vdso/math64.h b/include/vdso/math64.h
index 22ae212..c628d6c 100644
--- a/include/vdso/math64.h
+++ b/include/vdso/math64.h
@@ -2,15 +2,16 @@
#ifndef __VDSO_MATH64_H
#define __VDSO_MATH64_H
-static __always_inline u32
-__iter_div_u64_rem(u64 dividend, u32 divisor, u64 *remainder)
+static __always_inline u32 __iter_div_u64_rem(u64 dividend, u32 divisor, u64 *remainder)
{
u32 ret = 0;
while (dividend >= divisor) {
- /* The following asm() prevents the compiler from
- optimising this loop into a modulo operation. */
- asm("" : "+rm"(dividend));
+ /*
+ * Prevent the compiler from optimising this loop into a
+ * modulo operation.
+ */
+ OPTIMIZER_HIDE_VAR(dividend);
dividend -= divisor;
ret++;
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision
2026-09-16 2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
` (3 preceding siblings ...)
2026-09-16 2:32 ` [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask Zhan Xusheng
@ 2026-09-30 12:41 ` Thomas Gleixner
4 siblings, 0 replies; 11+ messages in thread
From: Thomas Gleixner @ 2026-09-30 12:41 UTC (permalink / raw)
To: Zhan Xusheng, thomas.weissschuh
Cc: luto, vincenzo.frascino, david.laight.linux, linux-kernel, zhanxusheng
On Wed, Sep 16 2026 at 10:32, Zhan Xusheng wrote:
> Testing on x86-64: checkpatch --strict, gcc W=1 and make LLVM=1 are all
> clean. Built and booted with kernel/configs/x86_debug.config plus
> panic_on_warn=1; no warning fired and the log holds no lockdep or
> debug-object complaint. A guest with an aux clock enabled saw no vDSO
> reading ahead of a later ktime_get_aux() in 100000 samples at each of
> offset 0, +5.123456789 and a negative offs_aux.
>
> The generated code in vsyscall.o, vdso64 and vdso32 is byte-identical to
> v4, so the reflowed lines changed nothing.
I've removed the series from tip timers/vdso due to the build fails
which were reported by 0-day.
Please address them and repost an updated series.
Thanks,
tgl
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2026-09-30 12:41 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-16 2:32 [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 1/4] vdso/math64: Use OPTIMIZER_HIDE_VAR() in __iter_div_u64_rem() Zhan Xusheng
2026-09-25 7:44 ` Thomas Weißschuh
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 2/4] vdso/math64: Add and use __iter_div64_u64_rem() Zhan Xusheng
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 3/4] vdso/vsyscall: Keep the CLOCK_AUX base scaled Zhan Xusheng
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-16 2:32 ` [PATCH v5 4/4] vdso/gettimeofday: Assert that the clockid fits into the u32 bitmask Zhan Xusheng
2026-09-29 19:27 ` [tip: timers/vdso] " tip-bot2 for Zhan Xusheng
2026-09-30 12:41 ` [PATCH v5 0/4] vdso: Keep the CLOCK_AUX base at full precision Thomas Gleixner
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®