From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pg1-f174.google.com (mail-pg1-f174.google.com [209.85.215.174]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 839183D9025 for ; Fri, 10 Apr 2026 15:38:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.174 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775835488; cv=none; b=AYB6vCLhpxWWCyU7/UfrAE24nRVWwPXgZR07grI2Ocwfb0DsbMRogyNZIQIE8YJb8aHE4u49kAv16TiPuCKz5Wq1EqFTLBrIpk/HMb0doWEvMysJoWJCeBWOohfrXf1/UXamn297afGrFQw1S51iEyakrTHEElB4q+yy/1//ro0= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1775835488; c=relaxed/simple; bh=H2hDwPRkXNGT8+fwYV9nKDab93kJxnQR/8ekGUq8y2o=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=n26hUWdRkR93F55zLt+4/HFzqt5Xot8MQvy9zQl6vtMb/ECQQe+bMLIkAKvIn9ZR+8JR0nlmsrzhiVMRwIYea9ZmJfG7/3OXCS7jBDvmNxGFFw2j1DBDwNpOvBXUPCV5zTSgszhUxefyZpq9MWT/bF18NDzUKobZZrGxXLtpFco= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=m1YHn6Qd; arc=none smtp.client-ip=209.85.215.174 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="m1YHn6Qd" Received: by mail-pg1-f174.google.com with SMTP id 41be03b00d2f7-c76c60c7502so880479a12.0 for ; Fri, 10 Apr 2026 08:38:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1775835487; x=1776440287; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to; bh=aWQ748YuyuRSs8B6opXusugQoA7Vn8Yq0zG+bC5yMpk=; b=m1YHn6QdcCjH//cV6T+dkMTAx8CRRNLwAMHsMZnMb8YVKq/pHpPBluknDU/r0im8UL Zd0cHnnA67xo1J1XSQx23m3pojSqf1oo6MP+oUR0tD7k7Z/GomlDSUlIm+McmlD4Yeyg LEuYjLhD9OjJiHh7pKhXUrbyVwmm1VmK6lg3EgZR499hIY/cVURtLybOUyge68RnR64s 5lLL4UsnPfIVlRk3PleCtQV+qKHV7H0N8CdCJJ2T1Ri6emgaQMvjH9D9xkYk/g1O/+uy 1ygeh/H8w5qZ25NSTtBL0g31+cAB16ROHmz1d1w6SQYKYN+lfH5OHfXzA5FfFxWfOcgp mZfg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1775835487; x=1776440287; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to; bh=aWQ748YuyuRSs8B6opXusugQoA7Vn8Yq0zG+bC5yMpk=; b=EVi/c5x6j05f9LW3hMLlZtgMKHg/L4Robqfmostm+mY9tt7sQtAJDyX3swP0JnHR2E yFz3JBLbzbvPWCOg3/iUQuf5YZMyKpdUIiFoSecSsteh36OHpQqiGFuAbqI8um0n/TvO g9Az51EKGy1nHzoyHMW452R5yxSgIw1gRsBgakELINKKkpAdGtvIp+h0lttEYfA8Gerh KnD38pWdV/1psOA81iS6X32jkNNNNYmVGF79967x9iKS+n2UEvRs8ZkIe2WGZBw0L1Io 1VXiJ3g9DoyHO9S+5bgFKt9Znf+0XZW7XMI9xRLdwPcKNGlXWIDHKMhScDFjE5ZEN8/T jBAA== X-Forwarded-Encrypted: i=1; AJvYcCW7UlZ2AKFjCFgHQkGGkkno3rvQ2GN5U3KRPtFq+3xJ58OMbv1bc4roePR5z+/h3LxeLNnYrdZsQOeIwWU=@vger.kernel.org X-Gm-Message-State: AOJu0YzhofA6PH/nM5QbsCtSeviU/WLigh2/JXY1NjYGhs7Uc8uyHiJw a2n8QdkKNXdHvdhuhHtLnqgS+mHwqnl/FEqMr4bIbqh3+rMIqGiu+FCr X-Gm-Gg: AeBDieuBdfR/uB5s9AdsiF5NprcHR/wsCEkPlFe/f2FxltvQ4XJfHHOXnzEYHy8AASJ cqj55uXTnCNTymJmVAJ1iq0XiARhutIlXBzQMvm3FypX9eV5SNQGW5NyMcDHXb4MKEOxRpEYyNS 0ti5zmTEGR0za/w2tjaIjmlq/vYK2N+mELJmGmBch4Eqc8TKqmO8LOqmS7N/Ieg1HefuGrJmth8 pspe63IkHE2VWfUiYDeK23R1FHOkGz4oSuivUpJ2i9mXjwcOLeTG8hAtV1ygxDoqDIgJfLp35vV UPf+6mOfxpttBES+c5q4SSYPUlBvAWaR/4wOzWeVv3uBqbATShzNSu8ol8gX+um3QQxQBOxgeHq RNoHpfEe91cIOxXvS8Xhc6KVup3C0L98svYSgUtz73WqXZko/eTw45mCCkDEabGGW87fMfsv5CL GkTo7R3gTTZKftUu4MqRLUvBkekjEeFR+eDJ5YJWvh5kk1/mDG51nIqtRqeU0jj9l4iPU= X-Received: by 2002:a05:6a20:4326:b0:39b:da1e:3fad with SMTP id adf61e73a8af0-39fc9309b47mr8585773637.9.1775835486864; Fri, 10 Apr 2026 08:38:06 -0700 (PDT) Received: from localhost.localdomain ([2404:7a85:2900:3f0:6514:327:b8bf:af2d]) by smtp.gmail.com with ESMTPSA id 41be03b00d2f7-c79219f55c4sm3093591a12.22.2026.04.10.08.38.04 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Fri, 10 Apr 2026 08:38:06 -0700 (PDT) From: Mitsumasa KONDO To: andres@anarazel.de Cc: abuehaze@amazon.com, alisaidi@amazon.com, bigeasy@linutronix.de, blakgeof@amazon.com, dipietro.salvatore@gmail.com, dipiets@amazon.it, linux-kernel@vger.kernel.org, mark.rutland@arm.com, peterz@infradead.org, tglx@kernel.org, vschneid@redhat.com, kondo.mitsumasa@gmail.com Subject: Re: [PATCH 0/1] sched: Restore PREEMPT_NONE as default Date: Sat, 11 Apr 2026 00:38:03 +0900 Message-ID: <20260410153803.74151-1-kondo.mitsumasa@gmail.com> X-Mailer: git-send-email 2.49.0 In-Reply-To: References: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Hi, Thank you Salvatore for the thorough cross-architecture benchmarks and pg_stat_activity data, and Andres for the insightful analysis of the spinlock and huge page behavior. Apologies for yet another hypothesis, but I suspect the regression may not be limited to spinlock contention alone -- it could be triggering a secondary feedback loop in the kernel's Buffered I/O throttling: PREEMPT_LAZY -> lock holders preempted during page faults -> spinlock contention (838/1024 backends stalled) -> dirty page generation rate drops -> bdi->write_bandwidth estimate converges to artificially low value (exponential smoothing makes recovery slow) -> balance_dirty_pages() over-throttles on next burst -> throughput cannot recover -> 0.51x Huge pages break this loop at the entry point: fewer TLB misses mean fewer page faults while holding spinlocks, so contention never escalates and the bandwidth estimator stays calibrated. Note that the rseq slice extension was validated with Oracle, which uses Direct I/O, bypassing the page cache and balance_dirty_pages() entirely. PostgreSQL uses Buffered I/O by default, so this class of regression would not have been caught in that validation. My worst-case concern is that this loop could affect any Buffered I/O workload where PREEMPT_LAZY disrupts dirty page generation patterns, not just PostgreSQL. Regards, -- Mitsumasa KONDO NTT Software Innovation Center