From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f45.google.com (mail-pj1-f45.google.com [209.85.216.45]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 99FE421C16E for ; Fri, 14 Nov 2025 13:36:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.45 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1763127393; cv=none; b=EUmTU7skUco0uSpodJmF8n5JHRuih1B4WKdyOx5mXxMzVQ2gE/+D4mmaFi4jFUaq2zMol7WwUwFtQKH5PLPDvXkJTwn6FrtosNwydR89ebg36obFqfM9n8Zbyko5HorAa0OyiHuae4FaXrqt3L5Xo4iSCf2LMPTUQIuwkWGMNyA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1763127393; c=relaxed/simple; bh=ZMRlT1qNtbqF0Veemigg10nBAhJkqiE7TfK5COiTJuU=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=byt7nL9wlGsa/AaIAtosrruGVQeW93BaBZDoR1pyAEuLrCmcNiskhtpX8u9z+RxCdBYqUeghDts7UNx9pheYyA+JC8zEu5d+tWcjsOs4dM4Z5/40+mTaaGG8/5dvceJAtbFioWLMDPNNrTG6lRowQn/BA+qXfwLrnNc0k0GNzdk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=rivosinc.com; spf=pass smtp.mailfrom=rivosinc.com; dkim=pass (2048-bit key) header.d=rivosinc.com header.i=@rivosinc.com header.b=P9RTHMhq; arc=none smtp.client-ip=209.85.216.45 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=rivosinc.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=rivosinc.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=rivosinc.com header.i=@rivosinc.com header.b="P9RTHMhq" Received: by mail-pj1-f45.google.com with SMTP id 98e67ed59e1d1-343774bd9b4so1894429a91.2 for ; Fri, 14 Nov 2025 05:36:30 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=rivosinc.com; s=google; t=1763127390; x=1763732190; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=DHEr0VGdvNGDZiR6SQYgH1KlLGPNA+7KAqiw0a2paZY=; b=P9RTHMhqfU/lHCsS3W0VxDuyRkwOPbe6dMHkrSUSIbbqOpieXORxeFwCxEJEUe25xl h1v4ED+kzNJAf5sgHg2FJXkLrQELL2fnY76tDzxS6AeFkV14Q0rvyPQBdG6/PdZGstsk ByX7V9Yw8vByzSxfYfd956Ex36erI4X498F68qKr/k0k2Xzdyhc7ayGosX6/N/pTL2rm TryCSkME3SUi/wB7bjUtP/Pu5z2vVVE/xbhpoYXegQ1krtKn8LJKEP2HlQRpK1U1iotM 6a6cPdDxYZOt/SsgOTvJtOnhxjRm+01tbza8XxROIn8WETo/+Lc5mrZhMXOY/IO3PEAX M9tw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1763127390; x=1763732190; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=DHEr0VGdvNGDZiR6SQYgH1KlLGPNA+7KAqiw0a2paZY=; b=rh2kPN9t1BKftg02K7Uzw/kvBB4R/ecMGoWXS4tzRwHaBBjqp/5mOGn7KHX7E4N7y+ qaQFNW9MiUvSNIi3hkj6i7HiSs/CXhx9+mfM12kDXxzIpU4ppDvr1XTXL6avj3N2YFW9 PKnLw3mzBOT+Pd37Up+de1v+weAVJtfDH/yWiETPDXB879c4CPLPbKpSlLrCKSyr0HpS hhE9nAUjfhqCI2DeESOImknRcrbBCR5tKYpA6GmdK/F5am83fqapOIRU733dtM93v+Nb kdLzkHkCRPQ+FlBS2PIuqJ42mi9m2LYHL+3xIYqoc0tBVLVb1wikSBfBo2r2R6OPi0ZT DrfA== X-Forwarded-Encrypted: i=1; AJvYcCV7ECa3nKXGcyyFwmGBxhCQYX7/CyritJeqBOsB5QacffRjp7zUZL/2mU90nhW/Z+Lvpyl+M5jOu9xf2xU=@vger.kernel.org X-Gm-Message-State: AOJu0YwRuBe7PDuR8YwVw3VOMFMOyCWTI/ixtpHVTL+0JB5Nh/SxvW+C NlHgv9m17lX9EqjH46qXESuNi1jB92qXybZNofem6WSgGl46mVXWByk8Xr19WJ8qiDc= X-Gm-Gg: ASbGncvpP/nt8d8JX/lYWu7wKYN8neoPP6DopFeI8s2pLGDrywvZ5nAbIqodEHfOwuO 1LyUVFOIwYKq/9YF/LIm0sNMeyyG1dGD7QT3NaeL3NIO94m6jRFnZd3caPwIQ78LZQ8W0ei0N6h g5g2pkLwGtMEzsmR82pK2mQbFbSAdSwNK4mXjlgjRT03OTLnw55wKmYudqEbrlVsPHr4cp4fUdC TFqtqNxw+8gi4ZNifEPo/7LLwab86afiXmIxz+tJBP4SsulTMbSicwYCdSg8bgxSOJIkdIx4e6B ovZqU+yAlbujq2aO9C/wSkg2uNC8h8mZ1KbhBEsuTq76UptkUM1hrvkCN9YvlkDrxuKsn/fyT4o e6is8DaTQWG3kD1uc02dWzOHNMReShjAvaDuOjr4AoSU6MIj7YplIJMIvFatkHT8ebXLHDVJgHQ sMZoJyMJ36yzj6iwXSYlCHyQt5BM7SBpZwVxM0STx8tZUVtQ== X-Google-Smtp-Source: AGHT+IGTlOsBEwW7dz/WZ/DHBPhbduYubtpAD4QTiOhGNCk+QaT1FYVrfJAMaEe5hmWNAFKDgI1qjg== X-Received: by 2002:a17:90b:540e:b0:341:88c5:2073 with SMTP id 98e67ed59e1d1-343f9d906dfmr3058440a91.2.1763127389658; Fri, 14 Nov 2025 05:36:29 -0800 (PST) Received: from ?IPV6:2a01:e0a:e17:9700:16d2:7456:6634:9626? ([2a01:e0a:e17:9700:16d2:7456:6634:9626]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-343eac7ec71sm2423873a91.11.2025.11.14.05.36.22 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Fri, 14 Nov 2025 05:36:28 -0800 (PST) Message-ID: Date: Fri, 14 Nov 2025 14:36:16 +0100 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: How to Avoid Starving the Kernel When Using SSE To: =?UTF-8?B?5byg5bGV6bmP?= Cc: Paul Walmsley , Palmer Dabbelt , "linux-riscv@lists.infradead.org" , "linux-kernel@vger.kernel.org" , "linux-arm-kernel@lists.infradead.org" , Himanshu Chauhan , Anup Patel , =?UTF-8?B?6Lev5pet?= , Atish Patra , =?UTF-8?B?QmrDtnJuIFTDtnBlbA==?= , =?UTF-8?B?5bSU6L+Q6L6J?= , =?UTF-8?B?5YWD56u5?= References: Content-Language: en-US From: =?UTF-8?B?Q2zDqW1lbnQgTMOpZ2Vy?= In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Hi Zhanpeng, On 11/14/25 11:24, 张展鹏 wrote: > Hi Clément, > > Lately, I've been thinking about how to avoid starving the kernel when > using SSE: > SSE is powered by M-mode irqs such as M-mode PMU irq for perf sampling and > M-mode IPI for inter-hart injection. Meanwhile, kernel is powered by S-mode > irqs, so the kernel may experience starvation when there is a flood of > M-mode irqs, and kernel may cause such flooding of M-mode irqs when using > SSE, either deliberately or inadvertently: > > 1. Malicious SSE handler: Kernel may deliberately register a bad SSE > handler, which triggers a new inter-hart SSE request via ecall. This will > cause an endless loop of SSE, rendering the kernel unresponsive. In this > case, the only thing SBI can do is to prevent the nesting of SSE in > `sbi_trap_handler` and ensure that SSE events are executed in priority > order. That seems quite convoluted. Anyone that can load a module can do worse than crashing the kernel :) > 2. Perf sampling: Kernel may inadvertently choose a bad parameter for > Perf, which causes PMU irqs to occur too frequently. Continuous PMU irqs > will leave the system with no time to respond to S-mode irqs. But this one concern however is valid ! > > Hence, I think we are supposed to improve the SSE framework to avoid > starving the kernel so easily. > > Here is a case study of perf sampling: > When using PMU-SSE for Perf sampling, the kernel may hang and become > unresponsive due to the PMU-SSE loop. Once we start to process a Perf > sampling using PMU-SSE, the kernel may fail to respond to `Ctrl+C` or fail > to exit after the timing of `sleep 1` completes (these are the two most > commonly used time-based sampling methods in perf). > > By default, perf uses a relatively high sampling frequency, namely > `perf_event_max_sample_rate`, and will adjust it on demand if sampling > takes too much time. If this frequency/period goes beyond what system can > handle, it will make SSE events connect end-to-end, and the system will get > stuck in an endless loop of "SSE → PMU interrupt → SSE". The kernel is then > starved (at this point, if you print the `sepc` of SSE completion, you will > find that the `sepc` remains unchanged each time, indicating that the > kernel is stuck), and the kernel can never escape from this loop of > PMU-SSE, because it can neither respond to Ctrl+C interrupts nor adjust the > sampling frequency. > > Current solution: The key to this problem is that every time we finish > sse_complete, there is already a new PMU irq pending. Then we resume the > kernel execution via mret, and the system will immediately trap back into > SSE. > > The PMU-SSE-Perf processing flow includes the following steps: `sse_inject` > (mret to SSE handler), `pmu_stop` (clear PMU pending bit), `pmu_start` (set > a new value for PMU counter), and `sse_complete` (resume execution to the > point where the kernel was interrupted). The reason why kernel traps right > after `sse_complete` is that there is a new PMU irq generated between > `pmu_start` and `sse_complete`. > > In order to address this issue, we propose to delay the procedure of > re-starting the overflowed PMU counter during PMU-SSE. When kernel triggers > an ecall to restart the overflowed PMU counters, SBI can check whether it > is SSE-powered PMU handling. If so, we temporarily modify mhpmevent CSR to > stop counting kernel events. In this process, M-mode events are always > inhibited, and U-mode code will not be executed during the > `pmu_sbi_ovf_handler`, so we only need to inhibit the counting of kernel > events. I'd rather let the kernel control the PMU SSE event delivery by masking it at the end of the SSE handler and reenabling it later. Additionally, that solution being in the SBI itself, it does not guarantee that all SBI implementation will actually do that correctly. What seems odd is that the perf_event_sample_took() call at each end of PMU event handler should actually allow perf subsystem to throttle the rate. I'll take another look at that part to make sure it works as expected and that we aren't missing any bits. Thanks, Clément > > In this way, we can ensure that `pmu_sbi_ovf_handler` will not be > re-entered by the new PMU-SSE, and minimize the modification of perf logic. > The price is that we gave up sampling a small portion of kernel code(from > `pmu_ctr_start` to the end of `pmu_sbi_ovf_handler`), and we probably need > a new parameter in `pmu_ctr_start`. > > Looking forward to your suggestions. Thanks! > > Best regards, > Zhanpeng Zhang >