From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm2-f13.google.com (mail-wm2-f13.google.com [74.125.225.141]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CA12443C07C for ; Mon, 21 Sep 2026 09:55:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.225.141 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789984523; cv=none; b=WG+5Lt8YVmht1+KtvqRDgdfqiUEwqpqr/UUODSr58B/vXoxYMYQ3ggX3iSwBBZl3nlqdhpB7hSf1mYvgJaesol1W4irRvqQu/SUpsbhpXlGCCK0XU+TiwGA47bY4Hs9h8MlsBaa5IEdCMDglvWnH3JWseEMAHeqZ09t6/+Vp+xg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1789984523; c=relaxed/simple; bh=7UPq8U/iRVN8XOda/8iZfxpKbzPPyb6X7CAqcWoNFFU=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=BqRQZDO9aP7aDUI5y6fjZuQ13w1VYr1Nn27qMSjQ/6Pm2zsBBbE90rAzVzKlWhTco0TiNFvx1+VK3qq64RsuRSjxJI6jDBlej8oOq/gd0NSioHepFK2WidNqFmEJMtUYddBODqqnzZMCoFMuclXCpnvlJtg4Fb76fp+TuQYoL+I= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=Nu+2xLyL; arc=none smtp.client-ip=74.125.225.141 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="Nu+2xLyL" Received: by mail-wm2-f13.google.com with SMTP id 5b1f17b1804b1-49e7d2bb404so8567855e9.1 for ; Mon, 21 Sep 2026 02:55:21 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1789984520; x=1790589320; darn=vger.kernel.org; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=JfR8P12V2dfyx7WlZJfhBwYaBNVarwlPbqKqcSq6ZrQ=; b=Nu+2xLyLn/kj0yTLZuoytUnKkMUFgfxij0tOX2NaCXvlyFUXyawb5gPdHkp9Hp+FEH Qdrbcby9PKezy8+EjB8Dfu2eQeS/vZYiAb7KbgwJdyM4ko5aF10ES3h1KSG2pC7wFVCw EEO19wFt9xoe5w/xnGMzVOjqPkkgGZK9/FbDZU9D6Ny60hZeETH1+xJHjonxoD29yvL3 C7IYLPae7dBSqeKtb+rMXt+Z/FjpqULgnSTumhVlz6gwGrW0rt+v3rwOAYM7SucjV4EQ WjvgPyOIvrGfcbLcAA/BguYh8RX6UV7oMr+K9UMPEZVWaVZW4oBnjzsIfFZ12rHVxiGH nsaA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1789984520; x=1790589320; h=content-transfer-encoding:content-type:mime-version:references :in-reply-to:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=JfR8P12V2dfyx7WlZJfhBwYaBNVarwlPbqKqcSq6ZrQ=; b=Tf/BlzwT7tauaR+mID+gDW8bAfGggvN7FuoBQiTE7Bhldhns1e6+vMBMMDJtDybQ8P eRKDKzaCFXadWK2UQMTvS9ixsDgrXBrhc6fnHnYQV9aGgPoUtBHm9LauO/L6bXfr+qX/ iDs3H/dzhlXE/kMV2wrAn0TVEoQQUyx8ubs4lvVFruWf5DSt4efoklVF/7Ta1L6g4XX7 NswPheaW9aq62xba8Bdj3CozG8FNgI2sCaJCpGXru5Sp5k0MBEKWfZJyOBeF5kfdxp6o R5eoDY22ceFvxuz65KlVAcKoAcl6e0caDCY/mt8LaX8IviZ6xzwgrhteHvMXj9HtOAan 1KUg== X-Forwarded-Encrypted: i=1; AKwUvBy4Kt7/4a4OSXNqA9ZobgDnUgrYvKZuWNr257hjtny3h+LrmsvWggvw2osF8PgvTSezPWNZLOPrKiX6AMA=@vger.kernel.org X-Gm-Message-State: AFuF++mkaVm+nJROf8n6JDJQZv7tgLejzkSwUHyItU9875rIDLjUW2HD a6Ro2BvIiiJWV3LDqZ5GtWCyF5GXkFxLUfB23I+qUK5jaRTdSrTLYeVu X-Gm-Gg: AYBFou1ohchBvP57Q8jiaQ3M67Z5c1hLvK1OIDmt2x6hNb5uLTHSOAJejT4/vRV6ETT gz1XdFtyoDyluwgLDz860dRiSwrwQJK7jVyR8X4BhN2i7XHIuF6/3ntaH2ph9iiQCTP7g4FWYuX cxG/zVA05526BmYhj5v9AWRpZqKIiU2Jt+LTaWF9WsszQvtQzI2mMnTwo9bk5yyfoEvmp2yhJN3 vrEOIfdn6vMAEOk8KsDkrMouitfq2hslknIn5gkOoCjJGHYLv9R0McUbMk46I/pY+ELMNStwYSY vG1rPNcdHCjdyiEIKE/zatU+LbzQSPF/JJ/MEcJuE0Z8x0tJcFBfEF+V3D5sRIBl5RpRdDHLCAY 7M/lJwM5DSpys9M9K4eg9/8Un6UiRdiIKNMlS2CJ/H1T+id68Xf3ZejXKgGuaSd6pWPFGS814fX yYcGtxVcYluqX1UlSpkU2qMoj665xj0B4xCIG9X7rDyVALmIgZ81p744hnNHYFIUn5r2jGJ2ftE Wu3n+GhwvJhj9PjmGszS//ObRorFgom19GNZBXfepeb+LQ= X-Received: by 2002:a05:600c:3f06:b0:49d:257c:a735 with SMTP id 5b1f17b1804b1-49fc4ff42b4mr158855435e9.11.1789984519697; Mon, 21 Sep 2026 02:55:19 -0700 (PDT) Received: from pumpkin (82-69-66-36.dsl.in-addr.zen.co.uk. [82.69.66.36]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49fcd0eabfcsm225054205e9.3.2026.09.21.02.55.19 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 02:55:19 -0700 (PDT) Date: Mon, 21 Sep 2026 10:55:18 +0100 From: David Laight To: Heiko Carstens Cc: Alexander Gordeev , Sven Schnelle , Vasily Gorbik , Christian Borntraeger , Mete Durlu , Peter Zijlstra , Mark Rutland , Juergen Christ , Ilya Leoshkevich , linux-kernel@vger.kernel.org, linux-s390@vger.kernel.org Subject: Re: [PATCH v3 11/11] s390/percpu: Rework to simplify percpu_entry() and percpu_exit() Message-ID: <20260921105518.26f3dcf7@pumpkin> In-Reply-To: <20260921084005.4022574-12-hca@linux.ibm.com> References: <20260921084005.4022574-1-hca@linux.ibm.com> <20260921084005.4022574-12-hca@linux.ibm.com> X-Mailer: Claws Mail 4.1.1 (GTK 3.24.38; arm-unknown-linux-gnueabihf) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Mon, 21 Sep 2026 10:40:05 +0200 Heiko Carstens wrote: > The percpu code section functionality uses a rather complex method to > figure out if the register, which contains the address of the current > cpu's percpu variable, needs to be adjusted. > > If an interrupt happens within a percpu code section (indicated by a > lowcore field), the instruction at the interrupted location is > checked. If it is not a specific AG instruction, the register needs to > be updated. This mechanism needs to take kprobes into account, and > enforces a specific instruction ordering. > > Mark Rutland provided a different solution for arm64 [1] which comes > without such limitations, but requires to use one more instruction, and > two more registers. Given that this simplifies percpu_entry() and > percpu_exit() it seems to be worth to go that route. > > Change s390 to implement a similar approach. This requires to encode > three register numbers into the "percpu_register" field, which is used > to indicate if a percpu code section is executed. > > The used mviy instruction can write only one byte, which allows to > encode only two register numbers. Use a register pair for the inline > assemblies, and only encode the even register number of the register > pair to work around this. > > The generated code changes like this for e.g. a simple this_cpu_inc(): > > Old: > > c0 20 00 00 00 00 larl %r2,c6 <-- load address of percpu var > b9 04 00 32 lgr %r3,%r2 <-- pointless copy of address > eb 03 03 c0 00 52 mviy 960,3 <-- start of percpu code section > - gpr 3 contains percpu var address > e3 30 03 b8 00 08 ag %r3,952 <-- add percpu offset > eb 01 30 00 00 7a agsi 0(%r3),1 <-- atomic inc > eb 00 03 c0 00 52 mviy 960,0 <-- end of percpu code section > > New: > > c0 10 00 00 00 00 larl %r1,c6 <-- load address of percpu var > eb 21 03 c0 00 52 mviy 960,33 <-- start of percpu code section > 33 == 0x21: > - gpr 1 contains percpu var address > - gpr 2 used for percpu offset > - gpr 2+1 == 3 used for current cpu's percpu var address > e3 20 03 b8 00 04 lg %r2,952 <-- load percpu offset > 41 32 10 00 la %r3,0(%r2,%r1) <-- generate current cpu's percpu var address > eb 01 30 00 00 7a agsi 0(%r3),1 <-- atomic inc > eb 00 03 c0 00 52 mviy 960,0 <-- end of percpu code section > > In the above "new" code example 33 (0x21) is used as indicator value to > mark that a percpu code section is executed. This value implies that > registers 2 and 3 will (only) hold the percpu offset and the current > cpu's percpu var address. Those registers will be updated by > percpu_exit() if the process was migrated to a different cpu to contain > the percpu offset and percpu var address of the new cpu. > > Note that because of the pointless lgr instruction in the "old" code > example the two code sequences have identical size, and it looks like > only one more register is used. However this is because of suboptimal > gcc code generation. In both cases the register dependency chain length is 3 (ignoring the pointless lgr) so the execution time is likely to be the same. ... > + regpcp = FIELD_GET(PCPU_REG_PCP, regval); > + regoff = FIELD_GET(PCPU_REG_OFF, regval); ... > +#define PCPU_REG_PCP_SHIFT 0 > +#define PCPU_REG_PCP GENMASK(3, 0) > +#define PCPU_REG_OFF_SHIFT 4 > +#define PCPU_REG_OFF GENMASK(7, 4) ... > +#define __PCPU_CALC_REGVAL(regpcp, regoff) \ > + "(" regpcp " << " __stringify(PCPU_REG_PCP_SHIFT) ") |" \ > + "(" regoff " << " __stringify(PCPU_REG_OFF_SHIFT) ")" I'm not a big fan of GENMASK() + FIELD_GET() and I'm not at all sure it really helps here. Maybe: regpcp = (regval >> PCPU_REG_PCP_SHIFT) & 15; regoff = (regval >> PCPU_REG_OFF_SHIFT) & 15; Which at least uses the same constants for encode and decode. But I might just comment that the pcp register is in the high nibble (twice) and remove the 'crud'. #define __PCPU_CALC_REGVAL(regpcp, regoff) "(" regpcp " << 4 ) | " regoff regpcp = regval >> 4; regoff = regval & 15; David