From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from fout-a5-smtp.messagingengine.com (fout-a5-smtp.messagingengine.com [103.168.172.148]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id D81973C3C; Mon, 3 Mar 2025 03:05:42 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=103.168.172.148 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1740971145; cv=none; b=piCSN7wY4D9xv4EddBGvKreJ6Ruxv2fxsCGQEXEUte53KKMRyVp/UhJe5dQGA4NASGjqd6QmgJ2lCzMB2EhOyZS+HEDBTKU8BttMhHLC/GwJ3Weyr8mbf/BcKPzcvIJVqIQTgvxiDrieAvjkp/9EiWD7xCj/p69jOaq162ixJXQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1740971145; c=relaxed/simple; bh=ijqRe3FoeguBXic06ZVu9GpNzUVfAF6KQAi5sws2rqk=; h=MIME-Version:Date:From:To:Cc:Message-Id:In-Reply-To:References: Subject:Content-Type; b=OJHoOTxFIuWngPqtChYMMI6Q2ELLmB3dtTsKUq09yzF3K9fblTLsC98ypyLMTQSWdxq2gSlOU+Jxcs6BV+UpUDf3y4/GZjVTkFjFYrVmfLt4YAFj0/uvbsdFoCJt+Cgx2Gca8npYQhYkO18FZ4GpgUP3aoTSKmN4F4mmhHp/sDA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=squebb.ca; spf=pass smtp.mailfrom=squebb.ca; dkim=pass (2048-bit key) header.d=squebb.ca header.i=@squebb.ca header.b=YV5qZcu3; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b=AYxtOy6f; arc=none smtp.client-ip=103.168.172.148 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=squebb.ca Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=squebb.ca Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=squebb.ca header.i=@squebb.ca header.b="YV5qZcu3"; dkim=pass (2048-bit key) header.d=messagingengine.com header.i=@messagingengine.com header.b="AYxtOy6f" Received: from phl-compute-06.internal (phl-compute-06.phl.internal [10.202.2.46]) by mailfout.phl.internal (Postfix) with ESMTP id BE9AD138113D; Sun, 2 Mar 2025 22:05:41 -0500 (EST) Received: from phl-imap-10 ([10.202.2.85]) by phl-compute-06.internal (MEProxy); Sun, 02 Mar 2025 22:05:41 -0500 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=squebb.ca; h=cc :cc:content-transfer-encoding:content-type:content-type:date :date:from:from:in-reply-to:in-reply-to:message-id:mime-version :references:reply-to:subject:subject:to:to; s=fm2; t=1740971141; x=1741057541; bh=iUaStrziO9Bkt9RaZInHU8giBomnwVHdGsgmyVuTGy4=; b= YV5qZcu3irfgq8ZdXuOVHoFboZR8iHfpQsDy5VrpN+QC/AP29daKaddGESccJcB9 LxjS9pditf8rWwGYnSdkuTkDLwv8uhc7AoM1+mcHERq1rNJT0/Y3O7TeyXCb8va9 JlpYHGUGw7WKabnr8ooYLsfeF4hqXJ7UPhcQe1TEATOU0sy5Jugw04gjty7h7Noy aEYsR8Xg8CV3Eqcbdf3XD1zqqPUFwapM6AKm9de6VTIgPfwLvL8BRMYGlgIh37jO NdnYr7Vk1apsvXWKtYN1TAP1vhw/RX3PxDzhKDluTFmIcZF20HaQjdPR+LISYDVI cerALL0FZa/qTloeOfMu4g== DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d= messagingengine.com; h=cc:cc:content-transfer-encoding :content-type:content-type:date:date:feedback-id:feedback-id :from:from:in-reply-to:in-reply-to:message-id:mime-version :references:reply-to:subject:subject:to:to:x-me-proxy :x-me-sender:x-me-sender:x-sasl-enc; s=fm1; t=1740971141; x= 1741057541; bh=iUaStrziO9Bkt9RaZInHU8giBomnwVHdGsgmyVuTGy4=; b=A YxtOy6f8WNfNbW+dkggvmYJgWjavTXQe9wfcTy4LaqklMpNfI6CAB4hWRmDZhRdz dc3R3v0V81gGrgVuAcBq6c0lbvBSF+Xz0Y5EV//6m27TTnetUhbaDCBnD+axVg03 apfKHlrXUKn1YFpADYmQBwxIC9C5dWXhm1DrA+CpDKgHYI7rPB42AQqPCsy2g6ip ixelY7rWRqoPK+DBy5klc1Fpw7t0/45GDpNMTGPU5owKZz7G1rEa7STrNobARncZ X2OpGpDTsTX5+EZHQMoPsjxXENXBuNTXXBH6ZYJS0i4Y/BgShhOy+UlJDlnJqLOh uwg0XcvGXDbasJ9Tcvl4A== X-ME-Sender: X-ME-Proxy-Cause: gggruggvucftvghtrhhoucdtuddrgeefvddrtddtgdeljeellecutefuodetggdotefrod ftvfcurfhrohhfihhlvgemucfhrghsthforghilhdpggftfghnshhusghstghrihgsvgdp uffrtefokffrpgfnqfghnecuuegrihhlohhuthemuceftddtnecusecvtfgvtghiphhivg hnthhsucdlqddutddtmdenucfjughrpefoggffhffvvefkjghfufgtgfesthhqredtredt jeenucfhrhhomhepfdforghrkhcurfgvrghrshhonhdfuceomhhpvggrrhhsohhnqdhlvg hnohhvohesshhquhgvsggsrdgtrgeqnecuggftrfgrthhtvghrnhepudeltdevfedufeff gfeggeehtdeigefghedugefhleetvdfggffghfehhfefteffnecuffhomhgrihhnpehinh htvghlrdgtohhmnecuvehluhhsthgvrhfuihiivgeptdenucfrrghrrghmpehmrghilhhf rhhomhepmhhpvggrrhhsohhnqdhlvghnohhvohesshhquhgvsggsrdgtrgdpnhgspghrtg hpthhtohepuddvpdhmohguvgepshhmthhpohhuthdprhgtphhtthhopegurghvvghmsegu rghvvghmlhhofhhtrdhnvghtpdhrtghpthhtohepvgguuhhmrgiivghtsehgohhoghhlvg drtghomhdprhgtphhtthhopegrnhhthhhonhihrdhlrdhnghhuhigvnhesihhnthgvlhdr tghomhdprhgtphhtthhopehprhiivghmhihslhgrfidrkhhithhsiigvlhesihhnthgvlh drtghomhdprhgtphhtthhopehvihhtrghlhidrlhhifhhshhhithhssehinhhtvghlrdgt ohhmpdhrtghpthhtohepkhhusggrsehkvghrnhgvlhdrohhrghdprhgtphhtthhopehinh htvghlqdifihhrvgguqdhlrghnsehlihhsthhsrdhoshhuohhslhdrohhrghdprhgtphht thhopegrnhgurhgvfidonhgvthguvghvsehluhhnnhdrtghhpdhrtghpthhtoheprghnug hrvgifsehluhhnnhdrtghh X-ME-Proxy: Feedback-ID: ibe194615:Fastmail Received: by mailuser.phl.internal (Postfix, from userid 501) id 7D0FE3C0066; Sun, 2 Mar 2025 22:05:40 -0500 (EST) X-Mailer: MessagingEngine.com Webmail Interface Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Date: Sun, 02 Mar 2025 22:05:19 -0500 From: "Mark Pearson" To: "Vitaly Lifshits" , "Andrew Lunn" Cc: anthony.l.nguyen@intel.com, przemyslaw.kitszel@intel.com, andrew+netdev@lunn.ch, davem@davemloft.net, edumazet@google.com, kuba@kernel.org, pabeni@redhat.com, intel-wired-lan@lists.osuosl.org, netdev@vger.kernel.org, linux-kernel@vger.kernel.org Message-Id: <77d8cbfa-93a2-4d49-b061-b8c59b767241@app.fastmail.com> In-Reply-To: <4b5b0f52-7ed8-7eef-2467-fa59ca5de937@intel.com> References: <20250226194422.1030419-1-mpearson-lenovo@squebb.ca> <36ae9886-8696-4f8a-a1e4-b93a9bd47b2f@lunn.ch> <50d86329-98b1-4579-9cf1-d974cf7a748d@app.fastmail.com> <1a4ed373-9d27-4f4b-9e75-9434b4f5cad9@lunn.ch> <9f460418-99c6-49f9-ac2c-7a957f781e17@app.fastmail.com> <4b5b0f52-7ed8-7eef-2467-fa59ca5de937@intel.com> Subject: Re: [Intel-wired-lan] [PATCH] e1000e: Link flap workaround option for false IRP events Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: quoted-printable Hi Vitaly, On Sun, Mar 2, 2025, at 8:09 AM, Lifshits, Vitaly wrote: > Hi Mark, > >> Hi Andrew >>=20 >> On Thu, Feb 27, 2025, at 11:07 AM, Andrew Lunn wrote: >>>>>> + e1e_rphy(hw, PHY_REG(772, 26), &phy_data); >>>>> >>>>> Please add some #define for these magic numbers, so we have some i= dea >>>>> what PHY register you are actually reading. That in itself might h= elp >>>>> explain how the workaround actually works. >>>>> >>>> >>>> I don't know what this register does I'm afraid - that's Intel know= ledge and has not been shared. >>> >>> What PHY is it? Often it is just a COTS PHY, and the datasheet might >>> be available. >>> >>> Given your setup description, pause seems like the obvious thing to >>> check. When trying to debug this, did you look at pause settings? >>> Knowing what this register is might also point towards pause, or >>> something totally different. >>> >>> Andrew >>=20 >> For the PHY - do you know a way of determining this easily? I can rea= ch out to the platform team but that will take some time. I'm not seeing= anything in the kernel logs, but if there's a recommended way of confir= ming that would be appreciated. > > The PHY is I219 PHY. > The datasheet is indeed accessible to the public:=20 > https://cdrdv2-public.intel.com/612523/ethernet-connection-i219-datash= eet.pdf > >>=20 >> We did look at at the pause pieces - which I agree seems like an obvi= ous candidate given the speed mismatch on the network. >> Experts on the Intel networking team did reproduce the issue in their= lab and looked at this for many weeks without determining root cause. I= wish it was as obvious as pause control configuration :) >>=20 >> Thanks >> Mark >>=20 > > Reading this register was suggested for debug purposes to understand i= f=20 > there is some misconfiguration. We did not find any misconfiguration. > The issue as we discovered was a link status change interrupt caused t= he=20 > CSME to reset the adapter causing the link flap. > > We were unable to determine what causes the link status change interru= pt=20 > in the first place. As stated in the comment, it was only ever observe= d=20 > on Lenovo P5/P7systems and we couldn't ever reproduce on other systems= .=20 > The reproduction in our lab was on a P5 system as well. > > > Regarding the suggested workaround, there isn=E2=80=99t a clear unders= tanding=20 > why it works. We suspect that reading a PHY register is probably=20 > prevents the CSME from resetting the PHY when it handles the LSC=20 > interrupt it gets. However, it can also be a matter of slight timing=20 > variations. > We communicated that this solution is not likely to be accepted to the=20 > kernel as is, and the initial responses on the mailing list demonstrat= e=20 > the pushback. We do understand the frustration of end-users that may=20 > experience the problem. A couple of suggestions that can make it look=20 > less =E2=80=9Cout-of-the-blue=E2=80=9D are: try a short delay instead = of the register=20 > read, or read a more common register like PHY STATUS instead. > On a different topic, I suggest removing the part of the comment below: > * Intel unable to determine root cause. Thank you for the details. I agree that the Intel networking team communicated to us that the solut= ion provided to us was not suitable for upstream (just adding in the rea= d directly) - I agreed with that and add the part about it being a modul= e option to make it, hopefully, more palatable. =20 The problem is that Intel have stated to us, in writing, that they would= not investigate further and would not work on an upstream appropriate s= olution (a requirement for us - out of tree patches are not suitable as = I'm sure we all agree). =20 I'm not going to go into the internal communication details/escalations/= etc here - but this is the only reason I'm doing this patch. I would muc= h rather Intel had continued the debug exercise. > The issue went through joint debug by Intel and Lenovo, and no obvious=20 > spec violations by either party were found. There doesn=E2=80=99t seem= to be=20 > value in including this information in the comments of upstream code. I'm OK to remove that statement, but I do want to highlight that it is f= actually correct. My intent with this patch was not to dump on Intel or cause offence, but= to provide context as to why a Lenovo employee is doing a somewhat pecu= liar workaround for an Intel networking device :) I was planning on updating the commit message so I will look to make it = less inflammatory. Thanks for chiming in and providing the details. I'm treading very caref= ully in what is an extremely awkward position to be put.=20 Mark