From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1755563AbdEHTER (ORCPT ); Mon, 8 May 2017 15:04:17 -0400 Received: from mail-yw0-f174.google.com ([209.85.161.174]:35249 "EHLO mail-yw0-f174.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1752983AbdEHTEO (ORCPT ); Mon, 8 May 2017 15:04:14 -0400 Date: Mon, 8 May 2017 15:04:11 -0400 From: Tejun Heo To: Pavel Machek Cc: Boris Brezillon , David Woodhouse , Hans de Goede , Ricard Wanderlof , linux-scsi@vger.kernel.org, linux-kernel@vger.kernel.org, linux-ide@vger.kernel.org, linux-mtd@lists.infradead.org, Henrique de Moraes Holschuh Subject: Re: Race to power off harming SATA SSDs Message-ID: <20170508190411.GB12079@htj.duckdns.org> References: <1494231215.6528.22.camel@infradead.org> <1494233673.6528.28.camel@infradead.org> <7ee84982-cfad-94d7-4e22-4edbeb852b32@redhat.com> <1494238390.6528.42.camel@infradead.org> <20170508135005.0b9b200b@bbrezillon> <20170508164322.GA9781@amd> <20170508174303.GA12079@htj.duckdns.org> <20170508185615.GA16268@amd> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20170508185615.GA16268@amd> User-Agent: Mutt/1.8.2 (2017-04-18) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hello, On Mon, May 08, 2017 at 08:56:15PM +0200, Pavel Machek wrote: > Well... the SMART counter tells us that the device was not shut down > correctly. Do we have reason to believe that it is _not_ telling us > truth? It is more than one device. It also finished power off command successfully. > SSDs die when you power them without warning: > http://lkcl.net/reports/ssd_analysis.html > > What kind of data would you like to see? "I have been using linux and > my SSD died"? We have had such reports. "I have killed 10 SSDs in a > week then I added one second delay, and this SSD survived 6 months"? Repeating shutdown cycles and showing that the device actually is in trouble would be great. It doesn't have to reach full-on device failure. Showing some sign of corruption would be enough - increase in CRC failure counts, bad block counts (a lot of devices report remaining reserve or lifetime in one way or the other) and so on. Right now, it might as well be just the SMART counter being funky. Thanks. -- tejun