From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from lgeamrelo13.lge.com (lgeamrelo13.lge.com [156.147.23.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B4E0D37757C for ; Fri, 7 Aug 2026 06:41:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=156.147.23.53 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786084895; cv=none; b=PLmjBF3G+4S+rh0AAmoBTdlyX/ND9YX0V/i5vmFeEF6ZlyPA6yrE4mPfo8XYxR4b/JwAlzS/XTd7EpaBLxUpTjxgxuDew6zkKQcD9tFaRxJKte1VfyIllcyMr97DjT2KHtxDivUFRCHwc3oOL2lu40QY8Q3+3AvOipH8rIyFmsY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786084895; c=relaxed/simple; bh=eJOKWi8b54m969cZ3KudjECyUGrEDgQYi/+iVKmyZvg=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Z68yz3XezwDF2YjUWI1xppWFvi0tRP9uwcRmzQTUaGsSf1TsGvLsxt9eFyJeyuWGIsd7Gp87Z/wqWyjvMcHRbyOXxiPEqjwY2KyObQEfC/rxauhZgOmenmFoD/JNc117jNZ+8/F8/9PhzQtodSiM7XA+a6qI+BjYV0lzta93gc8= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=lge.com; spf=pass smtp.mailfrom=lge.com; arc=none smtp.client-ip=156.147.23.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=lge.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=lge.com Received: from unknown (HELO lgemrelse7q.lge.com) (156.147.1.151) by 156.147.23.53 with ESMTP; 7 Aug 2026 15:41:22 +0900 X-Original-SENDERIP: 156.147.1.151 X-Original-MAILFROM: youngjun.park@lge.com Received: from unknown (HELO yjaykim-PowerEdge-T330) (10.177.112.156) by 156.147.1.151 with ESMTP; 7 Aug 2026 15:41:22 +0900 X-Original-SENDERIP: 10.177.112.156 X-Original-MAILFROM: youngjun.park@lge.com Date: Fri, 7 Aug 2026 15:41:22 +0900 From: Youngjun Park To: Andrew Morton Cc: Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , her0gyugyu@gmail.com, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH v3 0/2] mm/swap: skip empty clusters in the swapoff scan Message-ID: References: <20260806193228.458685-1-youngjun.park@lge.com> <20260806130655.420e16bf621b580cc5f08040@linux-foundation.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260806130655.420e16bf621b580cc5f08040@linux-foundation.org> On Thu, Aug 06, 2026 at 01:06:55PM -0700, Andrew Morton wrote: > On Fri, 7 Aug 2026 04:32:26 +0900 Youngjun Park wrote: > > > find_next_to_unuse() walks a swap device one offset at a time. Slot > > state now lives in a per cluster swap table, so patch 2 dismisses an > > empty cluster with one counter read instead of SWAPFILE_CLUSTER table > > reads. > > Thanks. > > Can you help us understand how significant this change is for users? > If "not very" then I'd prefer to defer consideraton of the series until > after 7.3-rc1. Hello Andrew "Not very" in the common case, though there is a case where the win is clear. No bug and no user report. For now I would rather defer to after 7.3-rc1. And for your reference, here is the details. Every swapoff does a little less work now, because the scan steps over an unused area one cluster at a time. But IMHO most of the swapoff time goes to unuse_mm() and to reading the pages back in. The gain shows on a large swap device that is almost empty, when the last pages still in use are near the end of it. The scan has to walk up to them, and today it looks at every slot on the way. Now the empty clusters in between are skipped in one step. I have no measured times yet, since that case has to be set up on purpose. What I did is the arithmetic for the case that skips best, For example 1T of swap with 256M slots with SWAPFILE_CLUSTER = 512 and everything free but the far end: - today: 256M table reads - with the skip: 512K counter reads That should be around half a second of scan saved. What I have checked is that empty clusters are skipped as intended, so the scan does less work. That work is a small part of swapoff, so depending on the situation it may be too small to see in clock time. Thanks, Youngjun