From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from lgeamrelo13.lge.com (lgeamrelo13.lge.com [156.147.23.53]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 18D5A38B7BA for ; Thu, 6 Aug 2026 05:22:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=156.147.23.53 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785993734; cv=none; b=IxFrQPLzSC3P2iAT1Hr7FalKSiAiP6FbUNIUg1ke2qBC5RpT38u+9PNUPSDdps0RgZUlMEjNfT1t80/jRbupO2ytHYHJcuCN/T4crXVbL3dd80Hbnb09PGhMtTUJVErZc4B3xiCK3vS+AzYLz5xiH0qQPqnB4kSK4578gThDvxs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785993734; c=relaxed/simple; bh=2Yj2ebN/GAdNmnMO2td+S2t21gZI4sCVqR4uhnpuJUo=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=Ucma54Md90dqp2JuU244ORR/mqq65fGdGpg7aSNwD3KDBwJ3R54lME8XzmL45ZI0QGKGVovSyWteU8FoUQ2cRrkczUXZEWX5J2LTxOhuO+aTW0U8lErmBaO2Uw2USrjmywzqXj9vRW9LF/M/7m7YgdBgXEVnkLJildJr1UotaJI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=lge.com; spf=pass smtp.mailfrom=lge.com; arc=none smtp.client-ip=156.147.23.53 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=lge.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=lge.com Received: from unknown (HELO lgemrelse6q.lge.com) (156.147.1.121) by 156.147.23.53 with ESMTP; 6 Aug 2026 14:22:07 +0900 X-Original-SENDERIP: 156.147.1.121 X-Original-MAILFROM: youngjun.park@lge.com Received: from unknown (HELO yjaykim-PowerEdge-T330) (10.177.112.156) by 156.147.1.121 with ESMTP; 6 Aug 2026 14:22:07 +0900 X-Original-SENDERIP: 10.177.112.156 X-Original-MAILFROM: youngjun.park@lge.com Date: Thu, 6 Aug 2026 14:22:06 +0900 From: Youngjun Park To: Barry Song Cc: Youngjun Park , Andrew Morton , Chris Li , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH v2 2/2] mm/swap: scan by cluster in find_next_to_unuse() Message-ID: References: <20260805141146.127776-1-youngjun.park@lge.com> <20260805141146.127776-3-youngjun.park@lge.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Thu, Aug 06, 2026 at 10:33:54AM +0800, Barry Song wrote: > Reviewed-by: Barry Song Hi Barry, Thanks for the review :) > [...] > > > + i = prev + 1; > > + while (i < si->max) { > > + ci = __swap_offset_to_cluster(si, i); > > + ci_off = i % SWAPFILE_CLUSTER; > > + end = i - ci_off + SWAPFILE_CLUSTER; > > + > > + /* > > + * An empty cluster has no slot in use, so skip it whole. > > + * A slot is uncounted only after its folio left the swap > > + * cache, so there is nothing here for try_to_unuse() to act on. > > + * Count only drops here, so a READ_ONCE() without ci->lock is > > + * enough, unlike in every other cluster_is_empty() caller. > > + */ > > + if (!READ_ONCE(ci->count)) { > > + i = end; > > cond_resched(); > > - } > > + continue; > > + } > > > > - if (i == si->max) > > - i = 0; > > + for (; i < end; ci_off++, i++) { > > + swp_tb = swap_table_get(ci, ci_off); > > + if (!swp_tb_is_null(swp_tb) && !swp_tb_is_bad(swp_tb)) > > + return i; > > You have the following in the changelog: > > " The inner loop runs to the end of the cluster rather than to si->max. > The swap table is always SWAPFILE_CLUSTER entries and swapon() masks > [si->max, round_up(si->max, SWAPFILE_CLUSTER)) as bad, so the tail of a > partial last cluster is rejected by swp_tb_is_bad() and never returned." > But I wonder whether this explanation should be part of the code > comment instead. Yeah right. If I remain the code as it is, I will move this changelog on to the code itself. > Otherwise, people may wonder why this is safe and ask for the below: > end = min_t(unsigned long, i - ci_off + SWAPFILE_CLUSTER, si->max); > How expensive is the min() operation? If it is cheap enough, maybe > we should just do the min() unconditionally? Not expensive. Kairui suggested keeping the end calculation simple(As I assume his intention?), so I dropped the min() in v1. But after thinking about the retry case, keeping the min_t() seems clearer and can also avoid walking the masked tail of the last cluster before retrying. So I think I will keep the min_t() version (inclding move ci_off calculation only to where it is needed) like below + ci = __swap_offset_to_cluster(si, i); + end = min_t(unsigned long, i - ci_off + SWAPFILE_CLUSTER, si->max); + /* + * An empty cluster has no slot in use, so skip it whole. + * A slot is uncounted only after its folio left the swap + * cache, so there is nothing here for try_to_unuse() to act on. + * Count only drops here, so a READ_ONCE() without ci->lock is + * enough, unlike in every other cluster_is_empty() caller. + */ + if (!READ_ONCE(ci->count)) { + i = end; cond_resched(); - } + continue; + } + + ci_off = i % SWAPFILE_CLUSTER; + for (; i < end; ci_off++, i++) { + swp_tb = swap_table_get(ci, ci_off); + if (!swp_tb_is_null(swp_tb) && !swp_tb_is_bad(swp_tb)) + return i; + } + cond_resched(); + } I think both of good enough. But, IMHO, remaining min_t calculation is my preference at now. Kairui and Barry how do you think? - Follow Barry's suggestion. remain min_t calculation. - Add comment why we don't need min_t calculation.(also barry's suggestion) Thanks Youngjun