From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f47.google.com (mail-pj1-f47.google.com [209.85.216.47]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 5932B280A56 for ; Mon, 8 Jun 2026 12:50:30 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.47 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780923031; cv=none; b=hdov+8/YTXs/5Xz96cb7VjBL5DUJjCAf1xgYdwQPJIKgYE4CC8fhc5gRsh3XlLYZTXlqtURsUjjBZwNaWwjFXdYyi0Mg2v5bwOLZwSEGOCCmqge8C3XR3OEiNjGYvadSHGa1OrXrFKGfBkMHpZXXXWl36yCn4P+R6REbWmwmsrg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780923031; c=relaxed/simple; bh=TG6r/3Lp+ODfq+/oUWl3Qo2qwI2yFx1wavbyGgxzkRg=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=DYh9OZGhAmVVhLhI617yCsnK/8HqF0DTENwbvsO1HFi/8amcZrjyZflNuicCrmNNWC7pX8GmTmWQbvXMH6v5RgLvKjIaK4cRVYvt3NMLK1fuohAYtVhp+qC2FXrrDyFBMM4/F6PRZ+zZ4/EFwFY51cv6MaILOWF9ekNd95KZDrI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=TDEEEbcw; arc=none smtp.client-ip=209.85.216.47 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="TDEEEbcw" Received: by mail-pj1-f47.google.com with SMTP id 98e67ed59e1d1-36bb3551f6eso3665630a91.1 for ; Mon, 08 Jun 2026 05:50:30 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1780923030; x=1781527830; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=DiIdK82JBJK4ZmshrTXuh1DETm09pI5fnq4TS+v0gOw=; b=TDEEEbcw5RoeiTZWI0FoBMax+eR8C2tcdBXsoC+PVkb2kta9/b0PKb+Cryf5bxVHk9 EXPPksQj/YtwtXgbDHTY3JQOJ2Nv6FluWu5JO5iPqok0/j/Ha3g8Kwnio4Qkf6wqdZoH F3Z+eqfxbLHS9xrmNQ6pM1dy8aSKtU4csj1lrP29ZHYowwpqmQnYnMYH0G2stf+oH2Ub RZ2uPCmsfX6MKtKxwBmj9eFS7ESqVI5Jk6eldnDE1yu+/W+Tkwv1nCqTl7Zsoa0yq9pQ 6L0dacSQlLRGULqGdfFEKRBg9atT7bR5DArD2+zguLh+IVU9a01pvi+daWLHuCxHX+vp io7g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1780923030; x=1781527830; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=DiIdK82JBJK4ZmshrTXuh1DETm09pI5fnq4TS+v0gOw=; b=YtNr0nMATTxGTw40KnPc0cGLUql+6az6ZQfbfScho7UdEoKnEsJAxJ9rWNBEtfym9E uWdQZ9sQx11099NjAKb2sfY2e9sVPnYh1WvuhyCIyVC5wcg23b7yVSSmmyZ7BL2MfmnC ERqTX0H1Sm1YJR9D1wcdByb/iV8vBoR7FIyyalmCRNzzoUDU77sWR87MwkXB8r+fNJyY 9FEh3MaKFaIwrjocHVo7rS9dEFWKOOCO14UNK9Z1yRPQ2mQMr0rkIJGAVyLcKt1XwggR 15a88sEnl/EKspZ0vovvmWXg/tU16J0akpI3dM3LsgGdvDbC+EtCnCnOrYI4y1SbqWZo Z5sA== X-Forwarded-Encrypted: i=1; AFNElJ9OSy6ZrHtsIU+XBnqtBtkgbV8bsJwaIELhKv8sXycalaeSJGfpS5d7NU1z41mkP3pSNJG86ioDDc4bL/w=@vger.kernel.org X-Gm-Message-State: AOJu0YzABzxkuIvV7ASSZHGXHj7CPDjwF2UFHLGYttbGttVC6VwmGx7n P+a6EZfbm0xItZv0U2D9lsg75ZMwgT0CsdpQgmIDTipmWukqD2Dd/MqC X-Gm-Gg: Acq92OGCqrmG2h125HLlsUWHzxWXimhN7HzWPvjAglPf50mt6mLuHdXSSJ8MzqYcbzg HmrqLs5icmXP+MrqdpniNQC3sHX1/vtuMQMnCOlJhxZCc5oXlcb+n0bUevewgoePBkKyBVx3Fbq eHrskIqSsnA2tqhYlOafWsUbDe132r+S6NOJ+fghd4tMMhGTTxpuDmOAlGMLBPrVM2OriHGsa05 IL0gjxNsoWwI7knaeLZ1CIhF0ONdq3KikB9MAFbD/jHOHvw29yp0dxflqVuyR0qlRHN/lGmxmv/ IbGOQ5i0XASZFK6upX/qjux0De7iiMYQiVIXEMuoavfoiryscMb3ib2GZvsTrNJtah+jUDOQHz9 X3wUkfdAe8miNw8dgLB9eQOg596Jfg8jAyH4vYWel1TzKgMvSGnjWyGfrBdG1koHtMECx+WCg12 OyMiK/vHOPJhjhOsg+3diMihJpS8lZ4f6MAyXd8f64IdsGcjhwrjtdRg== X-Received: by 2002:a17:90b:5345:b0:35c:30a8:330 with SMTP id 98e67ed59e1d1-370ebff342fmr15636109a91.0.1780923029551; Mon, 08 Jun 2026 05:50:29 -0700 (PDT) Received: from [10.125.192.72] ([210.184.73.204]) by smtp.gmail.com with ESMTPSA id d9443c01a7336-2c1663981basm179520225ad.67.2026.06.08.05.50.22 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Mon, 08 Jun 2026 05:50:29 -0700 (PDT) Message-ID: <90730fa7-62e7-d5f4-b638-23b22a8509f2@gmail.com> Date: Mon, 8 Jun 2026 20:50:19 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:102.0) Gecko/20100101 Thunderbird/102.15.0 Subject: Re: [PATCH v3 1/4] mm/zswap: Make shrink_worker writeback cursor per-memcg To: Nhat Pham Cc: Yosry Ahmed , akpm@linux-foundation.org, tj@kernel.org, hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org, mkoutny@suse.com, chengming.zhou@linux.dev, muchun.song@linux.dev, roman.gushchin@linux.dev, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, Hao Jia References: <20260526114601.67041-1-jiahao.kernel@gmail.com> <20260526114601.67041-2-jiahao.kernel@gmail.com> <8c0e60e1-5713-69f0-a687-088c87e75764@gmail.com> <9898f83d-fae9-e284-6b85-c7f4089840a0@gmail.com> From: Hao Jia In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit On 2026/6/5 01:23, Nhat Pham wrote: > On Thu, Jun 4, 2026 at 6:06 AM Hao Jia wrote: >> >> >> >> On 2026/6/4 13:34, Yosry Ahmed wrote: >>>>>> For instance, suppose a parent memcg has two children, memcg1 and memcg2, >>>>>> each with 200MB of zswap (100MB inactive). Triggering proactive writeback on >>>>>> the parent memcg will exhaust memcg1's inactive zswap pages. After that, >>>>>> even though memcg2 still has plenty of inactive zswap pages, it will >>>>>> continue to write back memcg1's active zswap pages. Writing back active >>>>>> zswap pages causes the user-space agent to prematurely abort the writeback >>>>>> because it detects that certain memcg metrics have exceeded predefined >>>>>> thresholds. >>>>> >>>>> This will only happen if the reclaim size is smaller than the batch >>>>> size, right? Otherwise the kernel should reclaim more or less equally >>>>> from both memcgs? >>>>> >>>> >>>> I gave it some thought. Not using a cursor could lead to unfairness >>>> issues with certain writeback sizes: >>>> >>>> - If the writeback size is an odd multiple of WB_BATCH (e.g., >>>> triggering a writeback of 3 * WB_BATCH), with 2 child cgroups, the >>>> writeback ratio might end up being 2:1. >>>> - If a memcg has 5 child cgroups and a writeback of 2 * WB_BATCH is >>>> triggered, it might repeatedly write back from only the first 2 child >>>> cgroups. >>>> >>>> Although setting a smaller WB_BATCH might mitigate this unfairness, it >>>> could hurt writeback efficiency. Let's just use per-memcg cursors to >>>> completely fix these corner cases. >>> >>> Exactly, the batch size should be small enough that any unfairness is >>> not a problem. I would honestly just do batching without a per-memcg >>> cursor, unless we have numbers to prove that the efficiency is >>> affected when we use a small batch size. Let's only introduce >>> complexity when needed please. > > I'm impartial towards the complexity of per-memcg cursor. I don't > think it's that big of a deal, but only if it's warranted. > > Hao, if you're convinced that doing small batch is not efficient, > could you run some experiments to show the improvement bigger batchign > and fairness? Maybe implement a small batch, no-memcg cursor first. > Then implement a patch on top of it to add per-memcg cursor, and show > how much performance win we can get from that patch on top of the > patch series? > Thanks for the suggestion! I ran some tests and found that neither the per-memcg cursor nor different batch sizes have a significant impact on proactive writeback performance. However, exactly as we suspected, without the per-memcg cursor, the writeback distribution among child memcgs is highly unfair. Test Setup: zswap config: 18G capacity, LZ4 compression. cgroup hierarchy: 1 parent test memcg with 10 child memcgs. Allocation: Allocated 1600MB of anonymous pages in each child memcg. To ensure compressibility, the first half of each page was filled with random data and the second half with zeros. Force to zswap: Ran echo "1600M" > memory.reclaim on each child memcg to squeeze all their memory into zswap. Trigger writeback: Ran echo " zswap_writeback_only" > memory.reclaim on the parent cgroup 200 times, with a 2-second interval between each run. Metric: Monitored the zswpwb_proactive metric in memory.stat to observe the writeback volume. **Note**: The size here refers to the uncompressed memory size. Also, since the second-chance algorithm would cause many writebacks to fall short of the target size, I **bypassed** it during these tests to avoid interference. Without cursor (size: 1M, batch: 32) child wb_pages wb_MB share% child0 6368 24.88 12.50 child1 6368 24.88 12.50 child2 6368 24.88 12.50 child3 6368 24.88 12.50 child4 6368 24.88 12.50 child5 6368 24.88 12.50 child6 6368 24.88 12.50 child7 6368 24.88 12.50 child8 0 0.00 0.00 child9 0 0.00 0.00 Without cursor (size: 1M, batch: 128) child wb_pages wb_MB share% child0 25472 99.50 50.00 child1 25472 99.50 50.00 child2 0 0.00 0.00 child3 0 0.00 0.00 child4 0 0.00 0.00 child5 0 0.00 0.00 child6 0 0.00 0.00 child7 0 0.00 0.00 child8 0 0.00 0.00 child9 0 0.00 0.00 Without cursor (size: 6M, batch: 128) child wb_pages wb_MB share% child0 51200 200.00 16.67 child1 51200 200.00 16.67 child2 25600 100.00 8.33 child3 25600 100.00 8.33 child4 25600 100.00 8.33 child5 25600 100.00 8.33 child6 25600 100.00 8.33 child7 25600 100.00 8.33 child8 25600 100.00 8.33 child9 25600 100.00 8.33 With cursor (size: 1M, batch: 32) child wb_pages wb_MB share% child0 5120 20.00 10.00 child1 5120 20.00 10.00 child2 5120 20.00 10.00 child3 5120 20.00 10.00 child4 5120 20.00 10.00 child5 5120 20.00 10.00 child6 5120 20.00 10.00 child7 5120 20.00 10.00 child8 5120 20.00 10.00 child9 5120 20.00 10.00 With cursor (size: 1M, batch: 128) child wb_pages wb_MB share% child0 5120 20.00 10.00 child1 5120 20.00 10.00 child2 5120 20.00 10.00 child3 5120 20.00 10.00 child4 5120 20.00 10.00 child5 5120 20.00 10.00 child6 5120 20.00 10.00 child7 5120 20.00 10.00 child8 5120 20.00 10.00 child9 5120 20.00 10.00 Thakns, Hao