From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f44.google.com (mail-pj1-f44.google.com [209.85.216.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E436F47DD48 for ; Thu, 4 Jun 2026 13:06:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.44 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780578417; cv=none; b=gR7l2DiBalGxrTkL90DBgLEkyLfCfYNOJ/pFoQN6jXBf1GFfeQQ1AwaVjPvmi1mnMefwQudXVVEZ+OoYR6I4cBNTAhU5itZ6C+NLR9xxyahzOyD9Sjg3wDcI36H5+K3xYSkvy4TBcKOmixk4VakNn2B9GSvL4EYFALemQRRN4tg= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780578417; c=relaxed/simple; bh=wRbCF69wlx67VwicjostQ42SihvIU2JiEg8O7U3Nwyg=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=uZ16/NFHFkmzbAMzYAOP7uN5LORAS5ltolao990GPwxrLoDiqbRxBuOkUawuYiBVLYuYI2S+M4FnEaI2ZmBfSLtKShUz7bj5YgnzuBSTJv8SfwfWgDs6zO4q3VMdWFOPcRaV5VI72+L529TKFTEJFkLfwXIz7uwtlXxuGUstenM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=J1ts5vak; arc=none smtp.client-ip=209.85.216.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="J1ts5vak" Received: by mail-pj1-f44.google.com with SMTP id 98e67ed59e1d1-36ad15213fbso491819a91.0 for ; Thu, 04 Jun 2026 06:06:55 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1780578415; x=1781183215; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=x/V6RPepsyqoQUt2D3lkrhZmC+vMiyf2WaK9hrkxCm4=; b=J1ts5vakIYSwW7lGar4RSiIYRA6yKQkJvTCaeGuiFHGWM235fn8ul3kNp/CpYx+Vh3 DTNsJxgCvd7ub8MJAHC8884ys5ZSTzxbq/hv7RJTHMwPxNg3WlIsFjaw327KeykrL2+A Rq3tNntFtFR1ZFS5HaU+h8TvooTvkW0Y9jsYAcmiV2Fjb1mxM1XOesRCd/OW+dojaUne 9MsTGgUiY5A6sEyRsprLph4EV8E9KEdmsSO/HwgdtyPo8Nqg4mR8V5D2tbsAq6fHFKH0 uFTZnORdn3VMa+oFkPbfIqduKHf2yk5gFN83wjDV3cQpBXHhvN9QkTiH9xPD0gydXTQw w84Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1780578415; x=1781183215; h=content-transfer-encoding:in-reply-to:from:references:cc:to:subject :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=x/V6RPepsyqoQUt2D3lkrhZmC+vMiyf2WaK9hrkxCm4=; b=eSIOgkDZpIBfzTyErABiypeC4q+Wnz81GaLiRoW4z1+aD+050ZDa2yf/8JTfVbQdlC OQID61O8CsMKvW+QRRd1BvR/sIPx8KZOh9FHcI22elPwMnNMF54qnRgi4LU3VjAsJynR Mg1Ltqdla3keDyIgsgrfKqD/aSMg6jzLB9Tvvaa1gKGPCcukTHg3fMJ8+/FSgGClzEVW 5P4wIIKFJ21+J8v3M9rrxQ0E2AP77lvRXTFmco/x4lN15YbztuLwlO4gqBze3Kr2UDQf ldqpIgqD0tW4Im6hauj8acUWLEb+qL46TgS/sTcyXPDhuO2kD/Ot0dQH0SYX131vxmx8 h+yg== X-Forwarded-Encrypted: i=1; AFNElJ/xa9J+wgHLP1aT6MewmD7so/YL5siK+/HaYQTrB+mDc6yqGMzHiRc/e2y8VzxmKQ3YHLYf3I+5rITAXvk=@vger.kernel.org X-Gm-Message-State: AOJu0Yws6j3H7dmitRUST89G8T6qQH9zh9zCO0oBosqopG4vwk5lU3n6 3doiPjopn1YR3GCoUvOQRnJadX2JtlGuh/XWx76ZueEnHEHoqJwJAcAiyCVoQg== X-Gm-Gg: Acq92OFfItO7GRGMovRHcX62/NgNK8Xg2D6EMfU1B8I/2SUno+lEKuMfAnSHEAHZDOv PizUP8dp9I1bomOzjvUNspUxo4ApRRzETa8iHdEXdeWiZzePjrf+Xbqy/Z/MpBgXARJKwxizpb/ wu8Ap/Nn/U3x1rOFQfRKiSCmXv3LZZ+kb4yqQhL5Z75WCwLcGh57mvK4uwR2FTvtpw5STsgXQwI cbT+dgGL/vOq/K/XD0S0UNMEm2q5zBeDq+/2t4SEIHuTmfe9AL0eRATtQEvW6mmzq16IcUufTU4 AaErJQGZuc8VNOFI/ag09xdBEBiiCxTMAFZfpIsWOQgca5k2KuZpoqQOa2U8/GnV8xf37EqtU0J zYKNRf1SG9HoGplkM6sYCO1J8T52ootpvxfvUe1yXMmzJUfXXJZgaZYwnyBd6uk/bKEumoOUHHD 1/nh4rggdnW5VcqmeXHDU2tz1wEIQff8y4r53p08TZsx1tpqoYMKzJuQ== X-Received: by 2002:a17:90b:2d8d:b0:36d:ae6a:22fe with SMTP id 98e67ed59e1d1-36e32288dc4mr8152975a91.16.1780578415000; Thu, 04 Jun 2026 06:06:55 -0700 (PDT) Received: from [10.125.192.75] ([210.184.73.204]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-36f6bf8284dsm3942710a91.4.2026.06.04.06.06.46 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Thu, 04 Jun 2026 06:06:54 -0700 (PDT) Message-ID: Date: Thu, 4 Jun 2026 21:06:43 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:102.0) Gecko/20100101 Thunderbird/102.15.0 Subject: Re: [PATCH v3 1/4] mm/zswap: Make shrink_worker writeback cursor per-memcg To: Yosry Ahmed Cc: akpm@linux-foundation.org, tj@kernel.org, hannes@cmpxchg.org, shakeel.butt@linux.dev, mhocko@kernel.org, mkoutny@suse.com, nphamcs@gmail.com, chengming.zhou@linux.dev, muchun.song@linux.dev, roman.gushchin@linux.dev, cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-doc@vger.kernel.org, Hao Jia References: <20260526114601.67041-1-jiahao.kernel@gmail.com> <20260526114601.67041-2-jiahao.kernel@gmail.com> <8c0e60e1-5713-69f0-a687-088c87e75764@gmail.com> <9898f83d-fae9-e284-6b85-c7f4089840a0@gmail.com> From: Hao Jia In-Reply-To: Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit On 2026/6/4 13:34, Yosry Ahmed wrote: >>>> For instance, suppose a parent memcg has two children, memcg1 and memcg2, >>>> each with 200MB of zswap (100MB inactive). Triggering proactive writeback on >>>> the parent memcg will exhaust memcg1's inactive zswap pages. After that, >>>> even though memcg2 still has plenty of inactive zswap pages, it will >>>> continue to write back memcg1's active zswap pages. Writing back active >>>> zswap pages causes the user-space agent to prematurely abort the writeback >>>> because it detects that certain memcg metrics have exceeded predefined >>>> thresholds. >>> >>> This will only happen if the reclaim size is smaller than the batch >>> size, right? Otherwise the kernel should reclaim more or less equally >>> from both memcgs? >>> >> >> I gave it some thought. Not using a cursor could lead to unfairness >> issues with certain writeback sizes: >> >> - If the writeback size is an odd multiple of WB_BATCH (e.g., >> triggering a writeback of 3 * WB_BATCH), with 2 child cgroups, the >> writeback ratio might end up being 2:1. >> - If a memcg has 5 child cgroups and a writeback of 2 * WB_BATCH is >> triggered, it might repeatedly write back from only the first 2 child >> cgroups. >> >> Although setting a smaller WB_BATCH might mitigate this unfairness, it >> could hurt writeback efficiency. Let's just use per-memcg cursors to >> completely fix these corner cases. > > Exactly, the batch size should be small enough that any unfairness is > not a problem. I would honestly just do batching without a per-memcg > cursor, unless we have numbers to prove that the efficiency is > affected when we use a small batch size. Let's only introduce > complexity when needed please. If you prefer not to use per-cgroup cursors, do we still need to keep the global cursor (i.e., the root cgroup's cursor) zswap_next_shrink? I found this part to be quite tricky when trying to reuse the main logic of shrink_worker() in zswap_proactive_writeback(). Of course, I think we could also keep zswap_next_shrink and write a small helper to check if it's the root cgroup, allowing us to use different memcg iteration methods. Thanks, Hao