From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qv1-f44.google.com (mail-qv1-f44.google.com [209.85.219.44]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 184C31531C4 for ; Wed, 8 Jan 2025 04:23:06 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.44 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736310188; cv=none; b=tNsmrcZGn3mt+J5A6nZg8fw3bTiR/QVowm7i7cw+j8794KhmOjuKzCNekOAkBQdTe7UeAo06wjk1I4KPpPkwhEHvzOHUOU3lnFN5h7OYCoPjbGv8s3qdgf9/vd2Lv6iaChuHu9/6mrAGz+2Owq6TKd7napLVXbhqqVOGqCK55IM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736310188; c=relaxed/simple; bh=/5i5hnmoOFLW4Y3LRZxMa9xXtGZM3tJfjJg8lSP+8Lg=; h=MIME-Version:References:In-Reply-To:From:Date:Message-ID:Subject: To:Cc:Content-Type; b=Q1/7OygMMp6hw9i6AHZw+fHMH2bzqvVsM7LsX6ElLN06bEZAe+V0sYVGPGe2pcXCvHNR7Eyc0r+Uow098fBhVObBrDGVxSGdQN24c6QDZGqx0f8PpewpU2hD8U743wqlxHwSbLbtPO6JzxyUIig8h5jmnFxguar3xYa1BcniRrU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=eUUGZGuv; arc=none smtp.client-ip=209.85.219.44 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="eUUGZGuv" Received: by mail-qv1-f44.google.com with SMTP id 6a1803df08f44-6dd049b5428so137058886d6.2 for ; Tue, 07 Jan 2025 20:23:06 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1736310186; x=1736914986; darn=vger.kernel.org; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:from:to:cc:subject:date:message-id:reply-to; bh=dGPpQ6VJIV9xnsCD1bMq8dKJH1L2LIGUbC/YN69ZewQ=; b=eUUGZGuvYZaZNRD6r1HPUQnK/vKwor4FyajafYbKXgG2Kq70HXFyr0mmc0ifKOruyb q6f+BmaChOk2xLXloUWZyLG0J81KMgBik85XZjyYJcl9+fQEWZdbUUe0GyKoF+w19Lpz 5+fhVqHEynPOHI22cSfGAtjT+g1sI8mvEhrHcktnga65Bc409Q3WiQrlJtzwpRpEY7Fe ZC9ck0mhnPGE0KEyE0JXRaYPhDXMSmEH/Rm8ZfXKMPpIrKPhp3YFQmlqs9s6dFOwgwCs akx3wuof/B1a2iVD7LhgGy8N0kiT8obWZBwzDjyZAeF/RpxpvRiEebew9xRTA+NQ9wLc U85g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1736310186; x=1736914986; h=cc:to:subject:message-id:date:from:in-reply-to:references :mime-version:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=dGPpQ6VJIV9xnsCD1bMq8dKJH1L2LIGUbC/YN69ZewQ=; b=Gt9/3StGpKmMlb5q8NtUGe6EVdhLMA6MrtzP1tLNm8Wl8svDR3Vfk4GSE3kDG2Wmg8 P6prBIP8ouIshLVlJyQg/Io4y3ei+iQR8JtDv8r+/ig1sZtqf8J0ESPR1kseRf/km5Pd p8wM6mCybbqj0zFBJD9T2zmtCrlfuUGeuX02gxYOz3vYGvRRM3ioKPv55N7nuso8Rvpw RiWi6ZQEs4t7XuJiAzSh1v3hGUJWFSSiz7UyI+H3owiUE80UCrLvPyuOPUyCJ22fMW+Z Ure5eIQTVWT1We3xv7V2ArmqF/mmalPRvqxqNUxyXW/TA+nlP2HP+sbUWwj8N3wADhrn RIFA== X-Gm-Message-State: AOJu0Yw/zgyN3Hn7lHoagWyTMTLzOMdYIpanpugp0YUyAyHFbYire6py accienBZKCdFxQ/vg0uUv0aaX8/v1w51vm3u3tMt/HGw9o1CJRyD2oUFC4P03IPwSql3FlOfMtV 2v7Uu96Ouhz0kvjmTgZ/vHWW+Cy4L8XzT+tpb X-Gm-Gg: ASbGncvJqMBBZiQJkForBRhuZnb2io3uz32Gpj6eKu767HZsqivgEBOOqtEdiz/dReX J+knKxOOU9PA9DTTDN/BRiQcvKqyPCR9YoBM= X-Google-Smtp-Source: AGHT+IFJk+LnjgzfAgjKMHkGNMZINBuze+9KKgJnWmwTidN5rr18hYT2maqgMy22Lp3yaIP9FHYwKKIkEhmdmvdMiN0= X-Received: by 2002:a05:6214:2263:b0:6d8:871d:49f1 with SMTP id 6a1803df08f44-6df9b2d875amr26384756d6.44.1736310185822; Tue, 07 Jan 2025 20:23:05 -0800 (PST) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 References: <20241221063119.29140-1-kanchana.p.sridhar@intel.com> <20241221063119.29140-12-kanchana.p.sridhar@intel.com> In-Reply-To: From: Yosry Ahmed Date: Tue, 7 Jan 2025 20:22:29 -0800 X-Gm-Features: AbW1kvahJqqOC4RLQWaNTlCxRD-xLkUmXDWJfgFp5dXM9_qBnCBSg5cYKQkB1W4 Message-ID: Subject: Re: [PATCH v5 11/12] mm: zswap: Restructure & simplify zswap_store() to make it amenable for batching. To: "Sridhar, Kanchana P" Cc: "linux-kernel@vger.kernel.org" , "linux-mm@kvack.org" , "hannes@cmpxchg.org" , "nphamcs@gmail.com" , "chengming.zhou@linux.dev" , "usamaarif642@gmail.com" , "ryan.roberts@arm.com" , "21cnbao@gmail.com" <21cnbao@gmail.com>, "akpm@linux-foundation.org" , "linux-crypto@vger.kernel.org" , "herbert@gondor.apana.org.au" , "davem@davemloft.net" , "clabbe@baylibre.com" , "ardb@kernel.org" , "ebiggers@google.com" , "surenb@google.com" , "Accardi, Kristen C" , "Feghali, Wajdi K" , "Gopal, Vinodh" Content-Type: text/plain; charset="UTF-8" [..] > > > diff --git a/mm/zswap.c b/mm/zswap.c > > > index 99cd78891fd0..1be0f1807bfc 100644 > > > --- a/mm/zswap.c > > > +++ b/mm/zswap.c > > > @@ -1467,77 +1467,129 @@ static void shrink_worker(struct work_struct > > *w) > > > * main API > > > **********************************/ > > > > > > -static ssize_t zswap_store_page(struct page *page, > > > - struct obj_cgroup *objcg, > > > - struct zswap_pool *pool) > > > +static bool zswap_compress_folio(struct folio *folio, > > > + struct zswap_entry *entries[], > > > + struct zswap_pool *pool) > > > { > > > - swp_entry_t page_swpentry = page_swap_entry(page); > > > - struct zswap_entry *entry, *old; > > > + long index, nr_pages = folio_nr_pages(folio); > > > > > > - /* allocate entry */ > > > - entry = zswap_entry_cache_alloc(GFP_KERNEL, page_to_nid(page)); > > > - if (!entry) { > > > - zswap_reject_kmemcache_fail++; > > > - return -EINVAL; > > > + for (index = 0; index < nr_pages; ++index) { > > > + struct page *page = folio_page(folio, index); > > > + > > > + if (!zswap_compress(page, entries[index], pool)) > > > + return false; > > > } > > > > > > - if (!zswap_compress(page, entry, pool)) > > > - goto compress_failed; > > > + return true; > > > +} > > > > > > - old = xa_store(swap_zswap_tree(page_swpentry), > > > - swp_offset(page_swpentry), > > > - entry, GFP_KERNEL); > > > - if (xa_is_err(old)) { > > > - int err = xa_err(old); > > > +/* > > > + * Store all pages in a folio. > > > + * > > > + * The error handling from all failure points is consolidated to the > > > + * "store_folio_failed" label, based on the initialization of the zswap > > entries' > > > + * handles to ERR_PTR(-EINVAL) at allocation time, and the fact that the > > > + * entry's handle is subsequently modified only upon a successful > > zpool_malloc() > > > + * after the page is compressed. > > > + */ > > > +static ssize_t zswap_store_folio(struct folio *folio, > > > + struct obj_cgroup *objcg, > > > + struct zswap_pool *pool) > > > +{ > > > + long index, nr_pages = folio_nr_pages(folio); > > > + struct zswap_entry **entries = NULL; > > > + int node_id = folio_nid(folio); > > > + size_t compressed_bytes = 0; > > > > > > - WARN_ONCE(err != -ENOMEM, "unexpected xarray error: %d\n", > > err); > > > - zswap_reject_alloc_fail++; > > > - goto store_failed; > > > + entries = kmalloc(nr_pages * sizeof(*entries), GFP_KERNEL); > > > > We can probably use kcalloc() here. > > I am a little worried about the latency penalty of kcalloc() in the reclaim path, > especially since I am not relying on zero-initialized memory for "entries".. Hmm good point, for a 2M THP we could be allocating an entire page here. [..] > > > @@ -1549,8 +1601,8 @@ bool zswap_store(struct folio *folio) > > > struct mem_cgroup *memcg = NULL; > > > struct zswap_pool *pool; > > > size_t compressed_bytes = 0; > > > + ssize_t bytes; > > > bool ret = false; > > > - long index; > > > > > > VM_WARN_ON_ONCE(!folio_test_locked(folio)); > > > VM_WARN_ON_ONCE(!folio_test_swapcache(folio)); > > > @@ -1584,15 +1636,11 @@ bool zswap_store(struct folio *folio) > > > mem_cgroup_put(memcg); > > > } > > > > > > - for (index = 0; index < nr_pages; ++index) { > > > - struct page *page = folio_page(folio, index); > > > - ssize_t bytes; > > > + bytes = zswap_store_folio(folio, objcg, pool); > > > + if (bytes < 0) > > > + goto put_pool; > > > > > > - bytes = zswap_store_page(page, objcg, pool); > > > - if (bytes < 0) > > > - goto put_pool; > > > - compressed_bytes += bytes; > > > - } > > > + compressed_bytes = bytes; > > > > What's the point of having both compressed_bytes and bytes now? > > The main reason was to cleanly handle a negative error value returned in "bytes" > (declared as ssize_t), as against a true total "compressed_bytes" (declared as size_t) > for the folio to use for objcg charging. This is similar to the current mainline > code where zswap_store() calls zswap_store_page(). I was hoping to avoid potential > issues with overflow/underflow, and for maintainability. Let me know if this is Ok. It makes sense in the current mainline because we store the return value of each call to zswap_store_page() in 'bytes', then check if it's an error value, then add it to 'compressed_bytes'. Now we have a single call to zswap_store_folio() and a single return value. AFAICT, there is currently no benefit to storing it in 'bytes', checking it, then moving it to 'compressed_bytes'. The compiler will probably optimize the variable away anyway, but it looks weird.