From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f169.google.com (mail-pl1-f169.google.com [209.85.214.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A3C621DBB13 for ; Thu, 16 Jan 2025 09:22:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.169 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737019381; cv=none; b=h0ytTgCnhi48M3Yu/Pu384woehTOarBOAX81rLHzSLnV8poAo8VyWHUkiZiFqx2YuYmDBm5SA1aIYnSR+MTDjR3ItwvzQdFI2zEM8xGuMcwuMGFkjpsmI9wNSDbhCsw/FfN7JnFQnVKE/4dZsqerj+P0KkeTo+ZRVxIEVz5ch5E= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737019381; c=relaxed/simple; bh=mfh+qIeYV+yER3hsK1uzccjIzoMLpGca7+fTQehlHC0=; h=From:To:Cc:Subject:Date:Message-Id:MIME-Version:Content-Type; b=oWTIeGR5t9NQwBNVJ3+GjAcYs5heZshFpspKcVpQrbR7oY3LN3GEFR/onSeqkipJCfoKukPujqV2lCKQkKvltlFvxQv9xVlHDEQJgV1z+cyPP2+6Fcuzqhi+DsLZsGRz+TTUjawp7S7h+WFOnvCQhl9MRF1tFAAlIp5y4RaC1Bw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=SoxJRjzP; arc=none smtp.client-ip=209.85.214.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="SoxJRjzP" Received: by mail-pl1-f169.google.com with SMTP id d9443c01a7336-219f8263ae0so10426635ad.0 for ; Thu, 16 Jan 2025 01:22:59 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1737019379; x=1737624179; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to; bh=7xRWpo+T7uUV2bVwHTa7TpcAURmustR4SiBpyzSbmGY=; b=SoxJRjzPgw2gBIlzeDY7xdWzQ7Jzvrs+4Ow15Qg3a9OE9wWfh69Fp6J8hW8y6kHz2K 34K+VrDuLUiK+G1SgNPC51mj0JP/x7pTMdlGQeZNKK0K89pKIcImlhUX2zmlZc+siTT2 Y8AahANJInJAfYaqs1iUnVuW30mMgDgcd2+lWmlbBmSKlIlrZ0FjehqjWATReotE2sFa /+G7I3UvVbgR8E4AdpRLeJLWfI8C1oZvTp0izg85hbhdVtlbGwY9HZEDcOlCzZWFWa+g 4NOgTmxsDMw8hUiVjDNqbTCsSQb0ysSw0mThkNmGw42LE1hV8j+paYM+uH2ifWZyaV1r 5uEw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1737019379; x=1737624179; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=7xRWpo+T7uUV2bVwHTa7TpcAURmustR4SiBpyzSbmGY=; b=CGdeSP8HNTuFeBy3CIWapwkIBSzUQso8XY4H8zeFau6i33xh6+M9xXly2yTiqtfL5z hdQyaOg5WXZ2PemVfe/OYrz+g/QJtFQyDD801RpsIQYgFoFeAVt2UIGclrzn2ViTQ3Lt RK2roUbMnZQEZVCtldYJrQNgC+byGhvXx6n264nxktpQEz7RZ0YnoXTNHId/Yh+xXnEO R6vjOKlFEPJD6lscYAShq7z7Z8ZJSSN1m5hHpY4x0PLKZVJYdz6xN91OP67a1ZbJ7+KL x5MDXuMvnNnFM7/8Vg+iQH56iRs27C/2gtUCKXDZ8zmCwY4NHqUHaTAGHOQZ3AF074Bb OAXQ== X-Forwarded-Encrypted: i=1; AJvYcCWofrP2VKqaT4tete9Qwh5NhXV8PmrbkXLFW1GCUcm4Aawx2Cosnk/9uVzfIN21X/eZ2l3qiGhKzleZGS0=@vger.kernel.org X-Gm-Message-State: AOJu0YwHvF8J125NfJtoUkJOeWMgBp9CqWC/rTpJZ0njij6TV0/E849F qiXshDSvgXX2w398HMUxwjYZyk74zl4U9fgmV8IQS7pM31NNfAAX X-Gm-Gg: ASbGnctMbeMqscuIsDdoXgnmyNZqjArDxLHc9jTqmsamhXHVoWKiKI1Qz1ALDOh03ue HeXY3LRbun1XmJz1xmr/JS8USGByONK9aNK178DVO8JDhNFkhwThog8xYXHyOfhYOqmsaBlC/34 EI8LcjxEPn9OrVVoxagCg931P1+kVWynbTCA7P4ZY2vtCQ2bvNWkIAJG/g6K6u9RrxA4Kl4c3if zNpmV9WkIOlhodunhs/NHX31l8bF/BCUDrQlVVcvoMPoco= X-Google-Smtp-Source: AGHT+IG+hGvwVWEKkEmMNIBmLVXYEOB7+s/MtGgovJ800P0ZGy4dP6w4jiNWq1tDZ3gCVZ2BRwyeMA== X-Received: by 2002:a05:6a00:3e01:b0:725:e37d:cd35 with SMTP id d2e1a72fcca58-72d21ff4af7mr50049961b3a.18.1737019378738; Thu, 16 Jan 2025 01:22:58 -0800 (PST) Received: from localhost ([1.54.215.56]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-72d4054942dsm10368393b3a.21.2025.01.16.01.22.58 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 16 Jan 2025 01:22:58 -0800 (PST) From: Nhat Pham To: lsf-pc@lists.linux-foundation.org, akpm@linux-foundation.org, hannes@cmpxchg.org Cc: ryncsn@gmail.com, chengming.zhou@linux.dev, yosryahmed@google.com, chrisl@kernel.org, linux-mm@kvack.org, kernel-team@meta.com, linux-kernel@vger.kernel.org, shakeel.butt@linux.dev, hch@infradead.org, hughd@google.com, 21cnbao@gmail.com, usamaarif642@gmail.com Subject: [LSF/MM/BPF TOPIC] Virtual Swap Space Date: Thu, 16 Jan 2025 01:22:54 -0800 Message-Id: <20250116092254.204549-1-nphamcs@gmail.com> X-Mailer: git-send-email 2.40.1 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit My apologies if I missed any interested party in the cc list - hopefully the mailing lists cc's suffice :) I'd like to (re-)propose the topic of swap abstraction layer for the conference, as a continuation of Yosry's proposals at LSFMMBPF 2023 (see [1], [2], [3]). (AFAICT, the same idea has been floated by Rik van Riel since at least 2011 - see [8]). I have a working(-ish) prototype, which hopefully will be submission-ready soon. For now, I'd like to give the motivation/context for the topic, as well as some high level design: I. Motivation Currently, when an anon page is swapped out, a slot in a backing swap device is allocated and stored in the page table entries that refer to the original page. This slot is also used as the "key" to find the swapped out content, as well as the index to swap data structures, such as the swap cache, or the swap cgroup mapping. Tying a swap entry to its backing slot in this way is performant and efficient when swap is purely just disk space, and swapoff is rare. However, the advent of many swap optimizations has exposed major drawbacks of this design. The first problem is that we occupy a physical slot in the swap space, even for pages that are NEVER expected to hit the disk: pages compressed and stored in the zswap pool, zero-filled pages, or pages rejected by both of these optimizations when zswap writeback is disabled. This is the arguably central shortcoming of zswap: * In deployments when no disk space can be afforded for swap (such as mobile and embedded devices), users cannot adopt zswap, and are forced to use zram. This is confusing for users, and creates extra burdens for developers, having to develop and maintain similar features for two separate swap backends (writeback, cgroup charging, THP support, etc.). For instance, see the discussion in [4]. * Resource-wise, it is hugely wasteful in terms of disk usage, and limits the memory saving potentials of these optimizations by the static size of the swapfile, especially in high memory systems that can have up to terabytes worth of memory. It also creates significant challenges for users who rely on swap utilization as an early OOM signal. Another motivation for a swap redesign is to simplify swapoff, which is complicated and expensive in the current design. Tight coupling between a swap entry and its backing storage means that it requires a whole page table walk to update all the page table entries that refer to this swap entry, as well as updating all the associated swap data structures (swap cache, etc.). II. High Level Design Overview To fix the aforementioned issues, we need an abstraction that separates a swap entry from its physical backing storage. IOW, we need to “virtualize” the swap space: swap clients will work with a virtual swap slot (that is dynamically allocated on-demand), storing it in page table entries, and using it to index into various swap-related data structures. The backing storage is decoupled from this slot, and the newly introduced layer will “resolve” the ID to the actual storage, as well as cooperating with the swap cache to handle all the required synchronization. This layer also manages other metadata of the swap entry, such as its lifetime information (swap count), via a dynamically allocated per-entry swap descriptor: struct swp_desc { swp_entry_t vswap; union { swp_slot_t slot; struct folio *folio; struct zswap_entry *zswap_entry; }; struct rcu_head rcu; rwlock_t lock; enum swap_type type; #ifdef CONFIG_MEMCG atomic_t memcgid; #endif atomic_t in_swapcache; struct kref refcnt; atomic_t swap_count; }; This design allows us to: * Decouple zswap (and zeromapped swap entry) from backing swapfile: simply associate the swap ID with one of the supported backends: a zswap entry, a zero-filled swap page, a slot on the swapfile, or a page in memory . * Simplify and optimize swapoff: we only have to fault the page in and have the swap ID points to the page instead of the on-disk swap slot. No need to perform any page table walking :) III. Future Use Cases Other than decoupling swap backends and optimizing swapoff, this new design allows us to implement the following more easily and efficiently: * Multi-tier swapping (as mentioned in [5]), with transparent transferring (promotion/demotion) of pages across tiers (see [8] and [9]). Similar to swapoff, with the old design we would need to perform the expensive page table walk. * Swapfile compaction to alleviate fragmentation (as proposed by Ying Huang in [6]). * Mixed backing THP swapin (see [7]): Once you have pinned down the backing store of THPs, then you can dispatch each range of subpages to appropriate pagein handler. [1]: https://lore.kernel.org/all/CAJD7tkbCnXJ95Qow_aOjNX6NOMU5ovMSHRC+95U4wtW6cM+puw@mail.gmail.com/ [2]: https://lwn.net/Articles/932077/ [3]: https://www.youtube.com/watch?v=Hwqw_TBGEhg [4]: https://lore.kernel.org/all/Zqe_Nab-Df1CN7iW@infradead.org/ [5]: https://lore.kernel.org/lkml/CAF8kJuN-4UE0skVHvjUzpGefavkLULMonjgkXUZSBVJrcGFXCA@mail.gmail.com/ [6]: https://lore.kernel.org/linux-mm/87o78mzp24.fsf@yhuang6-desk2.ccr.corp.intel.com/ [7]: https://lore.kernel.org/all/CAGsJ_4ysCN6f7qt=6gvee1x3ttbOnifGneqcRm9Hoeun=uFQ2w@mail.gmail.com/ [8]: https://lore.kernel.org/linux-mm/4DA25039.3020700@redhat.com/ [9]: https://lore.kernel.org/all/CA+ZsKJ7DCE8PMOSaVmsmYZL9poxK6rn0gvVXbjpqxMwxS2C9TQ@mail.gmail.com/