From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qv1-f50.google.com (mail-qv1-f50.google.com [209.85.219.50]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 9AF5E231836 for ; Mon, 24 Nov 2025 19:33:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.219.50 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1764012788; cv=none; b=nasCK1bOOaUr1T/UEIAPSDYXHuwLU7OBsx7y7ARyAszndCRaYfK0UA2WJyOZyPgS9xOUlb3NZ8vLmnSc1AR/bQmdP6q/X9gRCrUY0o0BUhCCkCLf+pe/CfAR7j/IGf+MHZPSh5ujODwJLKZC1FMM9iwZ+i4xSdGzPLyRj1IQpqE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1764012788; c=relaxed/simple; bh=AqwLCO/D8x2VP7dhel9IQNDZxwJZX0JTG7UXvS7ACQI=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=nOhc3VV2NpCDxIabderTQ3S6M31XpYZe6SXO0WbYUQCVvTfo1V+XBjPng2nrnIfzO/miIaK+8rf1ZkemG9WCyK7ER3STiFKjTTet3VHUgFf2ImwZioQEnHxqP4WCRiBUF+6x1JybrTutWhXgQx3hqfPnlg5mE3VeGTkCPbl3jb4= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b=bp9YU6KR; arc=none smtp.client-ip=209.85.219.50 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg.org header.i=@cmpxchg.org header.b="bp9YU6KR" Received: by mail-qv1-f50.google.com with SMTP id 6a1803df08f44-882399d60baso38838716d6.0 for ; Mon, 24 Nov 2025 11:33:04 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg.org; s=google; t=1764012783; x=1764617583; darn=vger.kernel.org; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:from:to :cc:subject:date:message-id:reply-to; bh=gsMhFDvKNXWpM7LYd4pBI2Bvu1L+GYJfurU1oXkgcDc=; b=bp9YU6KRV0sJfp3SK7yms8ZXtRCGXBEs5NA9V5jpXzw+mVGvn5MBg2kX2Wx/BRmA2N PL29HuNPTO0vPBLUbXcTNrUhGV225Wo7S9qSxcjqEeT7JDDfvQZ3O/CCv1nL63RnoCKC eEowJbW0oc0Z2s0ZoVY1L1QD3Qea90AJ5pFs0XWDY3LidUj/hrHzaUE9MutCof42/sfj gH/hXblMf4K+5sDcLkS4kQz/P5JKTpxaNv2Uwd1vOYc3OQvvSw2X6MxjRVRgRSz4e9O0 wQx56Y1Gw1xQIYI2mHmD7ZYYZuixb33PtsraAFT+OSgApqMfdU5dUeJIjRhdi6Ik7gjE K5KQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1764012783; x=1764617583; h=in-reply-to:content-transfer-encoding:content-disposition :mime-version:references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=gsMhFDvKNXWpM7LYd4pBI2Bvu1L+GYJfurU1oXkgcDc=; b=wH4MWnFfQ3PAPsBGsOmwz6zIipZ0muP8D59T34GpSL5DyN0J3tjuidFZ+1zFzILeeR 2Vv2kFvVp/rTrGmQcICMBxcTXgdrwAab7olQX4xgUiCiN4NoamtqXpoR0FLV1uKwn+T5 vIH+7e4yrio3xJaHQnyr1Kl0gfJyJTQt/MsGPh3XE9CCX+renT993cmKImRA5AzsOLtO ma7ETsvssCAi+tdNSOAuT6LxUVfGaH+iVG7m4gPzsS0q0ZwK2M2Yoh7GwCnsmuCXu7a7 yOXecOSy/QgmUu9zquB2hAVbxwIt90ojWKMef+t+B+qLwtKJm0YXvnPYQ9AkPy/3PcuZ vO/A== X-Forwarded-Encrypted: i=1; AJvYcCWlT9txs2de3yrcxp4SxwvMG96CGGnk93namjaZJupNz961fNJAJiQNvN0Mxn/HVz4hbL67aXAy/BRBONY=@vger.kernel.org X-Gm-Message-State: AOJu0YwN36PgPJ8NPWhOBl3IWlYujLMvKn5wzEULmYvj+0CYp2HM0ZZ0 3Fwne9aC1GcCGHLQUjiegBnrx62FrJVFTeg7TP/yotR2joPS18jvr9CSaQoMf+l1lKE= X-Gm-Gg: ASbGncuKra0/jucVJgkuxlDacJhGmAZciRISNF1W4WDzu4X3tcrfpEZc0F/5glKDBn0 k3x3b3OQVI9YWzr/okbyrmd48QyFt45fZdJP/VWEzwowOFep/lRiGRhgikcAvpxeKu3DXBiN4Wo 5Ug5/jk+rgbexvUWvMJpEC7bw+L6djGU7wJtoXnOJoOvHF6miTVL/CbY/V4nGnjbrHHRl0Ze1xD eXaAj6sitx9IxvPi5f3JxPsxvD6gvNL+1YFPpGsxmlOGPCVbHUwRVfWNvJMSq4M0bEK/Nyx5tqM PX6uYausYCeFYp1X1VIjlbmQnsy7D2aNydQZbmB8W6rGgzt8FgxxSMAxzkG9cQO+TVlDcC1Vcbo kOLLcIfatModEQRSCKqJoaahjLVt0kz7coFpW+fm7wIA7cv6u9RGY0kxVPVU1yvHtPXZR0ZmxJO 3HV0Q5N9L0oP+UJ+1OT1OI X-Google-Smtp-Source: AGHT+IGDlV7TeJ3rSOfQeqMqm+uAJsOndBpeBaxnYTwlv3dFlHP9QgivOyk0xR0KJb1NqyhU3n/p9g== X-Received: by 2002:a05:6214:3118:b0:70f:a4b0:1eb8 with SMTP id 6a1803df08f44-8863ae58273mr2050396d6.13.1764012783160; Mon, 24 Nov 2025 11:33:03 -0800 (PST) Received: from localhost ([2603:7000:c01:2716:e601:6a28:ae2e:9b22]) by smtp.gmail.com with ESMTPSA id 6a1803df08f44-8846e47161esm107276746d6.20.2025.11.24.11.33.02 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 24 Nov 2025 11:33:02 -0800 (PST) Date: Mon, 24 Nov 2025 14:32:58 -0500 From: Johannes Weiner To: Chris Li Cc: Andrew Morton , Kairui Song , Kemeng Shi , Nhat Pham , Baoquan He , Barry Song , Yosry Ahmed , Chengming Zhou , linux-mm@kvack.org, linux-kernel@vger.kernel.org, pratmal@google.com, sweettea@google.com, gthelen@google.com, weixugc@google.com Subject: Re: [PATCH RFC] mm: ghost swapfile support for zswap Message-ID: <20251124193258.GB476776@cmpxchg.org> References: <20251121-ghost-v1-1-cfc0efcf3855@kernel.org> <20251121114011.GA71307@cmpxchg.org> <20251124172717.GA476776@cmpxchg.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: On Mon, Nov 24, 2025 at 09:24:18PM +0300, Chris Li wrote: > On Mon, Nov 24, 2025 at 8:27 PM Johannes Weiner wrote: > > > > On Fri, Nov 21, 2025 at 05:52:09PM -0800, Chris Li wrote: > > > On Fri, Nov 21, 2025 at 3:40 AM Johannes Weiner wrote: > > > > > > > > On Fri, Nov 21, 2025 at 01:31:43AM -0800, Chris Li wrote: > > > > > The current zswap requires a backing swapfile. The swap slot used > > > > > by zswap is not able to be used by the swapfile. That waste swapfile > > > > > space. > > > > > > > > > > The ghost swapfile is a swapfile that only contains the swapfile header > > > > > for zswap. The swapfile header indicate the size of the swapfile. There > > > > > is no swap data section in the ghost swapfile, therefore, no waste of > > > > > swapfile space. As such, any write to a ghost swapfile will fail. To > > > > > prevents accidental read or write of ghost swapfile, bdev of > > > > > swap_info_struct is set to NULL. Ghost swapfile will also set the SSD > > > > > flag because there is no rotation disk access when using zswap. > > > > > > > > Zswap is primarily a compressed cache for real swap on secondary > > > > storage. It's indeed quite important that entries currently in zswap > > > > don't occupy disk slots; but for a solution to this to be acceptable, > > > > it has to work with the primary usecase and support disk writeback. > > > > > > Well, my plan is to support the writeback via swap.tiers. > > > > Do you have a link to that proposal? > > My 2024 LSF swap pony talk already has a mechanism to redirect page > cache swap entries to different physical locations. > That can also work for redirecting swap entries in different swapfiles. > > https://lore.kernel.org/linux-mm/CANeU7QnPsTouKxdK2QO8Opho6dh1qMGTox2e5kFOV8jKoEJwig@mail.gmail.com/ I looked through your slides and the LWN article, but it's very hard for me to find answers to my questions in there. In your proposal, let's say you have a swp_entry_t in the page table. What does it describe, and what are the data structures to get from this key to user data in the following scenarios: - Data is in a swapfile - Data is in zswap - Data is in being written from zswap to a swapfile - Data is back in memory due to a fault from another page table > > My understanding of swap tiers was about grouping different swapfiles > > and assigning them to cgroups. The issue with writeback is relocating > > the data that a swp_entry_t page table refers to - without having to > > find and update all the possible page tables. I'm not sure how > > swap.tiers solve this problem. > > swap.tiers is part of the picture. You are right the LPC topic mostly > covers the per cgroup portion. The VFS swap ops are my two slides of > the LPC 2023. You read from one swap file and write to another swap > file with a new swap entry allocated. Ok, and from what you wrote below, presumably at this point you would put a redirection pointer in the old location to point to the new one. This way you only have the indirection IF such a relocation actually happened, correct? But how do you store new data in the freed up old slot? > > As to your specific points - we use xarray lookups in the page cache > > fast path. It's a bold claim to say this would be too much overhead > > during swapins. > > Yes, we just get rid of xarray in swap cache lookup and get some > performance gain from it. > You are saying one extra xarray is no problem, can your team demo some > performance number of impact of the extra xarray lookup in VS? Just > run some swap benchmarks and share the result. Average and worst-case for all common usecases matter. There is no code on your side for the writeback case. (And it's exceedingly difficult to even get a mental model of how it would work from your responses and the slides you have linked). > > Two, it's not clear to me how you want to make writeback efficient > > *without* any sort of swap entry redirection. Walking all relevant > > page tables is expensive; and you have to be able to find them first. > > Swap cache can have a physical location redirection, see my 2024 LPC > slides. I have considered that way before the VS discussion. > https://lore.kernel.org/linux-mm/CANeU7QnPsTouKxdK2QO8Opho6dh1qMGTox2e5kFOV8jKoEJwig@mail.gmail.com/ There are no matches for "redir" in either the email or the slides.