From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f202.google.com (mail-pl1-f202.google.com [209.85.214.202]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id ED3E018FDBE for ; Fri, 5 Dec 2025 00:38:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.202 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1764895137; cv=none; b=jJgZp7DsM0of7SGI9//I46oE6+lCRzDViMYqBlW9yH6qJNtPLQCWh5CLcxuVkpErTRrHydxNhob2T/DB/TppBSxUZka+qEDGRHbS2xZ2uVr5nqkwBsHfecqj6wcF5Fn0zMIXAecKk/DukqFp5Sfr1sQx7O3BsPm3E7MMWGgKUxI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1764895137; c=relaxed/simple; bh=jRGGucjhuVnCsCWM106TPSms/6ZdG6m5JJbmww1ag3Y=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=lxHwTdNjcgihsyR5xJUj0Fu7OzWKUTpIM0CIIrLDZijaaypsVSqZgQ7fHCW5W8YZIqxyJguyegOmnkrF9BoAg5VChtQaFmvNNTEO4lL0qnICRPV8rpLPjOgG1HGE8WjIVSbudqO/E489C6kocHRmxdIz9uj9Ei1GVOmy+FegPnM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--ackerleytng.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=34DqmuTQ; arc=none smtp.client-ip=209.85.214.202 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--ackerleytng.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="34DqmuTQ" Received: by mail-pl1-f202.google.com with SMTP id d9443c01a7336-297df52c960so33793515ad.1 for ; Thu, 04 Dec 2025 16:38:55 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20230601; t=1764895135; x=1765499935; darn=vger.kernel.org; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:from:to:cc:subject:date:message-id:reply-to; bh=jRGGucjhuVnCsCWM106TPSms/6ZdG6m5JJbmww1ag3Y=; b=34DqmuTQAOYTwBw4QGQbiZFFYau25ZdxDBIuUKPDl29jr5DbkOdflQD5MBNBv7VJlh LBkUde0mBfwk1WIY3gP+ApUPEcyhQrk4rakNiMVY2tvSIc1LxNv336I8s4V9iHKsh7x6 c6gzxPs7zfmCkdVejY5cQwgI4RECFXVMCaUUF1zTX0qlkOT9Di3UCnBcecgwbRkbMQFj TzwFNn+xYmBoH8lcbExrgfW7U3t7N6OtRpFW/fZMRuTNavsSoVcRSXdTt8SA+kZy+t6J bjmZGF+UZ0oVwx1eK69MA0YVIQyF/YECJPqlU8RCi3a/Jpj/F09+SMYzubJ9zcYX5JAC Ch/Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1764895135; x=1765499935; h=cc:to:from:subject:message-id:references:mime-version:in-reply-to :date:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=jRGGucjhuVnCsCWM106TPSms/6ZdG6m5JJbmww1ag3Y=; b=kHIptRjOOSH0J0LQFCcYehd2NYZzyTB160QVX/rUvmyPqjn8HJuvXHgOMC4hXJRK/5 jIRv814AqDl0HlRvD+KiEt+ZpkkqBsgIsLdGBuVtVRG+wNH+NXWrtU950IlOSJUE3ee5 ve2KyBFreXP6j4npLniECKz0wIT7Tjp7mMZSchrpehtGHiBgdPCmK2/TIPUcFzyknEFR ijNV2MZM9BVNW+zTiiOe/sWZmn5u0WHp9yiDQjC1yAdrL/l83LAga6F1MCrqh7Gtwo5I CfJHidB/Ot1y9q/NT+UJp6Mg1QH/nU9hYVw50/PXyaQ0PVID+CZDWbapEdCKV1fne1JT psVw== X-Forwarded-Encrypted: i=1; AJvYcCUZSvNkT14cO/TqYjVyiHwZ9f5BFdkNWnFeK6p2EMAQOOzqK6o2cIAIPXga631UtXJRm2qOWvLgkndDBFM=@vger.kernel.org X-Gm-Message-State: AOJu0Yyc4J96CysH8k5puuzw+CLaMByP7fgPBLdPUiQ7PMXFuwYlbhqc WiCZuxoV5IvPLmMOFuVF2Qw/SXjmrIz06mAXkm1XSXorBSWfv/PVssy/9o8j5tCe32dtPXkDJAl XI4Wug95FI+d659p8z6rlglAcGA== X-Google-Smtp-Source: AGHT+IEI+4QWSaT65yNp/2TgxcZCaFaqgv6I/9Ib9/x1qCnTFXLErIvauCddCu2q9NT77PNc/mruEXOJV7Fmqs3GVw== X-Received: from pgjr9.prod.google.com ([2002:a63:ec49:0:b0:bac:a20:5f05]) (user=ackerleytng job=prod-delivery.src-stubby-dispatcher) by 2002:a05:6a21:3387:b0:342:352c:77e5 with SMTP id adf61e73a8af0-363f5ea890bmr9279671637.54.1764895135064; Thu, 04 Dec 2025 16:38:55 -0800 (PST) Date: Thu, 04 Dec 2025 16:38:53 -0800 In-Reply-To: Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20251117224701.1279139-1-ackerleytng@google.com> Message-ID: Subject: Re: [RFC PATCH 0/4] Extend xas_split* to support splitting arbitrarily large entries From: Ackerley Tng To: Matthew Wilcox Cc: akpm@linux-foundation.org, linux-fsdevel@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, david@redhat.com, michael.roth@amd.com, vannapurve@google.com Content-Type: text/plain; charset="UTF-8" Ackerley Tng writes: > Matthew Wilcox writes: > >> On Mon, Nov 17, 2025 at 02:46:57PM -0800, Ackerley Tng wrote: >>> guest_memfd is planning to store huge pages in the filemap, and >>> guest_memfd's use of huge pages involves splitting of huge pages into >>> individual pages. Splitting of huge pages also involves splitting of >>> the filemap entries for the pages being split. > >> >> Hm, I'm not most concerned about the number of nodes you're allocating. > > Thanks for reminding me, I left this out of the original message. > > Splitting the xarray entry for a 1G folio (in a shift-18 node for > order=18 on x86), assuming XA_CHUNK_SHIFT is 6, would involve > > + shift-18 node (the original node will be reused - no new allocations) > + shift-12 node: 1 node allocated > + shift-6 node : 64 nodes allocated > + shift-0 node : 64 * 64 = 4096 nodes allocated > > This brings the total number of allocated nodes to 4161 nodes. struct > xa_node is 576 bytes, so that's 2396736 bytes or 2.28 MB, so splitting a > 1G folio to 4K pages costs ~2.5 MB just in filemap (XArray) entry > splitting. The other large memory cost would be from undoing HVO for the > HugeTLB folio. > At the guest_memfd biweekly call this morning, we touched on this topic again. David pointed out that the ~2MB overhead to store a 1G folio in the filemap seems a little high. IIUC the above is correct, so even if we put aside splitting, without multi-index XArrays, storing a 1G folio in the filemap would incur this number of nodes in overheads. (Hence multi-index XArrays are great :)) >> I'm most concerned that, once we have memdescs, splitting a 1GB page >> into 512 * 512 4kB pages is going to involve allocating about 20MB >> of memory (80 bytes * 512 * 512). > > I definitely need to catch up on memdescs. What's the best place for me > to learn/get an overview of how memdescs will describe memory/replace > struct folios? > > I think there might be a better way to solve the original problem of > usage tracking with memdesc support, but this was intended to make > progress before memdescs. > >> Is this necessary to do all at once? > > The plan for guest_memfd was to first split from 1G to 4K, then optimize > on that by splitting in stages, from 1G to 2M as much as possible, then > to 4K only for the page ranges that the guest shared with the host. David asked if splitting from 1G to 2M would remove the need for this extension patch series. On the call, I wrongly agreed - looking at the code again, even though the existing code kind of takes input for the target order of the split though xas, it actually still does not split to the requested order. I think some workarounds could be possible, but for the introduction of guest_memfd HugeTLB with folio restructuring, taking a dependency on non-uniform splits (splitting 1G to 511 2M folios and 512 4K folios) is significant complexity for a single series. It is significant because in addition to having to deal with non-uniform splits of the folios, we'd also have to deal with non-uniform HugeTLB vmemmap optimization. Hence I'm hoping that I could get help reviewing these changes, so that guest_memfd HugeTLB with non-uniform splits could be handled in a later stage as an optimization. Besides, David says generalizing this could help unblock other things (I forgot the detail, maybe David can chime in here) :) Thanks!