From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qt1-f180.google.com (mail-qt1-f180.google.com [209.85.160.180]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 846D5214219 for ; Wed, 22 Jan 2025 14:42:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.180 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737556967; cv=none; b=hRG+1P6+WeiqsMRNlN6DiVX50jbc2XA40qu/jJcBzd4vxxVeSzYqfswSDQrX/itvSNHPqmkJItyaXoPpyPAc12IIjq91Dv0rfTSuxsFdh4tpQg3maSEvk12/zMID+mDER0uC/jJhCHQ3phJrlEyhd7KaSga4q+SGU1Yg1MxuBxo= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737556967; c=relaxed/simple; bh=it4tX9MiK7KF+YUEtM6iwhf6/3eF6Yt7QvezRW/BAs8=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=AVWboShw/7T8GHFxz99qZ30Qvvc6vpr/HKgzp73/fboHplDB6+djyRgzDQViWpUHPj0HggLmZsj1F6TZRvc5S3yq7BwXc1bc9XjJqTCwEfxCN2sGM3diV8/LO8Ap/RuxAFCcS4grlUu39J/IfIuNNd7d2ybeKrIosjyvMTK1ISo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org; spf=pass smtp.mailfrom=cmpxchg.org; dkim=pass (2048-bit key) header.d=cmpxchg-org.20230601.gappssmtp.com header.i=@cmpxchg-org.20230601.gappssmtp.com header.b=Li224FZv; arc=none smtp.client-ip=209.85.160.180 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=cmpxchg.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=cmpxchg-org.20230601.gappssmtp.com header.i=@cmpxchg-org.20230601.gappssmtp.com header.b="Li224FZv" Received: by mail-qt1-f180.google.com with SMTP id d75a77b69052e-46783d44db0so65836551cf.1 for ; Wed, 22 Jan 2025 06:42:43 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=cmpxchg-org.20230601.gappssmtp.com; s=20230601; t=1737556962; x=1738161762; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=jUTvI1zjodzKf4BKTdSmbVm2vNENh5QZ9jsPAlaP8rc=; b=Li224FZveQqIHruvgoZKQ9rIy8V7mEVsvWx0yh5+uY6b88rpjfC7HFs3qalWa66EjA Qjc3x+rdvRVH60EzxBUFsvLgeVzN35ZZTQA4fgKT0hwvj0f1qrhbjbTxxmA4KHmn0/p2 buohHCRNwQGNSXxjfIpi5VZ7grhg7xU/vRzoOYcLLloDQpCDksOuDbCyH935bB04vxCa JidSJvYeafrpXjPssMFMpibwcO0+OrMa+RLK4lVMb73vLDo73jEQaXnEWXOOuugoMw7x jZspV72UhsXpsn1uQMwQ5XtUmDe3EFiw5yHxnmUS/EZXyVrbM5NIAaMOThTry0E1xiB7 zNBw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1737556962; x=1738161762; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to; bh=jUTvI1zjodzKf4BKTdSmbVm2vNENh5QZ9jsPAlaP8rc=; b=okR+wwx6UHPk9Jb2upH8ahB2qlKpezo5AH3Kwp6x/LeIxWzQvIHod/rHiqGpygTSRH YGvQoD08Wcgy93HRbAis3ZIrLUp72T49VM8Gs9aGElppj3ZF6cIl58y1jN8vxLyiU8bL lGxUVlUvtPgHDAbPm4pu4WTS057T+Pm8nAERD0tOuBJk2tb5Gn9jh/ODy1Y4TCqHi+jH jgXvLoVlMo8JbO5+ZPAiUUQSRtS37ohgM/b6AE/o2yYAM70XEBiqBnCW1QC6VHFzPNJR 7ygqWIJRgXN5hzztca5gcnzKLXDnxbnFJnlXao+GcH9Qt/BC+CEvOakvu+kPu9SKMtbk UE2w== X-Forwarded-Encrypted: i=1; AJvYcCXiTPdyXcZNSA1uTCrq7B76Wa8xN0Co0kRjC0m/eFtMK6mWAGje1CyJt2ZLqL/04Yz0pDknSEMh3IrQX10=@vger.kernel.org X-Gm-Message-State: AOJu0Yzs5D5K6w/eaO4Skbc0BOoaLX/L5X2PdYcpP4AcWzdLst5rOnLu yv/dCAaao+TiF745l3UkyAQ5f7C+emwsnnBQX1BNkutEosuxciWJOCiRVy7oAOI= X-Gm-Gg: ASbGncuH8ef7fLWDFjoWrlFxgjxjb3MfAqINj0pDzb3mLqcO81YGD9fuWACTyIeX7qV u1iO57oMAWb+KGAROIgEqE58lpSg+J2b0+frInntMUOgY/iblfzpmTPj/b7j06ajEzF33uJxuT6 5T8z/jpHZ239eQWJr6XsQNQo3w9V4K4hrxS7V9CeqGB8s4ydX8yi8Wwv2UJklG1H1HSRia96tqZ trA7jDHVmCaQvGK+yuFbS50Pdf0imO2frDiUqw/wkWipdUW8skSLNRLWeDzrf4x4CMG X-Google-Smtp-Source: AGHT+IHj0fasPmM9mB40hLFoTkSprPapd9w8WXM3yi8YwYCjYzRwR0BeAsihTxBYsjOZcsaYzaW4zw== X-Received: by 2002:a05:622a:1a85:b0:467:61a5:1a85 with SMTP id d75a77b69052e-46e12aa4b98mr350905731cf.30.1737556962154; Wed, 22 Jan 2025 06:42:42 -0800 (PST) Received: from localhost ([2603:7000:c01:2716:da5e:d3ff:fee7:26e7]) by smtp.gmail.com with UTF8SMTPSA id d75a77b69052e-46e102ebfbdsm64601371cf.15.2025.01.22.06.42.41 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 22 Jan 2025 06:42:41 -0800 (PST) Date: Wed, 22 Jan 2025 09:42:40 -0500 From: Johannes Weiner To: yangge1116@126.com Cc: akpm@linux-foundation.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, 21cnbao@gmail.com, david@redhat.com, baolin.wang@linux.alibaba.com, vbabka@suse.cz, liuzixing@hygon.cn Subject: Re: [PATCH V2] mm: compaction: use the actual allocation context to determine the watermarks for costly order during async memory compaction Message-ID: <20250122144240.GA217180@cmpxchg.org> References: <1736991214-29069-1-git-send-email-yangge1116@126.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <1736991214-29069-1-git-send-email-yangge1116@126.com> On Thu, Jan 16, 2025 at 09:33:34AM +0800, yangge1116@126.com wrote: > From: yangge > > There are 4 NUMA nodes on my machine, and each NUMA node has 32GB > of memory. I have configured 16GB of CMA memory on each NUMA node, > and starting a 32GB virtual machine with device passthrough is > extremely slow, taking almost an hour. > > Long term GUP cannot allocate memory from CMA area, so a maximum of > 16 GB of no-CMA memory on a NUMA node can be used as virtual machine > memory. There is 16GB of free CMA memory on a NUMA node, which is > sufficient to pass the order-0 watermark check, causing the > __compaction_suitable() function to consistently return true. > > For costly allocations, if the __compaction_suitable() function always > returns true, it causes the __alloc_pages_slowpath() function to fail > to exit at the appropriate point. This prevents timely fallback to > allocating memory on other nodes, ultimately resulting in excessively > long virtual machine startup times. > Call trace: > __alloc_pages_slowpath > if (compact_result == COMPACT_SKIPPED || > compact_result == COMPACT_DEFERRED) > goto nopage; // should exit __alloc_pages_slowpath() from here > > We could use the real unmovable allocation context to have > __zone_watermark_unusable_free() subtract CMA pages, and thus we won't > pass the order-0 check anymore once the non-CMA part is exhausted. There > is some risk that in some different scenario the compaction could in > fact migrate pages from the exhausted non-CMA part of the zone to the > CMA part and succeed, and we'll skip it instead. But only __GFP_NORETRY > allocations should be affected in the immediate "goto nopage" when > compaction is skipped, others will attempt with DEF_COMPACT_PRIORITY > anyway and won't fail without trying to compact-migrate the non-CMA > pageblocks into CMA pageblocks first, so it should be fine. > > After this fix, it only takes a few tens of seconds to start a 32GB > virtual machine with device passthrough functionality. > > Link: https://lore.kernel.org/lkml/1736335854-548-1-git-send-email-yangge1116@126.com/ > Signed-off-by: yangge > Acked-by: Vlastimil Babka Acked-by: Johannes Weiner