From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 17DC93587AB for ; Fri, 16 Jan 2026 10:33:08 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768559590; cv=none; b=mgdX3UQPxS7aBd4L9PT4JlcIdjcsQr9z39VnvXpnuGApmXWvUQH8qZ47H2EqODWymc3byOoCfrcuA0L9aKmOd+TSgfXltVfjbo7auhAHHwmRrgIBYJpk+BRDnud74aWMGJx0tdzKVy+U87Qo1xTRQfQqAniBNByna0YSc9PBPbY= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1768559590; c=relaxed/simple; bh=RN0XrMg9lKA3d119PQWdhF3ZLiX7njIeNM578BBxQn4=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=OYy0zitXddy4cbxUKEgtJKKOy4IFRCs/3+BkZ4Xz9MHWzH53siW45JGYCIWotsybUqzwVzPB9yGTduptHE0FsslJEBR0zUiO5ccP65tYYKa7gpwnwVUu7yHrkVXUCIoN4ltkcC3OA3KJVLsxoSS1KKsOmQ9D8q+yA4oruenyZzA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=gTenCNl9; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=Om5GZj2U; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="gTenCNl9"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="Om5GZj2U" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1768559587; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=5YXd/m3+AVM4e1ysf2ifYY+anBmpbJHUpwkVBLh6Ky4=; b=gTenCNl9fYDmMVR31z1vozLacoB8t3MsR1qGbEj3iFSDqONVGEEEslsImQc2KudnZ89Rd3 0STRIeSBFPBkqyeayLJt8iyF1gk2IkVRUXITKl6uUcZhttUqm+Wh/L4en1nZK0iipRDfa3 lRJl7WajapBwNnXKNJwc7IXWQ7Acfu4= Received: from mail-lj1-f198.google.com (mail-lj1-f198.google.com [209.85.208.198]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-344-jDWCGl-tO76wAXI9P-JzRQ-1; Fri, 16 Jan 2026 05:33:04 -0500 X-MC-Unique: jDWCGl-tO76wAXI9P-JzRQ-1 X-Mimecast-MFC-AGG-ID: jDWCGl-tO76wAXI9P-JzRQ_1768559583 Received: by mail-lj1-f198.google.com with SMTP id 38308e7fff4ca-38317963123so11640611fa.2 for ; Fri, 16 Jan 2026 02:33:04 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1768559583; x=1769164383; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=5YXd/m3+AVM4e1ysf2ifYY+anBmpbJHUpwkVBLh6Ky4=; b=Om5GZj2U3Lp1YKwiH3VRsSfGpm8vbuJH4zbaxFXE5yhB5VSdynkQg6sm9gnXK3q3N1 CLEhiVRmxznvzBGr0Mw7VgUwskc8gALFff/pJh/Embo0VkZQyjxeU+z3NRknrGgjc07L LITCeARLMfNI950C45Zg0yogoWvj4UB002hHfoxS3eeBoQmmY2JGEA5UHjwXEaUzxuNd cvziKqUk5OCgf2w37SMllW3Hkgr4KAaX2gzsoMy8ijL8YN6aIC/mERnJppkympsQQxiG rwqCcRJvCfzeVS72ZR8AoPHLVffJHB6wA0IS/Le75+Suzra15U3x5pl5cPTptnyuQMc3 tbCw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1768559583; x=1769164383; h=content-transfer-encoding:in-reply-to:from:content-language :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=5YXd/m3+AVM4e1ysf2ifYY+anBmpbJHUpwkVBLh6Ky4=; b=hAM277Tzhyov/GfQSxvZGzKqyVSDiOMVHGVNWbpT5lWMdqipyxiVOtQRQDx8O7FfDI 9S8kdwC82KG75uvbnHyVeDycI9IvcR2CE7amryDldMMaSLBzcLM6K3sosJ6vGICQ+xOq WmJOFUpYr1P6MayqbEeRnOrg+gq7Acvc+JE+GQLO0IcXBp7f63PBQfcfgPgwE/7TlMQj 0OXmzgHFPGWLy/iPln1CZ/HVyLNsqiwB0MLAZr3e3T5PXD4fIUd0Y+XHbwIdciKR6x4x rnq5/+i3GwK1t48ycS4cBZvGzMAwAm1JUZ4c2Xjk6pNcOHDfQrAzem5P9m8UstbnOJEz Zc0A== X-Gm-Message-State: AOJu0Yx9dFmyTnPVz2o+dg10upFCYxmfJIMgGQefyk4YzDC1xe90bDsP rIG4S79yf/0CJFCCqGLx8UgQPCPImwmaNVhjXiIkmN5JArLtC7KXv5pVfs633KEvHU82EDK8aKb lmLKiT6TBFsO7ClNiusdxw4Rq8Mj2HWbtIaxjEKYogpNZIFz75KEeBv3i6D0Ii1wA X-Gm-Gg: AY/fxX4K8NuhiGqu8bu/U5GP9a9MEYdc2i+awk0uCsVl6ee16qO3pFfzOk5vSGymjDY x3UafdACu/0psW1HXsjMcRMJaxViiv0evt+/d/h4xptzcdHAMORwHjECbvG4PPHpx0T7sW8RATL xlleXqXPfERK5vO6aZQ3qoxVAxIIo9HjvrxOirjCQgIlDFtluIlmEnWYWDKdEGVsGGoaknfqRSG VQqQngEGl1sKecH267++BGY3+NXBNGDbqlhQAAsCi94nqzJPlonfb3K28qNnPTkfC1152bLXwBd cu61ZxzogpWVIo2IALFQafLGOn9cgGe2yPodi5uwjp9SeVuxo1vAlaLJLD+oKZlnxiZJTj+sld4 rwJNZrxP6iN348MvejbuQngpZW4hWzrAZ5Rw= X-Received: by 2002:a2e:be8f:0:b0:37b:b00b:799d with SMTP id 38308e7fff4ca-38384269cdemr6704051fa.24.1768559582714; Fri, 16 Jan 2026 02:33:02 -0800 (PST) X-Received: by 2002:a2e:be8f:0:b0:37b:b00b:799d with SMTP id 38308e7fff4ca-38384269cdemr6703981fa.24.1768559582254; Fri, 16 Jan 2026 02:33:02 -0800 (PST) Received: from [192.168.1.86] (85-23-51-1.bb.dnainternet.fi. [85.23.51.1]) by smtp.gmail.com with ESMTPSA id 38308e7fff4ca-38384fb9ab8sm6522041fa.48.2026.01.16.02.33.01 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Fri, 16 Jan 2026 02:33:01 -0800 (PST) Message-ID: <0a5e6358-9383-4605-bfb7-d2cc48335458@redhat.com> Date: Fri, 16 Jan 2026 12:33:01 +0200 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH 0/3] Migrate on fault for device pages To: Balbir Singh Cc: linux-kernel@vger.kernel.org, David Hildenbrand , Jason Gunthorpe , Leon Romanovsky , Alistair Popple , Zi Yan , Matthew Brost , linux-mm@kvack.org References: <20260114091923.3950465-1-mpenttil@redhat.com> <65cf6d71-9440-423a-9a70-b6b40622440e@nvidia.com> Content-Language: en-US From: =?UTF-8?Q?Mika_Penttil=C3=A4?= In-Reply-To: <65cf6d71-9440-423a-9a70-b6b40622440e@nvidia.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Hi Balbir! On 1/16/26 01:16, Balbir Singh wrote: > On 1/14/26 20:19, mpenttil@redhat.com wrote: >> From: Mika Penttilä >> >> Currently, the way device page faulting and migration works >> is not optimal, if you want to do both fault handling and >> migration at once. >> >> Being able to migrate not present pages (or pages mapped with incorrect >> permissions, eg. COW) to the GPU requires doing either of the >> following sequences: >> >> 1. hmm_range_fault() - fault in non-present pages with correct permissions, etc. >> 2. migrate_vma_*() - migrate the pages >> >> Or: >> >> 1. migrate_vma_*() - migrate present pages >> 2. If non-present pages detected by migrate_vma_*(): >> a) call hmm_range_fault() to fault pages in >> b) call migrate_vma_*() again to migrate now present pages >> >> The problem with the first sequence is that you always have to do two >> page walks even when most of the time the pages are present or zero page >> mappings so the common case takes a performance hit. >> >> The second sequence is better for the common case, but far worse if >> pages aren't present because now you have to walk the page tables three >> times (once to find the page is not present, once so hmm_range_fault() >> can find a non-present page to fault in and once again to setup the >> migration). It is also tricky to code correctly. One page table walk >> could costs over 1000 cpu cycles on X86-64, which is a significant hit. >> >> We should be able to walk the page table once, faulting >> pages in as required and replacing them with migration entries if >> requested. >> >> Add a new flag to HMM APIs, HMM_PFN_REQ_MIGRATE, >> which tells to prepare for migration also during fault handling. >> Also, for the migrate_vma_setup() call paths, a flags, MIGRATE_VMA_FAULT, >> is added to tell to add fault handling to migrate. >> >> Tested in X86-64 VM with HMM test device, passing the selftests. >> Tested also rebased on the >> "Remove device private pages from physical address space" series: >> https://lore.kernel.org/linux-mm/20260107091823.68974-1-jniethe@nvidia.com/ >> plus a small patch to adjust with no problems. >> >> Changes from RFC: >> - rebase on 6.19-rc5 >> - adjust for the device THP >> - changes from feedback >> >> Revisions: >> - RFC https://lore.kernel.org/linux-mm/20250814072045.3637192-1-mpenttil@redhat.com/ >> >> Cc: David Hildenbrand >> Cc: Jason Gunthorpe >> Cc: Leon Romanovsky >> Cc: Alistair Popple >> Cc: Balbir Singh >> Cc: Zi Yan >> Cc: Matthew Brost >> Suggested-by: Alistair Popple >> Signed-off-by: Mika Penttilä >> >> Mika Penttilä (3): >> mm: unified hmm fault and migrate device pagewalk paths >> mm: add new testcase for the migrate on fault case >> mm:/migrate_device.c: remove migrate_vma_collect_*() functions >> >> include/linux/hmm.h | 17 +- >> include/linux/migrate.h | 6 +- >> lib/test_hmm.c | 100 +++- >> lib/test_hmm_uapi.h | 19 +- >> mm/hmm.c | 657 +++++++++++++++++++++++-- >> mm/migrate_device.c | 589 +++------------------- >> tools/testing/selftests/mm/hmm-tests.c | 54 ++ >> 7 files changed, 869 insertions(+), 573 deletions(-) >> > > I see some kernel test robot failures, I assume there will be a new version > for review? > > Balbir > Yes I will fix those and send a new version. The test robot failures are basically missing guards for !MIGRATE configs. Thanks, Mika