From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1753522AbdAaBhT (ORCPT ); Mon, 30 Jan 2017 20:37:19 -0500 Received: from mga06.intel.com ([134.134.136.31]:11137 "EHLO mga06.intel.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751180AbdAaBhS (ORCPT ); Mon, 30 Jan 2017 20:37:18 -0500 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.33,312,1477983600"; d="scan'208";a="60011915" Subject: Re: [RFC V2 02/12] mm: Isolate HugeTLB allocations away from CDM nodes To: Anshuman Khandual , linux-kernel@vger.kernel.org, linux-mm@kvack.org References: <20170130033602.12275-1-khandual@linux.vnet.ibm.com> <20170130033602.12275-3-khandual@linux.vnet.ibm.com> <01671749-c649-e015-4f51-7acaa1fb5b80@intel.com> Cc: mhocko@suse.com, vbabka@suse.cz, mgorman@suse.de, minchan@kernel.org, aneesh.kumar@linux.vnet.ibm.com, bsingharora@gmail.com, srikar@linux.vnet.ibm.com, haren@linux.vnet.ibm.com, jglisse@redhat.com, dan.j.williams@intel.com From: Dave Hansen Message-ID: Date: Mon, 30 Jan 2017 17:37:09 -0800 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:45.0) Gecko/20100101 Thunderbird/45.5.1 MIME-Version: 1.0 In-Reply-To: Content-Type: text/plain; charset=windows-1252 Content-Transfer-Encoding: 7bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 01/30/2017 05:03 PM, Anshuman Khandual wrote: > On 01/30/2017 10:49 PM, Dave Hansen wrote: >> On 01/29/2017 07:35 PM, Anshuman Khandual wrote: >>> HugeTLB allocation/release/accounting currently spans across all the nodes >>> under N_MEMORY node mask. Coherent memory nodes should not be part of these >>> allocations. So use system_ram() call to fetch system RAM only nodes on the >>> platform which can then be used for HugeTLB allocation purpose instead of >>> N_MEMORY node mask. This isolates coherent device memory nodes from HugeTLB >>> allocations. >> >> Does this end up making it impossible to use hugetlbfs to access device >> memory? > > Right, thats the implementation at the moment. But going forward if we need > to have HugeTLB pages on the CDM node, then we can implement through the > sysfs interface from individual NUMA node paths instead of changing the > generic HugeTLB path. I wrote this up in the cover letter but should also > have mentioned in the comment section of this patch as well. Does this > approach look okay ? The cover letter is not the most approachable document I've ever seen. :) > "Now, we ensure complete HugeTLB allocation isolation from CDM nodes. Going > forward if we need to support HugeTLB allocation on CDM nodes on targeted > basis, then we would have to enable those allocations through the > /sys/devices/system/node/nodeN/hugepages/hugepages-16384kB/nr_hugepages > interface while still ensuring isolation from other generic sysctl and > /sys/kernel/mm/hugepages/hugepages-16384kB/nr_hugepages interfaces." That would be passable if that's the only way you can allocate hugetlbfs pages. But we also have the fault-based allocations that can pull stuff right out of the buddy allocator. This approach would break that path entirely. FWIW, I think you really need to separate the true "CDM" stuff that's *really* device-specific from the parts of this from which you really just want to implement isolation.