From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1758511AbdAECD6 (ORCPT ); Wed, 4 Jan 2017 21:03:58 -0500 Received: from mail-pf0-f178.google.com ([209.85.192.178]:34447 "EHLO mail-pf0-f178.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1753732AbdAECD4 (ORCPT ); Wed, 4 Jan 2017 21:03:56 -0500 Subject: Re: [PATCH v3] arm64: mm: Fix NOMAP page initialization To: Ard Biesheuvel , Robert Richter References: <20161216165437.21612-1-rrichter@cavium.com> Cc: Russell King , Catalin Marinas , Will Deacon , David Daney , Mark Rutland , James Morse , Yisheng Xie , "linux-arm-kernel@lists.infradead.org" , "linux-kernel@vger.kernel.org" , "linux-mm@kvack.org" From: Hanjun Guo Message-ID: Date: Thu, 5 Jan 2017 10:03:48 +0800 User-Agent: Mozilla/5.0 (Windows NT 6.1; WOW64; rv:45.0) Gecko/20100101 Thunderbird/45.5.1 MIME-Version: 1.0 In-Reply-To: Content-Type: text/plain; charset=utf-8; format=flowed Content-Transfer-Encoding: 7bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 2017/1/4 21:56, Ard Biesheuvel wrote: > On 16 December 2016 at 16:54, Robert Richter wrote: >> On ThunderX systems with certain memory configurations we see the >> following BUG_ON(): >> >> kernel BUG at mm/page_alloc.c:1848! >> >> This happens for some configs with 64k page size enabled. The BUG_ON() >> checks if start and end page of a memmap range belongs to the same >> zone. >> >> The BUG_ON() check fails if a memory zone contains NOMAP regions. In >> this case the node information of those pages is not initialized. This >> causes an inconsistency of the page links with wrong zone and node >> information for that pages. NOMAP pages from node 1 still point to the >> mem zone from node 0 and have the wrong nid assigned. >> >> The reason for the mis-configuration is a change in pfn_valid() which >> reports pages marked NOMAP as invalid: >> >> 68709f45385a arm64: only consider memblocks with NOMAP cleared for linear mapping >> >> This causes pages marked as nomap being no longer reassigned to the >> new zone in memmap_init_zone() by calling __init_single_pfn(). >> >> Fixing this by implementing an arm64 specific early_pfn_valid(). This >> causes all pages of sections with memory including NOMAP ranges to be >> initialized by __init_single_page() and ensures consistency of page >> links to zone, node and section. >> > > I like this solution a lot better than the first one, but I am still > somewhat uneasy about having the kernel reason about attributes of > pages it should not touch in the first place. But the fact that > early_pfn_valid() is only used a single time in the whole kernel does > give some confidence that we are not simply moving the problem > elsewhere. > > Given that you are touching arch/arm/ as well as arch/arm64, could you > explain why only arm64 needs this treatment? Is it simply because we > don't have NUMA support there? > > Considering that Hisilicon D05 suffered from the same issue, I would > like to get some coverage there as well. Hanjun, is this something you > can arrange? Thanks Sure, we will test this patch with LTP MM stress test (which triggers the bug on D05), and give the feedback. Thanks Hanjun