From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1751286AbdH1WeC (ORCPT ); Mon, 28 Aug 2017 18:34:02 -0400 Received: from mail.linuxfoundation.org ([140.211.169.12]:37440 "EHLO mail.linuxfoundation.org" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751194AbdH1WeB (ORCPT ); Mon, 28 Aug 2017 18:34:01 -0400 Date: Mon, 28 Aug 2017 15:33:59 -0700 From: Andrew Morton To: Michal Hocko Cc: Mel Gorman , Tejun Heo , , LKML , Michal Hocko Subject: Re: [PATCH] mm, memory_hotplug: do not back off draining pcp free pages from kworker context Message-Id: <20170828153359.f9b252f99647eebd339a3a89@linux-foundation.org> In-Reply-To: <20170828093341.26341-1-mhocko@kernel.org> References: <20170828093341.26341-1-mhocko@kernel.org> X-Mailer: Sylpheed 3.4.1 (GTK+ 2.24.23; x86_64-pc-linux-gnu) Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Mon, 28 Aug 2017 11:33:41 +0200 Michal Hocko wrote: > drain_all_pages backs off when called from a kworker context since > 0ccce3b924212 ("mm, page_alloc: drain per-cpu pages from workqueue > context") because the original IPI based pcp draining has been replaced > by a WQ based one and the check wanted to prevent from recursion and > inter workers dependencies. This has made some sense at the time > because the system WQ has been used and one worker holding the lock > could be blocked while waiting for new workers to emerge which can be a > problem under OOM conditions. > > Since then ce612879ddc7 ("mm: move pcp and lru-pcp draining into single > wq") has moved draining to a dedicated (mm_percpu_wq) WQ with a rescuer > so we shouldn't depend on any other WQ activity to make a forward > progress so calling drain_all_pages from a worker context is safe as > long as this doesn't happen from mm_percpu_wq itself which is not the > case because all workers are required to _not_ depend on any MM locks. > > Why is this a problem in the first place? ACPI driven memory hot-remove > (acpi_device_hotplug) is executed from the worker context. We end > up calling __offline_pages to free all the pages and that requires > both lru_add_drain_all_cpuslocked and drain_all_pages to do their job > otherwise we can have dangling pages on pcp lists and fail the offline > operation (__test_page_isolated_in_pageblock would see a page with 0 > ref. count but without PageBuddy set). > > Fix the issue by removing the worker check in drain_all_pages. > lru_add_drain_all_cpuslocked doesn't have this restriction so it works > as expected. > > Fixes: 0ccce3b924212 ("mm, page_alloc: drain per-cpu pages from workqueue context") > Signed-off-by: Michal Hocko No cc:stable?