From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1754032AbcFPKJA (ORCPT ); Thu, 16 Jun 2016 06:09:00 -0400 Received: from mail-lf0-f66.google.com ([209.85.215.66]:34038 "EHLO mail-lf0-f66.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751681AbcFPKI6 (ORCPT ); Thu, 16 Jun 2016 06:08:58 -0400 Date: Thu, 16 Jun 2016 13:08:54 +0300 From: "Kirill A. Shutemov" To: Hillf Danton Cc: "'Ebru Akagunduz'" , "Kirill A. Shutemov" , linux-kernel , linux-mm@kvack.org Subject: Re: [PATCHv9-rebased2 01/37] mm, thp: make swapin readahead under down_read of mmap_sem Message-ID: <20160616100854.GB18137@node.shutemov.name> References: <04f701d1c797$1ebe6b80$5c3b4280$@alibaba-inc.com> <04f801d1c79b$b46744a0$1d35cde0$@alibaba-inc.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <04f801d1c79b$b46744a0$1d35cde0$@alibaba-inc.com> User-Agent: Mutt/1.5.23.1 (2014-03-12) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Thu, Jun 16, 2016 at 02:52:52PM +0800, Hillf Danton wrote: > > > > From: Ebru Akagunduz > > > > Currently khugepaged makes swapin readahead under down_write. This patch > > supplies to make swapin readahead under down_read instead of down_write. > > > > The patch was tested with a test program that allocates 800MB of memory, > > writes to it, and then sleeps. The system was forced to swap out all. > > Afterwards, the test program touches the area by writing, it skips a page > > in each 20 pages of the area. > > > > Link: http://lkml.kernel.org/r/1464335964-6510-4-git-send-email-ebru.akagunduz@gmail.com > > Signed-off-by: Ebru Akagunduz > > Cc: Hugh Dickins > > Cc: Rik van Riel > > Cc: "Kirill A. Shutemov" > > Cc: Naoya Horiguchi > > Cc: Andrea Arcangeli > > Cc: Joonsoo Kim > > Cc: Cyrill Gorcunov > > Cc: Mel Gorman > > Cc: David Rientjes > > Cc: Vlastimil Babka > > Cc: Aneesh Kumar K.V > > Cc: Johannes Weiner > > Cc: Michal Hocko > > Cc: Minchan Kim > > Signed-off-by: Andrew Morton > > --- > > mm/huge_memory.c | 92 ++++++++++++++++++++++++++++++++++++++------------------ > > 1 file changed, 63 insertions(+), 29 deletions(-) > > > > diff --git a/mm/huge_memory.c b/mm/huge_memory.c > > index f2bc57c45d2f..96dfe3f09bf6 100644 > > --- a/mm/huge_memory.c > > +++ b/mm/huge_memory.c > > @@ -2378,6 +2378,35 @@ static bool hugepage_vma_check(struct vm_area_struct *vma) > > } > > > > /* > > + * If mmap_sem temporarily dropped, revalidate vma > > + * before taking mmap_sem. > > See below > > @@ -2401,11 +2430,18 @@ static void __collapse_huge_page_swapin(struct mm_struct *mm, > > continue; > > swapped_in++; > > ret = do_swap_page(mm, vma, _address, pte, pmd, > > - FAULT_FLAG_ALLOW_RETRY|FAULT_FLAG_RETRY_NOWAIT, > > + FAULT_FLAG_ALLOW_RETRY, > > Add a description in change log for it please. Ebru, would you address it? > > pteval); > > + /* do_swap_page returns VM_FAULT_RETRY with released mmap_sem */ > > + if (ret & VM_FAULT_RETRY) { > > + down_read(&mm->mmap_sem); > > + /* vma is no longer available, don't continue to swapin */ > > + if (hugepage_vma_revalidate(mm, vma, address)) > > + return false; > > Revalidate vma _after_ acquiring mmap_sem, but the above comment says _before_. Ditto. > > + if (!__collapse_huge_page_swapin(mm, vma, address, pmd)) { > > + up_read(&mm->mmap_sem); > > + goto out; > > Jump out with mmap_sem released, > > > + result = hugepage_vma_revalidate(mm, vma, address); > > + if (result) > > + goto out; > > but jump out again with mmap_sem held. > > They are cleaned up in subsequent darns? I didn't fold fixups for these > -- Kirill A. Shutemov