From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1422664AbcFGRuZ (ORCPT ); Tue, 7 Jun 2016 13:50:25 -0400 Received: from merlin.infradead.org ([205.233.59.134]:56735 "EHLO merlin.infradead.org" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S933143AbcFGRuW (ORCPT ); Tue, 7 Jun 2016 13:50:22 -0400 Date: Tue, 7 Jun 2016 19:50:17 +0200 From: Peter Zijlstra To: Mel Gorman Cc: Thomas Gleixner , Sebastian Andrzej Siewior , Mike Galbraith , Davidlohr Bueso , lkml Subject: Re: [PATCH] futex: Calculate the futex key based on a tail page for file-based futexes Message-ID: <20160607175017.GK30154@twins.programming.kicks-ass.net> References: <20160607123017.GJ2469@suse.de> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20160607123017.GJ2469@suse.de> User-Agent: Mutt/1.5.23.1 (2014-03-12) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Tue, Jun 07, 2016 at 01:30:17PM +0100, Mel Gorman wrote: > Mike Galbraith reported that the LTP test case futex_wake04 was broken > by commit 65d8fc777f6d ("futex: Remove requirement for lock_page() > in get_futex_key()"). > > This test case uses futexes backed by hugetlbfs pages and so there is an > associated inode with a futex stored on such pages. The problem is that > the key is being calculated based on the head page index of the hugetlbfs > page and not the tail page. > > Prior to the optimisation, the page lock was used to stabilise mappings and > pin the inode is file-backed which is overkill. If the page was a compound > page, the head page was automatically looked up as part of the page lock > operation but the tail page index was used to calculate the futex key. > > After the optimisation, the compound head is looked up early and the page > lock is only relied upon to identify truncated pages, special pages or a > shmem page moving to swapcache. The head page is looked up because without > the page lock, special care has to be taken to pin the inode correctly. > However, the tail page is still required to calculate the futex key so this > patch records the tail page. > > On vanilla 4.6, the output of the test case is; > > futex_wake04 0 TINFO : Hugepagesize 2097152 > futex_wake04 1 TFAIL : futex_wake04.c:126: Bug: wait_thread2 did not wake after 30 secs. > > With the patch applied > > futex_wake04 0 TINFO : Hugepagesize 2097152 > futex_wake04 1 TPASS : Hi hydra, thread2 awake! > > Reported-by: Mike Galbraith > Signed-off-by: Mel Gorman Acked-by: Peter Zijlstra (Intel) > --- > kernel/futex.c | 14 +++++++++++--- > 1 file changed, 11 insertions(+), 3 deletions(-) > > diff --git a/kernel/futex.c b/kernel/futex.c > index c20f06f38ef3..6555d5459e98 100644 > --- a/kernel/futex.c > +++ b/kernel/futex.c > @@ -469,7 +469,7 @@ get_futex_key(u32 __user *uaddr, int fshared, union futex_key *key, int rw) > { > unsigned long address = (unsigned long)uaddr; > struct mm_struct *mm = current->mm; > - struct page *page; > + struct page *page, *tail; > struct address_space *mapping; > int err, ro = 0; > > @@ -530,7 +530,15 @@ get_futex_key(u32 __user *uaddr, int fshared, union futex_key *key, int rw) > * considered here and page lock forces unnecessarily serialization > * From this point on, mapping will be re-verified if necessary and > * page lock will be acquired only if it is unavoidable > - */ > + * > + * Mapping checks require the head page for any compound page so the > + * head page and mapping is looked up now. For anonymous pages, it > + * does not matter if the page splits in the future as the key is > + * based on the address. For filesystem-backed pages, the tail is > + * required as the index of the page determines the key. For > + * base pages, there is no tail page and tail == page. > + */ > + tail = page; > page = compound_head(page); > mapping = READ_ONCE(page->mapping); > > @@ -654,7 +662,7 @@ get_futex_key(u32 __user *uaddr, int fshared, union futex_key *key, int rw) > > key->both.offset |= FUT_OFF_INODE; /* inode-based key */ > key->shared.inode = inode; > - key->shared.pgoff = basepage_index(page); > + key->shared.pgoff = basepage_index(tail); > rcu_read_unlock(); > } >