From: Peter Zijlstra <peterz@infradead.org>
To: torvalds@linux-foundation.org, will.deacon@arm.com,
oleg@redhat.com, paulmck@linux.vnet.ibm.com,
benh@kernel.crashing.org, mpe@ellerman.id.au, npiggin@gmail.com
Cc: linux-kernel@vger.kernel.org, mingo@kernel.org,
stern@rowland.harvard.edu, peterz@infradead.org,
Mel Gorman <mgorman@suse.de>, Rik van Riel <riel@redhat.com>
Subject: [RFC][PATCH 1/5] mm: Rework {set,clear,mm}_tlb_flush_pending()
Date: Wed, 07 Jun 2017 18:15:02 +0200 [thread overview]
Message-ID: <20170607162013.705678923@infradead.org> (raw)
In-Reply-To: <20170607161501.819948352@infradead.org>
[-- Attachment #1: peterz-mm_tlb_flush_pending.patch --]
[-- Type: text/plain, Size: 3126 bytes --]
Commit:
af2c1401e6f9 ("mm: numa: guarantee that tlb_flush_pending updates are visible before page table updates")
added smp_mb__before_spinlock() to set_tlb_flush_pending(). I think we
can solve the same problem without this barrier.
If instead we mandate that mm_tlb_flush_pending() is used while
holding the PTL we're guaranteed to observe prior
set_tlb_flush_pending() instances.
For this to work we need to rework migrate_misplaced_transhuge_page()
a little and move the test up into do_huge_pmd_numa_page().
Cc: Mel Gorman <mgorman@suse.de>
Cc: Rik van Riel <riel@redhat.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
---
--- a/include/linux/mm_types.h
+++ b/include/linux/mm_types.h
@@ -527,18 +527,16 @@ static inline cpumask_t *mm_cpumask(stru
*/
static inline bool mm_tlb_flush_pending(struct mm_struct *mm)
{
- barrier();
+ /*
+ * Must be called with PTL held; such that our PTL acquire will have
+ * observed the store from set_tlb_flush_pending().
+ */
return mm->tlb_flush_pending;
}
static inline void set_tlb_flush_pending(struct mm_struct *mm)
{
mm->tlb_flush_pending = true;
-
- /*
- * Guarantee that the tlb_flush_pending store does not leak into the
- * critical section updating the page tables
- */
- smp_mb__before_spinlock();
+ barrier();
}
/* Clearing is done after a TLB flush, which also provides a barrier. */
static inline void clear_tlb_flush_pending(struct mm_struct *mm)
--- a/mm/huge_memory.c
+++ b/mm/huge_memory.c
@@ -1410,6 +1410,7 @@ int do_huge_pmd_numa_page(struct vm_faul
unsigned long haddr = vmf->address & HPAGE_PMD_MASK;
int page_nid = -1, this_nid = numa_node_id();
int target_nid, last_cpupid = -1;
+ bool need_flush = false;
bool page_locked;
bool migrated = false;
bool was_writable;
@@ -1490,10 +1491,29 @@ int do_huge_pmd_numa_page(struct vm_faul
}
/*
+ * Since we took the NUMA fault, we must have observed the !accessible
+ * bit. Make sure all other CPUs agree with that, to avoid them
+ * modifying the page we're about to migrate.
+ *
+ * Must be done under PTL such that we'll observe the relevant
+ * set_tlb_flush_pending().
+ */
+ if (mm_tlb_flush_pending(mm))
+ need_flush = true;
+
+ /*
* Migrate the THP to the requested node, returns with page unlocked
* and access rights restored.
*/
spin_unlock(vmf->ptl);
+
+ /*
+ * We are not sure a pending tlb flush here is for a huge page
+ * mapping or not. Hence use the tlb range variant
+ */
+ if (need_flush)
+ flush_tlb_range(vma, haddr, haddr + HPAGE_PMD_SIZE);
+
migrated = migrate_misplaced_transhuge_page(vma->vm_mm, vma,
vmf->pmd, pmd, vmf->address, page, target_nid);
if (migrated) {
--- a/mm/migrate.c
+++ b/mm/migrate.c
@@ -1935,12 +1935,6 @@ int migrate_misplaced_transhuge_page(str
put_page(new_page);
goto out_fail;
}
- /*
- * We are not sure a pending tlb flush here is for a huge page
- * mapping or not. Hence use the tlb range variant
- */
- if (mm_tlb_flush_pending(mm))
- flush_tlb_range(vma, mmun_start, mmun_end);
/* Prepare a page as a migration target */
__SetPageLocked(new_page);
next prev parent reply other threads:[~2017-06-07 16:21 UTC|newest]
Thread overview: 41+ messages / expand[flat|nested] mbox.gz Atom feed top
2017-06-07 16:15 [RFC][PATCH 0/5] Getting rid of smp_mb__before_spinlock Peter Zijlstra
2017-06-07 16:15 ` Peter Zijlstra [this message]
2017-06-09 14:45 ` [RFC][PATCH 1/5] mm: Rework {set,clear,mm}_tlb_flush_pending() Will Deacon
2017-06-09 18:42 ` Peter Zijlstra
2017-07-28 17:45 ` Peter Zijlstra
2017-08-01 10:31 ` Will Deacon
2017-08-01 12:02 ` Benjamin Herrenschmidt
2017-08-01 12:14 ` Peter Zijlstra
2017-08-01 16:39 ` Peter Zijlstra
2017-08-01 16:44 ` Will Deacon
2017-08-01 16:48 ` Peter Zijlstra
2017-08-01 22:59 ` Peter Zijlstra
2017-08-02 1:23 ` Benjamin Herrenschmidt
2017-08-02 8:11 ` Peter Zijlstra
2017-08-02 8:15 ` Will Deacon
2017-08-02 8:43 ` Will Deacon
2017-08-02 8:51 ` Peter Zijlstra
2017-08-02 9:02 ` Will Deacon
2017-08-02 22:54 ` Benjamin Herrenschmidt
2017-08-02 8:45 ` Peter Zijlstra
2017-08-02 9:02 ` Will Deacon
2017-08-02 9:18 ` Peter Zijlstra
2017-08-02 13:57 ` Benjamin Herrenschmidt
2017-08-02 15:46 ` Peter Zijlstra
2017-08-02 0:17 ` Benjamin Herrenschmidt
2017-08-01 22:42 ` Benjamin Herrenschmidt
2017-06-07 16:15 ` [RFC][PATCH 2/5] locking: Introduce smp_mb__after_spinlock() Peter Zijlstra
2017-06-07 16:15 ` [RFC][PATCH 3/5] overlayfs: Remove smp_mb__before_spinlock() usage Peter Zijlstra
2017-06-07 16:15 ` [RFC][PATCH 4/5] locking: Remove smp_mb__before_spinlock() Peter Zijlstra
2017-06-07 16:15 ` [RFC][PATCH 5/5] powerpc: Remove SYNC from _switch Peter Zijlstra
2017-06-08 0:32 ` Nicholas Piggin
2017-06-08 6:54 ` Peter Zijlstra
2017-06-08 7:29 ` Nicholas Piggin
2017-06-08 7:57 ` Peter Zijlstra
2017-06-08 8:21 ` Nicholas Piggin
2017-06-08 9:54 ` Michael Ellerman
2017-06-08 10:00 ` Nicholas Piggin
2017-06-08 12:45 ` Peter Zijlstra
2017-06-08 13:18 ` Nicholas Piggin
2017-06-08 13:47 ` Peter Zijlstra
2017-06-09 14:49 ` [RFC][PATCH 0/5] Getting rid of smp_mb__before_spinlock Will Deacon
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20170607162013.705678923@infradead.org \
--to=peterz@infradead.org \
--cc=benh@kernel.crashing.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mgorman@suse.de \
--cc=mingo@kernel.org \
--cc=mpe@ellerman.id.au \
--cc=npiggin@gmail.com \
--cc=oleg@redhat.com \
--cc=paulmck@linux.vnet.ibm.com \
--cc=riel@redhat.com \
--cc=stern@rowland.harvard.edu \
--cc=torvalds@linux-foundation.org \
--cc=will.deacon@arm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome