From: Mel Gorman <mgorman@suse.de>
To: Jiri Slaby <jslaby@suse.cz>
Cc: Linux-Stable <stable@vger.kernel.org>,
LKML <linux-kernel@vger.kernel.org>, Mel Gorman <mgorman@suse.de>
Subject: [PATCH 23/97] swap: add a simple detector for inappropriate swapin readahead
Date: Thu, 28 Aug 2014 19:34:31 +0100 [thread overview]
Message-ID: <1409250945-30874-24-git-send-email-mgorman@suse.de> (raw)
In-Reply-To: <1409250945-30874-1-git-send-email-mgorman@suse.de>
From: Shaohua Li <shli@kernel.org>
commit 579f82901f6f41256642936d7e632f3979ad76d4 upstream.
This is a patch to improve swap readahead algorithm. It's from Hugh and
I slightly changed it.
Hugh's original changelog:
swapin readahead does a blind readahead, whether or not the swapin is
sequential. This may be ok on harddisk, because large reads have
relatively small costs, and if the readahead pages are unneeded they can
be reclaimed easily - though, what if their allocation forced reclaim of
useful pages? But on SSD devices large reads are more expensive than
small ones: if the readahead pages are unneeded, reading them in caused
significant overhead.
This patch adds very simplistic random read detection. Stealing the
PageReadahead technique from Konstantin Khlebnikov's patch, avoiding the
vma/anon_vma sophistications of Shaohua Li's patch, swapin_nr_pages()
simply looks at readahead's current success rate, and narrows or widens
its readahead window accordingly. There is little science to its
heuristic: it's about as stupid as can be whilst remaining effective.
The table below shows elapsed times (in centiseconds) when running a
single repetitive swapping load across a 1000MB mapping in 900MB ram
with 1GB swap (the harddisk tests had taken painfully too long when I
used mem=500M, but SSD shows similar results for that).
Vanilla is the 3.6-rc7 kernel on which I started; Shaohua denotes his
Sep 3 patch in mmotm and linux-next; HughOld denotes my Oct 1 patch
which Shaohua showed to be defective; HughNew this Nov 14 patch, with
page_cluster as usual at default of 3 (8-page reads); HughPC4 this same
patch with page_cluster 4 (16-page reads); HughPC0 with page_cluster 0
(1-page reads: no readahead).
HDD for swapping to harddisk, SSD for swapping to VertexII SSD. Seq for
sequential access to the mapping, cycling five times around; Rand for
the same number of random touches. Anon for a MAP_PRIVATE anon mapping;
Shmem for a MAP_SHARED anon mapping, equivalent to tmpfs.
One weakness of Shaohua's vma/anon_vma approach was that it did not
optimize Shmem: seen below. Konstantin's approach was perhaps mistuned,
50% slower on Seq: did not compete and is not shown below.
HDD Vanilla Shaohua HughOld HughNew HughPC4 HughPC0
Seq Anon 73921 76210 75611 76904 78191 121542
Seq Shmem 73601 73176 73855 72947 74543 118322
Rand Anon 895392 831243 871569 845197 846496 841680
Rand Shmem 1058375 1053486 827935 764955 764376 756489
SSD Vanilla Shaohua HughOld HughNew HughPC4 HughPC0
Seq Anon 24634 24198 24673 25107 21614 70018
Seq Shmem 24959 24932 25052 25703 22030 69678
Rand Anon 43014 26146 28075 25989 26935 25901
Rand Shmem 45349 45215 28249 24268 24138 24332
These tests are, of course, two extremes of a very simple case: under
heavier mixed loads I've not yet observed any consistent improvement or
degradation, and wider testing would be welcome.
Shaohua Li:
Test shows Vanilla is slightly better in sequential workload than Hugh's
patch. I observed with Hugh's patch sometimes the readahead size is
shrinked too fast (from 8 to 1 immediately) in sequential workload if
there is no hit. And in such case, continuing doing readahead is good
actually.
I don't prepare a sophisticated algorithm for the sequential workload
because so far we can't guarantee sequential accessed pages are swap out
sequentially. So I slightly change Hugh's heuristic - don't shrink
readahead size too fast.
Here is my test result (unit second, 3 runs average):
Vanilla Hugh New
Seq 356 370 360
Random 4525 2447 2444
Attached graph is the swapin/swapout throughput I collected with 'vmstat
2'. The first part is running a random workload (till around 1200 of
the x-axis) and the second part is running a sequential workload.
swapin and swapout throughput are almost identical in steady state in
both workloads. These are expected behavior. while in Vanilla, swapin
is much bigger than swapout especially in random workload (because wrong
readahead).
Original patches by: Shaohua Li and Konstantin Khlebnikov.
[fengguang.wu@intel.com: swapin_nr_pages() can be static]
Signed-off-by: Hugh Dickins <hughd@google.com>
Signed-off-by: Shaohua Li <shli@fusionio.com>
Signed-off-by: Fengguang Wu <fengguang.wu@intel.com>
Cc: Rik van Riel <riel@redhat.com>
Cc: Wu Fengguang <fengguang.wu@intel.com>
Cc: Minchan Kim <minchan@kernel.org>
Cc: Konstantin Khlebnikov <khlebnikov@openvz.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
Signed-off-by: Mel Gorman <mgorman@suse.de>
---
include/linux/page-flags.h | 4 +--
mm/swap_state.c | 63 +++++++++++++++++++++++++++++++++++++++++++---
2 files changed, 62 insertions(+), 5 deletions(-)
diff --git a/include/linux/page-flags.h b/include/linux/page-flags.h
index dd7d45b..67fc8a2 100644
--- a/include/linux/page-flags.h
+++ b/include/linux/page-flags.h
@@ -228,9 +228,9 @@ PAGEFLAG(OwnerPriv1, owner_priv_1) TESTCLEARFLAG(OwnerPriv1, owner_priv_1)
TESTPAGEFLAG(Writeback, writeback) TESTSCFLAG(Writeback, writeback)
PAGEFLAG(MappedToDisk, mappedtodisk)
-/* PG_readahead is only used for file reads; PG_reclaim is only for writes */
+/* PG_readahead is only used for reads; PG_reclaim is only for writes */
PAGEFLAG(Reclaim, reclaim) TESTCLEARFLAG(Reclaim, reclaim)
-PAGEFLAG(Readahead, reclaim) /* Reminder to do async read-ahead */
+PAGEFLAG(Readahead, reclaim) TESTCLEARFLAG(Readahead, reclaim)
#ifdef CONFIG_HIGHMEM
/*
diff --git a/mm/swap_state.c b/mm/swap_state.c
index e6f15f8..fdde6f9 100644
--- a/mm/swap_state.c
+++ b/mm/swap_state.c
@@ -63,6 +63,8 @@ unsigned long total_swapcache_pages(void)
return ret;
}
+static atomic_t swapin_readahead_hits = ATOMIC_INIT(4);
+
void show_swap_cache_info(void)
{
printk("%lu pages in swap cache\n", total_swapcache_pages());
@@ -286,8 +288,11 @@ struct page * lookup_swap_cache(swp_entry_t entry)
page = find_get_page(swap_address_space(entry), entry.val);
- if (page)
+ if (page) {
INC_CACHE_INFO(find_success);
+ if (TestClearPageReadahead(page))
+ atomic_inc(&swapin_readahead_hits);
+ }
INC_CACHE_INFO(find_total);
return page;
@@ -389,6 +394,50 @@ struct page *read_swap_cache_async(swp_entry_t entry, gfp_t gfp_mask,
return found_page;
}
+static unsigned long swapin_nr_pages(unsigned long offset)
+{
+ static unsigned long prev_offset;
+ unsigned int pages, max_pages, last_ra;
+ static atomic_t last_readahead_pages;
+
+ max_pages = 1 << ACCESS_ONCE(page_cluster);
+ if (max_pages <= 1)
+ return 1;
+
+ /*
+ * This heuristic has been found to work well on both sequential and
+ * random loads, swapping to hard disk or to SSD: please don't ask
+ * what the "+ 2" means, it just happens to work well, that's all.
+ */
+ pages = atomic_xchg(&swapin_readahead_hits, 0) + 2;
+ if (pages == 2) {
+ /*
+ * We can have no readahead hits to judge by: but must not get
+ * stuck here forever, so check for an adjacent offset instead
+ * (and don't even bother to check whether swap type is same).
+ */
+ if (offset != prev_offset + 1 && offset != prev_offset - 1)
+ pages = 1;
+ prev_offset = offset;
+ } else {
+ unsigned int roundup = 4;
+ while (roundup < pages)
+ roundup <<= 1;
+ pages = roundup;
+ }
+
+ if (pages > max_pages)
+ pages = max_pages;
+
+ /* Don't shrink readahead too fast */
+ last_ra = atomic_read(&last_readahead_pages) / 2;
+ if (pages < last_ra)
+ pages = last_ra;
+ atomic_set(&last_readahead_pages, pages);
+
+ return pages;
+}
+
/**
* swapin_readahead - swap in pages in hope we need them soon
* @entry: swap entry of this memory
@@ -412,11 +461,16 @@ struct page *swapin_readahead(swp_entry_t entry, gfp_t gfp_mask,
struct vm_area_struct *vma, unsigned long addr)
{
struct page *page;
- unsigned long offset = swp_offset(entry);
+ unsigned long entry_offset = swp_offset(entry);
+ unsigned long offset = entry_offset;
unsigned long start_offset, end_offset;
- unsigned long mask = (1UL << page_cluster) - 1;
+ unsigned long mask;
struct blk_plug plug;
+ mask = swapin_nr_pages(offset) - 1;
+ if (!mask)
+ goto skip;
+
/* Read a page_cluster sized and aligned cluster around offset. */
start_offset = offset & ~mask;
end_offset = offset | mask;
@@ -430,10 +484,13 @@ struct page *swapin_readahead(swp_entry_t entry, gfp_t gfp_mask,
gfp_mask, vma, addr);
if (!page)
continue;
+ if (offset != entry_offset)
+ SetPageReadahead(page);
page_cache_release(page);
}
blk_finish_plug(&plug);
lru_add_drain(); /* Push any new pages onto the LRU now */
+skip:
return read_swap_cache_async(entry, gfp_mask, vma, addr);
}
--
1.8.4.5
next prev parent reply other threads:[~2014-08-28 18:36 UTC|newest]
Thread overview: 101+ messages / expand[flat|nested] mbox.gz Atom feed top
2014-08-28 18:34 [PATCH 00/97] Misc series of functional/performance fixes for 3.12-stable Mel Gorman
2014-08-28 18:34 ` [PATCH 01/97] mm: thp: cleanup: mv alloc_hugepage to better place Mel Gorman
2014-08-28 18:34 ` [PATCH 02/97] mm: thp: khugepaged: add policy for finding target node Mel Gorman
2014-08-28 18:34 ` [PATCH 03/97] slab: correct pfmemalloc check Mel Gorman
2014-08-28 18:34 ` [PATCH 04/97] mm: prevent setting of a value less than 0 to min_free_kbytes Mel Gorman
2014-08-28 18:34 ` [PATCH 05/97] mm: fix bad rss-counter if remap_file_pages raced migration Mel Gorman
2014-08-28 18:34 ` [PATCH 06/97] hugetlb: ensure hugepage access is denied if hugepages are not supported Mel Gorman
2014-08-28 18:34 ` [PATCH 07/97] mm: exclude memoryless nodes from zone_reclaim Mel Gorman
2014-08-28 18:34 ` [PATCH 08/97] mm, thp: do not allow thp faults to avoid cpuset restrictions Mel Gorman
2014-09-26 9:53 ` Jiri Slaby
2014-10-13 15:17 ` Mel Gorman
2014-08-28 18:34 ` [PATCH 09/97] swap: change swap_info singly-linked list to list_head Mel Gorman
2014-08-28 18:34 ` [PATCH 10/97] lib/plist: add helper functions Mel Gorman
2014-08-28 18:34 ` [PATCH 11/97] lib/plist: add plist_requeue Mel Gorman
2014-08-28 18:34 ` [PATCH 12/97] swap: change swap_list_head to plist, add swap_avail_head Mel Gorman
2014-08-28 18:34 ` [PATCH 13/97] readahead: fix sequential read cache miss detection Mel Gorman
2014-08-28 18:34 ` [PATCH 14/97] mm: get rid of unnecessary overhead of trace_mm_page_alloc_extfrag() Mel Gorman
2014-08-28 18:34 ` [PATCH 15/97] mm: __rmqueue_fallback() should respect pageblock type Mel Gorman
2014-08-28 18:34 ` [PATCH 16/97] mm, x86: Account for TLB flushes only when debugging Mel Gorman
2014-08-28 18:34 ` [PATCH 17/97] x86/mm: Clean up inconsistencies when flushing TLB ranges Mel Gorman
2014-08-28 18:34 ` [PATCH 18/97] x86/mm: Eliminate redundant page table walk during TLB range flushing Mel Gorman
2014-08-28 18:34 ` [PATCH 19/97] mm: compaction: trace compaction begin and end Mel Gorman
2014-08-28 18:34 ` [PATCH 20/97] mm: compaction: encapsulate defer reset logic Mel Gorman
2014-08-28 18:34 ` [PATCH 21/97] mm: compaction: do not mark unmovable pageblocks as skipped in async compaction Mel Gorman
2014-08-28 18:34 ` [PATCH 22/97] mm: compaction: reset scanner positions immediately when they meet Mel Gorman
2014-08-28 18:34 ` Mel Gorman [this message]
2014-08-28 18:34 ` [PATCH 24/97] mm: vmscan: shrink all slab objects if tight on memory Mel Gorman
2014-08-28 18:34 ` [PATCH 25/97] mm: vmscan: call NUMA-unaware shrinkers irrespective of nodemask Mel Gorman
2014-08-28 18:34 ` [PATCH 26/97] mm: get rid of unnecessary pageblock scanning in setup_zone_migrate_reserve Mel Gorman
2014-08-28 18:34 ` [PATCH 27/97] mm, compaction: avoid isolating pinned pages Mel Gorman
2014-08-28 18:34 ` [PATCH 28/97] mm/compaction: disallow high-order page for migration target Mel Gorman
2014-08-28 18:34 ` [PATCH 29/97] mm/compaction: do not call suitable_migration_target() on every page Mel Gorman
2014-08-28 18:34 ` [PATCH 30/97] mm/compaction: change the timing to check to drop the spinlock Mel Gorman
2014-08-28 18:34 ` [PATCH 31/97] mm/compaction: check pageblock suitability once per pageblock Mel Gorman
2014-08-28 18:34 ` [PATCH 32/97] mm/compaction: clean-up code on success of ballon isolation Mel Gorman
2014-08-28 18:34 ` [PATCH 33/97] mm, compaction: determine isolation mode only once Mel Gorman
2014-08-28 18:34 ` [PATCH 34/97] mm, compaction: ignore pageblock skip when manually invoking compaction Mel Gorman
2014-08-28 18:34 ` [PATCH 35/97] mm/readahead.c: fix readahead failure for memoryless NUMA nodes and limit readahead pages Mel Gorman
2014-08-28 18:34 ` [PATCH 36/97] mm: optimize put_mems_allowed() usage Mel Gorman
2014-08-28 18:34 ` [PATCH 37/97] mm/filemap.c: avoid always dirtying mapping->flags on O_DIRECT Mel Gorman
2014-08-28 18:34 ` [PATCH 38/97] mm: vmscan: respect NUMA policy mask when shrinking slab on direct reclaim Mel Gorman
2014-08-28 18:34 ` [PATCH 39/97] mm: vmscan: shrink_slab: rename max_pass -> freeable Mel Gorman
2014-08-28 18:34 ` [PATCH 40/97] vmscan: reclaim_clean_pages_from_list() must use mod_zone_page_state() Mel Gorman
2014-08-28 18:34 ` [PATCH 41/97] mm: per-thread vma caching Mel Gorman
2014-08-28 18:34 ` [PATCH 42/97] mm: don't pointlessly use BUG_ON() for sanity check Mel Gorman
2014-08-28 18:34 ` [PATCH 43/97] lib: radix-tree: add radix_tree_delete_item() Mel Gorman
2014-08-28 18:34 ` [PATCH 44/97] mm: shmem: save one radix tree lookup when truncating swapped pages Mel Gorman
2014-08-28 18:34 ` [PATCH 45/97] mm: filemap: move radix tree hole searching here Mel Gorman
2014-08-28 18:34 ` [PATCH 46/97] mm + fs: prepare for non-page entries in page cache radix trees Mel Gorman
2014-08-28 18:34 ` [PATCH 47/97] mm: madvise: fix MADV_WILLNEED on shmem swapouts Mel Gorman
2014-08-28 18:34 ` [PATCH 48/97] mm: remove read_cache_page_async() Mel Gorman
2014-08-28 18:34 ` [PATCH 49/97] callers of iov_copy_from_user_atomic() don't need pagecache_disable() Mel Gorman
2014-08-28 18:34 ` [PATCH 50/97] mm/readahead.c: inline ra_submit Mel Gorman
2014-08-28 18:34 ` [PATCH 51/97] mm/compaction: clean up unused code lines Mel Gorman
2014-08-28 18:35 ` [PATCH 52/97] mm/compaction: cleanup isolate_freepages() Mel Gorman
2014-08-28 18:35 ` [PATCH 53/97] mm, migration: add destination page freeing callback Mel Gorman
2014-08-28 18:35 ` [PATCH 54/97] mm, compaction: return failed migration target pages back to freelist Mel Gorman
2014-08-28 18:35 ` [PATCH 55/97] mm, compaction: add per-zone migration pfn cache for async compaction Mel Gorman
2014-08-28 18:35 ` [PATCH 56/97] mm, compaction: embed migration mode in compact_control Mel Gorman
2014-08-28 18:35 ` [PATCH 57/97] mm, compaction: terminate async compaction when rescheduling Mel Gorman
2014-08-28 18:35 ` [PATCH 58/97] mm/compaction: do not count migratepages when unnecessary Mel Gorman
2014-08-28 18:35 ` [PATCH 59/97] mm/compaction: avoid rescanning pageblocks in isolate_freepages Mel Gorman
2014-08-28 18:35 ` [PATCH 60/97] mm, compaction: properly signal and act upon lock and need_sched() contention Mel Gorman
2014-08-28 18:35 ` [PATCH 61/97] x86/mm: In the PTE swapout page reclaim case clear the accessed bit instead of flushing the TLB Mel Gorman
2014-08-28 18:35 ` [PATCH 62/97] mm: fix direct reclaim writeback regression Mel Gorman
2014-08-28 18:35 ` [PATCH 63/97] fs/superblock: unregister sb shrinker before ->kill_sb() Mel Gorman
2014-08-28 18:35 ` [PATCH 64/97] fs/superblock: avoid locking counting inodes and dentries before reclaiming them Mel Gorman
2014-08-28 18:35 ` [PATCH 65/97] mm: vmscan: use proportional scanning during direct reclaim and full scan at DEF_PRIORITY Mel Gorman
2014-08-28 18:35 ` [PATCH 66/97] mm/page_alloc: prevent MIGRATE_RESERVE pages from being misplaced Mel Gorman
2014-08-28 18:35 ` [PATCH 67/97] mm/swap.c: clean up *lru_cache_add* functions Mel Gorman
2014-08-28 18:35 ` [PATCH 68/97] mm: page_alloc: do not update zlc unless the zlc is active Mel Gorman
2014-08-28 18:35 ` [PATCH 69/97] mm: page_alloc: do not treat a zone that cannot be used for dirty pages as "full" Mel Gorman
2014-08-28 18:35 ` [PATCH 70/97] include/linux/jump_label.h: expose the reference count Mel Gorman
2014-08-28 18:35 ` [PATCH 71/97] mm: page_alloc: use jump labels to avoid checking number_of_cpusets Mel Gorman
2014-08-28 18:35 ` [PATCH 72/97] mm: page_alloc: calculate classzone_idx once from the zonelist ref Mel Gorman
2014-08-28 18:35 ` [PATCH 73/97] mm: page_alloc: only check the zone id check if pages are buddies Mel Gorman
2014-08-28 18:35 ` [PATCH 74/97] mm: page_alloc: only check the alloc flags and gfp_mask for dirty once Mel Gorman
2014-08-28 18:35 ` [PATCH 75/97] mm: page_alloc: take the ALLOC_NO_WATERMARK check out of the fast path Mel Gorman
2014-08-28 18:35 ` [PATCH 76/97] mm: page_alloc: use unsigned int for order in more places Mel Gorman
2014-08-28 18:35 ` [PATCH 77/97] mm: page_alloc: reduce number of times page_to_pfn is called Mel Gorman
2014-08-28 18:35 ` [PATCH 78/97] mm: page_alloc: convert hot/cold parameter and immediate callers to bool Mel Gorman
2014-08-28 18:35 ` [PATCH 79/97] mm: page_alloc: lookup pageblock migratetype with IRQs enabled during free Mel Gorman
2014-08-28 18:35 ` [PATCH 80/97] mm: shmem: avoid atomic operation during shmem_getpage_gfp Mel Gorman
2014-08-28 18:35 ` [PATCH 81/97] mm: do not use atomic operations when releasing pages Mel Gorman
2014-08-28 18:35 ` [PATCH 82/97] mm: do not use unnecessary atomic operations when adding pages to the LRU Mel Gorman
2014-08-28 18:35 ` [PATCH 83/97] fs: buffer: do not use unnecessary atomic operations when discarding buffers Mel Gorman
2014-08-28 18:35 ` [PATCH 84/97] mm: non-atomically mark page accessed during page cache allocation where possible Mel Gorman
2014-08-28 18:35 ` [PATCH 85/97] mm: avoid unnecessary atomic operations during end_page_writeback() Mel Gorman
2014-08-28 18:35 ` [PATCH 86/97] shmem: fix init_page_accessed use to stop !PageLRU bug Mel Gorman
2014-08-28 18:35 ` [PATCH 87/97] mm/memory.c: use entry = ACCESS_ONCE(*pte) in handle_pte_fault() Mel Gorman
2014-08-28 18:35 ` [PATCH 88/97] mm, thp: only collapse hugepages to nodes with affinity for zone_reclaim_mode Mel Gorman
2014-08-28 18:35 ` [PATCH 89/97] mm: make copy_pte_range static again Mel Gorman
2014-08-28 18:35 ` [PATCH 90/97] vmalloc: use rcu list iterator to reduce vmap_area_lock contention Mel Gorman
2014-08-28 18:35 ` [PATCH 91/97] memcg, vmscan: Fix forced scan of anonymous pages Mel Gorman
2014-08-28 18:35 ` [PATCH 92/97] mm: pagemap: avoid unnecessary overhead when tracepoints are deactivated Mel Gorman
2014-08-28 18:35 ` [PATCH 93/97] mm: rearrange zone fields into read-only, page alloc, statistics and page reclaim lines Mel Gorman
2014-08-28 18:35 ` [PATCH 94/97] mm: move zone->pages_scanned into a vmstat counter Mel Gorman
2014-08-28 18:35 ` [PATCH 95/97] mm: vmscan: only update per-cpu thresholds for online CPU Mel Gorman
2014-08-28 18:35 ` [PATCH 96/97] mm: page_alloc: abort fair zone allocation policy when remotes nodes are encountered Mel Gorman
2014-08-28 18:35 ` [PATCH 97/97] mm: page_alloc: reduce cost of the fair zone allocation policy Mel Gorman
2015-01-28 1:13 ` [PATCH 00/97] Misc series of functional/performance fixes for 3.12-stable Greg KH
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1409250945-30874-24-git-send-email-mgorman@suse.de \
--to=mgorman@suse.de \
--cc=jslaby@suse.cz \
--cc=linux-kernel@vger.kernel.org \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®