2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-27 23:09:51 +02:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2010-09-27 23:09:51 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2010-09-27 23:09:51 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:02 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-27 23:09:51 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2009-10-13 15:02:11 +01:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2011-05-26 16:00:52 -04:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
include cleanup: Update gfp.h and slab.h includes to prepare for breaking implicit slab.h inclusion from percpu.h
percpu.h is included by sched.h and module.h and thus ends up being
included when building most .c files. percpu.h includes slab.h which
in turn includes gfp.h making everything defined by the two files
universally available and complicating inclusion dependencies.
percpu.h -> slab.h dependency is about to be removed. Prepare for
this change by updating users of gfp and slab facilities include those
headers directly instead of assuming availability. As this conversion
needs to touch large number of source files, the following script is
used as the basis of conversion.
http://userweb.kernel.org/~tj/misc/slabh-sweep.py
The script does the followings.
* Scan files for gfp and slab usages and update includes such that
only the necessary includes are there. ie. if only gfp is used,
gfp.h, if slab is used, slab.h.
* When the script inserts a new include, it looks at the include
blocks and try to put the new include such that its order conforms
to its surrounding. It's put in the include block which contains
core kernel includes, in the same order that the rest are ordered -
alphabetical, Christmas tree, rev-Xmas-tree or at the end if there
doesn't seem to be any matching order.
* If the script can't find a place to put a new include (mostly
because the file doesn't have fitting include block), it prints out
an error message indicating which .h file needs to be added to the
file.
The conversion was done in the following steps.
1. The initial automatic conversion of all .c files updated slightly
over 4000 files, deleting around 700 includes and adding ~480 gfp.h
and ~3000 slab.h inclusions. The script emitted errors for ~400
files.
2. Each error was manually checked. Some didn't need the inclusion,
some needed manual addition while adding it to implementation .h or
embedding .c file was more appropriate for others. This step added
inclusions to around 150 files.
3. The script was run again and the output was compared to the edits
from #2 to make sure no file was left behind.
4. Several build tests were done and a couple of problems were fixed.
e.g. lib/decompress_*.c used malloc/free() wrappers around slab
APIs requiring slab.h to be added manually.
5. The script was run on all .h files but without automatically
editing them as sprinkling gfp.h and slab.h inclusions around .h
files could easily lead to inclusion dependency hell. Most gfp.h
inclusion directives were ignored as stuff from gfp.h was usually
wildly available and often used in preprocessor macros. Each
slab.h inclusion directive was examined and added manually as
necessary.
6. percpu.h was updated not to include slab.h.
7. Build test were done on the following configurations and failures
were fixed. CONFIG_GCOV_KERNEL was turned off for all tests (as my
distributed build env didn't work with gcov compiles) and a few
more options had to be turned off depending on archs to make things
build (like ipr on powerpc/64 which failed due to missing writeq).
* x86 and x86_64 UP and SMP allmodconfig and a custom test config.
* powerpc and powerpc64 SMP allmodconfig
* sparc and sparc64 SMP allmodconfig
* ia64 SMP allmodconfig
* s390 SMP allmodconfig
* alpha SMP allmodconfig
* um on x86_64 SMP allmodconfig
8. percpu.h modifications were reverted so that it could be applied as
a separate patch and serve as bisection point.
Given the fact that I had only a couple of failures from tests on step
6, I'm fairly confident about the coverage of this conversion patch.
If there is a breakage, it's likely to be something in one of the arch
headers which should be easily discoverable easily on most builds of
the specific arch.
Signed-off-by: Tejun Heo <tj@kernel.org>
Guess-its-ok-by: Christoph Lameter <cl@linux-foundation.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Lee Schermerhorn <Lee.Schermerhorn@hp.com>
2010-03-24 17:04:11 +09:00
|
|
|
|
2010-05-31 14:28:19 +08:00
|
|
|
|
2010-05-28 09:29:17 +09:00
|
|
|
|
2010-12-02 14:31:19 -08:00
|
|
|
|
2011-06-15 15:08:48 -07:00
|
|
|
|
2011-07-13 13:14:27 +08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2015-06-24 16:57:36 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:02 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2009-12-21 19:56:42 +01:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-27 23:09:51 +02:00
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-07-31 16:43:02 -07:00
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-02-11 11:52:49 -05:00
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2014-09-19 16:29:31 +08:00
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2009-12-21 19:56:42 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:52 -07:00
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
2014-06-04 16:10:59 -07:00
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
2014-06-04 16:10:59 -07:00
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-12-10 15:43:10 -08:00
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
mm: vmscan: invoke slab shrinkers from shrink_zone()
The slab shrinkers are currently invoked from the zonelist walkers in
kswapd, direct reclaim, and zone reclaim, all of which roughly gauge the
eligible LRU pages and assemble a nodemask to pass to NUMA-aware
shrinkers, which then again have to walk over the nodemask. This is
redundant code, extra runtime work, and fairly inaccurate when it comes to
the estimation of actually scannable LRU pages. The code duplication will
only get worse when making the shrinkers cgroup-aware and requiring them
to have out-of-band cgroup hierarchy walks as well.
Instead, invoke the shrinkers from shrink_zone(), which is where all
reclaimers end up, to avoid this duplication.
Take the count for eligible LRU pages out of get_scan_count(), which
considers many more factors than just the availability of swap space, like
zone_reclaimable_pages() currently does. Accumulate the number over all
visited lruvecs to get the per-zone value.
Some nodes have multiple zones due to memory addressing restrictions. To
avoid putting too much pressure on the shrinkers, only invoke them once
for each such node, using the class zone of the allocation as the pivot
zone.
For now, this integrates the slab shrinking better into the reclaim logic
and gets rid of duplicative invocations from kswapd, direct reclaim, and
zone reclaim. It also prepares for cgroup-awareness, allowing
memcg-capable shrinkers to be added at the lruvec level without much
duplication of both code and runtime work.
This changes kswapd behavior, which used to invoke the shrinkers for each
zone, but with scan ratios gathered from the entire node, resulting in
meaningless pressure quantities on multi-zone nodes.
Zone reclaim behavior also changes. It used to shrink slabs until the
same amount of pages were shrunk as were reclaimed from the LRUs. Now it
merely invokes the shrinkers once with the zone's scan ratio, which makes
the shrinkers go easier on caches that implement aging and would prefer
feeding back pressure from recently used slab objects to unused LRU pages.
[vdavydov@parallels.com: assure class zone is populated]
Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
Cc: Dave Chinner <david@fromorbit.com>
Signed-off-by: Vladimir Davydov <vdavydov@parallels.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2014-12-12 16:56:13 -08:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
2015-02-12 14:58:54 -08:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-27 23:36:05 +02:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-27 23:31:30 +02:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-07-11 10:20:47 -07:00
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-07-11 10:20:47 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
tree-wide: fix assorted typos all over the place
That is "success", "unknown", "through", "performance", "[re|un]mapping"
, "access", "default", "reasonable", "[con]currently", "temperature"
, "channel", "[un]used", "application", "example","hierarchy", "therefore"
, "[over|under]flow", "contiguous", "threshold", "enough" and others.
Signed-off-by: André Goddard Rosa <andre.goddard@gmail.com>
Signed-off-by: Jiri Kosina <jkosina@suse.cz>
2009-11-14 13:09:05 -02:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
2014-06-04 16:11:01 -07:00
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:11:01 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
mm anon rmap: replace same_anon_vma linked list with an interval tree.
When a large VMA (anon or private file mapping) is first touched, which
will populate its anon_vma field, and then split into many regions through
the use of mprotect(), the original anon_vma ends up linking all of the
vmas on a linked list. This can cause rmap to become inefficient, as we
have to walk potentially thousands of irrelevent vmas before finding the
one a given anon page might fall into.
By replacing the same_anon_vma linked list with an interval tree (where
each avc's interval is determined by its vma's start and last pgoffs), we
can make rmap efficient for this use case again.
While the change is large, all of its pieces are fairly simple.
Most places that were walking the same_anon_vma list were looking for a
known pgoff, so they can just use the anon_vma_interval_tree_foreach()
interval tree iterator instead. The exception here is ksm, where the
page's index is not known. It would probably be possible to rework ksm so
that the index would be known, but for now I have decided to keep things
simple and just walk the entirety of the interval tree there.
When updating vma's that already have an anon_vma assigned, we must take
care to re-index the corresponding avc's on their interval tree. This is
done through the use of anon_vma_interval_tree_pre_update_vma() and
anon_vma_interval_tree_post_update_vma(), which remove the avc's from
their interval tree before the update and re-insert them after the update.
The anon_vma stays locked during the update, so there is no chance that
rmap would miss the vmas that are being updated.
Signed-off-by: Michel Lespinasse <walken@google.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Rik van Riel <riel@redhat.com>
Cc: Peter Zijlstra <a.p.zijlstra@chello.nl>
Cc: Daniel Santos <daniel.santos@pobox.com>
Cc: Hugh Dickins <hughd@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2012-10-08 16:31:39 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2012-12-02 19:56:50 +00:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2011-06-27 16:18:09 -07:00
|
|
|
|
|
|
|
|
|
2014-07-23 14:00:01 -07:00
|
|
|
|
2011-06-27 16:18:09 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
mm: change anon_vma linking to fix multi-process server scalability issue
The old anon_vma code can lead to scalability issues with heavily forking
workloads. Specifically, each anon_vma will be shared between the parent
process and all its child processes.
In a workload with 1000 child processes and a VMA with 1000 anonymous
pages per process that get COWed, this leads to a system with a million
anonymous pages in the same anon_vma, each of which is mapped in just one
of the 1000 processes. However, the current rmap code needs to walk them
all, leading to O(N) scanning complexity for each page.
This can result in systems where one CPU is walking the page tables of
1000 processes in page_referenced_one, while all other CPUs are stuck on
the anon_vma lock. This leads to catastrophic failure for a benchmark
like AIM7, where the total number of processes can reach in the tens of
thousands. Real workloads are still a factor 10 less process intensive
than AIM7, but they are catching up.
This patch changes the way anon_vmas and VMAs are linked, which allows us
to associate multiple anon_vmas with a VMA. At fork time, each child
process gets its own anon_vmas, in which its COWed pages will be
instantiated. The parents' anon_vma is also linked to the VMA, because
non-COWed pages could be present in any of the children.
This reduces rmap scanning complexity to O(1) for the pages of the 1000
child processes, with O(N) complexity for at most 1/N pages in the system.
This reduces the average scanning cost in heavily forking workloads from
O(N) to 2.
The only real complexity in this patch stems from the fact that linking a
VMA to anon_vmas now involves memory allocations. This means vma_adjust
can fail, if it needs to attach a VMA to anon_vma structures. This in
turn means error handling needs to be added to the calling functions.
A second source of complexity is that, because there can be multiple
anon_vmas, the anon_vma linking in vma_adjust can no longer be done under
"the" anon_vma lock. To prevent the rmap code from walking up an
incomplete VMA, this patch introduces the VM_LOCK_RMAP VMA flag. This bit
flag uses the same slot as the NOMMU VM_MAPPED_COPY, with an ifdef in mm.h
to make sure it is impossible to compile a kernel that needs both symbolic
values for the same bitflag.
Some test results:
Without the anon_vma changes, when AIM7 hits around 9.7k users (on a test
box with 16GB RAM and not quite enough IO), the system ends up running
>99% in system time, with every CPU on the same anon_vma lock in the
pageout code.
With these changes, AIM7 hits the cross-over point around 29.7k users.
This happens with ~99% IO wait time, there never seems to be any spike in
system time. The anon_vma lock contention appears to be resolved.
[akpm@linux-foundation.org: cleanups]
Signed-off-by: Rik van Riel <riel@redhat.com>
Cc: KOSAKI Motohiro <kosaki.motohiro@jp.fujitsu.com>
Cc: Larry Woodman <lwoodman@redhat.com>
Cc: Lee Schermerhorn <Lee.Schermerhorn@hp.com>
Cc: Minchan Kim <minchan.kim@gmail.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Hugh Dickins <hugh.dickins@tiscali.co.uk>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2010-03-05 13:42:07 -08:00
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
mm: change anon_vma linking to fix multi-process server scalability issue
The old anon_vma code can lead to scalability issues with heavily forking
workloads. Specifically, each anon_vma will be shared between the parent
process and all its child processes.
In a workload with 1000 child processes and a VMA with 1000 anonymous
pages per process that get COWed, this leads to a system with a million
anonymous pages in the same anon_vma, each of which is mapped in just one
of the 1000 processes. However, the current rmap code needs to walk them
all, leading to O(N) scanning complexity for each page.
This can result in systems where one CPU is walking the page tables of
1000 processes in page_referenced_one, while all other CPUs are stuck on
the anon_vma lock. This leads to catastrophic failure for a benchmark
like AIM7, where the total number of processes can reach in the tens of
thousands. Real workloads are still a factor 10 less process intensive
than AIM7, but they are catching up.
This patch changes the way anon_vmas and VMAs are linked, which allows us
to associate multiple anon_vmas with a VMA. At fork time, each child
process gets its own anon_vmas, in which its COWed pages will be
instantiated. The parents' anon_vma is also linked to the VMA, because
non-COWed pages could be present in any of the children.
This reduces rmap scanning complexity to O(1) for the pages of the 1000
child processes, with O(N) complexity for at most 1/N pages in the system.
This reduces the average scanning cost in heavily forking workloads from
O(N) to 2.
The only real complexity in this patch stems from the fact that linking a
VMA to anon_vmas now involves memory allocations. This means vma_adjust
can fail, if it needs to attach a VMA to anon_vma structures. This in
turn means error handling needs to be added to the calling functions.
A second source of complexity is that, because there can be multiple
anon_vmas, the anon_vma linking in vma_adjust can no longer be done under
"the" anon_vma lock. To prevent the rmap code from walking up an
incomplete VMA, this patch introduces the VM_LOCK_RMAP VMA flag. This bit
flag uses the same slot as the NOMMU VM_MAPPED_COPY, with an ifdef in mm.h
to make sure it is impossible to compile a kernel that needs both symbolic
values for the same bitflag.
Some test results:
Without the anon_vma changes, when AIM7 hits around 9.7k users (on a test
box with 16GB RAM and not quite enough IO), the system ends up running
>99% in system time, with every CPU on the same anon_vma lock in the
pageout code.
With these changes, AIM7 hits the cross-over point around 29.7k users.
This happens with ~99% IO wait time, there never seems to be any spike in
system time. The anon_vma lock contention appears to be resolved.
[akpm@linux-foundation.org: cleanups]
Signed-off-by: Rik van Riel <riel@redhat.com>
Cc: KOSAKI Motohiro <kosaki.motohiro@jp.fujitsu.com>
Cc: Larry Woodman <lwoodman@redhat.com>
Cc: Lee Schermerhorn <Lee.Schermerhorn@hp.com>
Cc: Minchan Kim <minchan.kim@gmail.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Hugh Dickins <hugh.dickins@tiscali.co.uk>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2010-03-05 13:42:07 -08:00
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
mm anon rmap: replace same_anon_vma linked list with an interval tree.
When a large VMA (anon or private file mapping) is first touched, which
will populate its anon_vma field, and then split into many regions through
the use of mprotect(), the original anon_vma ends up linking all of the
vmas on a linked list. This can cause rmap to become inefficient, as we
have to walk potentially thousands of irrelevent vmas before finding the
one a given anon page might fall into.
By replacing the same_anon_vma linked list with an interval tree (where
each avc's interval is determined by its vma's start and last pgoffs), we
can make rmap efficient for this use case again.
While the change is large, all of its pieces are fairly simple.
Most places that were walking the same_anon_vma list were looking for a
known pgoff, so they can just use the anon_vma_interval_tree_foreach()
interval tree iterator instead. The exception here is ksm, where the
page's index is not known. It would probably be possible to rework ksm so
that the index would be known, but for now I have decided to keep things
simple and just walk the entirety of the interval tree there.
When updating vma's that already have an anon_vma assigned, we must take
care to re-index the corresponding avc's on their interval tree. This is
done through the use of anon_vma_interval_tree_pre_update_vma() and
anon_vma_interval_tree_post_update_vma(), which remove the avc's from
their interval tree before the update and re-insert them after the update.
The anon_vma stays locked during the update, so there is no chance that
rmap would miss the vmas that are being updated.
Signed-off-by: Michel Lespinasse <walken@google.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Rik van Riel <riel@redhat.com>
Cc: Peter Zijlstra <a.p.zijlstra@chello.nl>
Cc: Daniel Santos <daniel.santos@pobox.com>
Cc: Hugh Dickins <hughd@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2012-10-08 16:31:39 -07:00
|
|
|
|
|
|
|
|
|
mm: change anon_vma linking to fix multi-process server scalability issue
The old anon_vma code can lead to scalability issues with heavily forking
workloads. Specifically, each anon_vma will be shared between the parent
process and all its child processes.
In a workload with 1000 child processes and a VMA with 1000 anonymous
pages per process that get COWed, this leads to a system with a million
anonymous pages in the same anon_vma, each of which is mapped in just one
of the 1000 processes. However, the current rmap code needs to walk them
all, leading to O(N) scanning complexity for each page.
This can result in systems where one CPU is walking the page tables of
1000 processes in page_referenced_one, while all other CPUs are stuck on
the anon_vma lock. This leads to catastrophic failure for a benchmark
like AIM7, where the total number of processes can reach in the tens of
thousands. Real workloads are still a factor 10 less process intensive
than AIM7, but they are catching up.
This patch changes the way anon_vmas and VMAs are linked, which allows us
to associate multiple anon_vmas with a VMA. At fork time, each child
process gets its own anon_vmas, in which its COWed pages will be
instantiated. The parents' anon_vma is also linked to the VMA, because
non-COWed pages could be present in any of the children.
This reduces rmap scanning complexity to O(1) for the pages of the 1000
child processes, with O(N) complexity for at most 1/N pages in the system.
This reduces the average scanning cost in heavily forking workloads from
O(N) to 2.
The only real complexity in this patch stems from the fact that linking a
VMA to anon_vmas now involves memory allocations. This means vma_adjust
can fail, if it needs to attach a VMA to anon_vma structures. This in
turn means error handling needs to be added to the calling functions.
A second source of complexity is that, because there can be multiple
anon_vmas, the anon_vma linking in vma_adjust can no longer be done under
"the" anon_vma lock. To prevent the rmap code from walking up an
incomplete VMA, this patch introduces the VM_LOCK_RMAP VMA flag. This bit
flag uses the same slot as the NOMMU VM_MAPPED_COPY, with an ifdef in mm.h
to make sure it is impossible to compile a kernel that needs both symbolic
values for the same bitflag.
Some test results:
Without the anon_vma changes, when AIM7 hits around 9.7k users (on a test
box with 16GB RAM and not quite enough IO), the system ends up running
>99% in system time, with every CPU on the same anon_vma lock in the
pageout code.
With these changes, AIM7 hits the cross-over point around 29.7k users.
This happens with ~99% IO wait time, there never seems to be any spike in
system time. The anon_vma lock contention appears to be resolved.
[akpm@linux-foundation.org: cleanups]
Signed-off-by: Rik van Riel <riel@redhat.com>
Cc: KOSAKI Motohiro <kosaki.motohiro@jp.fujitsu.com>
Cc: Larry Woodman <lwoodman@redhat.com>
Cc: Lee Schermerhorn <Lee.Schermerhorn@hp.com>
Cc: Minchan Kim <minchan.kim@gmail.com>
Cc: Andrea Arcangeli <aarcange@redhat.com>
Cc: Hugh Dickins <hugh.dickins@tiscali.co.uk>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2010-03-05 13:42:07 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-12-02 19:56:50 +00:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:11:01 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-12-12 16:54:36 -08:00
|
|
|
|
2011-06-27 16:18:09 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2014-07-23 14:00:01 -07:00
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2012-10-08 16:31:25 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:11:02 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-12-12 16:54:36 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:11:01 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:11:01 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2014-06-04 16:11:01 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-04-15 16:13:05 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-04-15 16:13:05 -07:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-27 23:31:30 +02:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-01-21 15:49:08 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-30 22:57:33 -03:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-29 13:16:20 +08:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-28 09:29:20 +09:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-08 10:19:36 +09:00
|
|
|
|
2010-05-28 09:29:20 +09:00
|
|
|
|
2015-06-24 16:56:53 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-28 09:29:20 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-08 10:19:36 +09:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2010-05-28 09:29:20 +09:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-30 22:57:33 -03:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
2013-02-22 16:35:53 -08:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-12-11 16:01:32 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:33 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2015-06-24 16:57:36 -07:00
|
|
|
|
|
|
|
|
|
2015-04-15 16:13:05 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2009-10-19 08:15:01 +02:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2009-10-19 08:15:01 +02:00
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2015-04-15 16:13:05 -07:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2015-04-15 16:13:05 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:48 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-08-06 15:47:04 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:48 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2014-01-23 15:53:14 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-07-11 10:20:47 -07:00
|
|
|
|
2014-01-23 15:53:14 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2014-07-30 16:08:28 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-28 09:29:17 +09:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
2014-07-30 16:08:30 -07:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2014-07-30 16:08:30 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2010-05-28 09:29:17 +09:00
|
|
|
|
2012-07-11 10:20:47 -07:00
|
|
|
|
2010-05-28 09:29:17 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2011-02-01 15:52:40 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-07-11 10:20:47 -07:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2012-07-11 10:20:47 -07:00
|
|
|
|
2011-12-13 09:27:58 -08:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2010-05-28 09:29:18 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:52 -07:00
|
|
|
|
2010-05-28 09:29:18 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:52 -07:00
|
|
|
|
2010-05-28 09:29:18 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-12-15 10:48:12 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-28 09:29:17 +09:00
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2010-05-28 09:29:19 +09:00
|
|
|
|
2013-02-22 16:35:51 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:57 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:05 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:02 -08:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-08 10:19:38 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:48 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2010-09-08 10:19:38 +09:00
|
|
|
|
|
|
|
|
|
2014-05-22 11:54:15 -07:00
|
|
|
|
2010-09-08 10:19:38 +09:00
|
|
|
|
2011-03-10 08:52:07 +01:00
|
|
|
|
2014-05-22 11:54:15 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-08 10:19:38 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
2010-09-08 10:19:38 +09:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2015-08-14 15:35:08 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:48 -07:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2015-06-24 16:56:48 -07:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-29 13:16:20 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-05-05 16:23:35 -07:00
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-02-01 15:52:41 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-09-30 13:45:23 -07:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2013-09-30 13:45:23 -07:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
|
|
|
|
|
2011-02-01 15:52:41 -08:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
2009-09-29 13:16:20 +08:00
|
|
|
|
|
|
|
|
|
2011-03-10 08:52:07 +01:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2014-08-06 16:06:49 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:45 -07:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2014-08-06 16:06:49 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:35:51 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2014-05-22 11:54:21 -07:00
|
|
|
|
2015-08-06 15:46:58 -07:00
|
|
|
|
2014-05-22 11:54:21 -07:00
|
|
|
|
2015-08-06 15:46:58 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:02 -08:00
|
|
|
|
2010-05-28 09:29:17 +09:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:59 +01:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
hwpoison: fix the handling path of the victimized page frame that belong to non-LRU
Until now, the kernel has the same policy to handle victimized page
frames that belong to kernel-space(reserved/slab-subsystem) or
non-LRU(unknown page state). In other word, the result of handling
either of these victimized page frames is (IGNORED | FAILED), and the
return value of memory_failure() is -EBUSY.
This patch is to avoid that memory_failure() returns very soon due to
the "true" value of (!PageLRU(p)), and it also ensures that
action_result() can report more precise information("reserved kernel",
"kernel slab", and "unknown page state") instead of "non LRU",
especially for memory errors which are detected by memory-scrubbing.
Andi said:
: While running the mcelog test suite on 3.14 I hit the following VM_BUG_ON:
:
: soft_offline: 0x56d4: unknown non LRU page type 3ffff800008000
: page:ffffea000015b400 count:3 mapcount:2097169 mapping: (null) index:0xffff8800056d7000
: page flags: 0x3ffff800004081(locked|slab|head)
: ------------[ cut here ]------------
: kernel BUG at mm/rmap.c:1495!
:
: I think what happened is that a LRU page turned into a slab page in
: parallel with offlining. memory_failure initially tests for this case,
: but doesn't retest later after the page has been locked.
:
: ...
:
: I ran this patch in a loop over night with some stress plus
: the mcelog test suite running in a loop. I cannot guarantee it hit it,
: but it should have given it a good beating.
:
: The kernel survived with no messages, although the mcelog test suite
: got killed at some point because it couldn't fork anymore. Probably
: some unrelated problem.
:
: So the patch is ok for me for .16.
Signed-off-by: Chen Yucong <slaoub@gmail.com>
Acked-by: Naoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Reported-by: Andi Kleen <andi@firstfloor.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2014-07-02 15:22:37 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-28 09:29:18 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-02-01 15:52:40 -08:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2010-05-28 09:29:18 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:10:35 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-22 16:32:44 -07:00
|
|
|
|
2014-01-23 15:53:14 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2014-01-23 15:53:14 -08:00
|
|
|
|
|
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2015-06-24 16:57:30 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
hwpoison: fix the handling path of the victimized page frame that belong to non-LRU
Until now, the kernel has the same policy to handle victimized page
frames that belong to kernel-space(reserved/slab-subsystem) or
non-LRU(unknown page state). In other word, the result of handling
either of these victimized page frames is (IGNORED | FAILED), and the
return value of memory_failure() is -EBUSY.
This patch is to avoid that memory_failure() returns very soon due to
the "true" value of (!PageLRU(p)), and it also ensures that
action_result() can report more precise information("reserved kernel",
"kernel slab", and "unknown page state") instead of "non LRU",
especially for memory errors which are detected by memory-scrubbing.
Andi said:
: While running the mcelog test suite on 3.14 I hit the following VM_BUG_ON:
:
: soft_offline: 0x56d4: unknown non LRU page type 3ffff800008000
: page:ffffea000015b400 count:3 mapcount:2097169 mapping: (null) index:0xffff8800056d7000
: page flags: 0x3ffff800004081(locked|slab|head)
: ------------[ cut here ]------------
: kernel BUG at mm/rmap.c:1495!
:
: I think what happened is that a LRU page turned into a slab page in
: parallel with offlining. memory_failure initially tests for this case,
: but doesn't retest later after the page has been locked.
:
: ...
:
: I ran this patch in a loop over night with some stress plus
: the mcelog test suite running in a loop. I cannot guarantee it hit it,
: but it should have given it a good beating.
:
: The kernel survived with no messages, although the mcelog test suite
: got killed at some point because it couldn't fork anymore. Probably
: some unrelated problem.
:
: So the patch is ok for me for .16.
Signed-off-by: Chen Yucong <slaoub@gmail.com>
Acked-by: Naoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Reported-by: Andi Kleen <andi@firstfloor.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2014-07-02 15:22:37 -07:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2013-02-22 16:35:51 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
mm/hwpoison: fix loss of PG_dirty for errors on mlocked pages
memory_failure() store the page flag of the error page before doing unmap,
and (only) if the first check with page flags at the time decided the
error page is unknown, it do the second check with the stored page flag
since memory_failure() does unmapping of the error pages before doing
page_action(). This unmapping changes the page state, especially
page_remove_rmap() (called from try_to_unmap_one()) clears PG_mlocked, so
page_action() can't catch mlocked pages after that.
However, memory_failure() can't handle memory errors on dirty mlocked
pages correctly. try_to_unmap_one will move the dirty bit from pte to the
physical page, the second check lose it since it check the stored page
flag. This patch fix it by restore PG_dirty flag to stored page flag if
the page is dirty.
Testcase:
#define _GNU_SOURCE
#include <stdlib.h>
#include <stdio.h>
#include <sys/mman.h>
#include <sys/types.h>
#include <errno.h>
#define PAGES_TO_TEST 2
#define PAGE_SIZE 4096
int main(void)
{
char *mem;
int i;
mem = mmap(NULL, PAGES_TO_TEST * PAGE_SIZE,
PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_LOCKED, 0, 0);
for (i = 0; i < PAGES_TO_TEST; i++)
mem[i * PAGE_SIZE] = 'a';
if (madvise(mem, PAGES_TO_TEST * PAGE_SIZE, MADV_HWPOISON) == -1)
return -1;
return 0;
}
Before patch:
[ 912.839247] Injecting memory failure for page 7dfb8 at 7f6b4e37b000
[ 912.839257] MCE 0x7dfb8: clean mlocked LRU page recovery: Recovered
[ 912.845550] MCE 0x7dfb8: clean mlocked LRU page still referenced by 1 users
[ 912.852586] Injecting memory failure for page 7e6aa at 7f6b4e37c000
[ 912.852594] MCE 0x7e6aa: clean mlocked LRU page recovery: Recovered
[ 912.858936] MCE 0x7e6aa: clean mlocked LRU page still referenced by 1 users
After patch:
[ 163.590225] Injecting memory failure for page 91bc2f at 7f9f5b0e5000
[ 163.590264] MCE 0x91bc2f: dirty mlocked LRU page recovery: Recovered
[ 163.596680] MCE 0x91bc2f: dirty mlocked LRU page still referenced by 1 users
[ 163.603831] Injecting memory failure for page 91cdd3 at 7f9f5b0e6000
[ 163.603852] MCE 0x91cdd3: dirty mlocked LRU page recovery: Recovered
[ 163.610305] MCE 0x91cdd3: dirty mlocked LRU page still referenced by 1 users
Signed-off-by: Wanpeng Li <liwanp@linux.vnet.ibm.com>
Reviewed-by: Naoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Cc: Andi Kleen <andi@firstfloor.org>
Cc: Tony Luck <tony.luck@intel.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2013-09-11 14:22:50 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:35:51 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
2010-05-28 09:29:17 +09:00
|
|
|
|
2009-09-16 11:50:15 +02:00
|
|
|
|
|
|
|
|
|
2011-12-15 10:48:12 -08:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2011-07-13 13:14:27 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-11-14 14:32:17 -08:00
|
|
|
|
2011-07-13 13:14:27 +08:00
|
|
|
|
|
|
|
|
|
2013-07-25 11:53:25 -07:00
|
|
|
|
2011-07-13 13:14:27 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:07:56 -07:00
|
|
|
|
2011-07-13 13:14:27 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-07-10 14:57:01 +05:30
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-07-13 13:14:27 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-28 09:29:19 +09:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-27 23:31:30 +02:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:53 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-09-30 13:45:22 -07:00
|
|
|
|
2013-09-11 14:22:53 -07:00
|
|
|
|
2015-06-24 16:56:48 -07:00
|
|
|
|
2013-09-11 14:22:53 -07:00
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:52 -07:00
|
|
|
|
2010-05-28 09:29:19 +09:00
|
|
|
|
2015-06-24 16:56:48 -07:00
|
|
|
|
2010-09-08 10:19:38 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-10-31 17:09:04 -07:00
|
|
|
|
2010-09-08 10:19:38 +09:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2013-09-11 14:22:54 -07:00
|
|
|
|
2010-09-27 23:31:30 +02:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-10 08:52:07 +01:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-28 09:29:19 +09:00
|
|
|
|
2010-09-27 23:31:30 +02:00
|
|
|
|
2013-02-22 16:34:02 -08:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
2010-09-08 10:19:40 +09:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
mm/memory-failure.c: fix bug triggered by unpoisoning empty zero page
Injecting memory failure for page 0x19d0 at 0xb77d2000
MCE 0x19d0: non LRU page recovery: Ignored
MCE: Software-unpoisoned page 0x19d0
BUG: Bad page state in process bash pfn:019d0
page:f3461a00 count:0 mapcount:0 mapping: (null) index:0x0
page flags: 0x40000404(referenced|reserved)
Modules linked in: nfsd auth_rpcgss i915 nfs_acl nfs lockd video drm_kms_helper drm bnep rfcomm sunrpc bluetooth psmouse parport_pc ppdev lp serio_raw fscache parport gpio_ich lpc_ich mac_hid i2c_algo_bit tpm_tis wmi usb_storage hid_generic usbhid hid e1000e firewire_ohci firewire_core ahci ptp libahci pps_core crc_itu_t
CPU: 3 PID: 2123 Comm: bash Not tainted 3.11.0-rc6+ #12
Hardware name: LENOVO 7034DD7/ , BIOS 9HKT47AUS 01//2012
00000000 00000000 e9625ea0 c15ec49b f3461a00 e9625eb8 c15ea119 c17cbf18
ef084314 000019d0 f3461a00 e9625ed8 c110dc8a f3461a00 00000001 00000000
f3461a00 40000404 00000000 e9625ef8 c110dcc1 f3461a00 f3461a00 000019d0
Call Trace:
dump_stack+0x41/0x52
bad_page+0xcf/0xeb
free_pages_prepare+0x12a/0x140
free_hot_cold_page+0x21/0x110
__put_single_page+0x21/0x30
put_page+0x25/0x40
unpoison_memory+0x107/0x200
hwpoison_unpoison+0x20/0x30
simple_attr_write+0xb6/0xd0
vfs_write+0xa0/0x1b0
SyS_write+0x4f/0x90
sysenter_do_call+0x12/0x22
Disabling lock debugging due to kernel taint
Testcase:
#define _GNU_SOURCE
#include <stdlib.h>
#include <stdio.h>
#include <sys/mman.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/types.h>
#include <errno.h>
#define PAGES_TO_TEST 1
#define PAGE_SIZE 4096
int main(void)
{
char *mem;
mem = mmap(NULL, PAGES_TO_TEST * PAGE_SIZE,
PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS, 0, 0);
if (madvise(mem, PAGES_TO_TEST * PAGE_SIZE, MADV_HWPOISON) == -1)
return -1;
munmap(mem, PAGES_TO_TEST * PAGE_SIZE);
return 0;
}
There is one page reference count for default empty zero page,
madvise_hwpoison add another one by get_user_pages_fast. memory_hwpoison
reduce one page reference count since it's a non LRU page.
unpoison_memory release the last page reference count and free empty zero
page to buddy system which is not correct since empty zero page has
PG_reserved flag. This patch fix it by don't reduce the page reference
count under 1 against empty zero page.
Signed-off-by: Wanpeng Li <liwanp@linux.vnet.ibm.com>
Reviewed-by: Naoya Horiguchi <n-horiguchi@ah.jp.nec.com>
Cc: Andi Kleen <andi@firstfloor.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2013-09-11 14:23:01 -07:00
|
|
|
|
2009-12-16 12:19:58 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:01 +01:00
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-06-24 16:56:48 -07:00
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
2012-05-29 15:06:16 -07:00
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
2012-05-29 15:06:16 -07:00
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
2012-05-29 15:06:16 -07:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-08-14 15:34:56 -07:00
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:01 -07:00
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:33:59 -08:00
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
|
|
|
|
|
2013-02-22 16:33:59 -08:00
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
2013-02-22 16:33:59 -08:00
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
2015-04-15 16:14:38 -07:00
|
|
|
|
2015-08-14 15:34:59 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-04-15 16:14:38 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-06-04 16:08:25 -07:00
|
|
|
|
2013-09-11 14:22:01 -07:00
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
2011-10-31 17:09:04 -07:00
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:01 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
2013-12-18 17:08:54 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-08 10:19:39 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
2013-02-22 16:33:59 -08:00
|
|
|
|
|
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-05-24 17:12:20 -07:00
|
|
|
|
2010-09-27 23:31:30 +02:00
|
|
|
|
2013-02-22 16:34:03 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-05-24 17:12:20 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
2011-06-15 15:08:48 -07:00
|
|
|
|
2013-02-22 16:35:14 -08:00
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
2015-08-06 15:47:11 -07:00
|
|
|
|
|
|
|
|
|
2014-06-04 16:08:25 -07:00
|
|
|
|
2013-02-22 16:35:14 -08:00
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
2014-01-21 15:51:17 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-27 23:31:30 +02:00
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-08-06 15:47:11 -07:00
|
|
|
|
|
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
2010-09-27 23:31:30 +02:00
|
|
|
|
2011-10-31 17:09:04 -07:00
|
|
|
|
2009-12-16 12:20:00 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:56 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-03-03 15:38:18 -08:00
|
|
|
|
2013-09-11 14:22:56 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
mem-hotplug: implement get/put_online_mems
kmem_cache_{create,destroy,shrink} need to get a stable value of
cpu/node online mask, because they init/destroy/access per-cpu/node
kmem_cache parts, which can be allocated or destroyed on cpu/mem
hotplug. To protect against cpu hotplug, these functions use
{get,put}_online_cpus. However, they do nothing to synchronize with
memory hotplug - taking the slab_mutex does not eliminate the
possibility of race as described in patch 2.
What we need there is something like get_online_cpus, but for memory.
We already have lock_memory_hotplug, which serves for the purpose, but
it's a bit of a hammer right now, because it's backed by a mutex. As a
result, it imposes some limitations to locking order, which are not
desirable, and can't be used just like get_online_cpus. That's why in
patch 1 I substitute it with get/put_online_mems, which work exactly
like get/put_online_cpus except they block not cpu, but memory hotplug.
[ v1 can be found at https://lkml.org/lkml/2014/4/6/68. I NAK'ed it by
myself, because it used an rw semaphore for get/put_online_mems,
making them dead lock prune. ]
This patch (of 2):
{un}lock_memory_hotplug, which is used to synchronize against memory
hotplug, is currently backed by a mutex, which makes it a bit of a
hammer - threads that only want to get a stable value of online nodes
mask won't be able to proceed concurrently. Also, it imposes some
strong locking ordering rules on it, which narrows down the set of its
usage scenarios.
This patch introduces get/put_online_mems, which are the same as
get/put_online_cpus, but for memory hotplug, i.e. executing a code
inside a get/put_online_mems section will guarantee a stable value of
online nodes, present pages, etc.
lock_memory_hotplug()/unlock_memory_hotplug() are removed altogether.
Signed-off-by: Vladimir Davydov <vdavydov@parallels.com>
Cc: Christoph Lameter <cl@linux.com>
Cc: Pekka Enberg <penberg@kernel.org>
Cc: Tang Chen <tangchen@cn.fujitsu.com>
Cc: Zhang Yanfei <zhangyanfei@cn.fujitsu.com>
Cc: Toshi Kani <toshi.kani@hp.com>
Cc: Xishi Qiu <qiuxishi@huawei.com>
Cc: Jiang Liu <liuj97@gmail.com>
Cc: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
Cc: David Rientjes <rientjes@google.com>
Cc: Wen Congyang <wency@cn.fujitsu.com>
Cc: Yasuaki Ishimatsu <isimatu.yasuaki@jp.fujitsu.com>
Cc: Lai Jiangshan <laijs@cn.fujitsu.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2014-06-04 16:07:18 -07:00
|
|
|
|
2013-11-12 15:07:26 -08:00
|
|
|
|
2013-09-11 14:22:56 -07:00
|
|
|
|
mem-hotplug: implement get/put_online_mems
kmem_cache_{create,destroy,shrink} need to get a stable value of
cpu/node online mask, because they init/destroy/access per-cpu/node
kmem_cache parts, which can be allocated or destroyed on cpu/mem
hotplug. To protect against cpu hotplug, these functions use
{get,put}_online_cpus. However, they do nothing to synchronize with
memory hotplug - taking the slab_mutex does not eliminate the
possibility of race as described in patch 2.
What we need there is something like get_online_cpus, but for memory.
We already have lock_memory_hotplug, which serves for the purpose, but
it's a bit of a hammer right now, because it's backed by a mutex. As a
result, it imposes some limitations to locking order, which are not
desirable, and can't be used just like get_online_cpus. That's why in
patch 1 I substitute it with get/put_online_mems, which work exactly
like get/put_online_cpus except they block not cpu, but memory hotplug.
[ v1 can be found at https://lkml.org/lkml/2014/4/6/68. I NAK'ed it by
myself, because it used an rw semaphore for get/put_online_mems,
making them dead lock prune. ]
This patch (of 2):
{un}lock_memory_hotplug, which is used to synchronize against memory
hotplug, is currently backed by a mutex, which makes it a bit of a
hammer - threads that only want to get a stable value of online nodes
mask won't be able to proceed concurrently. Also, it imposes some
strong locking ordering rules on it, which narrows down the set of its
usage scenarios.
This patch introduces get/put_online_mems, which are the same as
get/put_online_cpus, but for memory hotplug, i.e. executing a code
inside a get/put_online_mems section will guarantee a stable value of
online nodes, present pages, etc.
lock_memory_hotplug()/unlock_memory_hotplug() are removed altogether.
Signed-off-by: Vladimir Davydov <vdavydov@parallels.com>
Cc: Christoph Lameter <cl@linux.com>
Cc: Pekka Enberg <penberg@kernel.org>
Cc: Tang Chen <tangchen@cn.fujitsu.com>
Cc: Zhang Yanfei <zhangyanfei@cn.fujitsu.com>
Cc: Toshi Kani <toshi.kani@hp.com>
Cc: Xishi Qiu <qiuxishi@huawei.com>
Cc: Jiang Liu <liuj97@gmail.com>
Cc: Rafael J. Wysocki <rafael.j.wysocki@intel.com>
Cc: David Rientjes <rientjes@google.com>
Cc: Wen Congyang <wency@cn.fujitsu.com>
Cc: Yasuaki Ishimatsu <isimatu.yasuaki@jp.fujitsu.com>
Cc: Lai Jiangshan <laijs@cn.fujitsu.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2014-06-04 16:07:18 -07:00
|
|
|
|
2013-11-12 15:07:26 -08:00
|
|
|
|
2013-09-11 14:22:56 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-11-12 15:07:26 -08:00
|
|
|
|
2013-09-11 14:22:56 -07:00
|
|
|
|
|
|
|
|
|
2015-05-05 16:23:46 -07:00
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:56 -07:00
|
|
|
|
|
|
|
|
|
2015-05-05 16:23:46 -07:00
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:56 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|