2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2008-10-15 22:01:59 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-11-16 23:57:37 -05:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
include cleanup: Update gfp.h and slab.h includes to prepare for breaking implicit slab.h inclusion from percpu.h
percpu.h is included by sched.h and module.h and thus ends up being
included when building most .c files. percpu.h includes slab.h which
in turn includes gfp.h making everything defined by the two files
universally available and complicating inclusion dependencies.
percpu.h -> slab.h dependency is about to be removed. Prepare for
this change by updating users of gfp and slab facilities include those
headers directly instead of assuming availability. As this conversion
needs to touch large number of source files, the following script is
used as the basis of conversion.
http://userweb.kernel.org/~tj/misc/slabh-sweep.py
The script does the followings.
* Scan files for gfp and slab usages and update includes such that
only the necessary includes are there. ie. if only gfp is used,
gfp.h, if slab is used, slab.h.
* When the script inserts a new include, it looks at the include
blocks and try to put the new include such that its order conforms
to its surrounding. It's put in the include block which contains
core kernel includes, in the same order that the rest are ordered -
alphabetical, Christmas tree, rev-Xmas-tree or at the end if there
doesn't seem to be any matching order.
* If the script can't find a place to put a new include (mostly
because the file doesn't have fitting include block), it prints out
an error message indicating which .h file needs to be added to the
file.
The conversion was done in the following steps.
1. The initial automatic conversion of all .c files updated slightly
over 4000 files, deleting around 700 includes and adding ~480 gfp.h
and ~3000 slab.h inclusions. The script emitted errors for ~400
files.
2. Each error was manually checked. Some didn't need the inclusion,
some needed manual addition while adding it to implementation .h or
embedding .c file was more appropriate for others. This step added
inclusions to around 150 files.
3. The script was run again and the output was compared to the edits
from #2 to make sure no file was left behind.
4. Several build tests were done and a couple of problems were fixed.
e.g. lib/decompress_*.c used malloc/free() wrappers around slab
APIs requiring slab.h to be added manually.
5. The script was run on all .h files but without automatically
editing them as sprinkling gfp.h and slab.h inclusions around .h
files could easily lead to inclusion dependency hell. Most gfp.h
inclusion directives were ignored as stuff from gfp.h was usually
wildly available and often used in preprocessor macros. Each
slab.h inclusion directive was examined and added manually as
necessary.
6. percpu.h was updated not to include slab.h.
7. Build test were done on the following configurations and failures
were fixed. CONFIG_GCOV_KERNEL was turned off for all tests (as my
distributed build env didn't work with gcov compiles) and a few
more options had to be turned off depending on archs to make things
build (like ipr on powerpc/64 which failed due to missing writeq).
* x86 and x86_64 UP and SMP allmodconfig and a custom test config.
* powerpc and powerpc64 SMP allmodconfig
* sparc and sparc64 SMP allmodconfig
* ia64 SMP allmodconfig
* s390 SMP allmodconfig
* alpha SMP allmodconfig
* um on x86_64 SMP allmodconfig
8. percpu.h modifications were reverted so that it could be applied as
a separate patch and serve as bisection point.
Given the fact that I had only a couple of failures from tests on step
6, I'm fairly confident about the coverage of this conversion patch.
If there is a breakage, it's likely to be something in one of the arch
headers which should be easily discoverable easily on most builds of
the specific arch.
Signed-off-by: Tejun Heo <tj@kernel.org>
Guess-its-ok-by: Christoph Lameter <cl@linux-foundation.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Lee Schermerhorn <Lee.Schermerhorn@hp.com>
2010-03-24 17:04:11 +09:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-01-07 20:41:55 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
2013-09-29 11:24:49 -04:00
|
|
|
|
2006-09-30 20:52:18 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2012-01-07 20:41:55 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
|
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2010-06-06 10:38:15 -06:00
|
|
|
|
2010-04-01 20:36:30 -05:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-07-02 22:38:35 +10:00
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2009-09-15 20:04:57 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2008-04-29 00:58:56 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2008-04-29 00:58:56 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:44 -07:00
|
|
|
|
2008-04-29 00:58:56 -07:00
|
|
|
|
ext4: fix potential deadlock in ext4_nonda_switch()
In ext4_nonda_switch(), if the file system is getting full we used to
call writeback_inodes_sb_if_idle(). The problem is that we can be
holding i_mutex already, and this causes a potential deadlock when
writeback_inodes_sb_if_idle() when it tries to take s_umount. (See
lockdep output below).
As it turns out we don't need need to hold s_umount; the fact that we
are in the middle of the write(2) system call will keep the superblock
pinned. Unfortunately writeback_inodes_sb() checks to make sure
s_umount is taken, and the VFS uses a different mechanism for making
sure the file system doesn't get unmounted out from under us. The
simplest way of dealing with this is to just simply grab s_umount
using a trylock, and skip kicking the writeback flusher thread in the
very unlikely case that we can't take a read lock on s_umount without
blocking.
Also, we now check the cirteria for kicking the writeback thread
before we decide to whether to fall back to non-delayed writeback, so
if there are any outstanding delayed allocation writes, we try to get
them resolved as soon as possible.
[ INFO: possible circular locking dependency detected ]
3.6.0-rc1-00042-gce894ca #367 Not tainted
-------------------------------------------------------
dd/8298 is trying to acquire lock:
(&type->s_umount_key#18){++++..}, at: [<c02277d4>] writeback_inodes_sb_if_idle+0x28/0x46
but task is already holding lock:
(&sb->s_type->i_mutex_key#8){+.+...}, at: [<c01ddcce>] generic_file_aio_write+0x5f/0xd3
which lock already depends on the new lock.
2 locks held by dd/8298:
#0: (sb_writers#2){.+.+.+}, at: [<c01ddcc5>] generic_file_aio_write+0x56/0xd3
#1: (&sb->s_type->i_mutex_key#8){+.+...}, at: [<c01ddcce>] generic_file_aio_write+0x5f/0xd3
stack backtrace:
Pid: 8298, comm: dd Not tainted 3.6.0-rc1-00042-gce894ca #367
Call Trace:
[<c015b79c>] ? console_unlock+0x345/0x372
[<c06d62a1>] print_circular_bug+0x190/0x19d
[<c019906c>] __lock_acquire+0x86d/0xb6c
[<c01999db>] ? mark_held_locks+0x5c/0x7b
[<c0199724>] lock_acquire+0x66/0xb9
[<c02277d4>] ? writeback_inodes_sb_if_idle+0x28/0x46
[<c06db935>] down_read+0x28/0x58
[<c02277d4>] ? writeback_inodes_sb_if_idle+0x28/0x46
[<c02277d4>] writeback_inodes_sb_if_idle+0x28/0x46
[<c026f3b2>] ext4_nonda_switch+0xe1/0xf4
[<c0271ece>] ext4_da_write_begin+0x27/0x193
[<c01dcdb0>] generic_file_buffered_write+0xc8/0x1bb
[<c01ddc47>] __generic_file_aio_write+0x1dd/0x205
[<c01ddce7>] generic_file_aio_write+0x78/0xd3
[<c026d336>] ext4_file_write+0x480/0x4a6
[<c0198c1d>] ? __lock_acquire+0x41e/0xb6c
[<c0180944>] ? sched_clock_cpu+0x11a/0x13e
[<c01967e9>] ? trace_hardirqs_off+0xb/0xd
[<c018099f>] ? local_clock+0x37/0x4e
[<c0209f2c>] do_sync_write+0x67/0x9d
[<c0209ec5>] ? wait_on_retry_sync_kiocb+0x44/0x44
[<c020a7b9>] vfs_write+0x7b/0xe6
[<c020a9a6>] sys_write+0x3b/0x64
[<c06dd4bd>] syscall_call+0x7/0xb
Signed-off-by: "Theodore Ts'o" <tytso@mit.edu>
Cc: stable@vger.kernel.org
2012-09-19 22:42:36 -04:00
|
|
|
|
2008-04-29 00:58:56 -07:00
|
|
|
|
2010-09-21 11:51:01 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-07-09 22:36:45 +08:00
|
|
|
|
2010-10-04 14:25:33 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-21 11:51:01 +02:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-01-17 11:18:56 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-02-06 15:47:47 +00:00
|
|
|
|
|
|
|
|
|
2014-04-03 14:46:23 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:44 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-04-03 14:46:23 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:44 -08:00
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
2014-04-03 14:46:23 -07:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-13 20:07:36 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
2014-04-03 14:46:23 -07:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-13 20:07:36 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-11-23 20:56:45 +08:00
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-30 22:57:33 -03:00
|
|
|
|
2010-06-01 11:08:43 +02:00
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2010-06-08 18:15:15 +02:00
|
|
|
|
2009-09-23 20:33:40 +08:00
|
|
|
|
2010-06-08 18:15:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:44 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-08 18:15:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:44 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:46 -08:00
|
|
|
|
2014-04-03 14:46:23 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:41 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:41 +11:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2011-03-22 22:23:41 +11:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:34 -07:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2007-10-16 23:30:34 -07:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2007-10-16 23:30:34 -07:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2007-10-16 23:30:34 -07:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:47:55 +02:00
|
|
|
|
2012-11-26 16:29:51 -08:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:47:55 +02:00
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-04-02 16:56:37 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-23 19:37:09 +02:00
|
|
|
|
2009-04-02 16:56:37 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2012-09-11 08:28:18 +08:00
|
|
|
|
2012-03-09 07:26:22 -08:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2011-04-23 12:27:27 -06:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2011-10-07 21:51:56 -06:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
|
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
2011-04-23 12:27:27 -06:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2013-07-09 22:36:45 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
|
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-04-23 12:27:27 -06:00
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
|
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2011-04-23 12:27:27 -06:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:42 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2011-10-07 21:51:56 -06:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2011-04-23 12:27:27 -06:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2010-08-11 14:17:42 -07:00
|
|
|
|
2011-10-07 21:51:56 -06:00
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:37 +01:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2013-01-11 13:06:37 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
|
|
|
|
|
sched: Remove proliferation of wait_on_bit() action functions
The current "wait_on_bit" interface requires an 'action'
function to be provided which does the actual waiting.
There are over 20 such functions, many of them identical.
Most cases can be satisfied by one of just two functions, one
which uses io_schedule() and one which just uses schedule().
So:
Rename wait_on_bit and wait_on_bit_lock to
wait_on_bit_action and wait_on_bit_lock_action
to make it explicit that they need an action function.
Introduce new wait_on_bit{,_lock} and wait_on_bit{,_lock}_io
which are *not* given an action function but implicitly use
a standard one.
The decision to error-out if a signal is pending is now made
based on the 'mode' argument rather than being encoded in the action
function.
All instances of the old wait_on_bit and wait_on_bit_lock which
can use the new version have been changed accordingly and their
action functions have been discarded.
wait_on_bit{_lock} does not return any specific error code in the
event of a signal so the caller must check for non-zero and
interpolate their own error code as appropriate.
The wait_on_bit() call in __fscache_wait_on_invalidate() was
ambiguous as it specified TASK_UNINTERRUPTIBLE but used
fscache_wait_bit_interruptible as an action function.
David Howells confirms this should be uniformly
"uninterruptible"
The main remaining user of wait_on_bit{,_lock}_action is NFS
which needs to use a freezer-aware schedule() call.
A comment in fs/gfs2/glock.c notes that having multiple 'action'
functions is useful as they display differently in the 'wchan'
field of 'ps'. (and /proc/$PID/wchan).
As the new bit_wait{,_io} functions are tagged "__sched", they
will not show up at all, but something higher in the stack. So
the distinction will still be visible, only with different
function names (gds2_glock_wait versus gfs2_glock_dq_wait in the
gfs2/glock.c case).
Since first version of this patch (against 3.15) two new action
functions appeared, on in NFS and one in CIFS. CIFS also now
uses an action function that makes the same freezer aware
schedule call as NFS.
Signed-off-by: NeilBrown <neilb@suse.de>
Acked-by: David Howells <dhowells@redhat.com> (fscache, keys)
Acked-by: Steven Whitehouse <swhiteho@redhat.com> (gfs2)
Acked-by: Peter Zijlstra <peterz@infradead.org>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Steve French <sfrench@samba.org>
Cc: Linus Torvalds <torvalds@linux-foundation.org>
Link: http://lkml.kernel.org/r/20140707051603.28027.72349.stgit@notabene.brown
Signed-off-by: Ingo Molnar <mingo@kernel.org>
2014-07-07 15:16:04 +10:00
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2010-05-24 14:32:38 -07:00
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:47:58 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:47:58 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2012-10-08 16:33:45 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2010-12-01 17:33:37 -06:00
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2013-01-11 13:06:37 -08:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:21 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-07-02 22:38:35 +10:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:21 +01:00
|
|
|
|
2013-07-02 22:38:35 +10:00
|
|
|
|
2010-03-05 09:21:21 +01:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-07 13:35:44 +04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2012-05-03 14:47:57 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-07 13:35:44 +04:00
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2010-03-05 09:21:21 +01:00
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:37 +01:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-12-14 04:21:26 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
2013-12-14 04:21:26 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-10-08 16:33:45 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-29 13:28:09 -06:00
|
|
|
|
|
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-29 13:28:09 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-07-02 22:38:35 +10:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-24 19:40:46 +02:00
|
|
|
|
2012-06-09 11:10:55 +08:00
|
|
|
|
|
|
|
|
|
2010-10-24 19:40:46 +02:00
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2010-10-24 19:40:46 +02:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2011-07-11 23:08:50 -07:00
|
|
|
|
fs: new inode i_state corruption fix
There was a report of a data corruption
http://lkml.org/lkml/2008/11/14/121. There is a script included to
reproduce the problem.
During testing, I encountered a number of strange things with ext3, so I
tried ext2 to attempt to reduce complexity of the problem. I found that
fsstress would quickly hang in wait_on_inode, waiting for I_LOCK to be
cleared, even though instrumentation showed that unlock_new_inode had
already been called for that inode. This points to memory scribble, or
synchronisation problme.
i_state of I_NEW inodes is not protected by inode_lock because other
processes are not supposed to touch them until I_LOCK (and I_NEW) is
cleared. Adding WARN_ON(inode->i_state & I_NEW) to sites where we modify
i_state revealed that generic_sync_sb_inodes is picking up new inodes from
the inode lists and passing them to __writeback_single_inode without
waiting for I_NEW. Subsequently modifying i_state causes corruption. In
my case it would look like this:
CPU0 CPU1
unlock_new_inode() __sync_single_inode()
reg <- inode->i_state
reg -> reg & ~(I_LOCK|I_NEW) reg <- inode->i_state
reg -> inode->i_state reg -> reg | I_SYNC
reg -> inode->i_state
Non-atomic RMW on CPU1 overwrites CPU0 store and sets I_LOCK|I_NEW again.
Fix for this is rather than wait for I_NEW inodes, just skip over them:
inodes concurrently being created are not subject to data integrity
operations, and should not significantly contribute to dirty memory
either.
After this change, I'm unable to reproduce any of the added warnings or
hangs after ~1hour of running. Previously, the new warnings would start
immediately and hang would happen in under 5 minutes.
I'm also testing on ext3 now, and so far no problems there either. I
don't know whether this fixes the problem reported above, but it fixes a
real problem for me.
Cc: "Jorge Boncompte [DTI2]" <jorge@dti2.net>
Reported-by: Adrian Hunter <ext-adrian.hunter@nokia.com>
Cc: Jan Kara <jack@suse.cz>
Cc: <stable@kernel.org>
Signed-off-by: Nick Piggin <npiggin@suse.de>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2009-03-12 14:31:38 -07:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:47:56 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:47:59 +02:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-06-08 17:07:36 +02:00
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
2010-08-29 13:28:09 -06:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-10-08 16:33:45 -07:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:43 +11:00
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
writeback: speed up writeback of big dirty files
After making dirty a 100M file, the normal behavior is to start the
writeback for all data after 30s delays. But sometimes the following
happens instead:
- after 30s: ~4M
- after 5s: ~4M
- after 5s: all remaining 92M
Some analyze shows that the internal io dispatch queues goes like this:
s_io s_more_io
-------------------------
1) 100M,1K 0
2) 1K 96M
3) 0 96M
1) initial state with a 100M file and a 1K file
2) 4M written, nr_to_write <= 0, so write more
3) 1K written, nr_to_write > 0, no more writes(BUG)
nr_to_write > 0 in (3) fools the upper layer to think that data have all
been written out. The big dirty file is actually still sitting in
s_more_io. We cannot simply splice s_more_io back to s_io as soon as s_io
becomes empty, and let the loop in generic_sync_sb_inodes() continue: this
may starve newly expired inodes in s_dirty. It is also not an option to
draw inodes from both s_more_io and s_dirty, an let the loop go on: this
might lead to live locks, and might also starve other superblocks in sync
time(well kupdate may still starve some superblocks, that's another bug).
We have to return when a full scan of s_io completes. So nr_to_write > 0
does not necessarily mean that "all data are written". This patch
introduces a flag writeback_control.more_io to indicate that more io should
be done. With it the big dirty file no longer has to wait for the next
kupdate invokation 5s later.
In sync_sb_inodes() we only set more_io on super_blocks we actually
visited. This avoids the interaction between two pdflush deamons.
Also in __sync_single_inode() we don't blindly keep requeuing the io if the
filesystem cannot progress. Failing to do so may lead to 100% iowait.
Tested-by: Mike Snitzer <snitzer@gmail.com>
Signed-off-by: Fengguang Wu <wfg@mail.ustc.edu.cn>
Cc: Michael Rubin <mrubin@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-04 22:29:36 -08:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2009-09-24 15:25:11 +02:00
|
|
|
|
2011-07-08 14:14:41 +10:00
|
|
|
|
2011-07-29 22:14:35 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
|
|
|
|
|
2013-09-11 14:22:40 -07:00
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
writeback: refill b_io iff empty
There is no point to carry different refill policies between for_kupdate
and other type of works. Use a consistent "refill b_io iff empty" policy
which can guarantee fairness in an easy to understand way.
A b_io refill will setup a _fixed_ work set with all currently eligible
inodes and start a new round of walk through b_io. The "fixed" work set
means no new inodes will be added to the work set during the walk.
Only when a complete walk over b_io is done, new inodes that are
eligible at the time will be enqueued and the walk be started over.
This procedure provides fairness among the inodes because it guarantees
each inode to be synced once and only once at each round. So all inodes
will be free from starvations.
This change relies on wb_writeback() to keep retrying as long as we made
some progress on cleaning some pages and/or inodes. Without that ability,
the old logic on background works relies on aggressively queuing all
eligible inodes into b_io at every time. But that's not a guarantee.
The below test script completes a slightly faster now:
2.6.39-rc3 2.6.39-rc3-dyn-expire+
------------------------------------------------
all elapsed 256.043 252.367
stddev 24.381 12.530
tar elapsed 30.097 28.808
dd elapsed 13.214 11.782
#!/bin/zsh
cp /c/linux-2.6.38.3.tar.bz2 /dev/shm/
umount /dev/sda7
mkfs.xfs -f /dev/sda7
mount /dev/sda7 /fs
echo 3 > /proc/sys/vm/drop_caches
tic=$(cat /proc/uptime|cut -d' ' -f2)
cd /fs
time tar jxf /dev/shm/linux-2.6.38.3.tar.bz2 &
time dd if=/dev/zero of=/fs/zero bs=1M count=1000 &
wait
sync
tac=$(cat /proc/uptime|cut -d' ' -f2)
echo elapsed: $((tac - tic))
It maintains roughly the same small vs. large file writeout shares, and
offers large files better chances to be written in nice 4M chunks.
Analyzes from Dave Chinner in great details:
Let's say we have lots of inodes with 100 dirty pages being created,
and one large writeback going on. We expire 8 new inodes for every
1024 pages we write back.
With the old code, we do:
b_more_io (large inode) -> b_io (1l)
8 newly expired inodes -> b_io (1l, 8s)
writeback large inode 1024 pages -> b_more_io
b_more_io (large inode) -> b_io (8s, 1l)
8 newly expired inodes -> b_io (8s, 1l, 8s)
writeback 8 small inodes 800 pages
1 large inode 224 pages -> b_more_io
b_more_io (large inode) -> b_io (8s, 1l)
8 newly expired inodes -> b_io (8s, 1l, 8s)
.....
Your new code:
b_more_io (large inode) -> b_io (1l)
8 newly expired inodes -> b_io (1l, 8s)
writeback large inode 1024 pages -> b_more_io
(b_io == 8s)
writeback 8 small inodes 800 pages
b_io empty: (1800 pages written)
b_more_io (large inode) -> b_io (1l)
14 newly expired inodes -> b_io (1l, 14s)
writeback large inode 1024 pages -> b_more_io
(b_io == 14s)
writeback 10 small inodes 1000 pages
1 small inode 24 pages -> b_more_io (1l, 1s(24))
writeback 5 small inodes 500 pages
b_io empty: (2548 pages written)
b_more_io (large inode) -> b_io (1l, 1s(24))
20 newly expired inodes -> b_io (1l, 1s(24), 20s)
......
Rough progression of pages written at b_io refill:
Old code:
total large file % of writeback
1024 224 21.9% (fixed)
New code:
total large file % of writeback
1800 1024 ~55%
2550 1024 ~40%
3050 1024 ~33%
3500 1024 ~29%
3950 1024 ~26%
4250 1024 ~24%
4500 1024 ~22.7%
4700 1024 ~21.7%
4800 1024 ~21.3%
4800 1024 ~21.3%
(pretty much steady state from here)
Ok, so the steady state is reached with a similar percentage of
writeback to the large file as the existing code. Ok, that's good,
but providing some evidence that is doesn't change the shared of
writeback to the large should be in the commit message ;)
The other advantage to this is that we always write 1024 page chunks
to the large file, rather than smaller "whatever remains" chunks.
CC: Jan Kara <jack@suse.cz>
Acked-by: Mel Gorman <mel@csn.ul.ie>
Signed-off-by: Wu Fengguang <fengguang.wu@intel.com>
2010-07-21 20:11:53 -06:00
|
|
|
|
2011-10-07 21:51:56 -06:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-11-18 14:38:33 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:39 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-11-18 14:38:33 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-08-29 11:22:30 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-10-03 20:46:17 -06:00
|
|
|
|
2010-08-29 11:22:30 -06:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2010-08-29 11:22:30 -06:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
2009-09-16 19:22:48 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
|
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2011-04-21 12:06:32 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2009-09-23 20:33:40 +08:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2011-01-13 15:45:47 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2009-09-23 20:33:40 +08:00
|
|
|
|
|
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2010-11-18 14:38:33 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2011-10-19 11:44:41 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-21 20:32:30 -06:00
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
2010-07-21 20:32:30 -06:00
|
|
|
|
2011-10-19 11:44:41 +02:00
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
2010-07-07 13:24:07 +10:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2011-04-21 12:06:32 -06:00
|
|
|
|
2011-10-07 21:51:56 -06:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:07 +10:00
|
|
|
|
2010-08-29 11:22:30 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
writeback: try more writeback as long as something was written
writeback_inodes_wb()/__writeback_inodes_sb() are not aggressive in that
they only populate possibly a subset of eligible inodes into b_io at
entrance time. When the queued set of inodes are all synced, they just
return, possibly with all queued inode pages written but still
wbc.nr_to_write > 0.
For kupdate and background writeback, there may be more eligible inodes
sitting in b_dirty when the current set of b_io inodes are completed. So
it is necessary to try another round of writeback as long as we made some
progress in this round. When there are no more eligible inodes, no more
inodes will be enqueued in queue_io(), hence nothing could/will be
synced and we may safely bail.
For example, imagine 100 inodes
i0, i1, i2, ..., i90, i91, i99
At queue_io() time, i90-i99 happen to be expired and moved to s_io for
IO. When finished successfully, if their total size is less than
MAX_WRITEBACK_PAGES, nr_to_write will be > 0. Then wb_writeback() will
quit the background work (w/o this patch) while it's still over
background threshold. This will be a fairly normal/frequent case I guess.
Now that we do tagged sync and update inode->dirtied_when after the sync,
this change won't livelock sync(1). I actually tried to write 1 page
per 1ms with this command
write-and-fsync -n10000 -S 1000 -c 4096 /fs/test
and do sync(1) at the same time. The sync completes quickly on ext4,
xfs, btrfs.
Acked-by: Jan Kara <jack@suse.cz>
Signed-off-by: Wu Fengguang <fengguang.wu@intel.com>
2010-07-22 10:23:44 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2009-09-23 19:32:26 +02:00
|
|
|
|
|
|
|
|
|
writeback: try more writeback as long as something was written
writeback_inodes_wb()/__writeback_inodes_sb() are not aggressive in that
they only populate possibly a subset of eligible inodes into b_io at
entrance time. When the queued set of inodes are all synced, they just
return, possibly with all queued inode pages written but still
wbc.nr_to_write > 0.
For kupdate and background writeback, there may be more eligible inodes
sitting in b_dirty when the current set of b_io inodes are completed. So
it is necessary to try another round of writeback as long as we made some
progress in this round. When there are no more eligible inodes, no more
inodes will be enqueued in queue_io(), hence nothing could/will be
synced and we may safely bail.
For example, imagine 100 inodes
i0, i1, i2, ..., i90, i91, i99
At queue_io() time, i90-i99 happen to be expired and moved to s_io for
IO. When finished successfully, if their total size is less than
MAX_WRITEBACK_PAGES, nr_to_write will be > 0. Then wb_writeback() will
quit the background work (w/o this patch) while it's still over
background threshold. This will be a fairly normal/frequent case I guess.
Now that we do tagged sync and update inode->dirtied_when after the sync,
this change won't livelock sync(1). I actually tried to write 1 page
per 1ms with this command
write-and-fsync -n10000 -S 1000 -c 4096 /fs/test
and do sync(1) at the same time. The sync completes quickly on ext4,
xfs, btrfs.
Acked-by: Jan Kara <jack@suse.cz>
Signed-off-by: Wu Fengguang <fengguang.wu@intel.com>
2010-07-22 10:23:44 -06:00
|
|
|
|
2009-09-23 19:32:26 +02:00
|
|
|
|
2010-07-21 22:19:51 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-23 19:32:26 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2012-05-03 14:47:59 +02:00
|
|
|
|
2012-05-03 14:48:03 +02:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:47:59 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2011-04-21 12:06:32 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-05-04 19:54:37 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2010-08-03 12:51:16 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-30 08:55:52 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:44 -08:00
|
|
|
|
|
|
|
|
|
2010-11-18 14:38:33 -06:00
|
|
|
|
2011-01-13 15:45:44 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2011-01-13 15:45:44 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-17 12:51:03 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-30 08:55:52 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-07-08 16:00:14 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-08-11 14:17:44 -07:00
|
|
|
|
2010-08-03 12:51:16 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:44 -08:00
|
|
|
|
2010-08-11 14:17:44 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
|
|
|
|
|
2010-06-19 23:08:22 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2013-04-30 15:27:24 -07:00
|
|
|
|
2010-10-26 14:22:45 -07:00
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
2014-04-03 14:46:23 -07:00
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
2013-07-08 16:00:14 -07:00
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2014-04-03 14:46:22 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
writeback: replace custom worker pool implementation with unbound workqueue
Writeback implements its own worker pool - each bdi can be associated
with a worker thread which is created and destroyed dynamically. The
worker thread for the default bdi is always present and serves as the
"forker" thread which forks off worker threads for other bdis.
there's no reason for writeback to implement its own worker pool when
using unbound workqueue instead is much simpler and more efficient.
This patch replaces custom worker pool implementation in writeback
with an unbound workqueue.
The conversion isn't too complicated but the followings are worth
mentioning.
* bdi_writeback->last_active, task and wakeup_timer are removed.
delayed_work ->dwork is added instead. Explicit timer handling is
no longer necessary. Everything works by either queueing / modding
/ flushing / canceling the delayed_work item.
* bdi_writeback_thread() becomes bdi_writeback_workfn() which runs off
bdi_writeback->dwork. On each execution, it processes
bdi->work_list and reschedules itself if there are more things to
do.
The function also handles low-mem condition, which used to be
handled by the forker thread. If the function is running off a
rescuer thread, it only writes out limited number of pages so that
the rescuer can serve other bdis too. This preserves the flusher
creation failure behavior of the forker thread.
* INIT_LIST_HEAD(&bdi->bdi_list) is used to tell
bdi_writeback_workfn() about on-going bdi unregistration so that it
always drains work_list even if it's running off the rescuer. Note
that the original code was broken in this regard. Under memory
pressure, a bdi could finish unregistration with non-empty
work_list.
* The default bdi is no longer special. It now is treated the same as
any other bdi and bdi_cap_flush_forker() is removed.
* BDI_pending is no longer used. Removed.
* Some tracepoints become non-applicable. The following TPs are
removed - writeback_nothread, writeback_wake_thread,
writeback_wake_forker_thread, writeback_thread_start,
writeback_thread_stop.
Everything, including devices coming and going away and rescuer
operation under simulated memory pressure, seems to work fine in my
test setup.
Signed-off-by: Tejun Heo <tj@kernel.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Cc: Jens Axboe <axboe@kernel.dk>
Cc: Fengguang Wu <fengguang.wu@intel.com>
Cc: Jeff Moyer <jmoyer@redhat.com>
2013-04-01 19:08:06 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-08 18:15:07 +02:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-08 18:15:07 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2013-09-11 14:22:22 -07:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-08 18:15:07 +02:00
|
|
|
|
2009-09-14 13:12:40 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-14 13:12:40 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2013-01-11 13:06:37 -08:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-05-27 06:53:02 -04:00
|
|
|
|
2013-01-11 13:06:37 -08:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-23 15:19:20 -04:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-02 17:38:30 -04:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:41 +11:00
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
|
|
|
|
|
2013-09-11 14:23:04 -07:00
|
|
|
|
|
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:10:25 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2011-03-22 22:23:41 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
|
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-03-22 22:23:40 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-03-22 22:23:36 +11:00
|
|
|
|
2011-03-22 22:23:40 +11:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-03-22 22:23:40 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-03-22 22:23:40 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2011-03-22 22:23:40 +11:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2011-11-23 20:56:45 +08:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
|
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-06-06 10:38:15 -06:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2010-06-08 18:14:43 +02:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2012-07-03 16:45:27 +02:00
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:51 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-05-17 12:55:07 +02:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-11-23 20:56:45 +08:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2010-06-01 11:08:43 +02:00
|
|
|
|
2010-05-17 12:55:07 +02:00
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
|
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
|
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
|
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2011-11-23 20:56:45 +08:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
|
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2013-01-10 13:47:57 +08:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
|
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
|
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2014-02-21 11:19:04 +01:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:43 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2011-10-07 21:54:10 -06:00
|
|
|
|
2013-07-02 22:38:35 +10:00
|
|
|
|
2010-06-08 18:14:43 +02:00
|
|
|
|
|
|
|
|
|
2012-07-03 16:45:27 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:51 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
[PATCH] fix nr_unused accounting, and avoid recursing in iput with I_WILL_FREE set
list_move(&inode->i_list, &inode_in_use);
} else {
list_move(&inode->i_list, &inode_unused);
+ inodes_stat.nr_unused++;
}
}
wake_up_inode(inode);
Are you sure the above diff is correct? It was added somewhere between
2.6.5 and 2.6.8. I think it's wrong.
The only way I can imagine the i_count to be zero in the above path, is
that I_WILL_FREE is set. And if I_WILL_FREE is set, then we must not
increase nr_unused. So I believe the above change is buggy and it will
definitely overstate the number of unused inodes and it should be backed
out.
Note that __writeback_single_inode before calling __sync_single_inode, can
drop the spinlock and we can have both the dirty and locked bitflags clear
here:
spin_unlock(&inode_lock);
__wait_on_inode(inode);
iput(inode);
XXXXXXX
spin_lock(&inode_lock);
}
use inode again here
a construct like the above makes zero sense from a reference counting
standpoint.
Either we don't ever use the inode again after the iput, or the
inode_lock should be taken _before_ executing the iput (i.e. a __iput
would be required). Taking the inode_lock after iput means the iget was
useless if we keep using the inode after the iput.
So the only chance the 2.6 was safe to call __writeback_single_inode
with the i_count == 0, is that I_WILL_FREE is set (I_WILL_FREE will
prevent the VM to free the inode in XXXXX).
Potentially calling the above iput with I_WILL_FREE was also wrong
because it would recurse in iput_final (the second mainline bug).
The below (untested) patch fixes the nr_unused accounting, avoids recursing
in iput when I_WILL_FREE is set and makes sure (with the BUG_ON) that we
don't corrupt memory and that all holders that don't set I_WILL_FREE, keeps
a reference on the inode!
Signed-off-by: Andrea Arcangeli <andrea@suse.de>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-30 15:03:05 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
[PATCH] fix nr_unused accounting, and avoid recursing in iput with I_WILL_FREE set
list_move(&inode->i_list, &inode_in_use);
} else {
list_move(&inode->i_list, &inode_unused);
+ inodes_stat.nr_unused++;
}
}
wake_up_inode(inode);
Are you sure the above diff is correct? It was added somewhere between
2.6.5 and 2.6.8. I think it's wrong.
The only way I can imagine the i_count to be zero in the above path, is
that I_WILL_FREE is set. And if I_WILL_FREE is set, then we must not
increase nr_unused. So I believe the above change is buggy and it will
definitely overstate the number of unused inodes and it should be backed
out.
Note that __writeback_single_inode before calling __sync_single_inode, can
drop the spinlock and we can have both the dirty and locked bitflags clear
here:
spin_unlock(&inode_lock);
__wait_on_inode(inode);
iput(inode);
XXXXXXX
spin_lock(&inode_lock);
}
use inode again here
a construct like the above makes zero sense from a reference counting
standpoint.
Either we don't ever use the inode again after the iput, or the
inode_lock should be taken _before_ executing the iput (i.e. a __iput
would be required). Taking the inode_lock after iput means the iget was
useless if we keep using the inode after the iput.
So the only chance the 2.6 was safe to call __writeback_single_inode
with the i_count == 0, is that I_WILL_FREE is set (I_WILL_FREE will
prevent the VM to free the inode in XXXXX).
Potentially calling the above iput with I_WILL_FREE was also wrong
because it would recurse in iput_final (the second mainline bug).
The below (untested) patch fixes the nr_unused accounting, avoids recursing
in iput when I_WILL_FREE is set and makes sure (with the BUG_ON) that we
don't corrupt memory and that all holders that don't set I_WILL_FREE, keeps
a reference on the inode!
Signed-off-by: Andrea Arcangeli <andrea@suse.de>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-30 15:03:05 -08:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-04-21 18:19:44 -06:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2008-02-08 04:20:23 -08:00
|
|
|
|
[PATCH] writeback: fix range handling
When a writeback_control's `start' and `end' fields are used to
indicate a one-byte-range starting at file offset zero, the required
values of .start=0,.end=0 mean that the ->writepages() implementation
has no way of telling that it is being asked to perform a range
request. Because we're currently overloading (start == 0 && end == 0)
to mean "this is not a write-a-range request".
To make all this sane, the patch changes range of writeback_control.
So caller does: If it is calling ->writepages() to write pages, it
sets range (range_start/end or range_cyclic) always.
And if range_cyclic is true, ->writepages() thinks the range is
cyclic, otherwise it just uses range_start and range_end.
This patch does,
- Add LLONG_MAX, LLONG_MIN, ULLONG_MAX to include/linux/kernel.h
-1 is usually ok for range_end (type is long long). But, if someone did,
range_end += val; range_end is "val - 1"
u64val = range_end >> bits; u64val is "~(0ULL)"
or something, they are wrong. So, this adds LLONG_MAX to avoid nasty
things, and uses LLONG_MAX for range_end.
- All callers of ->writepages() sets range_start/end or range_cyclic.
- Fix updates of ->writeback_index. It seems already bit strange.
If it starts at 0 and ended by check of nr_to_write, this last
index may reduce chance to scan end of file. So, this updates
->writeback_index only if range_cyclic is true or whole-file is
scanned.
Signed-off-by: OGAWA Hirofumi <hirofumi@mail.parknet.co.jp>
Cc: Nathan Scott <nathans@sgi.com>
Cc: Anton Altaparmakov <aia21@cantab.net>
Cc: Steven French <sfrench@us.ibm.com>
Cc: "Vladimir V. Saveliev" <vs@namesys.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2006-06-23 02:03:26 -07:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-11-07 00:59:15 -08:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2012-05-03 14:48:00 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2010-10-06 10:48:20 +02:00
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:48 -08:00
|
|
|
|
2010-10-06 10:48:20 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2011-01-13 15:45:48 -08:00
|
|
|
|
2010-10-06 10:48:20 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|