2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2008-10-15 22:01:59 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2007-09-21 09:19:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
include cleanup: Update gfp.h and slab.h includes to prepare for breaking implicit slab.h inclusion from percpu.h
percpu.h is included by sched.h and module.h and thus ends up being
included when building most .c files. percpu.h includes slab.h which
in turn includes gfp.h making everything defined by the two files
universally available and complicating inclusion dependencies.
percpu.h -> slab.h dependency is about to be removed. Prepare for
this change by updating users of gfp and slab facilities include those
headers directly instead of assuming availability. As this conversion
needs to touch large number of source files, the following script is
used as the basis of conversion.
http://userweb.kernel.org/~tj/misc/slabh-sweep.py
The script does the followings.
* Scan files for gfp and slab usages and update includes such that
only the necessary includes are there. ie. if only gfp is used,
gfp.h, if slab is used, slab.h.
* When the script inserts a new include, it looks at the include
blocks and try to put the new include such that its order conforms
to its surrounding. It's put in the include block which contains
core kernel includes, in the same order that the rest are ordered -
alphabetical, Christmas tree, rev-Xmas-tree or at the end if there
doesn't seem to be any matching order.
* If the script can't find a place to put a new include (mostly
because the file doesn't have fitting include block), it prints out
an error message indicating which .h file needs to be added to the
file.
The conversion was done in the following steps.
1. The initial automatic conversion of all .c files updated slightly
over 4000 files, deleting around 700 includes and adding ~480 gfp.h
and ~3000 slab.h inclusions. The script emitted errors for ~400
files.
2. Each error was manually checked. Some didn't need the inclusion,
some needed manual addition while adding it to implementation .h or
embedding .c file was more appropriate for others. This step added
inclusions to around 150 files.
3. The script was run again and the output was compared to the edits
from #2 to make sure no file was left behind.
4. Several build tests were done and a couple of problems were fixed.
e.g. lib/decompress_*.c used malloc/free() wrappers around slab
APIs requiring slab.h to be added manually.
5. The script was run on all .h files but without automatically
editing them as sprinkling gfp.h and slab.h inclusions around .h
files could easily lead to inclusion dependency hell. Most gfp.h
inclusion directives were ignored as stuff from gfp.h was usually
wildly available and often used in preprocessor macros. Each
slab.h inclusion directive was examined and added manually as
necessary.
6. percpu.h was updated not to include slab.h.
7. Build test were done on the following configurations and failures
were fixed. CONFIG_GCOV_KERNEL was turned off for all tests (as my
distributed build env didn't work with gcov compiles) and a few
more options had to be turned off depending on archs to make things
build (like ipr on powerpc/64 which failed due to missing writeq).
* x86 and x86_64 UP and SMP allmodconfig and a custom test config.
* powerpc and powerpc64 SMP allmodconfig
* sparc and sparc64 SMP allmodconfig
* ia64 SMP allmodconfig
* s390 SMP allmodconfig
* alpha SMP allmodconfig
* um on x86_64 SMP allmodconfig
8. percpu.h modifications were reverted so that it could be applied as
a separate patch and serve as bisection point.
Given the fact that I had only a couple of failures from tests on step
6, I'm fairly confident about the coverage of this conversion patch.
If there is a breakage, it's likely to be something in one of the arch
headers which should be easily discoverable easily on most builds of
the specific arch.
Signed-off-by: Tejun Heo <tj@kernel.org>
Guess-its-ok-by: Christoph Lameter <cl@linux-foundation.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Lee Schermerhorn <Lee.Schermerhorn@hp.com>
2010-03-24 17:04:11 +09:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
2006-09-30 20:52:18 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-04-01 20:36:30 -05:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2009-09-15 20:04:57 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2008-04-29 00:58:56 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2008-04-29 00:58:56 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:44 -07:00
|
|
|
|
2008-04-29 00:58:56 -07:00
|
|
|
|
|
|
|
|
|
2010-09-21 11:51:01 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-04 14:25:33 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-09-21 11:51:01 +02:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
writeback: move bdi threads exiting logic to the forker thread
Currently, bdi threads can decide to exit if there were no useful activities
for 5 minutes. However, this causes nasty races: we can easily oops in the
'bdi_queue_work()' if the bdi thread decides to exit while we are waking it up.
And even if we do not oops, but the bdi tread exits immediately after we wake
it up, we'd lose the wake-up event and have an unnecessary delay (up to 5 secs)
in the bdi work processing.
This patch makes the forker thread to be the central place which not only
creates bdi threads, but also kills them if they were inactive long enough.
This better design-wise.
Another reason why this change was done is to prepare for the further changes
which will prevent the bdi threads from waking up every 5 sec and wasting
power. Indeed, when the task does not wake up periodically anymore, it won't be
able to exit either.
This patch also moves the the 'wake_up_bit()' call from the bdi thread to the
forker thread as well. So now the forker thread sets the BDI_pending bit, then
forks the task or kills it, then clears the bit and wakes up the waiting
process.
The only process which may wain on the bit is 'bdi_wb_shutdown()'. This
function was changed as well - now it first removes the bdi from the
'bdi_list', then waits on the 'BDI_pending' bit. Once it wakes up, it is
guaranteed that the forker thread won't race with it, because the bdi is not
visible. Note, the forker thread sets the 'BDI_pending' bit under the
'bdi->wb_lock' which is essential for proper serialization.
And additionally, when we change 'bdi->wb.task', we now take the
'bdi->work_lock', to make sure that we do not lose wake-ups which we otherwise
would when raced with, say, 'bdi_queue_work()'.
Signed-off-by: Artem Bityutskiy <Artem.Bityutskiy@nokia.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <jaxboe@fusionio.com>
2010-07-25 14:29:20 +03:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-13 20:07:36 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-13 20:07:36 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-01 11:08:43 +02:00
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
|
|
|
|
|
2010-06-08 18:15:15 +02:00
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2010-06-08 18:15:15 +02:00
|
|
|
|
2009-09-23 20:33:40 +08:00
|
|
|
|
2010-06-08 18:15:15 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2007-10-16 23:30:32 -07:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:34 -07:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2007-10-16 23:30:34 -07:00
|
|
|
|
2007-10-16 23:30:38 -07:00
|
|
|
|
2007-10-16 23:30:34 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2007-10-16 23:30:34 -07:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-04-02 16:56:37 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-23 19:37:09 +02:00
|
|
|
|
2009-04-02 16:56:37 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
|
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2009-04-02 16:56:37 -07:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
|
|
|
|
|
2009-09-24 15:12:57 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2009-09-24 14:42:33 +02:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:42 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2010-08-11 14:17:42 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:37 +01:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-03-05 09:21:37 +01:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2007-10-16 23:30:39 -07:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-24 14:32:38 -07:00
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-24 14:32:38 -07:00
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2010-03-05 09:21:37 +01:00
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-05-07 13:35:44 +04:00
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
2010-05-07 13:35:44 +04:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:21 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:37 +01:00
|
|
|
|
2010-03-05 09:21:21 +01:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-07 13:35:44 +04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:21 +01:00
|
|
|
|
|
|
|
|
|
2010-03-05 09:21:37 +01:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
2010-06-02 17:38:30 -04:00
|
|
|
|
2010-08-11 14:17:41 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:43 -07:00
|
|
|
|
2007-10-16 23:30:35 -07:00
|
|
|
|
2010-08-11 14:17:43 -07:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-08-11 14:17:43 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-08-11 14:17:43 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:43 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2007-10-16 23:30:35 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-08-11 14:17:41 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2010-10-23 06:55:17 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-08 18:14:58 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:58 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-06-09 15:31:01 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-09 15:31:01 +02:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-09 15:31:01 +02:00
|
|
|
|
2010-06-08 18:14:58 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-06-09 15:31:01 +02:00
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:58 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-24 19:40:46 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
fs: new inode i_state corruption fix
There was a report of a data corruption
http://lkml.org/lkml/2008/11/14/121. There is a script included to
reproduce the problem.
During testing, I encountered a number of strange things with ext3, so I
tried ext2 to attempt to reduce complexity of the problem. I found that
fsstress would quickly hang in wait_on_inode, waiting for I_LOCK to be
cleared, even though instrumentation showed that unlock_new_inode had
already been called for that inode. This points to memory scribble, or
synchronisation problme.
i_state of I_NEW inodes is not protected by inode_lock because other
processes are not supposed to touch them until I_LOCK (and I_NEW) is
cleared. Adding WARN_ON(inode->i_state & I_NEW) to sites where we modify
i_state revealed that generic_sync_sb_inodes is picking up new inodes from
the inode lists and passing them to __writeback_single_inode without
waiting for I_NEW. Subsequently modifying i_state causes corruption. In
my case it would look like this:
CPU0 CPU1
unlock_new_inode() __sync_single_inode()
reg <- inode->i_state
reg -> reg & ~(I_LOCK|I_NEW) reg <- inode->i_state
reg -> inode->i_state reg -> reg | I_SYNC
reg -> inode->i_state
Non-atomic RMW on CPU1 overwrites CPU0 store and sets I_LOCK|I_NEW again.
Fix for this is rather than wait for I_NEW inodes, just skip over them:
inodes concurrently being created are not subject to data integrity
operations, and should not significantly contribute to dirty memory
either.
After this change, I'm unable to reproduce any of the added warnings or
hangs after ~1hour of running. Previously, the new warnings would start
immediately and hang would happen in under 5 minutes.
I'm also testing on ext3 now, and so far no problems there either. I
don't know whether this fixes the problem reported above, but it fixes a
real problem for me.
Cc: "Jorge Boncompte [DTI2]" <jorge@dti2.net>
Reported-by: Adrian Hunter <ext-adrian.hunter@nokia.com>
Cc: Jan Kara <jack@suse.cz>
Cc: <stable@kernel.org>
Signed-off-by: Nick Piggin <npiggin@suse.de>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2009-03-12 14:31:38 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-24 19:40:46 +02:00
|
|
|
|
2009-04-02 16:56:37 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
writeback: fix time ordering of the per superblock dirty inode lists 3
While writeback is working against a dirty inode it does a check after trying
to write some of the inode's pages:
"did the lower layers skip some of the inode's dirty pages because they were
locked (or under writeback, or whatever)"
If this turns out to be true, we must move the inode back onto s_dirty and
redirty it. The reason for doing this is that fsync() and friends only check
the s_dirty list, and those functions want to know about those pages which
were locked, so they can be waited upon and, if necessary, rewritten.
Problem is, that redirtying was putting the inode onto the tail of s_dirty
without updating its timestamp. This causes a violation of s_dirty ordering.
Fix this by updating inode->dirtied_when when moving the inode onto s_dirty.
But the code is still a bit buggy? If the inode was _already_ dirty then we
don't need to move it at all. Oh well, hopefully it doesn't matter too much,
as that was a redirtying, which was very recent anwyay.
Cc: Mike Waychison <mikew@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2007-10-16 23:30:34 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2006-03-25 03:07:44 -08:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
writeback: speed up writeback of big dirty files
After making dirty a 100M file, the normal behavior is to start the
writeback for all data after 30s delays. But sometimes the following
happens instead:
- after 30s: ~4M
- after 5s: ~4M
- after 5s: all remaining 92M
Some analyze shows that the internal io dispatch queues goes like this:
s_io s_more_io
-------------------------
1) 100M,1K 0
2) 1K 96M
3) 0 96M
1) initial state with a 100M file and a 1K file
2) 4M written, nr_to_write <= 0, so write more
3) 1K written, nr_to_write > 0, no more writes(BUG)
nr_to_write > 0 in (3) fools the upper layer to think that data have all
been written out. The big dirty file is actually still sitting in
s_more_io. We cannot simply splice s_more_io back to s_io as soon as s_io
becomes empty, and let the loop in generic_sync_sb_inodes() continue: this
may starve newly expired inodes in s_dirty. It is also not an option to
draw inodes from both s_more_io and s_dirty, an let the loop go on: this
might lead to live locks, and might also starve other superblocks in sync
time(well kupdate may still starve some superblocks, that's another bug).
We have to return when a full scan of s_io completes. So nr_to_write > 0
does not necessarily mean that "all data are written". This patch
introduces a flag writeback_control.more_io to indicate that more io should
be done. With it the big dirty file no longer has to wait for the next
kupdate invokation 5s later.
In sync_sb_inodes() we only set more_io on super_blocks we actually
visited. This avoids the interaction between two pdflush deamons.
Also in __sync_single_inode() we don't blindly keep requeuing the io if the
filesystem cannot progress. Failing to do so may lead to 100% iowait.
Tested-by: Mike Snitzer <snitzer@gmail.com>
Signed-off-by: Fengguang Wu <wfg@mail.ustc.edu.cn>
Cc: Michael Rubin <mrubin@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-04 22:29:36 -08:00
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
writeback: speed up writeback of big dirty files
After making dirty a 100M file, the normal behavior is to start the
writeback for all data after 30s delays. But sometimes the following
happens instead:
- after 30s: ~4M
- after 5s: ~4M
- after 5s: all remaining 92M
Some analyze shows that the internal io dispatch queues goes like this:
s_io s_more_io
-------------------------
1) 100M,1K 0
2) 1K 96M
3) 0 96M
1) initial state with a 100M file and a 1K file
2) 4M written, nr_to_write <= 0, so write more
3) 1K written, nr_to_write > 0, no more writes(BUG)
nr_to_write > 0 in (3) fools the upper layer to think that data have all
been written out. The big dirty file is actually still sitting in
s_more_io. We cannot simply splice s_more_io back to s_io as soon as s_io
becomes empty, and let the loop in generic_sync_sb_inodes() continue: this
may starve newly expired inodes in s_dirty. It is also not an option to
draw inodes from both s_more_io and s_dirty, an let the loop go on: this
might lead to live locks, and might also starve other superblocks in sync
time(well kupdate may still starve some superblocks, that's another bug).
We have to return when a full scan of s_io completes. So nr_to_write > 0
does not necessarily mean that "all data are written". This patch
introduces a flag writeback_control.more_io to indicate that more io should
be done. With it the big dirty file no longer has to wait for the next
kupdate invokation 5s later.
In sync_sb_inodes() we only set more_io on super_blocks we actually
visited. This avoids the interaction between two pdflush deamons.
Also in __sync_single_inode() we don't blindly keep requeuing the io if the
filesystem cannot progress. Failing to do so may lead to 100% iowait.
Tested-by: Mike Snitzer <snitzer@gmail.com>
Signed-off-by: Fengguang Wu <wfg@mail.ustc.edu.cn>
Cc: Michael Rubin <mrubin@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-04 22:29:36 -08:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
writeback: speed up writeback of big dirty files
After making dirty a 100M file, the normal behavior is to start the
writeback for all data after 30s delays. But sometimes the following
happens instead:
- after 30s: ~4M
- after 5s: ~4M
- after 5s: all remaining 92M
Some analyze shows that the internal io dispatch queues goes like this:
s_io s_more_io
-------------------------
1) 100M,1K 0
2) 1K 96M
3) 0 96M
1) initial state with a 100M file and a 1K file
2) 4M written, nr_to_write <= 0, so write more
3) 1K written, nr_to_write > 0, no more writes(BUG)
nr_to_write > 0 in (3) fools the upper layer to think that data have all
been written out. The big dirty file is actually still sitting in
s_more_io. We cannot simply splice s_more_io back to s_io as soon as s_io
becomes empty, and let the loop in generic_sync_sb_inodes() continue: this
may starve newly expired inodes in s_dirty. It is also not an option to
draw inodes from both s_more_io and s_dirty, an let the loop go on: this
might lead to live locks, and might also starve other superblocks in sync
time(well kupdate may still starve some superblocks, that's another bug).
We have to return when a full scan of s_io completes. So nr_to_write > 0
does not necessarily mean that "all data are written". This patch
introduces a flag writeback_control.more_io to indicate that more io should
be done. With it the big dirty file no longer has to wait for the next
kupdate invokation 5s later.
In sync_sb_inodes() we only set more_io on super_blocks we actually
visited. This avoids the interaction between two pdflush deamons.
Also in __sync_single_inode() we don't blindly keep requeuing the io if the
filesystem cannot progress. Failing to do so may lead to 100% iowait.
Tested-by: Mike Snitzer <snitzer@gmail.com>
Signed-off-by: Fengguang Wu <wfg@mail.ustc.edu.cn>
Cc: Michael Rubin <mrubin@google.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2008-02-04 22:29:36 -08:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-10 12:07:27 +02:00
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-09 17:20:03 -07:00
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2009-09-24 15:25:11 +02:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
2010-03-11 14:09:47 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:39 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-26 14:21:45 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 19:22:48 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2010-08-09 17:20:03 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2009-09-23 20:33:40 +08:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-02 09:19:46 +02:00
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2009-09-23 20:33:40 +08:00
|
|
|
|
|
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-01-06 14:40:25 -08:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:07 +10:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-06-10 12:07:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:07 +10:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-23 19:32:26 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-23 19:32:26 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-23 19:32:26 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2010-07-07 13:24:07 +10:00
|
|
|
|
2009-09-23 19:32:26 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-23 19:32:26 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2010-08-03 12:51:16 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-30 08:55:52 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-05-17 12:51:03 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-30 08:55:52 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-16 15:18:25 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-08-11 14:17:44 -07:00
|
|
|
|
2010-08-03 12:51:16 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-08-11 14:17:44 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-19 23:08:22 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-19 23:08:22 +02:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-26 14:22:45 -07:00
|
|
|
|
2010-06-19 23:08:22 +02:00
|
|
|
|
2010-07-25 14:29:18 +03:00
|
|
|
|
2010-06-19 23:08:22 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:18 +03:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:15 +03:00
|
|
|
|
2010-08-28 08:52:10 +02:00
|
|
|
|
2010-05-18 14:31:45 +02:00
|
|
|
|
2010-07-25 14:29:15 +03:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
writeback: move bdi threads exiting logic to the forker thread
Currently, bdi threads can decide to exit if there were no useful activities
for 5 minutes. However, this causes nasty races: we can easily oops in the
'bdi_queue_work()' if the bdi thread decides to exit while we are waking it up.
And even if we do not oops, but the bdi tread exits immediately after we wake
it up, we'd lose the wake-up event and have an unnecessary delay (up to 5 secs)
in the bdi work processing.
This patch makes the forker thread to be the central place which not only
creates bdi threads, but also kills them if they were inactive long enough.
This better design-wise.
Another reason why this change was done is to prepare for the further changes
which will prevent the bdi threads from waking up every 5 sec and wasting
power. Indeed, when the task does not wake up periodically anymore, it won't be
able to exit either.
This patch also moves the the 'wake_up_bit()' call from the bdi thread to the
forker thread as well. So now the forker thread sets the BDI_pending bit, then
forks the task or kills it, then clears the bit and wakes up the waiting
process.
The only process which may wain on the bit is 'bdi_wb_shutdown()'. This
function was changed as well - now it first removes the bdi from the
'bdi_list', then waits on the 'BDI_pending' bit. Once it wakes up, it is
guaranteed that the forker thread won't race with it, because the bdi is not
visible. Note, the forker thread sets the 'BDI_pending' bit under the
'bdi->wb_lock' which is essential for proper serialization.
And additionally, when we change 'bdi->wb.task', we now take the
'bdi->work_lock', to make sure that we do not lose wake-ups which we otherwise
would when raced with, say, 'bdi_queue_work()'.
Signed-off-by: Artem Bityutskiy <Artem.Bityutskiy@nokia.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <jaxboe@fusionio.com>
2010-07-25 14:29:20 +03:00
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-25 14:29:15 +03:00
|
|
|
|
2010-05-18 14:31:45 +02:00
|
|
|
|
2010-05-17 12:51:03 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
writeback: move bdi threads exiting logic to the forker thread
Currently, bdi threads can decide to exit if there were no useful activities
for 5 minutes. However, this causes nasty races: we can easily oops in the
'bdi_queue_work()' if the bdi thread decides to exit while we are waking it up.
And even if we do not oops, but the bdi tread exits immediately after we wake
it up, we'd lose the wake-up event and have an unnecessary delay (up to 5 secs)
in the bdi work processing.
This patch makes the forker thread to be the central place which not only
creates bdi threads, but also kills them if they were inactive long enough.
This better design-wise.
Another reason why this change was done is to prepare for the further changes
which will prevent the bdi threads from waking up every 5 sec and wasting
power. Indeed, when the task does not wake up periodically anymore, it won't be
able to exit either.
This patch also moves the the 'wake_up_bit()' call from the bdi thread to the
forker thread as well. So now the forker thread sets the BDI_pending bit, then
forks the task or kills it, then clears the bit and wakes up the waiting
process.
The only process which may wain on the bit is 'bdi_wb_shutdown()'. This
function was changed as well - now it first removes the bdi from the
'bdi_list', then waits on the 'BDI_pending' bit. Once it wakes up, it is
guaranteed that the forker thread won't race with it, because the bdi is not
visible. Note, the forker thread sets the 'BDI_pending' bit under the
'bdi->wb_lock' which is essential for proper serialization.
And additionally, when we change 'bdi->wb.task', we now take the
'bdi->work_lock', to make sure that we do not lose wake-ups which we otherwise
would when raced with, say, 'bdi_queue_work()'.
Signed-off-by: Artem Bityutskiy <Artem.Bityutskiy@nokia.com>
Reviewed-by: Christoph Hellwig <hch@lst.de>
Signed-off-by: Jens Axboe <jaxboe@fusionio.com>
2010-07-25 14:29:20 +03:00
|
|
|
|
2010-06-19 23:08:22 +02:00
|
|
|
|
|
|
|
|
|
2010-07-07 13:24:06 +10:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-19 23:08:22 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-08 18:15:07 +02:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-08 18:15:07 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-08 18:15:07 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-06-08 18:15:07 +02:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-06-08 18:15:07 +02:00
|
|
|
|
2009-09-14 13:12:40 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2009-09-14 13:12:40 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-23 15:19:20 -04:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-06-02 17:38:30 -04:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-09 09:10:25 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-10-21 11:49:30 +11:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-07-25 14:29:21 +03:00
|
|
|
|
|
|
|
|
|
2010-07-25 14:29:22 +03:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
2010-06-02 17:38:30 -04:00
|
|
|
|
2009-09-09 09:08:54 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
|
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:43 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2010-06-08 18:14:43 +02:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2010-06-08 18:14:51 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-05-17 12:55:07 +02:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-30 09:05:48 -07:00
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
2010-06-01 11:08:43 +02:00
|
|
|
|
2010-05-17 12:55:07 +02:00
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:51 +02:00
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
2010-06-08 18:14:51 +02:00
|
|
|
|
2009-12-23 07:57:07 -05:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-29 11:16:17 -04:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:43 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
2010-06-08 18:14:43 +02:00
|
|
|
|
|
|
|
|
|
2010-06-08 18:14:51 +02:00
|
|
|
|
|
|
|
|
|
2010-07-06 08:59:53 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-09-16 15:13:54 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
2009-09-02 12:34:32 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
[PATCH] fix nr_unused accounting, and avoid recursing in iput with I_WILL_FREE set
list_move(&inode->i_list, &inode_in_use);
} else {
list_move(&inode->i_list, &inode_unused);
+ inodes_stat.nr_unused++;
}
}
wake_up_inode(inode);
Are you sure the above diff is correct? It was added somewhere between
2.6.5 and 2.6.8. I think it's wrong.
The only way I can imagine the i_count to be zero in the above path, is
that I_WILL_FREE is set. And if I_WILL_FREE is set, then we must not
increase nr_unused. So I believe the above change is buggy and it will
definitely overstate the number of unused inodes and it should be backed
out.
Note that __writeback_single_inode before calling __sync_single_inode, can
drop the spinlock and we can have both the dirty and locked bitflags clear
here:
spin_unlock(&inode_lock);
__wait_on_inode(inode);
iput(inode);
XXXXXXX
spin_lock(&inode_lock);
}
use inode again here
a construct like the above makes zero sense from a reference counting
standpoint.
Either we don't ever use the inode again after the iput, or the
inode_lock should be taken _before_ executing the iput (i.e. a __iput
would be required). Taking the inode_lock after iput means the iget was
useless if we keep using the inode after the iput.
So the only chance the 2.6 was safe to call __writeback_single_inode
with the i_count == 0, is that I_WILL_FREE is set (I_WILL_FREE will
prevent the VM to free the inode in XXXXX).
Potentially calling the above iput with I_WILL_FREE was also wrong
because it would recurse in iput_final (the second mainline bug).
The below (untested) patch fixes the nr_unused accounting, avoids recursing
in iput when I_WILL_FREE is set and makes sure (with the BUG_ON) that we
don't corrupt memory and that all holders that don't set I_WILL_FREE, keeps
a reference on the inode!
Signed-off-by: Andrea Arcangeli <andrea@suse.de>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-30 15:03:05 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
[PATCH] fix nr_unused accounting, and avoid recursing in iput with I_WILL_FREE set
list_move(&inode->i_list, &inode_in_use);
} else {
list_move(&inode->i_list, &inode_unused);
+ inodes_stat.nr_unused++;
}
}
wake_up_inode(inode);
Are you sure the above diff is correct? It was added somewhere between
2.6.5 and 2.6.8. I think it's wrong.
The only way I can imagine the i_count to be zero in the above path, is
that I_WILL_FREE is set. And if I_WILL_FREE is set, then we must not
increase nr_unused. So I believe the above change is buggy and it will
definitely overstate the number of unused inodes and it should be backed
out.
Note that __writeback_single_inode before calling __sync_single_inode, can
drop the spinlock and we can have both the dirty and locked bitflags clear
here:
spin_unlock(&inode_lock);
__wait_on_inode(inode);
iput(inode);
XXXXXXX
spin_lock(&inode_lock);
}
use inode again here
a construct like the above makes zero sense from a reference counting
standpoint.
Either we don't ever use the inode again after the iput, or the
inode_lock should be taken _before_ executing the iput (i.e. a __iput
would be required). Taking the inode_lock after iput means the iget was
useless if we keep using the inode after the iput.
So the only chance the 2.6 was safe to call __writeback_single_inode
with the i_count == 0, is that I_WILL_FREE is set (I_WILL_FREE will
prevent the VM to free the inode in XXXXX).
Potentially calling the above iput with I_WILL_FREE was also wrong
because it would recurse in iput_final (the second mainline bug).
The below (untested) patch fixes the nr_unused accounting, avoids recursing
in iput when I_WILL_FREE is set and makes sure (with the BUG_ON) that we
don't corrupt memory and that all holders that don't set I_WILL_FREE, keeps
a reference on the inode!
Signed-off-by: Andrea Arcangeli <andrea@suse.de>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2005-10-30 15:03:05 -08:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2008-02-08 04:20:23 -08:00
|
|
|
|
[PATCH] writeback: fix range handling
When a writeback_control's `start' and `end' fields are used to
indicate a one-byte-range starting at file offset zero, the required
values of .start=0,.end=0 mean that the ->writepages() implementation
has no way of telling that it is being asked to perform a range
request. Because we're currently overloading (start == 0 && end == 0)
to mean "this is not a write-a-range request".
To make all this sane, the patch changes range of writeback_control.
So caller does: If it is calling ->writepages() to write pages, it
sets range (range_start/end or range_cyclic) always.
And if range_cyclic is true, ->writepages() thinks the range is
cyclic, otherwise it just uses range_start and range_end.
This patch does,
- Add LLONG_MAX, LLONG_MIN, ULLONG_MAX to include/linux/kernel.h
-1 is usually ok for range_end (type is long long). But, if someone did,
range_end += val; range_end is "val - 1"
u64val = range_end >> bits; u64val is "~(0ULL)"
or something, they are wrong. So, this adds LLONG_MAX to avoid nasty
things, and uses LLONG_MAX for range_end.
- All callers of ->writepages() sets range_start/end or range_cyclic.
- Fix updates of ->writeback_index. It seems already bit strange.
If it starts at 0 and ended by check of nr_to_write, this last
index may reduce chance to scan end of file. So, this updates
->writeback_index only if range_cyclic is true or whole-file is
scanned.
Signed-off-by: OGAWA Hirofumi <hirofumi@mail.parknet.co.jp>
Cc: Nathan Scott <nathans@sgi.com>
Cc: Anton Altaparmakov <aia21@cantab.net>
Cc: Steven French <sfrench@us.ibm.com>
Cc: "Vladimir V. Saveliev" <vs@namesys.com>
Signed-off-by: Andrew Morton <akpm@osdl.org>
Signed-off-by: Linus Torvalds <torvalds@osdl.org>
2006-06-23 02:03:26 -07:00
|
|
|
|
|
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2005-11-07 00:59:15 -08:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
2007-10-16 23:30:44 -07:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2009-06-08 13:35:40 +02:00
|
|
|
|
2005-04-16 15:20:36 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2010-10-06 10:48:20 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|