2019-06-04 10:11:33 +02:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2015-02-05 02:03:32 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-05-15 11:30:47 +08:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-26 07:20:34 -08:00
|
|
|
|
2017-02-04 01:27:20 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-12-10 16:33:11 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-02-09 14:04:03 +11:00
|
|
|
|
2015-06-06 22:07:23 +02:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-03-18 20:01:16 +11:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
2015-03-10 09:27:55 +11:00
|
|
|
|
2015-01-02 23:00:14 +01:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-02-05 02:03:34 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
rhashtable: use BIT(0) for locking.
As reported by Guenter Roeck, the new bit-locking using
BIT(1) doesn't work on the m68k architecture. m68k only requires
2-byte alignment for words and longwords, so there is only one
unused bit in pointers to structs - We current use two, one for the
NULLS marker at the end of the linked list, and one for the bit-lock
in the head of the list.
The two uses don't need to conflict as we never need the head of the
list to be a NULLS marker - the marker is only needed to check if an
object has moved to a different table, and the bucket head cannot
move. The NULLS marker is only needed in a ->next pointer.
As we already have different types for the bucket head pointer (struct
rhash_lock_head) and the ->next pointers (struct rhash_head), it is
fairly easy to treat the lsb differently in each.
So: Initialize buckets heads to NULL, and use the lsb for locking.
When loading the pointer from the bucket head, if it is NULL (ignoring
the lock big), report as being the expected NULLS marker.
When storing a value into a bucket head, if it is a NULLS marker,
store NULL instead.
And convert all places that used bit 1 for locking, to use bit 0.
Fixes: 8f0db018006a ("rhashtable: use bit_spin_locks to protect hash bucket.")
Reported-by: Guenter Roeck <linux@roeck-us.net>
Tested-by: Guenter Roeck <linux@roeck-us.net>
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-12 11:52:08 +10:00
|
|
|
|
2015-02-05 02:03:34 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2020-06-03 18:12:43 +10:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2020-06-03 18:12:43 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2020-06-03 18:12:43 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-14 13:57:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
2019-05-16 15:19:48 +08:00
|
|
|
|
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-03-24 00:50:27 +11:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-02-20 00:53:38 +01:00
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
2015-01-02 23:00:21 +01:00
|
|
|
|
2019-04-02 10:07:45 +11:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2019-04-11 18:43:06 -05:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-08-21 22:01:48 -07:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-08-21 22:01:48 -07:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
|
|
|
|
|
2017-06-07 22:47:13 -04:00
|
|
|
|
2015-03-14 13:57:22 +11:00
|
|
|
|
2015-01-02 23:00:21 +01:00
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2015-01-02 23:00:21 +01:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
2015-02-05 02:03:32 +01:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
2019-04-12 11:52:07 +10:00
|
|
|
|
2015-03-24 14:18:17 +01:00
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-04-12 11:52:08 +10:00
|
|
|
|
|
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-05 02:03:32 +01:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
2015-02-05 02:03:32 +01:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2019-04-12 11:52:08 +10:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-09-22 10:51:52 +02:00
|
|
|
|
2015-02-05 02:03:32 +01:00
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-04-12 11:52:08 +10:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2015-03-24 14:18:17 +01:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
2015-02-05 02:03:32 +01:00
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
|
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
2019-05-16 15:19:48 +08:00
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 14:18:17 +01:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-03-31 12:58:48 -07:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 09:53:17 +11:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-11 09:43:48 +11:00
|
|
|
|
2015-03-14 13:57:23 +11:00
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:25 +11:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2016-08-12 20:10:44 +02:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2016-08-12 20:10:44 +02:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:25 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
rhashtable: Fix race in rhashtable_destroy() and use regular work_struct
When we put our declared work task in the global workqueue with
schedule_delayed_work(), its delay parameter is always zero.
Therefore, we should define a regular work in rhashtable structure
instead of a delayed work.
By the way, we add a condition to check whether resizing functions
are NULL before cancelling the work, avoiding to cancel an
uninitialized work.
Lastly, while we wait for all work items we submitted before to run
to completion with cancel_delayed_work(), ht->mutex has been taken in
rhashtable_destroy(). Moreover, cancel_delayed_work() doesn't return
until all work items are accomplished, and when work items are
scheduled, the work's function - rht_deferred_worker() will be called.
However, as rht_deferred_worker() also needs to acquire the lock,
deadlock might happen at the moment as the lock is already held before.
So if the cancel work function is moved out of the lock covered scope,
this will avoid the deadlock.
Fixes: 97defe1 ("rhashtable: Per bucket locks & deferred expansion/shrinking")
Signed-off-by: Ying Xue <ying.xue@windriver.com>
Cc: Thomas Graf <tgraf@suug.ch>
Acked-by: Thomas Graf <tgraf@suug.ch>
Signed-off-by: David S. Miller <davem@davemloft.net>
2015-01-16 11:13:09 +08:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-02-04 07:33:22 +11:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-03-12 15:28:40 +01:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2015-03-24 20:42:19 +00:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
2019-03-21 09:39:52 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-03-24 00:50:26 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:28 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-12-03 20:41:29 +08:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:28 +11:00
|
|
|
|
|
|
|
|
|
2015-04-22 09:41:46 +02:00
|
|
|
|
|
|
|
|
|
2015-12-03 20:41:29 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:28 +11:00
|
|
|
|
2018-08-21 22:01:45 -07:00
|
|
|
|
2015-12-03 20:41:29 +08:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:28 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-12-03 20:41:29 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2015-12-03 20:41:29 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:28 +11:00
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-04-12 11:52:07 +10:00
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2017-04-16 02:55:09 +02:00
|
|
|
|
2019-04-12 11:52:08 +10:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-03-04 17:29:48 +02:00
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
2018-03-04 17:29:48 +02:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-04-12 11:52:08 +10:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
2016-08-24 12:31:31 +02:00
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-05-15 11:30:47 +08:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:28 +11:00
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2019-04-12 11:52:08 +10:00
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-04-12 11:52:08 +10:00
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2022-12-06 11:36:32 -10:00
|
|
|
|
rhashtable: use bit_spin_locks to protect hash bucket.
This patch changes rhashtables to use a bit_spin_lock on BIT(1) of the
bucket pointer to lock the hash chain for that bucket.
The benefits of a bit spin_lock are:
- no need to allocate a separate array of locks.
- no need to have a configuration option to guide the
choice of the size of this array
- locking cost is often a single test-and-set in a cache line
that will have to be loaded anyway. When inserting at, or removing
from, the head of the chain, the unlock is free - writing the new
address in the bucket head implicitly clears the lock bit.
For __rhashtable_insert_fast() we ensure this always happens
when adding a new key.
- even when lockings costs 2 updates (lock and unlock), they are
in a cacheline that needs to be read anyway.
The cost of using a bit spin_lock is a little bit of code complexity,
which I think is quite manageable.
Bit spin_locks are sometimes inappropriate because they are not fair -
if multiple CPUs repeatedly contend of the same lock, one CPU can
easily be starved. This is not a credible situation with rhashtable.
Multiple CPUs may want to repeatedly add or remove objects, but they
will typically do so at different buckets, so they will attempt to
acquire different locks.
As we have more bit-locks than we previously had spinlocks (by at
least a factor of two) we can expect slightly less contention to
go with the slightly better cache behavior and reduced memory
consumption.
To enhance type checking, a new struct is introduced to represent the
pointer plus lock-bit
that is stored in the bucket-table. This is "struct rhash_lock_head"
and is empty. A pointer to this needs to be cast to either an
unsigned lock, or a "struct rhash_head *" to be useful.
Variables of this type are most often called "bkt".
Previously "pprev" would sometimes point to a bucket, and sometimes a
->next pointer in an rhash_head. As these are now different types,
pprev is NULL when it would have pointed to the bucket. In that case,
'blk' is used, together with correct locking protocol.
Signed-off-by: NeilBrown <neilb@suse.com>
Signed-off-by: David S. Miller <davem@davemloft.net>
2019-04-02 10:07:45 +11:00
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-04-24 08:29:13 +10:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2015-12-16 16:45:54 +08:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2015-12-19 10:45:28 +08:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2015-12-16 16:45:54 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-02-14 22:03:27 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-12-16 16:45:54 +08:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
|
|
|
|
|
2015-12-16 16:45:54 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-12-04 10:31:41 -08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2017-09-19 12:41:37 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2021-07-07 18:07:31 -07:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2017-12-04 10:31:41 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2017-12-04 10:31:41 -08:00
|
|
|
|
2015-03-16 10:42:27 +01:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
2018-04-24 08:29:13 +10:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
2015-12-16 16:45:54 +08:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
2015-12-16 16:45:54 +08:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
|
|
|
|
|
2015-12-16 16:45:54 +08:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
2018-04-24 08:29:13 +10:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2018-04-24 08:29:13 +10:00
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-04-24 08:29:13 +10:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-07-02 09:35:34 -07:00
|
|
|
|
2018-04-24 08:29:13 +10:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2017-12-04 10:31:41 -08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-07-06 15:51:20 +02:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:19 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2015-05-05 02:22:53 +02:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2017-12-04 10:31:42 -08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-09-19 12:41:37 +02:00
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
2015-03-16 10:42:27 +01:00
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
|
|
|
|
|
2015-03-15 21:12:04 +11:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 09:53:17 +11:00
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
|
|
|
|
|
2016-08-18 16:50:56 +08:00
|
|
|
|
2019-03-21 14:42:40 +11:00
|
|
|
|
|
|
|
|
|
2015-03-24 09:53:17 +11:00
|
|
|
|
2015-03-14 13:57:20 +11:00
|
|
|
|
2015-03-15 21:12:04 +11:00
|
|
|
|
|
|
|
|
|
2015-02-04 07:33:23 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-20 21:56:59 +11:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2018-07-16 13:26:13 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:21 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-12-10 16:33:11 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-25 13:07:45 +00:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-12-10 16:33:11 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-20 21:56:59 +11:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:21 +11:00
|
|
|
|
2015-03-20 21:57:00 +11:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
2015-03-24 09:53:17 +11:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
|
|
|
|
|
2015-03-19 22:31:13 +00:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-04-27 13:44:51 +08:00
|
|
|
|
|
|
|
|
|
2017-04-28 14:10:48 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-04-27 13:44:51 +08:00
|
|
|
|
2017-05-01 22:18:01 +02:00
|
|
|
|
2015-03-19 22:31:13 +00:00
|
|
|
|
2018-07-16 13:26:13 -07:00
|
|
|
|
2015-12-16 18:13:14 +08:00
|
|
|
|
2015-03-24 00:50:21 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-08-21 22:01:48 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-03-24 00:50:27 +11:00
|
|
|
|
2018-08-21 22:01:48 -07:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-01-07 13:41:57 +08:00
|
|
|
|
2015-03-12 15:28:40 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
2015-02-25 16:31:54 +01:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
rhashtable: add restart routine in rhashtable_free_and_destroy()
rhashtable_free_and_destroy() cancels re-hash deferred work
then walks and destroys elements. at this moment, some elements can be
still in future_tbl. that elements are not destroyed.
test case:
nft_rhash_destroy() calls rhashtable_free_and_destroy() to destroy
all elements of sets before destroying sets and chains.
But rhashtable_free_and_destroy() doesn't destroy elements of future_tbl.
so that splat occurred.
test script:
%cat test.nft
table ip aa {
map map1 {
type ipv4_addr : verdict;
elements = {
0 : jump a0,
1 : jump a0,
2 : jump a0,
3 : jump a0,
4 : jump a0,
5 : jump a0,
6 : jump a0,
7 : jump a0,
8 : jump a0,
9 : jump a0,
}
}
chain a0 {
}
}
flush ruleset
table ip aa {
map map1 {
type ipv4_addr : verdict;
elements = {
0 : jump a0,
1 : jump a0,
2 : jump a0,
3 : jump a0,
4 : jump a0,
5 : jump a0,
6 : jump a0,
7 : jump a0,
8 : jump a0,
9 : jump a0,
}
}
chain a0 {
}
}
flush ruleset
%while :; do nft -f test.nft; done
Splat looks like:
[ 200.795603] kernel BUG at net/netfilter/nf_tables_api.c:1363!
[ 200.806944] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC KASAN PTI
[ 200.812253] CPU: 1 PID: 1582 Comm: nft Not tainted 4.17.0+ #24
[ 200.820297] Hardware name: To be filled by O.E.M. To be filled by O.E.M./Aptio CRB, BIOS 5.6.5 07/08/2015
[ 200.830309] RIP: 0010:nf_tables_chain_destroy.isra.34+0x62/0x240 [nf_tables]
[ 200.838317] Code: 43 50 85 c0 74 26 48 8b 45 00 48 8b 4d 08 ba 54 05 00 00 48 c7 c6 60 6d 29 c0 48 c7 c7 c0 65 29 c0 4c 8b 40 08 e8 58 e5 fd f8 <0f> 0b 48 89 da 48 b8 00 00 00 00 00 fc ff
[ 200.860366] RSP: 0000:ffff880118dbf4d0 EFLAGS: 00010282
[ 200.866354] RAX: 0000000000000061 RBX: ffff88010cdeaf08 RCX: 0000000000000000
[ 200.874355] RDX: 0000000000000061 RSI: 0000000000000008 RDI: ffffed00231b7e90
[ 200.882361] RBP: ffff880118dbf4e8 R08: ffffed002373bcfb R09: ffffed002373bcfa
[ 200.890354] R10: 0000000000000000 R11: ffffed002373bcfb R12: dead000000000200
[ 200.898356] R13: dead000000000100 R14: ffffffffbb62af38 R15: dffffc0000000000
[ 200.906354] FS: 00007fefc31fd700(0000) GS:ffff88011b800000(0000) knlGS:0000000000000000
[ 200.915533] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 200.922355] CR2: 0000557f1c8e9128 CR3: 0000000106880000 CR4: 00000000001006e0
[ 200.930353] Call Trace:
[ 200.932351] ? nf_tables_commit+0x26f6/0x2c60 [nf_tables]
[ 200.939525] ? nf_tables_setelem_notify.constprop.49+0x1a0/0x1a0 [nf_tables]
[ 200.947525] ? nf_tables_delchain+0x6e0/0x6e0 [nf_tables]
[ 200.952383] ? nft_add_set_elem+0x1700/0x1700 [nf_tables]
[ 200.959532] ? nla_parse+0xab/0x230
[ 200.963529] ? nfnetlink_rcv_batch+0xd06/0x10d0 [nfnetlink]
[ 200.968384] ? nfnetlink_net_init+0x130/0x130 [nfnetlink]
[ 200.975525] ? debug_show_all_locks+0x290/0x290
[ 200.980363] ? debug_show_all_locks+0x290/0x290
[ 200.986356] ? sched_clock_cpu+0x132/0x170
[ 200.990352] ? find_held_lock+0x39/0x1b0
[ 200.994355] ? sched_clock_local+0x10d/0x130
[ 200.999531] ? memset+0x1f/0x40
V2:
- free all tables requested by Herbert Xu
Signed-off-by: Taehee Yoo <ap420073@gmail.com>
Acked-by: Herbert Xu <herbert@gondor.apana.org.au>
Signed-off-by: David S. Miller <davem@davemloft.net>
2018-07-08 11:55:51 +09:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2015-02-25 16:31:54 +01:00
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
rhashtable: Fix race in rhashtable_destroy() and use regular work_struct
When we put our declared work task in the global workqueue with
schedule_delayed_work(), its delay parameter is always zero.
Therefore, we should define a regular work in rhashtable structure
instead of a delayed work.
By the way, we add a condition to check whether resizing functions
are NULL before cancelling the work, avoiding to cancel an
uninitialized work.
Lastly, while we wait for all work items we submitted before to run
to completion with cancel_delayed_work(), ht->mutex has been taken in
rhashtable_destroy(). Moreover, cancel_delayed_work() doesn't return
until all work items are accomplished, and when work items are
scheduled, the work's function - rht_deferred_worker() will be called.
However, as rht_deferred_worker() also needs to acquire the lock,
deadlock might happen at the moment as the lock is already held before.
So if the cancel work function is moved out of the lock covered scope,
this will avoid the deadlock.
Fixes: 97defe1 ("rhashtable: Per bucket locks & deferred expansion/shrinking")
Signed-off-by: Ying Xue <ying.xue@windriver.com>
Cc: Thomas Graf <tgraf@suug.ch>
Acked-by: Thomas Graf <tgraf@suug.ch>
Signed-off-by: David S. Miller <davem@davemloft.net>
2015-01-16 11:13:09 +08:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
rhashtable: add restart routine in rhashtable_free_and_destroy()
rhashtable_free_and_destroy() cancels re-hash deferred work
then walks and destroys elements. at this moment, some elements can be
still in future_tbl. that elements are not destroyed.
test case:
nft_rhash_destroy() calls rhashtable_free_and_destroy() to destroy
all elements of sets before destroying sets and chains.
But rhashtable_free_and_destroy() doesn't destroy elements of future_tbl.
so that splat occurred.
test script:
%cat test.nft
table ip aa {
map map1 {
type ipv4_addr : verdict;
elements = {
0 : jump a0,
1 : jump a0,
2 : jump a0,
3 : jump a0,
4 : jump a0,
5 : jump a0,
6 : jump a0,
7 : jump a0,
8 : jump a0,
9 : jump a0,
}
}
chain a0 {
}
}
flush ruleset
table ip aa {
map map1 {
type ipv4_addr : verdict;
elements = {
0 : jump a0,
1 : jump a0,
2 : jump a0,
3 : jump a0,
4 : jump a0,
5 : jump a0,
6 : jump a0,
7 : jump a0,
8 : jump a0,
9 : jump a0,
}
}
chain a0 {
}
}
flush ruleset
%while :; do nft -f test.nft; done
Splat looks like:
[ 200.795603] kernel BUG at net/netfilter/nf_tables_api.c:1363!
[ 200.806944] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC KASAN PTI
[ 200.812253] CPU: 1 PID: 1582 Comm: nft Not tainted 4.17.0+ #24
[ 200.820297] Hardware name: To be filled by O.E.M. To be filled by O.E.M./Aptio CRB, BIOS 5.6.5 07/08/2015
[ 200.830309] RIP: 0010:nf_tables_chain_destroy.isra.34+0x62/0x240 [nf_tables]
[ 200.838317] Code: 43 50 85 c0 74 26 48 8b 45 00 48 8b 4d 08 ba 54 05 00 00 48 c7 c6 60 6d 29 c0 48 c7 c7 c0 65 29 c0 4c 8b 40 08 e8 58 e5 fd f8 <0f> 0b 48 89 da 48 b8 00 00 00 00 00 fc ff
[ 200.860366] RSP: 0000:ffff880118dbf4d0 EFLAGS: 00010282
[ 200.866354] RAX: 0000000000000061 RBX: ffff88010cdeaf08 RCX: 0000000000000000
[ 200.874355] RDX: 0000000000000061 RSI: 0000000000000008 RDI: ffffed00231b7e90
[ 200.882361] RBP: ffff880118dbf4e8 R08: ffffed002373bcfb R09: ffffed002373bcfa
[ 200.890354] R10: 0000000000000000 R11: ffffed002373bcfb R12: dead000000000200
[ 200.898356] R13: dead000000000100 R14: ffffffffbb62af38 R15: dffffc0000000000
[ 200.906354] FS: 00007fefc31fd700(0000) GS:ffff88011b800000(0000) knlGS:0000000000000000
[ 200.915533] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 200.922355] CR2: 0000557f1c8e9128 CR3: 0000000106880000 CR4: 00000000001006e0
[ 200.930353] Call Trace:
[ 200.932351] ? nf_tables_commit+0x26f6/0x2c60 [nf_tables]
[ 200.939525] ? nf_tables_setelem_notify.constprop.49+0x1a0/0x1a0 [nf_tables]
[ 200.947525] ? nf_tables_delchain+0x6e0/0x6e0 [nf_tables]
[ 200.952383] ? nft_add_set_elem+0x1700/0x1700 [nf_tables]
[ 200.959532] ? nla_parse+0xab/0x230
[ 200.963529] ? nfnetlink_rcv_batch+0xd06/0x10d0 [nfnetlink]
[ 200.968384] ? nfnetlink_net_init+0x130/0x130 [nfnetlink]
[ 200.975525] ? debug_show_all_locks+0x290/0x290
[ 200.980363] ? debug_show_all_locks+0x290/0x290
[ 200.986356] ? sched_clock_cpu+0x132/0x170
[ 200.990352] ? find_held_lock+0x39/0x1b0
[ 200.994355] ? sched_clock_local+0x10d/0x130
[ 200.999531] ? memset+0x1f/0x40
V2:
- free all tables requested by Herbert Xu
Signed-off-by: Taehee Yoo <ap420073@gmail.com>
Acked-by: Herbert Xu <herbert@gondor.apana.org.au>
Signed-off-by: David S. Miller <davem@davemloft.net>
2018-07-08 11:55:51 +09:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-03-31 12:58:48 -07:00
|
|
|
|
2019-04-12 11:52:08 +10:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2016-09-19 19:00:09 +08:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
rhashtable: add restart routine in rhashtable_free_and_destroy()
rhashtable_free_and_destroy() cancels re-hash deferred work
then walks and destroys elements. at this moment, some elements can be
still in future_tbl. that elements are not destroyed.
test case:
nft_rhash_destroy() calls rhashtable_free_and_destroy() to destroy
all elements of sets before destroying sets and chains.
But rhashtable_free_and_destroy() doesn't destroy elements of future_tbl.
so that splat occurred.
test script:
%cat test.nft
table ip aa {
map map1 {
type ipv4_addr : verdict;
elements = {
0 : jump a0,
1 : jump a0,
2 : jump a0,
3 : jump a0,
4 : jump a0,
5 : jump a0,
6 : jump a0,
7 : jump a0,
8 : jump a0,
9 : jump a0,
}
}
chain a0 {
}
}
flush ruleset
table ip aa {
map map1 {
type ipv4_addr : verdict;
elements = {
0 : jump a0,
1 : jump a0,
2 : jump a0,
3 : jump a0,
4 : jump a0,
5 : jump a0,
6 : jump a0,
7 : jump a0,
8 : jump a0,
9 : jump a0,
}
}
chain a0 {
}
}
flush ruleset
%while :; do nft -f test.nft; done
Splat looks like:
[ 200.795603] kernel BUG at net/netfilter/nf_tables_api.c:1363!
[ 200.806944] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC KASAN PTI
[ 200.812253] CPU: 1 PID: 1582 Comm: nft Not tainted 4.17.0+ #24
[ 200.820297] Hardware name: To be filled by O.E.M. To be filled by O.E.M./Aptio CRB, BIOS 5.6.5 07/08/2015
[ 200.830309] RIP: 0010:nf_tables_chain_destroy.isra.34+0x62/0x240 [nf_tables]
[ 200.838317] Code: 43 50 85 c0 74 26 48 8b 45 00 48 8b 4d 08 ba 54 05 00 00 48 c7 c6 60 6d 29 c0 48 c7 c7 c0 65 29 c0 4c 8b 40 08 e8 58 e5 fd f8 <0f> 0b 48 89 da 48 b8 00 00 00 00 00 fc ff
[ 200.860366] RSP: 0000:ffff880118dbf4d0 EFLAGS: 00010282
[ 200.866354] RAX: 0000000000000061 RBX: ffff88010cdeaf08 RCX: 0000000000000000
[ 200.874355] RDX: 0000000000000061 RSI: 0000000000000008 RDI: ffffed00231b7e90
[ 200.882361] RBP: ffff880118dbf4e8 R08: ffffed002373bcfb R09: ffffed002373bcfa
[ 200.890354] R10: 0000000000000000 R11: ffffed002373bcfb R12: dead000000000200
[ 200.898356] R13: dead000000000100 R14: ffffffffbb62af38 R15: dffffc0000000000
[ 200.906354] FS: 00007fefc31fd700(0000) GS:ffff88011b800000(0000) knlGS:0000000000000000
[ 200.915533] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 200.922355] CR2: 0000557f1c8e9128 CR3: 0000000106880000 CR4: 00000000001006e0
[ 200.930353] Call Trace:
[ 200.932351] ? nf_tables_commit+0x26f6/0x2c60 [nf_tables]
[ 200.939525] ? nf_tables_setelem_notify.constprop.49+0x1a0/0x1a0 [nf_tables]
[ 200.947525] ? nf_tables_delchain+0x6e0/0x6e0 [nf_tables]
[ 200.952383] ? nft_add_set_elem+0x1700/0x1700 [nf_tables]
[ 200.959532] ? nla_parse+0xab/0x230
[ 200.963529] ? nfnetlink_rcv_batch+0xd06/0x10d0 [nfnetlink]
[ 200.968384] ? nfnetlink_net_init+0x130/0x130 [nfnetlink]
[ 200.975525] ? debug_show_all_locks+0x290/0x290
[ 200.980363] ? debug_show_all_locks+0x290/0x290
[ 200.986356] ? sched_clock_cpu+0x132/0x170
[ 200.990352] ? find_held_lock+0x39/0x1b0
[ 200.994355] ? sched_clock_local+0x10d/0x130
[ 200.999531] ? memset+0x1f/0x40
V2:
- free all tables requested by Herbert Xu
Signed-off-by: Taehee Yoo <ap420073@gmail.com>
Acked-by: Herbert Xu <herbert@gondor.apana.org.au>
Signed-off-by: David S. Miller <davem@davemloft.net>
2018-07-08 11:55:51 +09:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
rhashtable: add restart routine in rhashtable_free_and_destroy()
rhashtable_free_and_destroy() cancels re-hash deferred work
then walks and destroys elements. at this moment, some elements can be
still in future_tbl. that elements are not destroyed.
test case:
nft_rhash_destroy() calls rhashtable_free_and_destroy() to destroy
all elements of sets before destroying sets and chains.
But rhashtable_free_and_destroy() doesn't destroy elements of future_tbl.
so that splat occurred.
test script:
%cat test.nft
table ip aa {
map map1 {
type ipv4_addr : verdict;
elements = {
0 : jump a0,
1 : jump a0,
2 : jump a0,
3 : jump a0,
4 : jump a0,
5 : jump a0,
6 : jump a0,
7 : jump a0,
8 : jump a0,
9 : jump a0,
}
}
chain a0 {
}
}
flush ruleset
table ip aa {
map map1 {
type ipv4_addr : verdict;
elements = {
0 : jump a0,
1 : jump a0,
2 : jump a0,
3 : jump a0,
4 : jump a0,
5 : jump a0,
6 : jump a0,
7 : jump a0,
8 : jump a0,
9 : jump a0,
}
}
chain a0 {
}
}
flush ruleset
%while :; do nft -f test.nft; done
Splat looks like:
[ 200.795603] kernel BUG at net/netfilter/nf_tables_api.c:1363!
[ 200.806944] invalid opcode: 0000 [#1] SMP DEBUG_PAGEALLOC KASAN PTI
[ 200.812253] CPU: 1 PID: 1582 Comm: nft Not tainted 4.17.0+ #24
[ 200.820297] Hardware name: To be filled by O.E.M. To be filled by O.E.M./Aptio CRB, BIOS 5.6.5 07/08/2015
[ 200.830309] RIP: 0010:nf_tables_chain_destroy.isra.34+0x62/0x240 [nf_tables]
[ 200.838317] Code: 43 50 85 c0 74 26 48 8b 45 00 48 8b 4d 08 ba 54 05 00 00 48 c7 c6 60 6d 29 c0 48 c7 c7 c0 65 29 c0 4c 8b 40 08 e8 58 e5 fd f8 <0f> 0b 48 89 da 48 b8 00 00 00 00 00 fc ff
[ 200.860366] RSP: 0000:ffff880118dbf4d0 EFLAGS: 00010282
[ 200.866354] RAX: 0000000000000061 RBX: ffff88010cdeaf08 RCX: 0000000000000000
[ 200.874355] RDX: 0000000000000061 RSI: 0000000000000008 RDI: ffffed00231b7e90
[ 200.882361] RBP: ffff880118dbf4e8 R08: ffffed002373bcfb R09: ffffed002373bcfa
[ 200.890354] R10: 0000000000000000 R11: ffffed002373bcfb R12: dead000000000200
[ 200.898356] R13: dead000000000100 R14: ffffffffbb62af38 R15: dffffc0000000000
[ 200.906354] FS: 00007fefc31fd700(0000) GS:ffff88011b800000(0000) knlGS:0000000000000000
[ 200.915533] CS: 0010 DS: 0000 ES: 0000 CR0: 0000000080050033
[ 200.922355] CR2: 0000557f1c8e9128 CR3: 0000000106880000 CR4: 00000000001006e0
[ 200.930353] Call Trace:
[ 200.932351] ? nf_tables_commit+0x26f6/0x2c60 [nf_tables]
[ 200.939525] ? nf_tables_setelem_notify.constprop.49+0x1a0/0x1a0 [nf_tables]
[ 200.947525] ? nf_tables_delchain+0x6e0/0x6e0 [nf_tables]
[ 200.952383] ? nft_add_set_elem+0x1700/0x1700 [nf_tables]
[ 200.959532] ? nla_parse+0xab/0x230
[ 200.963529] ? nfnetlink_rcv_batch+0xd06/0x10d0 [nfnetlink]
[ 200.968384] ? nfnetlink_net_init+0x130/0x130 [nfnetlink]
[ 200.975525] ? debug_show_all_locks+0x290/0x290
[ 200.980363] ? debug_show_all_locks+0x290/0x290
[ 200.986356] ? sched_clock_cpu+0x132/0x170
[ 200.990352] ? find_held_lock+0x39/0x1b0
[ 200.994355] ? sched_clock_local+0x10d/0x130
[ 200.999531] ? memset+0x1f/0x40
V2:
- free all tables requested by Herbert Xu
Signed-off-by: Taehee Yoo <ap420073@gmail.com>
Acked-by: Herbert Xu <herbert@gondor.apana.org.au>
Signed-off-by: David S. Miller <davem@davemloft.net>
2018-07-08 11:55:51 +09:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2015-01-02 23:00:20 +01:00
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2015-03-24 14:18:20 +01:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2014-08-02 11:47:44 +02:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2020-06-03 18:12:43 +10:00
|
|
|
|
2017-02-25 22:39:50 +08:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-02-25 22:39:50 +08:00
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
|
|
|
|
|
2019-04-02 10:07:45 +11:00
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
2019-04-02 10:07:45 +11:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
2020-07-24 20:14:34 +10:00
|
|
|
|
|
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2020-06-03 18:12:43 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
2018-06-18 12:52:50 +10:00
|
|
|
|
2017-02-11 19:26:47 +08:00
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|