Patch 1/3 1) No more one rwlock_t protecting the 'curtain' One major bottleneck on SMP machines is the fact that one rwlock is taken in ipt_do_table() entry and exit. That 2 atomic operations are the killer, and even if multiple readers can work together on the same table (using read_lock_bh()), the cache line containing rwlock still must be taken exclusively by each CPU at entry/exit of ipt_do_table(). As existing code already use separate copies of the data for each cpu, it is very easy to convert the central rwlock to separate rwlocks, allocated percpu (and NUMA aware). When a cpu enters ipt_do_table(), it can locks its local copy, touching a cache line that is not used by other cpus. Further operations are done on 'local' copy of the data. When a sum of all counters must be done, we can write_lock each part at a time, instead of locking all the parts, reducing the lock contention. Note1 : I also optimized get_counters(), using a SET_COUNTER() for the first cpu, avoiding a memset() and ADD_COUNTER() if SMP on other cpus. Signed-off-by: Eric Dumazet