From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1754406AbbLIKgo (ORCPT ); Wed, 9 Dec 2015 05:36:44 -0500 Received: from mailout1.samsung.com ([203.254.224.24]:45619 "EHLO mailout1.samsung.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1753649AbbLIKgj (ORCPT ); Wed, 9 Dec 2015 05:36:39 -0500 X-AuditID: cbfee68d-f79646d000001355-35-5668043449fb From: Rahul Jain To: davem@davemloft.net, edumazet@google.com, ast@plumgrid.com, jiri@mellanox.com, daniel@iogearbox.net, makita.toshiaki@lab.ntt.co.jp, noureddine@arista.com, herbert@gondor.apana.org.au Cc: netdev@vger.kernel.org, linux-kernel@vger.kernel.org, k.ashutosh@samsung.com Subject: [PATCH] net : To avoid execution of extra instructions in NET RX path when rps_map is not set but rps_needed is true. Date: Wed, 09 Dec 2015 16:05:21 +0530 Message-id: <1449657321-3797-1-git-send-email-rahul.jain@samsung.com> X-Mailer: git-send-email 1.9.1 X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFnrHLMWRmVeSWpSXmKPExsWyRsSkWteEJSPMoOOIsEXf3R4Wi8ULvzFb zDnfwmLx9NgjdovuVzIWc16UWdxrncBmcXnXHDaL6T8fMlkcWyBmseDJW1YHbo/d05sYPbas vMnkse2AqseCTaUeXTcuMXs0v3jO4vFs+mEmj81HPzB69G1ZxejxeZNcAFcUl01Kak5mWWqR vl0CV8b8zjvsBRdVKpa397A3ML6R7mLk5JAQMJF4sHovG4QtJnHh3nogm4tDSGAFo8ScjZfZ YYreHzrIApGYxSixct5dqKrvjBLf5zxgAaliE9CUWHZ5IhNIQkRgH6NE67GVYHOZBQIkNl34 zwiSEBZoYJQ49u0nI0iCRUBVYs/SeUwgNq+Aq8S5xvOsEPvkJE4em8wK0iAhsI9douHpd1aI BgGJb5MPAa3jAErISmw6wAxRLylxcMUNlgmMggsYGVYxiqYWJBcUJ6UXGeoVJ+YWl+al6yXn 525iBEbD6X/Pencw3j5gfYhRgINRiYf3okt6mBBrYllxZe4hRlOgDROZpUST84Exl1cSb2hs ZmRhamJqbGRuaaYkzqso9TNYSCA9sSQ1OzW1ILUovqg0J7X4ECMTB6dUA6P21EgryUmlp1cY +3zmb+UoueGWlvtF+/2NH0mlpdJJUXUJOXrKnwpPmPnUd9n1lDj9WNnUJ2y/oelq1IWa9ZJR tfwGl1JqPlRkRN66oPlgo+2stdOWnd8qm/b020Ve3kXcDqYidVqRR76lloh0XVM8vu2++GqR 2WuzN5ed6Ik8fHznuYudTUosxRmJhlrMRcWJACiCbg2BAgAA X-Brightmail-Tracker: H4sIAAAAAAAAA+NgFtrNIsWRmVeSWpSXmKPExsVy+t9jQV0Tlowwg1s7uS367vawWCxe+I3Z Ys75FhaLp8cesVt0v5KxmPOizOJe6wQ2i8u75rBZTP/5kMni2AIxiwVP3rI6cHvsnt7E6LFl 5U0mj20HVD0WbCr16Lpxidmj+cVzFo9n0w8zeWw++oHRo2/LKkaPz5vkAriiGhhtMlITU1KL FFLzkvNTMvPSbZW8g+Od403NDAx1DS0tzJUU8hJzU22VXHwCdN0yc4DuVVIoS8wpBQoFJBYX K+nbYZoQGuKmawHTGKHrGxIE12NkgAYS1jBmzO+8w15wUaVieXsPewPjG+kuRk4OCQETifeH DrJA2GISF+6tZ+ti5OIQEpjFKLFy3l0o5zujxPc5D8Cq2AQ0JZZdnsgEkhAR2Mco0XpsJRtI glkgQGLThf+MIAlhgQZGiWPffjKCJFgEVCX2LJ3HBGLzCrhKnGs8zwqxT07i5LHJrBMYuRcw MqxilEgtSC4oTkrPNcxLLdcrTswtLs1L10vOz93ECI64Z1I7GA/ucj/EKMDBqMTDe8ElPUyI NbGsuDL3EKMEB7OSCK/tF6AQb0piZVVqUX58UWlOavEhRlOgAyYyS4km5wOTQV5JvKGxibmp samliYWJmaWSOG/tpcgwIYH0xJLU7NTUgtQimD4mDk6pBsZN6UE3O7kPzO2pOMFysYxPI3HJ yavNk05mzGUME1ErcPh/PVgg886z8tLMpQd7Vfety449cr6pMEFeoIlDIzo/5oSM+Qez769M HbwKbV9esDaysd3IySZywTks0FrwTrWMX8qaCom32xoOLnO/eirtjlng/QMvt7QvYuuwZL/K 7lrHyb3/rBJLcUaioRZzUXEiADry7onOAgAA DLP-Filter: Pass X-MTR: 20000000000000000@CPGS X-CFilter-Loop: Reflected Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org From: Ashutosh Kaushik The patch fixes the issues with check of global flag "rps_needed" in RX Path (which process packets in TCP/IP stack like netif_rx and netif_receive_skb functions) These functions have flag CONFIG RPS which is enabled default in kernel and to enter in RPS mode, it depends on variable rps_needed. This variable is updated whenever value in /sys/class/net//queues/rx-0/rps_cpus is being changed. There are 2 scenarios where it is executing extra piece of code even when results would be same every time:- 1) Suppose in system more than one networking devices are connected i.e. wired (eth0) and wireless (wlan0). If I enable RPS using above method for wlan0, then it sets atomic variable in rps_needed flag which is global. Now,whenever traffic uses wired network, then it will execute that extra piece of code in RX path because of only dependency on global rps_needed variable as below: #ifdef CONFIG_RPS if (static_key_false(&rps_needed)) { struct rps_dev_flow voidflow, *rflow = &voidflow; int cpu = get_rps_cpu(skb->dev, skb, &rflow); if (cpu >= 0) { ret = enqueue_to_backlog(skb, cpu, &rflow->last_qtail); rcu_read_unlock(); return ret; } } #endif while every time, value returned from get_rps_cpu will be < 0 as rps_map value is not set for this network device using /sys/class/net/device/queues/rx-0/rps_cpus. And it will every time execute get_rps_cpu function with fail case in IF condition which will be an extra overhead for packet processing. 2) Another scenario is as: Suppose that we have enable RPS for wireless device say wlan0 using above specified method which will set value of rps_needed. After doing test with set value of rps_cpus in sysfs, we do rmmod driver of wireless device. Next time again when we do insmod without rebooting system, it always hit below code with fail case: #ifdef CONFIG_RPS if (static_key_false(&rps_needed)) { struct rps_dev_flow voidflow, *rflow = &voidflow; int cpu; preempt_disable(); rcu_read_lock(); cpu = get_rps_cpu(skb->dev, skb, &rflow); if (cpu < 0) cpu = smp_processor_id(); ret = enqueue_to_backlog(skb, cpu, &rflow->last_qtail); rcu_read_unlock(); preempt_enable(); } else #endif The reason behind this overhead of hitting this code with false case is same because before doing rmmod, we enabled RPS which set rps_needed flag. Next time when we do insmod, it will just create entry for network device in sysfs with default value which is 0 for rps_cpus. It implies that RPS is disable for that device. But due to unchanged value of rps_needed variable, it goes into IF condition every time (which is failed in get_rps_cpu function) even rps_cpus is 0. Because if we do not enable RPS for that network device, rps_map is not set. The patch adds a check to these two RX functions which will check RPS availability locally for device specific. Signed-off-by: Ashutosh Kaushik --- net/core/dev.c | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/net/core/dev.c b/net/core/dev.c index 5df6cbc..1aa4402 100644 --- a/net/core/dev.c +++ b/net/core/dev.c @@ -3531,12 +3531,14 @@ drop: static int netif_rx_internal(struct sk_buff *skb) { int ret; + struct netdev_rx_queue *dev_rxqueue = skb->dev->_rx; net_timestamp_check(netdev_tstamp_prequeue, skb); trace_netif_rx(skb); #ifdef CONFIG_RPS - if (static_key_false(&rps_needed)) { + if (static_key_false(&rps_needed) && + dev_rxqueue->rps_map && dev_rxqueue->rps_map->len) { struct rps_dev_flow voidflow, *rflow = &voidflow; int cpu; @@ -3986,6 +3988,7 @@ static int __netif_receive_skb(struct sk_buff *skb) static int netif_receive_skb_internal(struct sk_buff *skb) { int ret; + struct netdev_rx_queue *dev_rxqueue = skb->dev->_rx; net_timestamp_check(netdev_tstamp_prequeue, skb); @@ -3995,7 +3998,8 @@ static int netif_receive_skb_internal(struct sk_buff *skb) rcu_read_lock(); #ifdef CONFIG_RPS - if (static_key_false(&rps_needed)) { + if (static_key_false(&rps_needed) && + dev_rxqueue->rps_map && dev_rxqueue->rps_map->len) { struct rps_dev_flow voidflow, *rflow = &voidflow; int cpu = get_rps_cpu(skb->dev, skb, &rflow); -- 1.9.1