From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752730Ab3KSVPT (ORCPT ); Tue, 19 Nov 2013 16:15:19 -0500 Received: from shards.monkeyblade.net ([149.20.54.216]:58190 "EHLO shards.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751967Ab3KSVPO (ORCPT ); Tue, 19 Nov 2013 16:15:14 -0500 Date: Tue, 19 Nov 2013 16:15:11 -0500 (EST) Message-Id: <20131119.161511.2201597699288568150.davem@davemloft.net> To: avagin@openvz.org Cc: netdev@vger.kernel.org, linux-kernel@vger.kernel.org, criu@openvz.org, xemul@parallels.com, edumazet@google.com, kuznet@ms2.inr.ac.ru, jmorris@namei.org, yoshfuji@linux-ipv6.org, kaber@trash.net Subject: Re: [PATCH] tcp: don't update snd_nxt, when a socket is switched from repair mode From: David Miller In-Reply-To: <1384884606-9978-1-git-send-email-avagin@openvz.org> References: <1384884606-9978-1-git-send-email-avagin@openvz.org> X-Mailer: Mew version 6.5 on Emacs 24.1 / Mule 6.0 (HANACHIRUSATO) Mime-Version: 1.0 Content-Type: Text/Plain; charset=us-ascii Content-Transfer-Encoding: 7bit X-Greylist: Sender succeeded SMTP AUTH, not delayed by milter-greylist-4.5.1 (shards.monkeyblade.net [0.0.0.0]); Tue, 19 Nov 2013 13:15:14 -0800 (PST) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org From: Andrey Vagin Date: Tue, 19 Nov 2013 22:10:06 +0400 > snd_nxt must be updated synchronously with sk_send_head. Otherwise > tp->packets_out may be updated incorrectly, what may bring a kernel panic. > > Here is a kernel panic from my host. > [ 103.043194] BUG: unable to handle kernel NULL pointer dereference at 0000000000000048 > [ 103.044025] IP: [] tcp_rearm_rto+0xcf/0x150 > ... > [ 146.301158] Call Trace: > [ 146.301158] [] tcp_ack+0xcc0/0x12c0 > > Before this panic a tcp socket was restored. This socket had sent and > unsent data in the write queue. Sent data was restored in repair mode, > then the socket was switched from reapair mode and unsent data was > restored. After that the socket was switched back into repair mode. > > In that moment we had a socket where write queue looks like this: > snd_una snd_nxt write_seq > |_________|________| > | > sk_send_head > > After a second switching from repair mode the state of socket was > changed: > > snd_una snd_nxt, write_seq > |_________ ________| > | > sk_send_head > > This state is inconsistent, because snd_nxt and sk_send_head are not > synchronized. > > Bellow you can find a call trace, how packets_out can be incremented > twice for one skb, if snd_nxt and sk_send_head are not synchronized. > In this case packets_out will be always positive, even when > sk_write_queue is empty. > > tcp_write_wakeup > skb = tcp_send_head(sk); > tcp_fragment > if (!before(tp->snd_nxt, TCP_SKB_CB(buff)->end_seq)) > tcp_adjust_pcount(sk, skb, diff); > tcp_event_new_data_sent > tp->packets_out += tcp_skb_pcount(skb); > > I think update of snd_nxt isn't required, when a socket is switched from > repair mode. Because it's initialized in tcp_connect_init. Then when a > write queue is restored, snd_nxt is incremented in tcp_event_new_data_sent, > so it's always is in consistent state. > > I have checked, that the bug is not reproduced with this patch and > all tests about restoring tcp connections work fine. > > Signed-off-by: Andrey Vagin Applied and queued up for -stable, thank you.