From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org X-Spam-Level: X-Spam-Status: No, score=-17.2 required=3.0 tests=BAYES_00, HEADER_FROM_DIFFERENT_DOMAINS,INCLUDES_CR_TRAILER,INCLUDES_PATCH, MAILING_LIST_MULTI,NICE_REPLY_A,SPF_HELO_NONE,SPF_PASS,UNPARSEABLE_RELAY, USER_AGENT_SANE_1 autolearn=ham autolearn_force=no version=3.4.0 Received: from mail.kernel.org (mail.kernel.org [198.145.29.99]) by smtp.lore.kernel.org (Postfix) with ESMTP id 5342DC433EF for ; Fri, 10 Sep 2021 01:54:05 +0000 (UTC) Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by mail.kernel.org (Postfix) with ESMTP id 33B5B6113E for ; Fri, 10 Sep 2021 01:54:05 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S229665AbhIJBzO (ORCPT ); Thu, 9 Sep 2021 21:55:14 -0400 Received: from out30-133.freemail.mail.aliyun.com ([115.124.30.133]:32992 "EHLO out30-133.freemail.mail.aliyun.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S229648AbhIJBzM (ORCPT ); Thu, 9 Sep 2021 21:55:12 -0400 X-Alimail-AntiSpam: AC=PASS;BC=-1|-1;BR=01201311R781e4;CH=green;DM=||false|;DS=||;FP=0|-1|-1|-1|0|-1|-1|-1;HT=e01e04426;MF=joseph.qi@linux.alibaba.com;NM=1;PH=DS;RN=8;SR=0;TI=SMTPD_---0Unqrd8t_1631238837; Received: from B-D1K7ML85-0059.local(mailfrom:joseph.qi@linux.alibaba.com fp:SMTPD_---0Unqrd8t_1631238837) by smtp.aliyun-inc.com(127.0.0.1); Fri, 10 Sep 2021 09:53:58 +0800 Subject: Re: [Ocfs2-devel] [PATCH v2] ocfs2: Fix handle refcount leak in two exception handling paths To: Wengang Wang Cc: Chenyuan Mi , akpm , Xin Tan , Xiyu Yang , "yuanxzhang@fudan.edu.cn" , "linux-kernel@vger.kernel.org" , "ocfs2-devel@oss.oracle.com" References: <20210908102055.10168-1-cymi20@fudan.edu.cn> <06d9e055-29b9-731c-5a36-d888f2c83188@linux.alibaba.com> <6018AF95-3613-4D43-A3E6-7BAA0E0BE009@oracle.com> From: Joseph Qi Message-ID: Date: Fri, 10 Sep 2021 09:53:57 +0800 User-Agent: Mozilla/5.0 (Macintosh; Intel Mac OS X 10.15; rv:78.0) Gecko/20100101 Thunderbird/78.13.0 MIME-Version: 1.0 In-Reply-To: Content-Type: text/plain; charset=utf-8 Content-Language: en-US Content-Transfer-Encoding: 8bit Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 9/10/21 1:48 AM, Wengang Wang wrote: > > > On Sep 9, 2021, at 4:07 AM, Joseph Qi > wrote: > > Hi Wengang, > > On 9/9/21 1:12 AM, Wengang Wang wrote: > Hi, > > Sorry for late involving, but this doesn’t look right to me. > > On Sep 8, 2021, at 3:51 AM, Joseph Qi > wrote: > > > > On 9/8/21 6:20 PM, Chenyuan Mi wrote: > The reference counting issue happens in two exception handling paths > of ocfs2_replay_truncate_records(). When executing these two exception > handling paths, the function forgets to decrease the refcount of handle > increased by ocfs2_start_trans(), causing a refcount leak. > > Fix this issue by using ocfs2_commit_trans() to decrease the refcount > of handle in two handling paths. > > Signed-off-by: Chenyuan Mi > > Signed-off-by: Xiyu Yang > > Signed-off-by: Xin Tan > > > Reviewed-by: Joseph Qi > > --- > fs/ocfs2/alloc.c | 2 ++ > 1 file changed, 2 insertions(+) > > diff --git a/fs/ocfs2/alloc.c b/fs/ocfs2/alloc.c > index f1cc8258d34a..b05fde7edc3a 100644 > --- a/fs/ocfs2/alloc.c > +++ b/fs/ocfs2/alloc.c > @@ -5940,6 +5940,7 @@ static int ocfs2_replay_truncate_records(struct ocfs2_super *osb, > status = ocfs2_journal_access_di(handle, INODE_CACHE(tl_inode), tl_bh, > OCFS2_JOURNAL_ACCESS_WRITE); > if (status < 0) { > + ocfs2_commit_trans(osb, handle); > mlog_errno(status); > goto bail; > } > @@ -5964,6 +5965,7 @@ static int ocfs2_replay_truncate_records(struct ocfs2_super *osb, > data_alloc_bh, start_blk, > num_clusters); > if (status < 0) { > + ocfs2_commit_trans(osb, handle); > > As a transaction, stuff expected to be in the same handle should be treated as atomic. > Here the stuff includes the tl_bh and other metadata block which will be modified in ocfs2_free_clusters(). > Coming here, some of related meta blocks may be in the handle but others are not due to the error happened. > If you do a commit, partial meta blocks are committed to log. — that breaks the atomic idea, it will cause FS inconsistency. > So what’s reason you want to commit the meta block changes, which is not all of expected, in this handle to journal log? > > Do you really see a hit on the failure? or just you detected the refcount leak by code review? > > You may want to look at ocfs2_journal_dirty() for the error handling part. > > > For the first error handling, since we don't call ocfs2_journal_dirty() > yet, so won't be a problem. > For the second error handling, I think we don't have a better way. Look > at other callers of ocfs2_free_clusters(), we simply ignore the error > code. > Anyway, we should commit transaction if starts, otherwise journal will > be abnormal. > > I don't think so. If error happened, we should fail ocfs2, rather than do a partial committing. > Umm... not exactly... Take ocfs2_free_clusters() for example, when it fails in case of EIO or ENOMEM, we can't just abort journal in such cases, because it is not so serious, only a bit blocks still occupied and they will recovery during the next mount. That's why we have "errors=continue" in most filesystems, we should always consider the business continuity first. Also you can look at ext4_free_blocks() for reference. Thanks, Joseph