From: Zhihao Cheng <chengzhihao1@huawei.com>
To: <yukuai@fygo.io>, <song@kernel.org>, <shli@kernel.org>,
<neil@brown.name>
Cc: <linux-raid@vger.kernel.org>, <linux-kernel@vger.kernel.org>,
<yangerkun@huawei.com>, <yi.zhang@huawei.com>
Subject: Re: [PATCH v3 1/4] md/raid5: Hide the origin mddev->thread before takeover
Date: Fri, 9 Oct 2026 15:12:32 +0800 [thread overview]
Message-ID: <aaa9d0ee-99ca-55d9-ced4-953d25190d97@huawei.com> (raw)
In-Reply-To: <1194977a-ef25-4ad3-9bec-df27b6adc4b4@fygo.io>
在 2026/10/9 12:31, yu kuai 写道:
> Hi,
>
> 在 2026/9/23 19:21, Zhihao Cheng 写道:
>> The raid5 takeover invokes setup_conf and allocates strip heads, but
>> it wakes up the wrong thread, which lefts strip heads in the list
>> 'conf->released_stripes' and not being processed. If raid5_run fails,
>> the strip heads won't be released, which triggers the following slab
>> warnings (CONFIG_SLUB_DEBUG):
>> BUG raid5-md0 (Not tainted): Objects remaining on __kmem_cache_shutdown()
>> Object 0x0000000062fad548 @offset=3968
>> Object 0x000000007f74683c @offset=4960
>> WARNING: mm/slub.c:1268 at __slab_err+0x31/0x40, CPU#0: bash/865
>> RIP: 0010:__slab_err+0x31
>> Call Trace:
>> __kmem_cache_shutdown.cold+0x15b
>> kmem_cache_destroy+0x71
>> free_conf+0xf8
>> raid5_run.cold+0x463
>> level_store+0x64e
>> md_attr_store+0xd7
>
> Is this still a problem with following patch?
>
> [PATCH] md/raid5: drain released_stripes before destroying the cache -
> Li Youhong
> <https://lore.kernel.org/all/20260929083916.50620-1-dayou5941@163.com/>
Hi,
Youhong's patch could fix the problem, please take it, his solution is
better.
>
>>
>> The detailed triggering process is as follows:
>> mdadm --create /dev/md0 --level=1 --raid-devices=2 /dev/sda /dev/sdb
>> --force --assume-clean # create raid1, mddev->thread is raid1d
>> echo 5 > /sys/block/md0/md/level
>> level_store
>> raid5_takeover_raid1
>> setup_conf
>> grow_stripes
>> grow_one_stripe
>> sh = alloc_stripe
>> raid5_release_stripe
>> md_wakeup_thread(conf->mddev->thread) // wakeup raid1d
>> raid5_run
>> ENOMEM = raid5_create_ctx_pool
>> free_conf
>> shrink_stripes
>> drop_one_stripe // no strips found from the conf->inactive_list
>> kmem_cache_destroy(conf->slab_cache)
>> __kmem_cache_shutdown
>> free_partial
>> list_slab_objects // some entries are not released !
>>
>> Fix it by hiding the origin mddev->thread before takeover, so that
>> new allocating strip heads can be put into 'conf->inactive_list',
>> which can be found by drop_one_stripe().
>>
>> Fixes: 773ca82fa1ee ("raid5: make release_stripe lockless")
>> Signed-off-by: Zhihao Cheng <chengzhihao1@huawei.com>
>> ---
>> drivers/md/raid5.c | 37 ++++++++++++++++++++++++++++---------
>> 1 file changed, 28 insertions(+), 9 deletions(-)
>>
>> diff --git a/drivers/md/raid5.c b/drivers/md/raid5.c
>> index c091bba95c31..7e87e8a60f5f 100644
>> --- a/drivers/md/raid5.c
>> +++ b/drivers/md/raid5.c
>> @@ -9039,19 +9039,38 @@ static void *raid5_takeover(struct mddev *mddev)
>> * raid4 - trivial - just use a raid4 layout.
>> * raid6 - Providing it is a *_6 layout
>> */
>> - if (mddev->level == 0)
>> - return raid45_takeover_raid0(mddev, 5);
>> - if (mddev->level == 1)
>> - return raid5_takeover_raid1(mddev);
>> - if (mddev->level == 4) {
>> + void *ret = ERR_PTR(-EINVAL);
>> + struct md_thread *thread;
>> +
>> + thread = rcu_dereference_protected(mddev->thread,
>> + lockdep_is_held(&mddev->reconfig_mutex));
>> + /*
>> + * Set mddev->thread to NULL before setup_conf() to avoid waking up
>> + * wrong thread(eg. raid1), which can prevent the strips from being
>> + * left unreleased in the error handling path(free_conf) of raid5_run.
>> + */
>> + rcu_assign_pointer(mddev->thread, NULL);
>> +
>> + switch (mddev->level) {
>> + case 0:
>> + ret = raid45_takeover_raid0(mddev, 5);
>> + break;
>> + case 1:
>> + ret = raid5_takeover_raid1(mddev);
>> + break;
>> + case 4:
>> mddev->new_layout = ALGORITHM_PARITY_N;
>> mddev->new_level = 5;
>> - return setup_conf(mddev);
>> + ret = setup_conf(mddev);
>> + break;
>> + case 6:
>> + ret = raid5_takeover_raid6(mddev);
>> + break;
>> }
>> - if (mddev->level == 6)
>> - return raid5_takeover_raid6(mddev);
>>
>> - return ERR_PTR(-EINVAL);
>> + rcu_assign_pointer(mddev->thread, thread);
>> +
>> + return ret;
>> }
>>
>> static void *raid4_takeover(struct mddev *mddev)
>
next prev parent reply other threads:[~2026-10-09 7:12 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-23 11:21 [PATCH v3 0/4] md: Fix two raid5/raid10 bugs Zhihao Cheng
2026-09-23 11:21 ` [PATCH v3 1/4] md/raid5: Hide the origin mddev->thread before takeover Zhihao Cheng
2026-10-09 4:31 ` yu kuai
2026-10-09 7:12 ` Zhihao Cheng [this message]
2026-09-23 11:21 ` [PATCH v3 2/4] md: Handle pers->run failure in level_store Zhihao Cheng
2026-10-09 5:05 ` yu kuai
2026-09-23 11:21 ` [PATCH v3 3/4] md/raid5: Don't free conf on raid5_run failure Zhihao Cheng
2026-09-23 11:21 ` [PATCH v3 4/4] md/raid10: Don't free conf on raid10_run failure Zhihao Cheng
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aaa9d0ee-99ca-55d9-ced4-953d25190d97@huawei.com \
--to=chengzhihao1@huawei.com \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-raid@vger.kernel.org \
--cc=neil@brown.name \
--cc=shli@kernel.org \
--cc=song@kernel.org \
--cc=yangerkun@huawei.com \
--cc=yi.zhang@huawei.com \
--cc=yukuai@fygo.io \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®