From: Ben Hutchings <ben@decadent.org.uk>
To: linux-kernel@vger.kernel.org, stable@vger.kernel.org
Cc: akpm@linux-foundation.org, "NeilBrown" <neilb@suse.com>
Subject: [PATCH 3.2 47/60] md/raid10: ensure device failure recorded before write request returns.
Date: Sun, 15 Nov 2015 01:45:45 +0000 [thread overview]
Message-ID: <lsq.1447551945.119881246@decadent.org.uk> (raw)
In-Reply-To: <lsq.1447551944.536641563@decadent.org.uk>
3.2.73-rc1 review patch. If anyone has any objections, please let me know.
------------------
From: NeilBrown <neilb@suse.com>
commit 95af587e95aacb9cfda4a9641069a5244a540dc8 upstream.
When a write to one of the legs of a RAID10 fails, the failure is
recorded in the metadata of the other legs so that after a restart
the data on the failed drive wont be trusted even if that drive seems
to be working again (maybe a cable was unplugged).
Currently there is no interlock between the write request completing
and the metadata update. So it is possible that the write will
complete, the app will confirm success in some way, and then the
machine will crash before the metadata update completes.
This is an extremely small hole for a racy to fit in, but it is
theoretically possible and so should be closed.
So:
- set MD_CHANGE_PENDING when requesting a metadata update for a
failed device, so we can know with certainty when it completes
- queue requests that experienced an error on a new queue which
is only processed after the metadata update completes
- call raid_end_bio_io() on bios in that queue when the time comes.
Signed-off-by: NeilBrown <neilb@suse.com>
[bwh: Backported to 3.2: adjust context]
Signed-off-by: Ben Hutchings <ben@decadent.org.uk>
---
drivers/md/raid10.c | 29 ++++++++++++++++++++++++++++-
drivers/md/raid10.h | 6 ++++++
2 files changed, 34 insertions(+), 1 deletion(-)
--- a/drivers/md/raid10.c
+++ b/drivers/md/raid10.c
@@ -1280,6 +1280,7 @@ static void error(struct mddev *mddev, s
set_bit(Blocked, &rdev->flags);
set_bit(Faulty, &rdev->flags);
set_bit(MD_CHANGE_DEVS, &mddev->flags);
+ set_bit(MD_CHANGE_PENDING, &mddev->flags);
printk(KERN_ALERT
"md/raid10:%s: Disk failure on %s, disabling device.\n"
"md/raid10:%s: Operation continuing on %d devices.\n",
@@ -2215,6 +2216,7 @@ static void handle_write_completed(struc
}
put_buf(r10_bio);
} else {
+ bool fail = false;
for (m = 0; m < conf->copies; m++) {
int dev = r10_bio->devs[m].devnum;
struct bio *bio = r10_bio->devs[m].bio;
@@ -2227,6 +2229,7 @@ static void handle_write_completed(struc
rdev_dec_pending(rdev, conf->mddev);
} else if (bio != NULL &&
!test_bit(BIO_UPTODATE, &bio->bi_flags)) {
+ fail = true;
if (!narrow_write_error(r10_bio, m)) {
md_error(conf->mddev, rdev);
set_bit(R10BIO_Degraded,
@@ -2238,7 +2241,13 @@ static void handle_write_completed(struc
if (test_bit(R10BIO_WriteError,
&r10_bio->state))
close_write(r10_bio);
- raid_end_bio_io(r10_bio);
+ if (fail) {
+ spin_lock_irq(&conf->device_lock);
+ list_add(&r10_bio->retry_list, &conf->bio_end_io_list);
+ spin_unlock_irq(&conf->device_lock);
+ md_wakeup_thread(conf->mddev->thread);
+ } else
+ raid_end_bio_io(r10_bio);
}
}
@@ -2252,6 +2261,23 @@ static void raid10d(struct mddev *mddev)
md_check_recovery(mddev);
+ if (!list_empty_careful(&conf->bio_end_io_list) &&
+ !test_bit(MD_CHANGE_PENDING, &mddev->flags)) {
+ LIST_HEAD(tmp);
+ spin_lock_irqsave(&conf->device_lock, flags);
+ if (!test_bit(MD_CHANGE_PENDING, &mddev->flags)) {
+ list_add(&tmp, &conf->bio_end_io_list);
+ list_del_init(&conf->bio_end_io_list);
+ }
+ spin_unlock_irqrestore(&conf->device_lock, flags);
+ while (!list_empty(&tmp)) {
+ r10_bio = list_first_entry(&conf->bio_end_io_list,
+ struct r10bio, retry_list);
+ list_del(&r10_bio->retry_list);
+ raid_end_bio_io(r10_bio);
+ }
+ }
+
blk_start_plug(&plug);
for (;;) {
@@ -2860,6 +2886,7 @@ static struct r10conf *setup_conf(struct
spin_lock_init(&conf->device_lock);
INIT_LIST_HEAD(&conf->retry_list);
+ INIT_LIST_HEAD(&conf->bio_end_io_list);
spin_lock_init(&conf->resync_lock);
init_waitqueue_head(&conf->wait_barrier);
--- a/drivers/md/raid10.h
+++ b/drivers/md/raid10.h
@@ -40,6 +40,12 @@ struct r10conf {
sector_t chunk_mask;
struct list_head retry_list;
+ /* A separate list of r1bio which just need raid_end_bio_io called.
+ * This mustn't happen for writes which had any errors if the superblock
+ * needs to be written.
+ */
+ struct list_head bio_end_io_list;
+
/* queue pending writes and submit them on unplug */
struct bio_list pending_bio_list;
int pending_count;
next prev parent reply other threads:[~2015-11-15 2:05 UTC|newest]
Thread overview: 64+ messages / expand[flat|nested] mbox.gz Atom feed top
2015-11-15 1:45 [PATCH 3.2 00/60] 3.2.73-rc1 review Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 48/60] md/raid10: don't clear bitmap bit when bad-block-list write fails Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 11/60] x86/process: Add proper bound checks in 64bit get_wchan() Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 18/60] iio: accel: sca3000: memory corruption in sca3000_read_first_n_hw_rb() Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 01/60] Revert "KVM: MMU: fix validation of mmio page fault" Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 46/60] md/raid1: don't clear bitmap bit when bad-block-list write fails Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 51/60] net: add length argument to skb_copy_and_csum_datagram_iovec Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 60/60] KEYS: Fix crash when attempt to garbage collect an uninstantiated keyring Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 43/60] dm btree remove: fix a bug when rebalancing nodes after removal Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 17/60] clocksource: Fix abs() usage w/ 64bit values Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 32/60] xhci: don't finish a TD if we get a short transfer event mid TD Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 56/60] asix: Do full reset during ax88772_bind Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 16/60] md/raid0: apply base queue limits *before* disk_stack_limits Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 52/60] skbuff: Fix skb checksum flag on skb pull Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 41/60] mm: make sendfile(2) killable Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 21/60] tty: fix stall caused by missing memory barrier in drivers/tty/n_tty.c Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 58/60] KVM: x86: work around infinite loop in microcode when #AC is delivered Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 36/60] crypto: api - Only abort operations on fatal signal Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 13/60] mm: hugetlbfs: skip shared VMAs when unmapping private pages to satisfy a fault Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 42/60] ppp: fix pppoe_dev deletion condition in pppoe_release() Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 24/60] iwlwifi: dvm: fix D3 firmware PN programming Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 15/60] md/raid0: update queue parameter in a safer location Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 04/60] regmap: debugfs: Don't bother actually printing when calculating max length Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 44/60] dm btree: fix leak of bufio-backed block in btree_split_beneath error path Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 09/60] UBI: return ENOSPC if no enough space available Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 14/60] [SMB3] Do not fall back to SMBWriteX in set_file_size error cases Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 20/60] usb: Add device quirk for Logitech PTZ cameras Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 26/60] sched/core: Fix TASK_DEAD race in finish_task_switch() Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 06/60] m68k: Define asmlinkage_protect Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 12/60] genirq: Fix race in register_irq_proc() Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 19/60] USB: Add reset-resume quirk for two Plantronics usb headphones Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 53/60] skbuff: Fix skb checksum partial check Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 50/60] sched: declare pid_alive as inline Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 25/60] ALSA: synth: Fix conflicting OSS device registration on AWE32 Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 45/60] md/raid1: ensure device failure recorded before write request returns Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 10/60] MIPS: dma-default: Fix 32-bit fall back to GFP_DMA Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 49/60] mvsas: Fix NULL pointer dereference in mvs_slot_task_free Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 39/60] drm/nouveau/gem: return only valid domain when there's only one Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 27/60] 3w-9xxx: don't unmap bounce buffered commands Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 37/60] ASoC: wm8904: Correct number of EQ registers Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 22/60] drivers/tty: require read access for controlling terminal Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 31/60] iommu/vt-d: fix range computation when making room for large pages Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 30/60] crypto: ahash - ensure statesize is non-zero Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 07/60] x86/xen: Do not clip xen_e820_map to xen_e820_map_entries when sanitizing map Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 35/60] xhci: Add spurious wakeup quirk for LynxPoint-LP controllers Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 34/60] xhci: Switch Intel Lynx Point LP ports to EHCI on shutdown Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 38/60] IB/cm: Fix rb-tree duplicate free and use-after-free Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 29/60] ALSA: hda - Fix inverted internal mic on Lenovo G50-80 Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 05/60] ath9k: declare required extra tx headroom Ben Hutchings
2015-11-15 1:45 ` Ben Hutchings [this message]
2015-11-15 1:45 ` [PATCH 3.2 33/60] xhci: handle no ping response error properly Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 40/60] powerpc/rtas: Validate rtas.entry before calling enter_rtas() Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 03/60] regmap: debugfs: Ensure we don't underflow when printing access masks Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 08/60] UBI: Validate data_size Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 02/60] module: Fix locking in symbol_put_addr() Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 59/60] KEYS: Fix race between key destruction and finding a keyring by name Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 23/60] ppp: don't override sk->sk_state in pppoe_flush_dev() Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 57/60] Failing to send a CLOSE if file is opened WRONLY and server reboots on a 4.x mount Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 28/60] xen-blkfront: check for null drvdata in blkback_changed (XenbusStateClosing) Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 55/60] asix: Don't reset PHY on if_up for ASIX 88772 Ben Hutchings
2015-11-15 1:45 ` [PATCH 3.2 54/60] ethtool: Use kcalloc instead of kmalloc for ethtool_get_strings Ben Hutchings
2015-11-15 2:29 ` [PATCH 3.2 00/60] 3.2.73-rc1 review Ben Hutchings
2015-11-15 13:42 ` Guenter Roeck
2015-11-16 11:11 ` Ben Hutchings
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=lsq.1447551945.119881246@decadent.org.uk \
--to=ben@decadent.org.uk \
--cc=akpm@linux-foundation.org \
--cc=linux-kernel@vger.kernel.org \
--cc=neilb@suse.com \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®