From: Philipp Reisner <philipp.reisner@linbit.com>
To: Jens Axboe <axboe@fb.com>, linux-kernel@vger.kernel.org
Cc: drbd-dev@lists.linbit.com
Subject: [PATCH 37/38] drbd: avoid potential deadlock during handshake
Date: Wed, 25 Nov 2015 11:54:10 +0100 [thread overview]
Message-ID: <1448448851-10343-38-git-send-email-philipp.reisner@linbit.com> (raw)
In-Reply-To: <1448448851-10343-1-git-send-email-philipp.reisner@linbit.com>
From: Lars Ellenberg <lars.ellenberg@linbit.com>
During handshake communication, we also reconsider our device size,
using drbd_determine_dev_size(). Just in case we need to change the
offsets or layout of our on-disk metadata, we lock out application
and other meta data IO, and wait for the activity log to be "idle"
(no more referenced extents).
If this handshake happens just after a connection loss, with a fencing
policy of "resource-and-stonith", we have frozen IO.
If, additionally, the activity log was "starving" (too many incoming
random writes at that point in time), it won't become idle, ever,
because of the frozen IO, and this would be a lockup of the receiver
thread, and consquentially of DRBD.
Previous logic (re-)initialized with a special "empty" transaction
block, which required the activity log to fully drain first.
Instead, write out some standard activity log transactions.
Using lc_try_lock_for_transaction() instead of lc_try_lock() does not
care about pending activity log references, avoiding the potential
deadlock.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
---
drivers/block/drbd/drbd_actlog.c | 19 +++++++++++--------
drivers/block/drbd/drbd_int.h | 2 +-
drivers/block/drbd/drbd_nl.c | 33 +++++++++++++++++++--------------
3 files changed, 31 insertions(+), 23 deletions(-)
diff --git a/drivers/block/drbd/drbd_actlog.c b/drivers/block/drbd/drbd_actlog.c
index 4b484ac..10459a1 100644
--- a/drivers/block/drbd/drbd_actlog.c
+++ b/drivers/block/drbd/drbd_actlog.c
@@ -614,21 +614,24 @@ void drbd_al_shrink(struct drbd_device *device)
wake_up(&device->al_wait);
}
-int drbd_initialize_al(struct drbd_device *device, void *buffer)
+int drbd_al_initialize(struct drbd_device *device, void *buffer)
{
struct al_transaction_on_disk *al = buffer;
struct drbd_md *md = &device->ldev->md;
- sector_t al_base = md->md_offset + md->al_offset;
int al_size_4k = md->al_stripes * md->al_stripe_size_4k;
int i;
- memset(al, 0, 4096);
- al->magic = cpu_to_be32(DRBD_AL_MAGIC);
- al->transaction_type = cpu_to_be16(AL_TR_INITIALIZED);
- al->crc32c = cpu_to_be32(crc32c(0, al, 4096));
+ __al_write_transaction(device, al);
+ /* There may or may not have been a pending transaction. */
+ spin_lock_irq(&device->al_lock);
+ lc_committed(device->act_log);
+ spin_unlock_irq(&device->al_lock);
- for (i = 0; i < al_size_4k; i++) {
- int err = drbd_md_sync_page_io(device, device->ldev, al_base + i * 8, WRITE);
+ /* The rest of the transactions will have an empty "updates" list, and
+ * are written out only to provide the context, and to initialize the
+ * on-disk ring buffer. */
+ for (i = 1; i < al_size_4k; i++) {
+ int err = __al_write_transaction(device, al);
if (err)
return err;
}
diff --git a/drivers/block/drbd/drbd_int.h b/drivers/block/drbd/drbd_int.h
index df3d89d..b6844fe 100644
--- a/drivers/block/drbd/drbd_int.h
+++ b/drivers/block/drbd/drbd_int.h
@@ -1667,7 +1667,7 @@ extern int __drbd_change_sync(struct drbd_device *device, sector_t sector, int s
#define drbd_rs_failed_io(device, sector, size) \
__drbd_change_sync(device, sector, size, RECORD_RS_FAILED)
extern void drbd_al_shrink(struct drbd_device *device);
-extern int drbd_initialize_al(struct drbd_device *, void *);
+extern int drbd_al_initialize(struct drbd_device *, void *);
/* drbd_nl.c */
/* state info broadcast */
diff --git a/drivers/block/drbd/drbd_nl.c b/drivers/block/drbd/drbd_nl.c
index c7cd3df..f4ca273 100644
--- a/drivers/block/drbd/drbd_nl.c
+++ b/drivers/block/drbd/drbd_nl.c
@@ -903,15 +903,14 @@ drbd_determine_dev_size(struct drbd_device *device, enum dds_flags flags, struct
int md_moved, la_size_changed;
enum determine_dev_size rv = DS_UNCHANGED;
- /* race:
- * application request passes inc_ap_bio,
- * but then cannot get an AL-reference.
- * this function later may wait on ap_bio_cnt == 0. -> deadlock.
+ /* We may change the on-disk offsets of our meta data below. Lock out
+ * anything that may cause meta data IO, to avoid acting on incomplete
+ * layout changes or scribbling over meta data that is in the process
+ * of being moved.
*
- * to avoid that:
- * Suspend IO right here.
- * still lock the act_log to not trigger ASSERTs there.
- */
+ * Move is not exactly correct, btw, currently we have all our meta
+ * data in core memory, to "move" it we just write it all out, there
+ * are no reads. */
drbd_suspend_io(device);
buffer = drbd_md_get_buffer(device, __func__); /* Lock meta-data IO */
if (!buffer) {
@@ -919,9 +918,6 @@ drbd_determine_dev_size(struct drbd_device *device, enum dds_flags flags, struct
return DS_ERROR;
}
- /* no wait necessary anymore, actually we could assert that */
- wait_event(device->al_wait, lc_try_lock(device->act_log));
-
prev_first_sect = drbd_md_first_sector(device->ldev);
prev_size = device->ldev->md.md_size_sect;
la_size_sect = device->ldev->md.la_size_sect;
@@ -997,20 +993,29 @@ drbd_determine_dev_size(struct drbd_device *device, enum dds_flags flags, struct
* Clear the timer, to avoid scary "timer expired!" messages,
* "Superblock" is written out at least twice below, anyways. */
del_timer(&device->md_sync_timer);
- drbd_al_shrink(device); /* All extents inactive. */
+ /* We won't change the "al-extents" setting, we just may need
+ * to move the on-disk location of the activity log ringbuffer.
+ * Lock for transaction is good enough, it may well be "dirty"
+ * or even "starving". */
+ wait_event(device->al_wait, lc_try_lock_for_transaction(device->act_log));
+
+ /* mark current on-disk bitmap and activity log as unreliable */
prev_flags = md->flags;
- md->flags &= ~MDF_PRIMARY_IND;
+ md->flags |= MDF_FULL_SYNC | MDF_AL_DISABLED;
drbd_md_write(device, buffer);
+ drbd_al_initialize(device, buffer);
+
drbd_info(device, "Writing the whole bitmap, %s\n",
la_size_changed && md_moved ? "size changed and md moved" :
la_size_changed ? "size changed" : "md moved");
/* next line implicitly does drbd_suspend_io()+drbd_resume_io() */
drbd_bitmap_io(device, md_moved ? &drbd_bm_write_all : &drbd_bm_write,
"size changed", BM_LOCKED_MASK);
- drbd_initialize_al(device, buffer);
+ /* on-disk bitmap and activity log is authoritative again
+ * (unless there was an IO error meanwhile...) */
md->flags = prev_flags;
drbd_md_write(device, buffer);
--
1.9.1
next prev parent reply other threads:[~2015-11-25 11:09 UTC|newest]
Thread overview: 40+ messages / expand[flat|nested] mbox.gz Atom feed top
2015-11-25 10:53 [PATCH 00/38] DRBD update Philipp Reisner
2015-11-25 10:53 ` [PATCH 01/38] MAINTAINERS: Updated information for DRBD DRIVER Philipp Reisner
2015-11-25 10:53 ` [PATCH 02/38] drbd: Remove pointless check Philipp Reisner
2015-11-25 10:53 ` [PATCH 03/38] drbd: De-inline drbd_should_do_remote() and drbd_should_send_out_of_sync() Philipp Reisner
2015-11-25 10:53 ` [PATCH 04/38] drbd: Get rid of some first_peer_device() calls Philipp Reisner
2015-11-25 10:53 ` [PATCH 05/38] drbd: Move enum write_ordering_e to drbd.h Philipp Reisner
2015-11-25 10:53 ` [PATCH 06/38] drbd: drbd_adm_attach(): Add missing drbd_resync_after_changed() Philipp Reisner
2015-11-25 10:53 ` [PATCH 07/38] drbd: Fix locking across all resources Philipp Reisner
2015-11-25 10:53 ` [PATCH 08/38] drbd: Backport the "events2" command Philipp Reisner
2015-11-25 10:53 ` [PATCH 09/38] drbd: Backport the "status" command Philipp Reisner
2015-11-25 10:53 ` [PATCH 10/38] drbd: Deletion of an unnecessary check before the function call "lc_destroy" Philipp Reisner
2015-11-25 10:53 ` [PATCH 11/38] drbd: Replace 0 with the more meaningful GFP_NOWAIT Philipp Reisner
2015-11-25 10:53 ` [PATCH 12/38] drbd: Fix spurious disk-timeout Philipp Reisner
2015-11-25 10:53 ` [PATCH 13/38] drbd: drop remnants of connector -- we don't use it anymore in drbd 8.4 Philipp Reisner
2015-11-25 10:53 ` [PATCH 14/38] drbd: drbdsetup detach of an unresponsive local disk should not block IO "forever" Philipp Reisner
2015-11-25 10:53 ` [PATCH 15/38] drbd: also bump UUIDs if a diskless primary connects Philipp Reisner
2015-11-25 10:53 ` [PATCH 16/38] drbd: add comment why we want to first call local-io-error, then send state Philipp Reisner
2015-11-25 10:53 ` [PATCH 17/38] drbd: drbd_panic_after_delayed_completion_of_aborted_request() Philipp Reisner
2015-11-25 10:53 ` [PATCH 18/38] drbd: improve network timeout detection Philipp Reisner
2015-11-25 10:53 ` [PATCH 19/38] drbd: fix NULL deref in remember_new_state Philipp Reisner
2015-11-25 10:53 ` [PATCH 20/38] drbd: fix refcount error during detach of an already failed disk Philipp Reisner
2015-11-25 10:53 ` [PATCH 21/38] drbd: Rename asender to ack_receiver Philipp Reisner
2015-11-25 10:53 ` [PATCH 22/38] drbd: Create a dedicated workqueue for sending acks on the control connection Philipp Reisner
2015-11-25 10:53 ` [PATCH 23/38] drbd: prevent NULL pointer deref when resuming diskless primary Philipp Reisner
2015-11-25 10:53 ` [PATCH 24/38] drbd: debugfs: expose ed_data_gen_id Philipp Reisner
2015-11-25 10:53 ` [PATCH 25/38] drbd: use resource name in workqueue Philipp Reisner
2015-11-25 10:53 ` [PATCH 26/38] drbd: avoid redefinition of BITS_PER_PAGE Philipp Reisner
2015-11-25 10:54 ` [PATCH 27/38] drbd: use bitmap_weight() helper, don't open code Philipp Reisner
2015-11-25 10:54 ` [PATCH 28/38] drbd: fix spurious alert level printk Philipp Reisner
2015-11-25 10:54 ` [PATCH 29/38] drbd: fix queue limit setup for discard Philipp Reisner
2015-11-25 10:54 ` [PATCH 30/38] drbd: make drbd known to lsblk: use bd_link_disk_holder Philipp Reisner
2015-11-25 10:54 ` [PATCH 31/38] lru_cache: Converted lc_seq_printf_status to return void Philipp Reisner
2015-11-25 10:54 ` [PATCH 32/38] drbd: don't block forever in disconnect during resync if fencing=r-a-stonith Philipp Reisner
2015-11-25 10:54 ` [PATCH 33/38] drbd: fix memory leak in drbd_adm_resize Philipp Reisner
2015-11-25 10:54 ` [PATCH 34/38] drbd: fix "endless" transfer log walk in protocol A Philipp Reisner
2015-11-25 10:54 ` [PATCH 35/38] drbd: make suspend_io() / resume_io() must be thread and recursion safe Philipp Reisner
2015-11-25 10:54 ` [PATCH 36/38] drbd: separate out __al_write_transaction helper function Philipp Reisner
2015-11-25 10:54 ` Philipp Reisner [this message]
2015-11-25 10:54 ` [PATCH 38/38] drbd: fix error path during resize Philipp Reisner
2015-11-25 18:01 ` [PATCH 00/38] DRBD update Jens Axboe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1448448851-10343-38-git-send-email-philipp.reisner@linbit.com \
--to=philipp.reisner@linbit.com \
--cc=axboe@fb.com \
--cc=drbd-dev@lists.linbit.com \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®