mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Philipp Reisner <philipp.reisner@linbit.com>
To: Jens Axboe <axboe@fb.com>, linux-kernel@vger.kernel.org
Cc: drbd-dev@lists.linbit.com
Subject: [PATCH 23/38] drbd: prevent NULL pointer deref when resuming diskless primary
Date: Wed, 25 Nov 2015 11:53:56 +0100	[thread overview]
Message-ID: <1448448851-10343-24-git-send-email-philipp.reisner@linbit.com> (raw)
In-Reply-To: <1448448851-10343-1-git-send-email-philipp.reisner@linbit.com>

From: Lars Ellenberg <lars.ellenberg@linbit.com>

In a multiple error scenario, we may end up with a "frozen" Primary,
that has no access to any data (no local disk, no replication link).

If we then resume-io, we try to generate a new data generation id,
which will fail if there is no longer a local disk.

Double check for available local data,
which prevents the NULL pointer deref.

If we are diskless, turn the resume-io in this situation
into the first stage of a "force down", by bumping the "effective" data
gen id, which will prevent later attach or connect to the former data
set without first being demoted (deconfigured).

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
---
 drivers/block/drbd/drbd_nl.c | 25 ++++++++++++++++++++++++-
 1 file changed, 24 insertions(+), 1 deletion(-)

diff --git a/drivers/block/drbd/drbd_nl.c b/drivers/block/drbd/drbd_nl.c
index f35cefb..5e4adff 100644
--- a/drivers/block/drbd/drbd_nl.c
+++ b/drivers/block/drbd/drbd_nl.c
@@ -2920,7 +2920,30 @@ int drbd_adm_resume_io(struct sk_buff *skb, struct genl_info *info)
 	mutex_lock(&adm_ctx.resource->adm_mutex);
 	device = adm_ctx.device;
 	if (test_bit(NEW_CUR_UUID, &device->flags)) {
-		drbd_uuid_new_current(device);
+		if (get_ldev_if_state(device, D_ATTACHING)) {
+			drbd_uuid_new_current(device);
+			put_ldev(device);
+		} else {
+			/* This is effectively a multi-stage "forced down".
+			 * The NEW_CUR_UUID bit is supposedly only set, if we
+			 * lost the replication connection, and are configured
+			 * to freeze IO and wait for some fence-peer handler.
+			 * So we still don't have a replication connection.
+			 * And now we don't have a local disk either.  After
+			 * resume, we will fail all pending and new IO, because
+			 * we don't have any data anymore.  Which means we will
+			 * eventually be able to terminate all users of this
+			 * device, and then take it down.  By bumping the
+			 * "effective" data uuid, we make sure that you really
+			 * need to tear down before you reconfigure, we will
+			 * the refuse to re-connect or re-attach (because no
+			 * matching real data uuid exists).
+			 */
+			u64 val;
+			get_random_bytes(&val, sizeof(u64));
+			drbd_set_ed_uuid(device, val);
+			drbd_warn(device, "Resumed without access to data; please tear down before attempting to re-configure.\n");
+		}
 		clear_bit(NEW_CUR_UUID, &device->flags);
 	}
 	drbd_suspend_io(device);
-- 
1.9.1


  parent reply	other threads:[~2015-11-25 11:13 UTC|newest]

Thread overview: 40+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2015-11-25 10:53 [PATCH 00/38] DRBD update Philipp Reisner
2015-11-25 10:53 ` [PATCH 01/38] MAINTAINERS: Updated information for DRBD DRIVER Philipp Reisner
2015-11-25 10:53 ` [PATCH 02/38] drbd: Remove pointless check Philipp Reisner
2015-11-25 10:53 ` [PATCH 03/38] drbd: De-inline drbd_should_do_remote() and drbd_should_send_out_of_sync() Philipp Reisner
2015-11-25 10:53 ` [PATCH 04/38] drbd: Get rid of some first_peer_device() calls Philipp Reisner
2015-11-25 10:53 ` [PATCH 05/38] drbd: Move enum write_ordering_e to drbd.h Philipp Reisner
2015-11-25 10:53 ` [PATCH 06/38] drbd: drbd_adm_attach(): Add missing drbd_resync_after_changed() Philipp Reisner
2015-11-25 10:53 ` [PATCH 07/38] drbd: Fix locking across all resources Philipp Reisner
2015-11-25 10:53 ` [PATCH 08/38] drbd: Backport the "events2" command Philipp Reisner
2015-11-25 10:53 ` [PATCH 09/38] drbd: Backport the "status" command Philipp Reisner
2015-11-25 10:53 ` [PATCH 10/38] drbd: Deletion of an unnecessary check before the function call "lc_destroy" Philipp Reisner
2015-11-25 10:53 ` [PATCH 11/38] drbd: Replace 0 with the more meaningful GFP_NOWAIT Philipp Reisner
2015-11-25 10:53 ` [PATCH 12/38] drbd: Fix spurious disk-timeout Philipp Reisner
2015-11-25 10:53 ` [PATCH 13/38] drbd: drop remnants of connector -- we don't use it anymore in drbd 8.4 Philipp Reisner
2015-11-25 10:53 ` [PATCH 14/38] drbd: drbdsetup detach of an unresponsive local disk should not block IO "forever" Philipp Reisner
2015-11-25 10:53 ` [PATCH 15/38] drbd: also bump UUIDs if a diskless primary connects Philipp Reisner
2015-11-25 10:53 ` [PATCH 16/38] drbd: add comment why we want to first call local-io-error, then send state Philipp Reisner
2015-11-25 10:53 ` [PATCH 17/38] drbd: drbd_panic_after_delayed_completion_of_aborted_request() Philipp Reisner
2015-11-25 10:53 ` [PATCH 18/38] drbd: improve network timeout detection Philipp Reisner
2015-11-25 10:53 ` [PATCH 19/38] drbd: fix NULL deref in remember_new_state Philipp Reisner
2015-11-25 10:53 ` [PATCH 20/38] drbd: fix refcount error during detach of an already failed disk Philipp Reisner
2015-11-25 10:53 ` [PATCH 21/38] drbd: Rename asender to ack_receiver Philipp Reisner
2015-11-25 10:53 ` [PATCH 22/38] drbd: Create a dedicated workqueue for sending acks on the control connection Philipp Reisner
2015-11-25 10:53 ` Philipp Reisner [this message]
2015-11-25 10:53 ` [PATCH 24/38] drbd: debugfs: expose ed_data_gen_id Philipp Reisner
2015-11-25 10:53 ` [PATCH 25/38] drbd: use resource name in workqueue Philipp Reisner
2015-11-25 10:53 ` [PATCH 26/38] drbd: avoid redefinition of BITS_PER_PAGE Philipp Reisner
2015-11-25 10:54 ` [PATCH 27/38] drbd: use bitmap_weight() helper, don't open code Philipp Reisner
2015-11-25 10:54 ` [PATCH 28/38] drbd: fix spurious alert level printk Philipp Reisner
2015-11-25 10:54 ` [PATCH 29/38] drbd: fix queue limit setup for discard Philipp Reisner
2015-11-25 10:54 ` [PATCH 30/38] drbd: make drbd known to lsblk: use bd_link_disk_holder Philipp Reisner
2015-11-25 10:54 ` [PATCH 31/38] lru_cache: Converted lc_seq_printf_status to return void Philipp Reisner
2015-11-25 10:54 ` [PATCH 32/38] drbd: don't block forever in disconnect during resync if fencing=r-a-stonith Philipp Reisner
2015-11-25 10:54 ` [PATCH 33/38] drbd: fix memory leak in drbd_adm_resize Philipp Reisner
2015-11-25 10:54 ` [PATCH 34/38] drbd: fix "endless" transfer log walk in protocol A Philipp Reisner
2015-11-25 10:54 ` [PATCH 35/38] drbd: make suspend_io() / resume_io() must be thread and recursion safe Philipp Reisner
2015-11-25 10:54 ` [PATCH 36/38] drbd: separate out __al_write_transaction helper function Philipp Reisner
2015-11-25 10:54 ` [PATCH 37/38] drbd: avoid potential deadlock during handshake Philipp Reisner
2015-11-25 10:54 ` [PATCH 38/38] drbd: fix error path during resize Philipp Reisner
2015-11-25 18:01 ` [PATCH 00/38] DRBD update Jens Axboe

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=1448448851-10343-24-git-send-email-philipp.reisner@linbit.com \
    --to=philipp.reisner@linbit.com \
    --cc=axboe@fb.com \
    --cc=drbd-dev@lists.linbit.com \
    --cc=linux-kernel@vger.kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®