From: Alasdair G Kergon <agk@redhat.com>
To: device-mapper development <dm-devel@redhat.com>
Cc: linux-kernel@vger.kernel.org
Subject: Re: [dm-devel] [PATCH] dm: noflush resizing (0/3)
Date: Thu, 25 Oct 2007 02:24:56 +0100 [thread overview]
Message-ID: <20071025012456.GK10006@agk.fab.redhat.com> (raw)
In-Reply-To: <471FB83D.4060307@ce.jp.nec.com>
On Wed, Oct 24, 2007 at 05:25:17PM -0400, Jun'ichi Nomura wrote:
> - For some device-mapper targets (multipath and mirror),
> the mapping table sometimes has to be replaced to cope with device
> failure.
> OTOH, device-mapper flushes all pending I/Os upon table replacement
> and may result in I/O errors, if there are device failures.
> 'noflush' suspend is used to let dm queue the pending I/Os
> instead of flushing them.
> Since it's not possible for user space program to tell whether
> the suspend could cause I/O error, they always use
> 'noflush' to suspend mirror/multipath targets.
>
> - Currently resizing is disabled for 'noflush' suspend.
> Resizing occurs in the course of table replacement.
> To resize the device under use, device-mapper needs to get its
> bdev inode. However, using bdget() in this case could cause deadlock
> by waiting for I_LOCK where an I/O process holding I_LOCK is
> waiting for completion of table replacement.
Before reviewing the details of the proposed workaround, I'd like to see
a deeper analysis of the problem to see that there isn't a cleaner way
to resolve this.
For example:
Question) What are the realistic situations we must support that lead to
a resize during table reload with I/O outstanding?
- The resize is the purpose of the reload; noflush is only set to avoid losing
I/O if a path should fail. So any outstanding I/O may be expected to be
consistent with both the old and new sizes of the device. E.g. If it's
beyond the end of a shrinking device and userspace cared about not
losing that I/O, it would have waited for that I/O to be flushed
*before* issuing the resize. If the I/O is beyond the end of the
existing device but within the new size, userspace would have waited for
the resize operation to complete before allowing the new I/O to be
issued.
=> Is it OK for device-mapper to handle the device size check
internally, rejecting any I/O that falls beyond the end of the table (it
already must do this lookup anyway), and to update the size recorded in
the inode later, after I/O is flowing through the device again, but (of
course) before reporting that the resize operation is complete?
I.e. does it eliminate deadlocks if the bdget() and i_size_write()
happen after the 'resume'?
Alasdair
--
agk@redhat.com
next prev parent reply other threads:[~2007-10-25 1:25 UTC|newest]
Thread overview: 7+ messages / expand[flat|nested] mbox.gz Atom feed top
2007-10-24 21:25 Jun'ichi Nomura
2007-10-25 1:24 ` Alasdair G Kergon [this message]
2007-10-25 14:18 ` [dm-devel] " Jun'ichi Nomura
2007-10-25 14:48 ` Alasdair G Kergon
2007-10-25 18:46 ` Jun'ichi Nomura
2007-10-25 23:51 ` Jun'ichi Nomura
2007-10-26 0:37 ` Alasdair G Kergon
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20071025012456.GK10006@agk.fab.redhat.com \
--to=agk@redhat.com \
--cc=dm-devel@redhat.com \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome