From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1763115AbXGVTF7 (ORCPT ); Sun, 22 Jul 2007 15:05:59 -0400 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S1760130AbXGVTFq (ORCPT ); Sun, 22 Jul 2007 15:05:46 -0400 Received: from [212.12.190.174] ([212.12.190.174]:33970 "EHLO raad.intranet" rhost-flags-FAIL-FAIL-OK-FAIL) by vger.kernel.org with ESMTP id S1759610AbXGVTFp (ORCPT ); Sun, 22 Jul 2007 15:05:45 -0400 From: Al Boldi To: linux-fsdevel@vger.kernel.org Subject: Re: [RFH] Partition table recovery Date: Sun, 22 Jul 2007 22:05:53 +0300 User-Agent: KMail/1.5 Cc: linux-raid@vger.kernel.org, linux-kernel@vger.kernel.org References: <200707200813.03553.a1426z@gawab.com> <200707220710.31402.a1426z@gawab.com> <20070722162802.GA20174@thunk.org> In-Reply-To: <20070722162802.GA20174@thunk.org> MIME-Version: 1.0 Content-Type: text/plain; charset="iso-8859-1" Content-Transfer-Encoding: 7bit Content-Disposition: inline Message-Id: <200707222205.53556.a1426z@gawab.com> Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org Theodore Tso wrote: > On Sun, Jul 22, 2007 at 07:10:31AM +0300, Al Boldi wrote: > > Sounds great, but it may be advisable to hook this into the partition > > modification routines instead of mkfs/fsck. Which would mean that the > > partition manager could ask the kernel to instruct its fs subsystem to > > update the backup partition table for each known fs-type that supports > > such a feature. > > Well, let's think about this a bit. What are the requirements? > > 1) The partition manager should be able explicitly request that a new > backup of the partition tables be stashed in each filesystem that has > room for such a backup. That way, when the user affirmatively makes a > partition table change, it can get backed up in all of the right > places automatically. > > 2) The fsck program should *only* stash a backup of the partition > table if there currently isn't one in the filesystem. It may be that > the partition table has been corrupted, and so merely doing an fsck > should not transfer a current copy of the partition table to the > filesystem-secpfic backup area. It could be that the partition table > was only partially recovered, and we don't want to overwrite the > previously existing backups except on an explicit request from the > system administrator. > > 3) The mkfs program should automatically create a backup of the > current partition table layout. That way we get a backup in the newly > created filesystem as soon as it is created. > > 4) The exact location of the backup may vary from filesystem to > filesystem. For ext2/3/4, bytes 512-1023 are always unused, and don't > interfere with the boot sector at bytes 0-511, so that's the obvious > location. Other filesystems may have that location in use, and some > other location might be a better place to store it. Ideally it will > be a well-known location, that isn't dependent on finding an inode > table, or some such, but that may not be possible for all filesystems. > > OK, so how about this as a solution that meets the above requirements? > > /sbin/partbackup [] > > Will scan (i.e., /dev/hda, /dev/sdb, etc.) and create > a 512 byte partition backup, using the format I've previously > described. If is specified on the command line, it > will use the blkid library to determine the filesystem type of > , and then attempt to execute > /dev/partbackupfs. to write the partition backup to > . If is '-', then it will write the 512 byte > partition table to stdout. If is not specified on > the command line, /sbin/partbackup will iterate over all > partitions in , use the blkid library to attempt to > determine the correct filesystem type, and then execute > /sbin/partbackupfs. if such a backup program exists. > > /sbin/partbackupfs. > > ... is a filesystem specific program for filesystem type > . It will assure that (i.e., /dev/hda1, > /dev/sdb3) is of an appropriate filesystem type, and then read > 512 bytes from stdin and write it out to to an > appropriate place for that filesystem. > > Partition managers will be encouraged to check to see if > /sbin/partbackup exists, and if so, after the partition table is > written, will check to see if /sbin/partbackup exists, and if so, to > call it with just one argument (i.e., /sbin/partbackup /dev/hdb). > They SHOULD provide an option for the user to suppress the backup from > happening, but the backup should be the default behavior. > > An /etc/mkfs. program is encouraged to run /sbin/partbackup > with two arguments (i.e., /sbin/partbackup /dev/hdb /dev/hdb3) when > creating a filesystem. > > An /etc/fsck. program is encouraged to check to see if a > partition backup exists (assuming the filesystem supports it), and if > not, call /sbin/partbackup with two arguments. > > A filesystem utility package for a particular filesystem type is > encouraged to make the above changes to its mkfs and fsck programs, as > well as provide an /sbin/partbackupfs. program. Great! > I would do this all in userspace, though. Is there any reason to get > the kernel involved? I don't think so. Yes, doing things in userspace, when possible, is much better. But, a change in the partition table has to be relayed to the kernel, and when that change happens to be on a mounted disk, then the partition manager complains of not being able to update the kernel's view. So how can this be addressed? Thanks! -- Al