From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261265AbVABQQ6 (ORCPT ); Sun, 2 Jan 2005 11:16:58 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261269AbVABQQ5 (ORCPT ); Sun, 2 Jan 2005 11:16:57 -0500 Received: from open.hands.com ([195.224.53.39]:51874 "EHLO open.hands.com") by vger.kernel.org with ESMTP id S261265AbVABQQx (ORCPT ); Sun, 2 Jan 2005 11:16:53 -0500 Date: Sun, 2 Jan 2005 16:26:52 +0000 From: Luke Kenneth Casson Leighton To: linux-kernel@vger.kernel.org, xen-devel@lists.sf.net Subject: [XEN] using shmfs for swapspace Message-ID: <20050102162652.GA12268@lkcl.net> Mime-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline User-Agent: Mutt/1.5.5.1+cvs20040105i X-hands-com-MailScanner: Found to be clean X-MailScanner-From: lkcl@lkcl.net Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org hi, am starting to play with XEN - the virtualisation project (http://xen.sf.net). i'll give some background first of all and then the question - at the bottom - will make sense [when posting to lkml i often get questions asked that are answered by the background material i also provide... *sigh*] each virtual machine requires (typically) its own physical ram (a chunk of the host's real memory) and some virtual memory - swapspace. xen uses 32mb for its shm guest OS inter-communication. so, in the case i'm setting up, that's 5 virtual machines (only one of which can get away with having only 32mb of ram, the rest require 64mb) so that's five lots of 256mbyte swap files. the memory usage is the major concern: i only have 256mb of ram and you've probably by now added up that the above comes to 320mbytes. so i started looking at ways to minimise the memory usage. first, reducing each machine to only having 32mb of ram, and secondly, on the host, creating a MASSIVE swap file (1gbyte), making a MASSIVE shmfs/tmpfs partition (1gbyte) and then creating swap files in the tmpfs partition!!! the reasoning behind doing this is quite straightforward: by placing the swapfiles in a tmpfs, presumably then when one of the guest OSes requires some memory, then RAM on the host OS will be used until such time as the amount of RAM requested exceeds the host OSes physical memory, and then it will go into swap-space. this is presumed to be infinitely better than forcing the swapspace to be always on disk, especially with the guests only being allocated 32mbyte of physical RAM. here's the problems: 1) tmpfs doesn't support sparse files 2) files created in tmpfs don't support block devices (???) 3) as a workaround i have to create a swap partition in a 256mb file, (dd if=/dev/zero of=/mnt/swapfile bs=1M count=256 and do mkswap on it) then copy the ENTIRE file into the tmpfs-mounted partition. on every boot-up. per swapfile needed. eeeuw, yuk. so, my question is a strategic one: * in what other ways could the same results be achieved? in other words, what other ways can i publish block devices from the master OS (and they must be block devices for XEN guest OSes to be able to see them) that can be used as swap space, that will be in RAM if possible, bearing in mind that they can be recreated at boot time, i.e. they don't need to be persistent. ta, l. -- -- http://lkcl.net -- From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261824AbVACSj5 (ORCPT ); Mon, 3 Jan 2005 13:39:57 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261784AbVACSf6 (ORCPT ); Mon, 3 Jan 2005 13:35:58 -0500 Received: from dhcp93115068.columbus.rr.com ([24.93.115.68]:6411 "EHLO nineveh.rivenstone.net") by vger.kernel.org with ESMTP id S261775AbVACSbk convert rfc822-to-8bit (ORCPT ); Mon, 3 Jan 2005 13:31:40 -0500 Date: Mon, 3 Jan 2005 13:31:34 -0500 From: Joseph Fannin To: Luke Kenneth Casson Leighton Cc: linux-kernel@vger.kernel.org, xen-devel@lists.sf.net Subject: Re: [XEN] using shmfs for swapspace Message-ID: <20050103183133.GA19081@samarkand.rivenstone.net> Mail-Followup-To: Luke Kenneth Casson Leighton , linux-kernel@vger.kernel.org, xen-devel@lists.sf.net References: <20050102162652.GA12268@lkcl.net> Mime-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline Content-Transfer-Encoding: 8BIT In-Reply-To: <20050102162652.GA12268@lkcl.net> User-Agent: Mutt/1.4i Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Sun, Jan 02, 2005 at 04:26:52PM +0000, Luke Kenneth Casson Leighton wrote: [...] > this is presumed to be infinitely better than forcing the swapspace to > be always on disk, especially with the guests only being allocated > 32mbyte of physical RAM. I'd be interested in knowing how a tmpfs that's gone far into swap performs compared to a more normal on-disk fs. I don't know if anyone has ever looked into it. Is it comparable, or is tmpfs's ability to swap more a last-resort escape hatch? This is the part where I would add something valuable to this conversation, if I were going to do that. (But no.) -- Joseph Fannin jhf@rivenstone.net From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261693AbVACUn3 (ORCPT ); Mon, 3 Jan 2005 15:43:29 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261760AbVACUn3 (ORCPT ); Mon, 3 Jan 2005 15:43:29 -0500 Received: from open.hands.com ([195.224.53.39]:59603 "EHLO open.hands.com") by vger.kernel.org with ESMTP id S261693AbVACUnY (ORCPT ); Mon, 3 Jan 2005 15:43:24 -0500 Date: Mon, 3 Jan 2005 20:53:18 +0000 From: Luke Kenneth Casson Leighton To: linux-kernel@vger.kernel.org, xen-devel@lists.sf.net Subject: Re: [XEN] using shmfs for swapspace Message-ID: <20050103205318.GD6631@lkcl.net> References: <20050102162652.GA12268@lkcl.net> <20050103183133.GA19081@samarkand.rivenstone.net> Mime-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20050103183133.GA19081@samarkand.rivenstone.net> User-Agent: Mutt/1.5.5.1+cvs20040105i X-hands-com-MailScanner: Found to be clean X-MailScanner-From: lkcl@lkcl.net Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Mon, Jan 03, 2005 at 01:31:34PM -0500, Joseph Fannin wrote: > On Sun, Jan 02, 2005 at 04:26:52PM +0000, Luke Kenneth Casson Leighton wrote: > [...] > > this is presumed to be infinitely better than forcing the swapspace to > > be always on disk, especially with the guests only being allocated > > 32mbyte of physical RAM. > > I'd be interested in knowing how a tmpfs that's gone far into swap > performs compared to a more normal on-disk fs. I don't know if anyone > has ever looked into it. Is it comparable, or is tmpfs's ability to > swap more a last-resort escape hatch? > > This is the part where I would add something valuable to this > conversation, if I were going to do that. (But no.) :) okay. some kind person from ibm pointed out that of course if you use a file-based swap file (in xen terminology, disk=['file:/xen/guest1-swapfile,/dev/sda2,rw'] which means "publish guest1-swapfile on the DOM0 VM as /dev/sda2 hard drive on the guest1 VM) then you of course end up using the linux filesystem cache on DOM0 which is of course RAM-based. so this tends to suggest a strategy where you allocate as much memory as you can afford to the DOM0 VM, and as little as you can afford to the guests, and make the guest swap files bigger to compensate. ... and i thought it was going to need some wacky wacko non-sharing shared-memory virtual-memory pseudo-tmpfs block-based filesystem driver. dang. l. From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261889AbVACVP1 (ORCPT ); Mon, 3 Jan 2005 16:15:27 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261886AbVACVOT (ORCPT ); Mon, 3 Jan 2005 16:14:19 -0500 Received: from brown.brainfood.com ([146.82.138.61]:16350 "EHLO gradall.private.brainfood.com") by vger.kernel.org with ESMTP id S261889AbVACVHy (ORCPT ); Mon, 3 Jan 2005 16:07:54 -0500 Date: Mon, 3 Jan 2005 15:07:42 -0600 (CST) From: Adam Heath X-X-Sender: adam@gradall.private.brainfood.com To: Luke Kenneth Casson Leighton cc: "linux-kernel@vger.kernel.org" , "xen-devel@lists.sf.net" Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace In-Reply-To: <20050103205318.GD6631@lkcl.net> Message-ID: References: <20050102162652.GA12268@lkcl.net> <20050103183133.GA19081@samarkand.rivenstone.net> <20050103205318.GD6631@lkcl.net> MIME-Version: 1.0 Content-Type: TEXT/PLAIN; charset=US-ASCII Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Mon, 3 Jan 2005, Luke Kenneth Casson Leighton wrote: > On Mon, Jan 03, 2005 at 01:31:34PM -0500, Joseph Fannin wrote: > > On Sun, Jan 02, 2005 at 04:26:52PM +0000, Luke Kenneth Casson Leighton wrote: > > [...] > > > this is presumed to be infinitely better than forcing the swapspace to > > > be always on disk, especially with the guests only being allocated > > > 32mbyte of physical RAM. > > > > I'd be interested in knowing how a tmpfs that's gone far into swap > > performs compared to a more normal on-disk fs. I don't know if anyone > > has ever looked into it. Is it comparable, or is tmpfs's ability to > > swap more a last-resort escape hatch? > > > > This is the part where I would add something valuable to this > > conversation, if I were going to do that. (But no.) > > :) > > okay. > > some kind person from ibm pointed out that of course if you use a > file-based swap file (in xen terminology, > disk=['file:/xen/guest1-swapfile,/dev/sda2,rw'] which means "publish > guest1-swapfile on the DOM0 VM as /dev/sda2 hard drive on the > guest1 VM) then you of course end up using the linux filesystem cache > on DOM0 which is of course RAM-based. > > so this tends to suggest a strategy where you allocate as > much memory as you can afford to the DOM0 VM, and as little > as you can afford to the guests, and make the guest swap > files bigger to compensate. But the guest kernels need real ram to run programs in. The problem with dom0 doing the caching, is that dom0 has no idea about the usage pattern for the swap. It's just a plain file to dom0. Only each guest kernel knows how to combine swap reads/writes correctly. From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261957AbVACWsF (ORCPT ); Mon, 3 Jan 2005 17:48:05 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261942AbVACWNw (ORCPT ); Mon, 3 Jan 2005 17:13:52 -0500 Received: from clock-tower.bc.nu ([81.2.110.250]:6066 "EHLO localhost.localdomain") by vger.kernel.org with ESMTP id S261938AbVACWKi (ORCPT ); Mon, 3 Jan 2005 17:10:38 -0500 Subject: Re: [XEN] using shmfs for swapspace From: Alan Cox To: Luke Kenneth Casson Leighton Cc: Linux Kernel Mailing List , xen-devel@lists.sourceforge.net In-Reply-To: <20050103205318.GD6631@lkcl.net> References: <20050102162652.GA12268@lkcl.net> <20050103183133.GA19081@samarkand.rivenstone.net> <20050103205318.GD6631@lkcl.net> Content-Type: text/plain Content-Transfer-Encoding: 7bit Message-Id: <1104785749.13302.26.camel@localhost.localdomain> Mime-Version: 1.0 X-Mailer: Ximian Evolution 1.4.6 (1.4.6-2) Date: Mon, 03 Jan 2005 21:06:20 +0000 Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org > so this tends to suggest a strategy where you allocate as > much memory as you can afford to the DOM0 VM, and as little > as you can afford to the guests, and make the guest swap > files bigger to compensate. This is essentially what the mainframe folks are already doing and have been doing for some time because the kernel VM has no external inputs for saying "you are virtualised so be nice" for doing opportunistic page recycling ("I dont need this page but when I ask for it back please tell me if you trashed the content") and for hinting to the underlying VM what pages are best blasted out of existance first and how to communicate so we dont page them back in scanning them. From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261982AbVADDLR (ORCPT ); Mon, 3 Jan 2005 22:11:17 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261997AbVADDLR (ORCPT ); Mon, 3 Jan 2005 22:11:17 -0500 Received: from ppsw-2.csi.cam.ac.uk ([131.111.8.132]:51667 "EHLO ppsw-2.csi.cam.ac.uk") by vger.kernel.org with ESMTP id S261982AbVADDLO (ORCPT ); Mon, 3 Jan 2005 22:11:14 -0500 From: Mark Williamson To: xen-devel@lists.sourceforge.net Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace Date: Tue, 4 Jan 2005 03:04:09 +0000 User-Agent: KMail/1.7.1 Cc: Alan Cox , Luke Kenneth Casson Leighton , Linux Kernel Mailing List References: <20050102162652.GA12268@lkcl.net> <20050103205318.GD6631@lkcl.net> <1104785749.13302.26.camel@localhost.localdomain> In-Reply-To: <1104785749.13302.26.camel@localhost.localdomain> MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Content-Disposition: inline Message-Id: <200501040304.10128.maw48@cl.cam.ac.uk> X-Cam-ScannerInfo: http://www.cam.ac.uk/cs/email/scanner/ X-Cam-AntiVirus: No virus found X-Cam-SpamDetails: Not scanned Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org > for doing opportunistic page recycling ("I dont need this page but when > I ask for it back please tell me if you trashed the content") We've talked about doing this but AFAIK nobody has gotten round to it yet because there hasn't been a pressing need (IIRC, it was on the todo list when Xen 1.0 came out). IMHO, it doesn't look terribly difficult but would require (hopefully small) modifications to the architecture independent code, plus a little bit of support code in Xen. I'd quite like to look at this one fine day but I suspect there are more useful things I should do first... Cheers, Mark From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261569AbVADJVA (ORCPT ); Tue, 4 Jan 2005 04:21:00 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261577AbVADJVA (ORCPT ); Tue, 4 Jan 2005 04:21:00 -0500 Received: from open.hands.com ([195.224.53.39]:55177 "EHLO open.hands.com") by vger.kernel.org with ESMTP id S261569AbVADJU4 (ORCPT ); Tue, 4 Jan 2005 04:20:56 -0500 Date: Tue, 4 Jan 2005 09:30:44 +0000 From: Luke Kenneth Casson Leighton To: Adam Heath Cc: "linux-kernel@vger.kernel.org" , "xen-devel@lists.sf.net" Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace Message-ID: <20050104093044.GC10906@lkcl.net> References: <20050102162652.GA12268@lkcl.net> <20050103183133.GA19081@samarkand.rivenstone.net> <20050103205318.GD6631@lkcl.net> Mime-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: User-Agent: Mutt/1.5.5.1+cvs20040105i X-hands-com-MailScanner: Found to be clean X-MailScanner-From: lkcl@lkcl.net Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Mon, Jan 03, 2005 at 03:07:42PM -0600, Adam Heath wrote: > > so this tends to suggest a strategy where you allocate as > > much memory as you can afford to the DOM0 VM, and as little > > as you can afford to the guests, and make the guest swap > > files bigger to compensate. > > But the guest kernels need real ram to run programs in. > > The problem with dom0 doing the caching, is that dom0 has no idea about the > usage pattern for the swap. It's just a plain file to dom0. Only each guest > kernel knows how to combine swap reads/writes correctly. ... hmm... then that tends to suggest that this is an issue that should really be dealt with by XEN. that there needs to be coordination of swap management between the virtual machines. l. -- -- http://lkcl.net -- From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261647AbVADOGF (ORCPT ); Tue, 4 Jan 2005 09:06:05 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261651AbVADOGF (ORCPT ); Tue, 4 Jan 2005 09:06:05 -0500 Received: from mx1.redhat.com ([66.187.233.31]:25832 "EHLO mx1.redhat.com") by vger.kernel.org with ESMTP id S261647AbVADOF4 (ORCPT ); Tue, 4 Jan 2005 09:05:56 -0500 Date: Tue, 4 Jan 2005 09:05:13 -0500 (EST) From: Rik van Riel X-X-Sender: riel@chimarrao.boston.redhat.com To: Mark Williamson cc: xen-devel@lists.sourceforge.net, Alan Cox , Luke Kenneth Casson Leighton , Linux Kernel Mailing List Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace In-Reply-To: <200501040304.10128.maw48@cl.cam.ac.uk> Message-ID: References: <20050102162652.GA12268@lkcl.net> <20050103205318.GD6631@lkcl.net> <1104785749.13302.26.camel@localhost.localdomain> <200501040304.10128.maw48@cl.cam.ac.uk> MIME-Version: 1.0 Content-Type: TEXT/PLAIN; charset=US-ASCII; format=flowed Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Tue, 4 Jan 2005, Mark Williamson wrote: >> for doing opportunistic page recycling ("I dont need this page but when >> I ask for it back please tell me if you trashed the content") > > We've talked about doing this but AFAIK nobody has gotten round to it > yet because there hasn't been a pressing need (IIRC, it was on the todo > list when Xen 1.0 came out). > > IMHO, it doesn't look terribly difficult but would require (hopefully > small) modifications to the architecture independent code, plus a little > bit of support code in Xen. The architecture independant changes are fine, since they're also useful for S390(x), PPC64 and UML... > I'd quite like to look at this one fine day but I suspect there are more > useful things I should do first... I wonder if the same effect could be achieved by just measuring the VM pressure inside the guests and ballooning the guests as required, letting them grow and shrink with their workloads. That wouldn't need many kernel changes, maybe just a few extra statistics, or maybe all the needed stats already exist. It would also allow more complex policy to be done in userspace, eg. dealing with Xen guests of different priority... -- "Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it." - Brian W. Kernighan From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S261657AbVADOHp (ORCPT ); Tue, 4 Jan 2005 09:07:45 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S261656AbVADOHg (ORCPT ); Tue, 4 Jan 2005 09:07:36 -0500 Received: from mx1.redhat.com ([66.187.233.31]:233 "EHLO mx1.redhat.com") by vger.kernel.org with ESMTP id S261651AbVADOGm (ORCPT ); Tue, 4 Jan 2005 09:06:42 -0500 Date: Tue, 4 Jan 2005 09:06:24 -0500 (EST) From: Rik van Riel X-X-Sender: riel@chimarrao.boston.redhat.com To: Luke Kenneth Casson Leighton cc: Adam Heath , "linux-kernel@vger.kernel.org" , "xen-devel@lists.sf.net" Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace In-Reply-To: <20050104093044.GC10906@lkcl.net> Message-ID: References: <20050102162652.GA12268@lkcl.net> <20050103183133.GA19081@samarkand.rivenstone.net> <20050103205318.GD6631@lkcl.net> <20050104093044.GC10906@lkcl.net> MIME-Version: 1.0 Content-Type: TEXT/PLAIN; charset=US-ASCII; format=flowed Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Tue, 4 Jan 2005, Luke Kenneth Casson Leighton wrote: > then that tends to suggest that this is an issue that should > really be dealt with by XEN. Probably. > that there needs to be coordination of swap management between the > virtual machines. I'd like to see the maximum security separation possible between the unprivileged guests, though... -- "Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it." - Brian W. Kernighan From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S262108AbVAEAX3 (ORCPT ); Tue, 4 Jan 2005 19:23:29 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S262154AbVAEAUr (ORCPT ); Tue, 4 Jan 2005 19:20:47 -0500 Received: from natpreptil.rzone.de ([81.169.145.163]:63114 "EHLO natpreptil.rzone.de") by vger.kernel.org with ESMTP id S261869AbVAEATm (ORCPT ); Tue, 4 Jan 2005 19:19:42 -0500 From: Arnd Bergmann To: Mark Williamson Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace Date: Wed, 5 Jan 2005 01:11:54 +0100 User-Agent: KMail/1.6.2 Cc: xen-devel@lists.sourceforge.net, Alan Cox , Luke Kenneth Casson Leighton , Linux Kernel Mailing List References: <20050102162652.GA12268@lkcl.net> <1104785749.13302.26.camel@localhost.localdomain> <200501040304.10128.maw48@cl.cam.ac.uk> In-Reply-To: <200501040304.10128.maw48@cl.cam.ac.uk> MIME-Version: 1.0 Content-Type: multipart/signed; protocol="application/pgp-signature"; micalg=pgp-sha1; boundary="Boundary-02=_ODz2BcJHSHjhPM7"; charset="iso-8859-15" Content-Transfer-Encoding: 7bit Message-Id: <200501050111.59072.arnd@arndb.de> Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org --Boundary-02=_ODz2BcJHSHjhPM7 Content-Type: text/plain; charset="iso-8859-15" Content-Transfer-Encoding: quoted-printable Content-Disposition: inline On Dinsdag 04 Januar 2005 04:04, Mark Williamson wrote: > > for doing opportunistic page recycling ("I dont need this page but when > > I ask for it back please tell me if you trashed the content") >=20 > We've talked about doing this but AFAIK nobody has gotten round to it yet= =20 > because there hasn't been a pressing need (IIRC, it was on the todo list = when=20 > Xen 1.0 came out). >=20 > IMHO, it doesn't look terribly difficult but would require (hopefully sma= ll)=20 > modifications to the architecture independent code, plus a little bit of= =20 > support code in Xen. >=20 > I'd quite like to look at this one fine day but I suspect there are more= =20 > useful things I should do first... There are two other alternatives that are already used on s390 for making multi-level paging a little more pleasant: =2D Pseudo faults: When Linux accesses a page that it believes to be present but is actually swapped out in z/VM, the VM hypervisor causes a special PFAULT exception. Linux can then choose to either ignore this exception and continue, which will force VM to swap the page back in. Or it can do a task switch and wait for the page to come back. At the point where VM has read the page back from its swap device, it causes another exception, after which Linux wakes up the sleeping process. see arch/s390/mm/fault.c =2D Ballooning:=20 z/VM has an interface (DIAG 10) for the OS to tell it about a page that is currently unused. The kernel uses get_free_page to reserve a number of pages, then calls DIAG10 to give it to z/VM. The amount of pages to give back to the hypervisor is determined by a system wide workload manager. see arch/s390/mm/cmm.c When you want to introduce some interface in Xen, you probably want something more powerful than these, but it probably makes sense to see them as a base line of what can be done with practically no common code changes (if you don't do similar stuff already). Arnd <>< --Boundary-02=_ODz2BcJHSHjhPM7 Content-Type: application/pgp-signature Content-Description: signature -----BEGIN PGP SIGNATURE----- Version: GnuPG v1.2.4 (GNU/Linux) iD8DBQBB2zDO5t5GS2LDRf4RAuZLAJ40ZmXuYmHUA45RzjT/ykwcJCkuAgCgggtY KE4Qu15FknpsI11MLjIDbVc= =SCle -----END PGP SIGNATURE----- --Boundary-02=_ODz2BcJHSHjhPM7-- From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S262580AbVAFL15 (ORCPT ); Thu, 6 Jan 2005 06:27:57 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S262581AbVAFL15 (ORCPT ); Thu, 6 Jan 2005 06:27:57 -0500 Received: from open.hands.com ([195.224.53.39]:17304 "EHLO open.hands.com") by vger.kernel.org with ESMTP id S262580AbVAFL1y (ORCPT ); Thu, 6 Jan 2005 06:27:54 -0500 Date: Thu, 6 Jan 2005 11:38:27 +0000 From: Luke Kenneth Casson Leighton To: Rik van Riel Cc: Mark Williamson , xen-devel@lists.sourceforge.net, Alan Cox , Linux Kernel Mailing List Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace Message-ID: <20050106113827.GM6440@lkcl.net> References: <20050102162652.GA12268@lkcl.net> <20050103205318.GD6631@lkcl.net> <1104785749.13302.26.camel@localhost.localdomain> <200501040304.10128.maw48@cl.cam.ac.uk> Mime-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: User-Agent: Mutt/1.5.5.1+cvs20040105i X-hands-com-MailScanner: Found to be clean X-MailScanner-From: lkcl@lkcl.net Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Tue, Jan 04, 2005 at 09:05:13AM -0500, Rik van Riel wrote: > On Tue, 4 Jan 2005, Mark Williamson wrote: > > >>for doing opportunistic page recycling ("I dont need this page but when > >>I ask for it back please tell me if you trashed the content") > > > >We've talked about doing this but AFAIK nobody has gotten round to it > >yet because there hasn't been a pressing need (IIRC, it was on the todo > >list when Xen 1.0 came out). > > > >IMHO, it doesn't look terribly difficult but would require (hopefully > >small) modifications to the architecture independent code, plus a little > >bit of support code in Xen. > > The architecture independant changes are fine, since > they're also useful for S390(x), PPC64 and UML... > > >I'd quite like to look at this one fine day but I suspect there are more > >useful things I should do first... > > I wonder if the same effect could be achieved by just > measuring the VM pressure inside the guests and > ballooning the guests as required, letting them grow > and shrink with their workloads. mem = 64M-128M target = 64M "if needed, grow me to 128mb but if not, whittle down to 64". mem=64M-128 target=128M "if you absolutely have to, steal some of my memory, but don't nick any more than 64M". i'm probably going to have to "manually" implement something like this. l. From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S262515AbVAUViu (ORCPT ); Fri, 21 Jan 2005 16:38:50 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S262522AbVAUVit (ORCPT ); Fri, 21 Jan 2005 16:38:49 -0500 Received: from mx1.redhat.com ([66.187.233.31]:46238 "EHLO mx1.redhat.com") by vger.kernel.org with ESMTP id S262515AbVAUVhW (ORCPT ); Fri, 21 Jan 2005 16:37:22 -0500 Date: Fri, 21 Jan 2005 16:37:09 -0500 (EST) From: Rik van Riel X-X-Sender: riel@chimarrao.boston.redhat.com To: Arnd Bergmann cc: Mark Williamson , xen-devel@lists.sourceforge.net, Alan Cox , Luke Kenneth Casson Leighton , Linux Kernel Mailing List Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace In-Reply-To: <200501050111.59072.arnd@arndb.de> Message-ID: References: <20050102162652.GA12268@lkcl.net> <1104785749.13302.26.camel@localhost.localdomain> <200501040304.10128.maw48@cl.cam.ac.uk> <200501050111.59072.arnd@arndb.de> MIME-Version: 1.0 Content-Type: TEXT/PLAIN; charset=US-ASCII; format=flowed Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org On Wed, 5 Jan 2005, Arnd Bergmann wrote: > - Pseudo faults: These are a problem, because they turn what would be a single pageout into a pageout, a pagein, and another pageout, in effect tripling the amount of IO that needs to be done. > - Ballooning: Xen already has this. I wonder if it makes sense to consolidate the various balloon approaches into a single driver, and keep the amount of ballooned memory into account when reporting statistics in /proc/meminfo. > When you want to introduce some interface in Xen, you probably want > something more powerful than these, Xen has a nice balloon driver, that can also be controlled from outside the guest domain. -- "Debugging is twice as hard as writing the code in the first place. Therefore, if you write the code as cleverly as possible, you are, by definition, not smart enough to debug it." - Brian W. Kernighan From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S262465AbVA0BTu (ORCPT ); Wed, 26 Jan 2005 20:19:50 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S262420AbVA0AEv (ORCPT ); Wed, 26 Jan 2005 19:04:51 -0500 Received: from ppsw-4.csi.cam.ac.uk ([131.111.8.134]:62337 "EHLO ppsw-4.csi.cam.ac.uk") by vger.kernel.org with ESMTP id S262421AbVAZVES (ORCPT ); Wed, 26 Jan 2005 16:04:18 -0500 From: Mark Williamson To: Rik van Riel Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace Date: Wed, 26 Jan 2005 20:56:50 +0000 User-Agent: KMail/1.7.1 Cc: Arnd Bergmann , Mark Williamson , xen-devel@lists.sourceforge.net, Alan Cox , Luke Kenneth Casson Leighton , Linux Kernel Mailing List References: <20050102162652.GA12268@lkcl.net> <200501050111.59072.arnd@arndb.de> In-Reply-To: MIME-Version: 1.0 Content-Type: text/plain; charset="iso-8859-1" Content-Transfer-Encoding: 7bit Content-Disposition: inline Message-Id: <200501262056.50981.maw48@cl.cam.ac.uk> X-Cam-ScannerInfo: http://www.cam.ac.uk/cs/email/scanner/ X-Cam-AntiVirus: No virus found X-Cam-SpamDetails: Not scanned Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org > > - Pseudo faults: > > These are a problem, because they turn what would be a single > pageout into a pageout, a pagein, and another pageout, in > effect tripling the amount of IO that needs to be done. The Disco VMM tackled this by detecting attempts to double-page using a special virtual swap disk. Perhaps it would be possible to find some cleaner way to avoid wasteful double-paging by adding some more hooks for virtualised architectures... In any case, for now Xen guests are not swapped onto disk storage at runtime - they retain their physical memory reservation unless they alter it using the balloon driver. > Xen already has this. I wonder if it makes sense to > consolidate the various balloon approaches into a single > driver, and keep the amount of ballooned memory into > account when reporting statistics in /proc/meminfo. If multiple platforms want to do this, we could refactor the code so that the core of the balloon driver can be used in multiple archs. We could have an arch_release/request_memory() that the core balloon driver can call into to actually return memory to the VMM. > > When you want to introduce some interface in Xen, you probably want > > something more powerful than these, > > Xen has a nice balloon driver, that can also be > controlled from outside the guest domain. The Xen control interface made this fairly trivial to implement. Again, the balloon driver core could be plumbed into whatever the preferred virtual machine control interface for the platform is (I don't know if / how other platforms tackle this). Cheers, Mark From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S262573AbVA0KsV (ORCPT ); Thu, 27 Jan 2005 05:48:21 -0500 Received: (majordomo@vger.kernel.org) by vger.kernel.org id S262557AbVA0KmL (ORCPT ); Thu, 27 Jan 2005 05:42:11 -0500 Received: from main.gmane.org ([80.91.229.2]:33001 "EHLO main.gmane.org") by vger.kernel.org with ESMTP id S262563AbVA0Kd3 (ORCPT ); Thu, 27 Jan 2005 05:33:29 -0500 X-Injected-Via-Gmane: http://gmane.org/ To: linux-kernel@vger.kernel.org From: Nuutti Kotivuori Subject: Re: [Xen-devel] Re: [XEN] using shmfs for swapspace Date: Thu, 27 Jan 2005 12:33:10 +0200 Organization: Ye 'Ol Disorganized NNTPCache groupie Message-ID: <87brbb9n6h.fsf@aka.i.naked.iki.fi> References: <20050102162652.GA12268@lkcl.net> <200501050111.59072.arnd@arndb.de> <200501262056.50981.maw48@cl.cam.ac.uk> Mime-Version: 1.0 Content-Type: text/plain; charset=us-ascii X-Complaints-To: usenet@sea.gmane.org X-Gmane-NNTP-Posting-Host: naked.iki.fi User-Agent: Gnus/5.1007 (Gnus v5.10.7) XEmacs/21.4 (Corporate Culture, linux) Cancel-Lock: sha1:yAMwwyOpYkm41O5APJRt7d90230= Cache-Post-Path: aka.i.naked.iki.fi!unknown@aka.i.naked.iki.fi X-Cache: nntpcache 3.0.1 (see http://www.nntpcache.org/) Cc: xen-devel@lists.sourceforge.net Sender: linux-kernel-owner@vger.kernel.org X-Mailing-List: linux-kernel@vger.kernel.org Mark Williamson wrote: > If multiple platforms want to do this, we could refactor the code so > that the core of the balloon driver can be used in multiple archs. > We could have an arch_release/request_memory() that the core balloon > driver can call into to actually return memory to the VMM. This is also a thing that the UML project would probably be interested in. As a generalization though, what is needed is hot-pluggable memory in Linux kernel. Satisfies Xen, UML and any physical implementations at some point. -- Naked