From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752920AbcB2Q2k (ORCPT ); Mon, 29 Feb 2016 11:28:40 -0500 Received: from mail2-relais-roc.national.inria.fr ([192.134.164.83]:18514 "EHLO mail2-relais-roc.national.inria.fr" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1750843AbcB2Q2h (ORCPT ); Mon, 29 Feb 2016 11:28:37 -0500 X-IronPort-AV: E=Sophos;i="5.22,521,1449529200"; d="scan'208";a="205092582" Date: Mon, 29 Feb 2016 17:28:35 +0100 From: Samuel Thibault To: linux-kernel@vger.kernel.org, linux-mm@kvack.org Subject: Costless huge virtual memory? /dev/same, /dev/null? Message-ID: <20160229162835.GA2816@var.bordeaux.inria.fr> Mail-Followup-To: Samuel Thibault , linux-kernel@vger.kernel.org, linux-mm@kvack.org MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline User-Agent: Mutt/1.5.21+34 (58baf7c9f32f) (2010-12-30) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hello, I'm wondering whether we could introduce a /dev/same device to allow costless huge virtual memory. The use case is the simulation of the execution of a big irregular HPC application, to provision memory usage, cpu time, etc. We know how much time each computation loop takes, and it's easy to replace them with a mere accounting. We'd however like to avoid having to revamp the rest of the code, which does allocation/memcpys/etc., by just replacing the allocation calls with virtual allocations, i.e. allocations which return addresses of buffers that one can read/write, but the values you read are not necessarily what you wrote, i.e. the data is not actually properly stored (since we don't do the actual computations that's not a problem). The way we currently do this is by some folding: we map the same normal file several times contiguously to form the virtual allocation. By using a small 1MiB file, this limits memory consumption to 1MiB plus the page table (and fits the dumb data in a typical cache). This however creates one VMA per file mapping, we get limited by the 65535 VMA limit, and VMA lookup becomes slow. The way I could see is to have a /dev/same device: when you open it, it allocates one page. When you mmap it, it maps the same page over the whole resulting single VMA. This is a quite specific use case, but it seems to be easy to implement, and it seems to me that it could be integrated mainline. Actually I was thinking that /dev/null itself could be providing that service? (currently it returns ENODEV) What do people think? Is there perhaps another solution to achieve this that I didn't think about? Samuel