From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0b-0031df01.pphosted.com (mx0b-0031df01.pphosted.com [205.220.180.131]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DC4BF399CFD for ; Fri, 26 Jun 2026 10:47:21 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=205.220.180.131 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782470843; cv=none; b=l57buV7jVKV4hwJAOpjebIvA0D3S0shJoqHeARwGSzHCjoHH2jWk8IKX4MdTCVqhw1ygz+Mtkzd5lmcPv3Prt2thjnnLHoFzbMEYxh51GaLlVPhdGaFNP5Ittyu620IklwU3gr1g9liRg1rxMuF4wouBL3ExVtp1Mhcj9+cqsZs= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782470843; c=relaxed/simple; bh=SIf4Wby4OinN2UZ1ibS707vwvEHsMXIll5sjBwiWY4E=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=Jq6CBvjuc/nt7LbJYbYeHtaYGfqptmI1T8buWRTvACvL0f0QfeCFx+vrCZaeMp18tneifC+eseuQ/CiNZawQN+ER744BeEGzIR1dVxSTD+P73b/tm+6GWFxH0zqr+Ll1KQYyndR5nYiGnoRhChhmmEEGEsCs/MIRqKmo7r4WheQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oss.qualcomm.com; spf=pass smtp.mailfrom=oss.qualcomm.com; dkim=pass (2048-bit key) header.d=qualcomm.com header.i=@qualcomm.com header.b=iD7nxymC; dkim=pass (2048-bit key) header.d=oss.qualcomm.com header.i=@oss.qualcomm.com header.b=V7Dza1B1; arc=none smtp.client-ip=205.220.180.131 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oss.qualcomm.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=oss.qualcomm.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=qualcomm.com header.i=@qualcomm.com header.b="iD7nxymC"; dkim=pass (2048-bit key) header.d=oss.qualcomm.com header.i=@oss.qualcomm.com header.b="V7Dza1B1" Received: from pps.filterd (m0279871.ppops.net [127.0.0.1]) by mx0a-0031df01.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 65QActaY1069994 for ; Fri, 26 Jun 2026 10:47:21 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=qualcomm.com; h= cc:content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s=qcppdkim1; bh= PChBsdJTdy39hlcQc2xrVoLyCX7474bgwLOt0FMf6Ls=; b=iD7nxymCDFJjKejM yeACrs8LlvKBV3jswNLTh9psxaiW925eHUDLfBLMJfzSiROAU3dbupadpJwMvfLa h2JHcUBRcIL8wvD9gun5NNueROnQ/3zkjJ6g1y8vo/nseg2QDL+0GXjQ97XyRfMP bsZrUWOeb6z+nZ0oiTuF6juHHQOHsjTBD5GERvNKQac86Lojqjd0fFuZdM7HUn7M fv9N0L1X39Y6EgYGYMvanQeNvLQlTURU5+7TezrCFUYYyDs0UBJ8xA+A1w6UdTNm WcvliyHGhaxwusIwekR63QIMEPE/otIMFHKdN9UN7hr8tcP6TL6bWCJee12h+wrV YZX6bQ== Received: from mail-dy1-f197.google.com (mail-dy1-f197.google.com [74.125.82.197]) by mx0a-0031df01.pphosted.com (PPS) with ESMTPS id 4f1fcta06y-1 (version=TLSv1.3 cipher=TLS_AES_128_GCM_SHA256 bits=128 verify=NOT) for ; Fri, 26 Jun 2026 10:47:20 +0000 (GMT) Received: by mail-dy1-f197.google.com with SMTP id 5a478bee46e88-30ba395b047so2352790eec.0 for ; Fri, 26 Jun 2026 03:47:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oss.qualcomm.com; s=google; t=1782470840; x=1783075640; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to :content-language:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=PChBsdJTdy39hlcQc2xrVoLyCX7474bgwLOt0FMf6Ls=; b=V7Dza1B1J0iNLME01bocYO6FAT9fVNw8FI7GaTS2YvGe1Xre3N3Lh+NyOj7YGl6M+4 S2YNOZ2azpME6EEJVDfYPhWmNKAw05larcRzV4kjIh8WJTG6gAF8aMt/+WWvj+NiHe3U 7DP2yjCp1DnLqCbfkyvwzfvBHXpRKruY7qiPNVUl98Tr4kc6dDeqEgeVttXb7HbavKT2 e78yHpCvqESt+TXQAUIsdwJj6NC9k0lcNEvStKx0/rw1ShKfqJ13J7WPYxDJbSlTRbOj uI0NjbpO4H4hSOl4xYVel0nfcP+Qj67BTqV1duWmjx2Qbo1Qemi0A/x7jmUop+CzCtKG IBCQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1782470840; x=1783075640; h=content-transfer-encoding:in-reply-to:from:references:cc:to :content-language:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=PChBsdJTdy39hlcQc2xrVoLyCX7474bgwLOt0FMf6Ls=; b=m0J7OLX5tvdmUaNxVN8hujnYlwYrbfX0DLatwjodBqRHb/YpAWvuF3R9wq8BOef6jU ev01XtUU9TsDRKzn7fbO3LdKvVknZ6aTbjPmeupJSKMlgnusUocqVULo5akytc/kRcL3 aTzQ0Dpb+rN7lCS1SO7R6gs/GQf90SPgd5yZP6s2e348f4Uyty9JMd5nfoHQfyn13Cnv xOVFuFWHdFhvIi+as0vuMRdf5skP+VPuwMtC2fGtruVGLhWw8f61qUeUQs70BLO942Qv p1XlCYcGhjo60+GgnmvO+89yexgK4+Pi2nGozlLdl07iuh7DhcQXodUGhStEn8oWfk0h kXuw== X-Forwarded-Encrypted: i=1; AHgh+RrufM2ZbdZ0fNBqXSruxXJOt3yBMD54Pb4M5OoIxhY+XNrpmKUJixTEI/hS0PPHf2lSaTgghJpRvUI2Z9M=@vger.kernel.org X-Gm-Message-State: AOJu0Yx6qUewCkZmcoVq/q5uvLGA50J7dxSdmSMfZJlWGx4lmnThhWVj LM0yNea9ZoO72R1q+f7A14YnN2CTBzN8bM4ABK0XmDeJjJN91xuhM0aapbr78RYFUUZf3cyXbgw DnlhcmfV3V96okFlKKH4sjK1d+zccsiambkmkRXbsGRrorx5KSLzYcEQwj++B5lcM1uc= X-Gm-Gg: AfdE7cnfWM3/MwUVQDkfs5x7UPNjMKTi2p2Xx7jEHR2x/pzZJBsPROtVyLRIuIAFOtN I8Uo2EjnbDHZpPVSoVieNsiykS6f+KWVcI3N5JdARzaShWSluMUfGAIuo9RzD0j3y2+0qnB76jT yPIZx5rghw/pJ7NUHgPeQLFr1vhjCXNJIAFmuSDoVHaOFXJsmmXypBuTv1xsPNNrx9Dg6XSrdXd 9n4eXX7f2+5uMS2FkLVGhli18To+tlJjyvl8ZTNHdnvZkbgl1iPxrvyjwuWL2GiiQFGwThk7MLT 9D3mmKXpb1di49INpV1k1y69HjlcQF6azUqOaULHC0uQAG0x1gL+3pw188+v0kyNCpPZFclqHX/ Sz6uspe+Hv1IOSZ+p8wO41wdvsCJN/bcLJ17L9cpT X-Received: by 2002:a05:7301:37c4:b0:304:ab8:f899 with SMTP id 5a478bee46e88-30c84bcdac4mr6511062eec.8.1782470839608; Fri, 26 Jun 2026 03:47:19 -0700 (PDT) X-Received: by 2002:a05:7301:37c4:b0:304:ab8:f899 with SMTP id 5a478bee46e88-30c84bcdac4mr6510992eec.8.1782470838306; Fri, 26 Jun 2026 03:47:18 -0700 (PDT) Received: from [10.218.25.225] ([202.46.22.19]) by smtp.gmail.com with ESMTPSA id 5a478bee46e88-30c7c4ca240sm17658216eec.4.2026.06.26.03.47.09 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Fri, 26 Jun 2026 03:47:17 -0700 (PDT) Message-ID: Date: Fri, 26 Jun 2026 16:17:08 +0530 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH RFC 00/12] mm/vmalloc: migrate vmap_area indexing from rb-tree to maple-tree Content-Language: en-US To: Uladzislau Rezki , Matthew Wilcox Cc: Andrew Morton , "Liam R. Howlett" , Alice Ryhl , Andrew Ballance , linux-arm-msm@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, maple-tree@lists.infradead.org, Lorenzo Stoakes , Pranjal Shrivastava , Will Deacon , Suzuki K Poulose , Neil Armstrong , Mostafa Saleh , Balbir Singh , Suren Baghdasaryan , Marco Elver , Dmitry Vyukov , Alexander Potapenko , Shuah Khan , Dev Jain , Brendan Jackman , Puranjay Mohan , Santosh Shukla , Wyes Karny , Sudeep Holla References: <20260613-vmalloc_maple-v1-0-0aa740bb944b@oss.qualcomm.com> From: Pranjal Arya In-Reply-To: Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-Proofpoint-GUID: tzGWKXSu6O7hM9NdS4newBAWgMqbL0bH X-Proofpoint-ORIG-GUID: tzGWKXSu6O7hM9NdS4newBAWgMqbL0bH X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwNjI2MDA4NyBTYWx0ZWRfX55kCz08a6tHW t+r7943h7J0STPQmLPTOHetZA1BXKKFeXjZXIsRQ72HZHv//jM82cOEzwdYcKOI2qrAqBO2hUVN zM/YdfAUJ7/cG9cNeWMAGQlcoZ0zD3J/yZa6cy1g7a1YjGA8H3x4p938TaDWoXVWvSoucDXPvwC C8UX4watZ7eXJIzoJYOELvifD9gziYXZSsaEz/hH662Dcz79+quBL9xGmrfMKwbQobPVqxm7d+U sWeRs1av3JZ9J4R/33VKNsA9SBydzKvpjEPtBxVD1m0FEOmWOff3moDe/eGUca0RN3fRkA+OMo0 pUnF01Tc4DvkLzB10JB3rxNEl0EmaF7XAI8S4MdpiNh29G0oG19MxivrrpCpJkvIz1jmSs13OED fbGP/BNTQotiPDNo3MZ8iyeUwGxI6EWzIe02GtA21kofB6P9gQlCKdrUVVG9RvLlk10SJrHSil2 4ibiTyJHfz76yHJYBUg== X-Authority-Analysis: v=2.4 cv=FPkrAeos c=1 sm=1 tr=0 ts=6a3e58b8 cx=c_pps a=Uww141gWH0fZj/3QKPojxA==:117 a=fChuTYTh2wq5r3m49p7fHw==:17 a=IkcTkHD0fZMA:10 a=FelO9ux0wxsA:10 a=s4-Qcg_JpJYA:10 a=VkNPw1HP01LnGYTKEx00:22 a=u7WPNUs3qKkmUXheDGA7:22 a=3WHJM1ZQz_JShphwDgj5:22 a=O-PE4fd2i0mDks678lIA:9 a=QEXdDO2ut3YA:10 a=PxkB5W3o20Ba91AHUih5:22 X-Proofpoint-Spam-Info: AW1haW4tMjYwNjI2MDA4NyBTYWx0ZWRfXy2lb+rXHYd7d 5/Y7PvgkACfUddl4qn2NjiqfemLD9g4XqwsCjfvck9RggiXM9+R79/P81B5PLwqcdsSS2iomiLH KfoSUDBVmqHhA7M5rQ9VkYlu0sY/mJ4= X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1143,Hydra:6.1.125,FMLib:17.12.100.49 definitions=2026-06-26_03,2026-06-24_01,2025-10-01_01 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 clxscore=1015 priorityscore=1501 malwarescore=0 adultscore=0 bulkscore=0 impostorscore=0 phishscore=0 suspectscore=0 lowpriorityscore=0 classifier=typeunknown authscore=0 authtc= authcc= route=outbound adjust=0 reason=mlx scancount=1 engine=8.22.0-2606150000 definitions=main-2606260087 On 6/18/2026 3:36 PM, Uladzislau Rezki wrote: > On Tue, Jun 16, 2026 at 07:07:57PM +0100, Matthew Wilcox wrote: >> On Mon, Jun 15, 2026 at 11:52:22AM +0200, Uladzislau Rezki wrote: >>> On Sun, Jun 14, 2026 at 12:15:28AM +0100, Matthew Wilcox wrote: >>>> What I don't understand is why you maintain a separate "free" tree. >>>> It should not be necessary any more, but maybe you tried removing it >>>> already and found a performance problem? >>> >>> We maintain it in order to split several entities. That prevents >>> interfering between allocated data and vmap-free-space manager. >>> So in that case one context can easily access allocated data, for >>> example vread iterator, etc., whereas another can do an allocation. >>> >>> So by splitting parts i minimize lock-contention. >> >> Sure, but there are many ways to reduce lock contention. One is to not >> take locks at all; the maple tree is RCU-safe, so you can read the tree >> holding only the RCU read lock, as long as you obey the RCU rules. >> >> Specifically: >> - Write side has to RCU-free the objects that are stored in the tree >> - Read side has to trylock the objects it finds (and retry the walk >> if the trylock fails) >> - Read side can see a mixture of objects if the tree is changed while >> it is reading, but for any given index in the tree it is guaranteed >> to see one of the objects which has been referred to by that index. >> That is, if the write side overwrites an index that referred to >> object A with object B, the reader will see either object A or B. >> It will not see NULL and it will not see any other object. >> - If the write side stores both object C and object D in the tree, >> the read side may see neither, both, only C or only D. >> > Some thoughts about it. > > Having the tree which is RCU safe is good for sure. We can benefit from > at least in the: vmallocinfo scanning/dumping, possibly in the vread_iter() > when access to /proc/kcore and other places(which i need to check carefully). > But this is for read-only traversal. > I agree. The RCU safe busy tree is a foundation that enables lockless read only traversal. Will implement it in next patchset. > Switching to gap-based approach requires quite a bit of refactoring and it > should be a full switch without any hybrid schemes or mixes. I expect that > we remove more code then adding because of some parts will become hidden > like lookups/reserving range/erase, etc which is good. > > - replacing free_vmap_area to maple-tree gap based approach; > - rewriting pcpu-allocator which lives in the end of vmalloc space; > - refactoring per-cpu allocator which is also part of vmalloc space; > - vread iterator; > - vmalloc dump path; > - vmap_node logic(use gap-reserve to minimize contention); > - and more... > > To me such rewrite makes sense if we end up in something structural not > just because maple tree exists. The criteria i would go with are: at least > same performance level, remove more then add, the design stays at least in > same good shape. > > There are some drawback i am thinking of. One of them is maple insert path, > mas_store_gfp()? First we need to find an empty area, then set-range and do > mas_store_gfp() that uses gfp flag for its internal allocation. If we are > under spin-lock sleeping is not possible, using NOWAIT or ATOMIC is not a > case thus we should somehow pre-allocate outside the lock and store the range > without any allocation. > I am planning to have following approach on this: 1. mas_preallocate(GFP_NOWAIT | __GFP_NOWARN) + mas_store_prealloc will be fast path. The preallocate attempt will be non sleeping and, if it succeeds, the subsequent store won't require allocation. 2. mas_store_gfp(GFP_ATOMIC | __GFP_NOWARN) fallback: if preallocate fails (rare but possible under memory pressure), GFP_ATOMIC will make a non sleeping allocation attempt inline. 3. vmap_retry_list recovery queue: if both above fail, the VA will be added to a non indexed retry list. The allocator will scan the list on subsequent calls, and purge workers will drain it. This will avoid any leak or panic under sustained slab pressure. Neither GFP_NOWAIT nor GFP_ATOMIC can sleep under a spinlock. The retry will queue provide the correctness backstop to avoid GFP_KERNEL blocking inside the lock. > The allocator operation: > - finds an empty range; > - publishes VA that blocks that range. > > those two have to be serialized among other writes. Otherwise two CPUs can use > same empty range and both try to reserve them. If preallocate outside the lock, > the "alloc" side has to validate that a selected range is still empty and only > then store VA to block the range. > Sure. Both will be called under a lock. Holding both under the same lock is the better serialization approach. No two concurrent allocators will observe the same gap and both succeed. In the next patchset, I'll explicitly add a comment on the lock declaration describing this approach in commit message. > I think it is worth to prototype something to see how it would go. I may be > missing something for sure. > > Thank you for your input! > > -- > Uladzislau Rezki Thank you for the detailed design questions. This will make upcoming patchset substantially cleaner than the RFC. BR, Pranjal