From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from frasgout.his.huawei.com (frasgout.his.huawei.com [185.176.79.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 32C7B392FE4 for ; Fri, 19 Dec 2025 16:19:04 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=185.176.79.56 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766161148; cv=none; b=i8z9nJOcdhjr0xjgXs3ImNAMxwzHTUoFVUpjgj1RqZqjyitNPMI8XGC1OfThzDoa2f7mv3EFxFe3X+j7yj5o5TohJFN/tl4KTQnV/zYlHTjVpNPKFsxpKAJOaYcXlyuk/tfnVdVk9ueHKjVfeWYJU8nXYClfkJ6u6KEBBFtb5W8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1766161148; c=relaxed/simple; bh=/SlQj+UmTcJhTzB/LkiNQpcEdW+IdssAmrm7ofOcs1Y=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=s+BYg1tD3XWJWap7foCis5uZRuxXL9+MAgaGT1vg751iIJsRabNw267Au46FW8644ClC2qrmCYUTOEiwUHxMR2DWrsX7GoF7B9QEe2ZcvwigPJHCp+ChD8bjNvJOUaIrfIZjMuOFKQEYyDF30SDGIVpCnjl7DtgRPTOfyDoqTAQ= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=h-partners.com; spf=pass smtp.mailfrom=h-partners.com; arc=none smtp.client-ip=185.176.79.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=h-partners.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=h-partners.com Received: from mail.maildlp.com (unknown [172.18.224.150]) by frasgout.his.huawei.com (SkyGuard) with ESMTPS id 4dXt4j3NmNzJ467y; Sat, 20 Dec 2025 00:18:29 +0800 (CST) Received: from mscpeml500003.china.huawei.com (unknown [7.188.49.51]) by mail.maildlp.com (Postfix) with ESMTPS id 4D02B40539; Sat, 20 Dec 2025 00:19:02 +0800 (CST) Received: from [10.123.123.67] (10.123.123.67) by mscpeml500003.china.huawei.com (7.188.49.51) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.11; Fri, 19 Dec 2025 19:18:58 +0300 Message-ID: <9822c658-c2f0-4b1c-9eef-9ffa865e44f7@h-partners.com> Date: Fri, 19 Dec 2025 19:18:53 +0300 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [RFC PATCH 2/2] mm: implement page refcount locking via dedicated bit To: Kiryl Shutsemau CC: , , , , , , , , , , , , , , , , , , , , , , , , References: <81e3c45f49bdac231e831ec7ba09ef42fbb77930.1766145604.git.gladyshev.ilya1@h-partners.com> Content-Language: en-US From: Gladyshev Ilya In-Reply-To: Content-Type: text/plain; charset="UTF-8"; format=flowed Content-Transfer-Encoding: 7bit X-ClientProxiedBy: lhrpeml100010.china.huawei.com (7.191.174.197) To mscpeml500003.china.huawei.com (7.188.49.51) On 12/19/2025 5:50 PM, Kiryl Shutsemau wrote: > On Fri, Dec 19, 2025 at 12:46:39PM +0000, Gladyshev Ilya wrote: >> The current atomic-based page refcount implementation treats zero >> counter as dead and requires a compare-and-swap loop in folio_try_get() >> to prevent incrementing a dead refcount. This CAS loop acts as a >> serialization point and can become a significant bottleneck during >> high-frequency file read operations. >> >> This patch introduces FOLIO_LOCKED_BIT to distinguish between a > > s/FOLIO_LOCKED_BIT/PAGEREF_LOCKED_BIT/ Ack, thanks >> (temporary) zero refcount and a locked (dead/frozen) state. Because now >> incrementing counter doesn't affect it's locked/unlocked state, it is >> possible to use an optimistic atomic_fetch_add() in >> page_ref_add_unless_zero() that operates independently of the locked bit. >> The locked state is handled after the increment attempt, eliminating the >> need for the CAS loop. > > I don't think I follow. > > Your trick with the PAGEREF_LOCKED_BIT helps with serialization against > page_ref_freeze(), but I don't think it does anything to serialize > against freeing the page under you. > > Like, if the page in the process of freeing, page allocator sets its > refcount to zero and your version of page_ref_add_unless_zero() > successfully acquirees reference for the freed page. > > How is it safe? Page is freed only after a successful page_ref_dec_and_test() call, which will set LOCKED_BIT. This bit will persist until set_page_count(1) is called somewhere in the allocation path [alloc_pages()], and effectively block any "use after free" users.