From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pl1-f176.google.com (mail-pl1-f176.google.com [209.85.214.176]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8343C3CF21E for ; Tue, 25 Aug 2026 09:17:03 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.214.176 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649425; cv=none; b=RaHTPzR1L5eOR9/fgwu0VBCDI5fvXAIVzRB7rVT8qwnEkxQ5Qc1gk287Ac0rnKdEHvLKDQo62vdhLPvkJXxKD9rlttUrIzQKRCcez+npCRMlrllHYEBGmLVzEk4H3OLujNacsGBRbL7DYSNvgiqxUnlozUPoh5ZBlaLwuqKB/gA= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787649425; c=relaxed/simple; bh=n39d1Y3WRz91AbqgzkVngKPips7SwY3qmP+NR2SNYK0=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=Etf5vGGQQI6MzaSGEhspe0OUMugnaBiDZ08+DKfpwrGKBMvgINVlY8cibfE36BfYk/V3kV1k7812U17ZjS7y1yexc16NGSS9H5VsTrGozCh/qO3no1kOa4CvhsrDUzB/9tCP33WdGRsY3j1idEWOvmWb5WbS87b5z0L77hAFI/w= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=qIjadi2l; arc=none smtp.client-ip=209.85.214.176 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="qIjadi2l" Received: by mail-pl1-f176.google.com with SMTP id d9443c01a7336-2d6f624c323so3567375ad.0 for ; Tue, 25 Aug 2026 02:17:03 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787649423; x=1788254223; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=pJOMnBC8+UOkkBmi9TvRXsVyiFLnHyyDV0TZJc/rZ/c=; b=qIjadi2lB0cGs4DLatwKeh0Vq+THJeya3EzuocQLp4P8eDATEuAeyluHng7EKP0rxa QmRYtEhflDD0EqqdQQ3yKQNyzMhi2xbbV9XOTnR4Ij4/V2/3OxDJvvzTLNYa/k3GPMN0 wWJPYwgXg21MdZ90Q5xp550vRvQ+JtsL9SGzfbwUEOuhl671y4X54iwBpqSPeU9koygy I4MEAkw0HQjA4DDuL/iKJRypq4AnNldStNu3+NrtSJ0K90TplC/VeV39+hr49aGTurSC 0BZMqLQhHezcvmifUSep0bhe6BIscP3/UQ8dXEK2GRVzwE67YARij1MLQ/16rR0gSvGh ATZQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787649423; x=1788254223; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:cc:to:from:x-gm-gg:x-gm-message-state:from :to:cc:subject:date:message-id:reply-to:content-type; bh=pJOMnBC8+UOkkBmi9TvRXsVyiFLnHyyDV0TZJc/rZ/c=; b=Ga/zodTzOzATfTIZUyW9MwkfZNLpSeSbMuPTKOxWI3qQH22trV9ls6e80HreMIbs4m 7J7a0Xp66CflxqgUHvtwe8cS3S5F2VkIRN3uXK5nH571mZE1tuBM/UvNTNIWfWGaDHY4 C3CCEMo0G0LbjXgCPEJ0TeuMj0c7Z414YTVTj0XxxAYqX5JPII2nRhaa/O+/i6UXwhsM yslggRHSEHWQQpIWfcBrkZWL5Zv1yvchA1XbywFXuYEJPZ6rqzNH94E7UY2bEXwaXXb6 6BzMDkIRBgB5Ym/4VQqqmF5m9uCgZ3oC2hPOBdXwewklgvPdH2la5Yw5e1yh0GaR32ln qWtA== X-Forwarded-Encrypted: i=1; AHgh+RrqD1Csf0Ao9HVIk5oIJzVDEWqdQHcYJgGrChmTiv7I64tfPtXJWbJz/WP3XhX130EqFTpHNZEHpnhZ2d4=@vger.kernel.org X-Gm-Message-State: AFuF++kHSmfzX++OmSYxoruBFzTKSDoqUSLsz8mg9Ta96FketREQk/d4 w+gGWmAHFnIO0hUQqqEUoEZW4i9rjQkQ4jBVAUse+uiJ3elDSa2uj4L6 X-Gm-Gg: AR+sD13Wn82L87C0TwsHOtWHsocyFl5QXk2JKl9fwJq9+w1f1ONlVBGpogeiBRwr8LA kXbe+75Eddbfv4Us76XckNR79HYhqfqT4H30FvgVfGAQ36WpEKPFuV/vpoLxSb47/extcHLvq2S dYBd8k5BZhNZf3nZrRy2d2ocdPVY+PQF2QoqY+SwcxZPzDrqhYhjdUCF4vNw6lBxHxxGUgx8vq7 e+IqRtAZulyyeK+R6YVx6tzGAW2oiBnl+fe01rFPOmsBmhvo1yHVSrf19Frupn6rUnDHdbqpm4h XRl3/80CKQpKf7S2Ky7n/fRRQYPruKWpO0yPrdlBZEJQxIBIyqF+K0DhjT3wExWkI7FxySzlXXL wd9J44ohDKfvx7eOJ2tSJFYLNxXYSMrAKydZb38qZEeonuUOmPU1ZWfRgQVYz7DFlj+TwAxQmUv bnRLpCKN8ZzlXdBYcENj5uKbvPsYuqgLu1JA4wGI1rwZWp5J0y4hTVZPx7LfZVr7xRvh2Uq2kGH n2/9AhDKVkACDJuDKM= X-Received: by 2002:a17:90b:4b02:b0:36b:bec8:94c5 with SMTP id 98e67ed59e1d1-395df2595femr42727078a91.10.1787649422635; Tue, 25 Aug 2026 02:17:02 -0700 (PDT) Received: from localhost.localdomain ([103.120.31.178]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-141860fe08csm52122658c88.8.2026.08.25.02.16.58 (version=TLS1_3 cipher=TLS_CHACHA20_POLY1305_SHA256 bits=256/256); Tue, 25 Aug 2026 02:17:02 -0700 (PDT) From: Khawar Ahemad To: bpf@vger.kernel.org, linux-kernel@vger.kernel.org Cc: ast@kernel.org, daniel@iogearbox.net, andrii@kernel.org, eddyz87@gmail.com, jiayuan.chen@linux.dev, emil@etsalapatis.com Subject: [PATCH v6 2/4] bpf: arena: allocate the fault-in page outside the lock Date: Tue, 25 Aug 2026 14:46:45 +0530 Message-ID: <20260825091647.81632-3-ahemadkhawar123@gmail.com> X-Mailer: git-send-email 2.54.0 In-Reply-To: <20260825091647.81632-1-ahemadkhawar123@gmail.com> References: <20260825091647.81632-1-ahemadkhawar123@gmail.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit From: Jiayuan Chen arena_vm_fault() allocated the page while holding arena->spinlock, so it could only use the non-blocking allocator. Once the memcg is at memory.max that allocation just fails, the fault turns into VM_FAULT_SIGSEGV, and the process gets a SIGSEGV on a perfectly valid arena address. Hitting memory.max is routine (e.g. page cache from reading a big file), so this kills innocent processes. Rework the fault handler: - Preallocate the page before taking the lock, like do_anonymous_page() does, so it can sleep and go through reclaim and the memcg OOM killer, instead of turning a routine memory.max into a fake segfault. - On allocation failure return VM_FAULT_SIGBUS. The allocation already ran reclaim and the OOM killer, so the failure is non-recoverable. For a task faulting its own arena this changes nothing: the OOM killer already picked it inside the allocation and it dies by SIGKILL, the SIGBUS is shadowed by the pending fatal signal, and the memcg OOM is still reported. VM_FAULT_OOM would instead be retried by the fault path and can livelock when the charged memcg is not the faulting task's (e.g. a shared arena) and its OOM killer cannot reach it. - A lockless probe skips that preallocation when a page is already mapped (e.g. allocated by the bpf program), so the common case wastes no allocation. The rare race where such a page is freed before we take the lock falls back to the non-blocking allocator under the lock. - Return VM_FAULT_SIGBUS for the other non-recoverable errors (lock failure, range-tree and page-table failures) instead of VM_FAULT_SIGSEGV; only BPF_F_SEGV_ON_FAULT, and a scratch-page hole under that flag, is a real user addressing error and keeps VM_FAULT_SIGSEGV. - Tidy up the error labels. Reviewed-by: Emil Tsalapatis Signed-off-by: Jiayuan Chen --- kernel/bpf/arena.c | 90 +++++++++++++++++++++++++++++++++++----------- 1 file changed, 69 insertions(+), 21 deletions(-) diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c index 7b6847200b..fa462a0ff1 100644 --- a/kernel/bpf/arena.c +++ b/kernel/bpf/arena.c @@ -481,7 +481,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf) struct bpf_map *map = vmf->vma->vm_file->private_data; struct bpf_arena *arena = container_of(map, struct bpf_arena, map); struct mem_cgroup *new_memcg, *old_memcg; - struct page *page; + struct page *page, *new_page = NULL; + vm_fault_t fault_ret; long kbase, kaddr; unsigned long flags; int ret; @@ -489,59 +490,106 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf) kbase = bpf_arena_get_kern_vm_start(arena); kaddr = kbase + (u32)(vmf->address); - if (raw_res_spin_lock_irqsave(&arena->spinlock, flags)) + page = vmalloc_to_page((void *)kaddr); + if (!page && !(arena->map.map_flags & BPF_F_SEGV_ON_FAULT)) { + /* + * Preallocate outside the lock with a sleepable allocator so it + * can reclaim and run the memcg OOM killer, which the + * non-blocking allocator under arena->spinlock cannot. A NULL + * return is non-recoverable, so fail with VM_FAULT_SIGBUS; + * VM_FAULT_OOM would be retried by the fault path and can + * livelock when the charged memcg is not the faulting task's. + */ + bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg); + new_page = bpf_map_alloc_page_sleepable(map); + bpf_map_memcg_exit(old_memcg, new_memcg); + if (!new_page) + return VM_FAULT_SIGBUS; + } + + if (raw_res_spin_lock_irqsave(&arena->spinlock, flags)) { /* * A failed lock means a possible deadlock was detected. Don't * return VM_FAULT_RETRY: this handler never took mmap_lock, but * the fault path would re-take it on retry and deadlock. Fail. */ + if (new_page) + free_pages_nolock(new_page, 0); return VM_FAULT_SIGBUS; + } page = vmalloc_to_page((void *)kaddr); if (page) { - if (page == arena->scratch_page) - /* BPF triggered scratch here; don't lazy-alloc over it */ - goto out_sigsegv; + if (page == arena->scratch_page) { + /* + * A scratch page marks a hole. Segfault only if the user + * asked for it; otherwise we could lazy-allocate but + * choose not to over a hole, so report a bus error. + */ + fault_ret = (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) ? + VM_FAULT_SIGSEGV : VM_FAULT_SIGBUS; + goto out_err_locked; + } /* already have a page vmap-ed */ goto out; } + if (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) { + /* User space requested to segfault when page is not allocated by bpf prog */ + fault_ret = VM_FAULT_SIGSEGV; + goto out_err_locked; + } + bpf_map_memcg_enter(&arena->map, &old_memcg, &new_memcg); - if (arena->map.map_flags & BPF_F_SEGV_ON_FAULT) - /* User space requested to segfault when page is not allocated by bpf prog */ - goto out_sigsegv_memcg; + if (!new_page) { + /* + * Very rare race: the bpf program had allocated a page here, so + * the lockless probe saw it and we skipped preallocation, but it + * freed the page before we took the lock. Now we do need one; + * sleeping is not allowed here, so fall back to the non-blocking + * allocator and give up if it fails. + */ + ret = bpf_map_alloc_pages(map, map->numa_node, 1, &new_page); + if (ret) { + fault_ret = VM_FAULT_SIGBUS; + goto out_err_locked_memcg; + } + } ret = range_tree_clear(&arena->rt, vmf->pgoff, 1); - if (ret) - goto out_sigsegv_memcg; - - struct apply_range_data data = { .arena = arena, .pages = &page, .i = 0 }; - /* Account into memcg of the process that created bpf_arena */ - ret = bpf_map_alloc_pages(map, NUMA_NO_NODE, 1, &page); if (ret) { - range_tree_set(&arena->rt, vmf->pgoff, 1); - goto out_sigsegv_memcg; + fault_ret = VM_FAULT_SIGBUS; + goto out_err_locked_memcg; } + struct apply_range_data data = { .arena = arena, .pages = &new_page, .i = 0 }; ret = apply_to_page_range(&init_mm, kaddr, PAGE_SIZE, apply_range_set_cb, &data); if (ret) { range_tree_set(&arena->rt, vmf->pgoff, 1); - free_pages_nolock(page, 0); - goto out_sigsegv_memcg; + fault_ret = VM_FAULT_SIGBUS; + goto out_err_locked_memcg; } flush_vmap_cache(kaddr, PAGE_SIZE); bpf_map_memcg_exit(old_memcg, new_memcg); + /* new_page was consumed */ + page = new_page; + new_page = NULL; out: page_ref_add(page, 1); raw_res_spin_unlock_irqrestore(&arena->spinlock, flags); + if (new_page) + free_pages_nolock(new_page, 0); vmf->page = page; return 0; -out_sigsegv_memcg: + +out_err_locked_memcg: bpf_map_memcg_exit(old_memcg, new_memcg); -out_sigsegv: +out_err_locked: raw_res_spin_unlock_irqrestore(&arena->spinlock, flags); - return VM_FAULT_SIGSEGV; + if (new_page) + free_pages_nolock(new_page, 0); + return fault_ret; } static const struct vm_operations_struct arena_vm_ops = { -- 2.54.0 (Apple Git-157)