From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-qk1-f181.google.com (mail-qk1-f181.google.com [209.85.222.181]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 4C09B35DA69 for ; Wed, 17 Jun 2026 18:04:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.222.181 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781719463; cv=none; b=X4+Lv84al4ava8XpybLs2T4xzlyebW99rG5iaDffSf0MKiVKgXgyAQQwGqtjgcxNZuPCGFxdYqDKQN1OO5cjl0eMVxquZ3vZHb62FIRV42I3ZSRL+yO7e8jTuKpD59QYwQfHA9eqmNR2YCgDpf5+y+lU/WYbgFLnZlC9FwEUREM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781719463; c=relaxed/simple; bh=nW/o6mN/VLj0TH3DDkLkpNNnQ7XaRbBINRNcg8MuK0M=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=RNpi0iGXfFjrxR1MftfG/OUUhCgtmFeZtbrbnh/i/LakhCOEHR7a/LgKwwm30cR5y/vmkk0XRbIR0nUTkISEk3+Z1p5uzV3cl2XgQV5hbb4cr1SBAI+jAJNC2sXN/+Cx4N5U/58bTL03U4Be1pe8BuUGH6NNjzMnzuyOy9+MQmA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ziepe.ca; spf=pass smtp.mailfrom=ziepe.ca; dkim=pass (2048-bit key) header.d=ziepe.ca header.i=@ziepe.ca header.b=N3i5JS8W; arc=none smtp.client-ip=209.85.222.181 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=ziepe.ca Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=ziepe.ca Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=ziepe.ca header.i=@ziepe.ca header.b="N3i5JS8W" Received: by mail-qk1-f181.google.com with SMTP id af79cd13be357-915c36e32abso9435185a.2 for ; Wed, 17 Jun 2026 11:04:22 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=ziepe.ca; s=google; t=1781719461; x=1782324261; darn=vger.kernel.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=43aClsIewUmRyeXaG13jiRHX9eRfmYOa6P/4ggqgHgg=; b=N3i5JS8WiPW9wMtDmP+tLncBT1v2nChl5x+U4qaXrL1oLwC7yGfPQECKqWVn4WvJCu 8MIuuCTf3v0ir2b+tiShNupP3Y8xf6B0zZFvas47hWcBSdmMHGeD6IDd+HLMi67QZeSw ArKaG/BDlbyuFFMydtgRQiYd0qFhQRCY51QreuD6ZjLwji/VnUgDXeXMWPtWnjyai74z eSQcxWOtxmXlUdIS25HoxCj73Na8ZVO/GRmvoxb4oLG8yuH3AyPbWPuLtLD/wcaxI/h1 yXQWG0+xtLOcN5wITN/YLnDvrLVZy6BHhaZepx6kkkn5L0rP+2UGGH5uWdXDZT2XEb/A IkXw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781719461; x=1782324261; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=43aClsIewUmRyeXaG13jiRHX9eRfmYOa6P/4ggqgHgg=; b=mZxOFn8raOaw+3zyzdSltfK3mSVVCdj0n61S5EwxPfonXkqdF7MQzdBkuFeClAmMvJ RZWC+1jHrbWp8v8APktnI1btcYjrBC7d2XpAOdER+jnTdQCz68xmcKsYyUnPZP4cm4n6 uvxllHl1PTYFy0EtOie0gu/Y0tWsoiebyz1It3kpd4q07SUKsU3M0Y7n/75xK7KQjcwd PRC51nHvEVNZJmI0F8GNe7Oxnb/iuq2R9+grONqrR9+nax1kckurenJSp6Y4brPl0RQo GrbC+CQ5K69RJjcvpUVyu5CmQDOaafN3yDpAbxDCDFPe03CLuDAv0NvgejBfiaRkCWSx /UHQ== X-Forwarded-Encrypted: i=1; AFNElJ9p09wL8nNMsRdBabO3Y0Zltwn+JfIcTg0LY3iiqzczSt9CGtH+OcoVQh5pLER4jNngS9mxLtJHP29EOt4=@vger.kernel.org X-Gm-Message-State: AOJu0YxPz3VF2eKWUsOZGhvnHThSTr4nB/hte5dSGcsQqC+WOi6+RreW EjEqs3vbwPI/OG9YNLvora4sn1YOtzBkd1U7Nj39OMA6QNdI41WcApZ5GcVNjEPJzJw= X-Gm-Gg: Acq92OF39WNIIL2hAlhkOB64C9mBLo4Z8mTLu9rP0qxf3l9aOduRF4F4+KwCh3/Vh1l onWFDz5v+4vtrRcgVHy7KOyVMCznhzIKXLxwetS8r4a8ihks8byd0a7fDtwps35sUH30bgeQr/L ujMZ/EL3g36KxMuowtpdHUvPw4rKECk/SYKYrs4AaltNP8KqTKW14qMwusohFMJ/qAafqfLoo5i ITqaQNo8D/LjIZ9OxtCS7VgzDcGI6BlbF7e2q8Pwj9Juul1bzTf0dI5+W0s+PNjG6RfQdczzNVj oV8m+2HITTDK/7xqzwUNNLbHIMz2Yh+/HX1lE9oL7dAQUYvW0XNDYvuP75XGWgUNiR+HD2yMCHk /nlP7UJiFHLMnD4V9IrgOa6lSLDK4qxaF6Rmohbuf6bvGI0QIEUFeTWhRTrzAfPcumzTrtin+VG 3/ZL8KQg2LquTVSK0ddNR7RHSDKm9emp0gAApPdKd+j1P6cPiLOS+g1ZdBzhqtgLeJ+xMvBCotF ura9Q== X-Received: by 2002:a05:620a:2685:b0:915:c365:ffba with SMTP id af79cd13be357-91d8bfad7e6mr871504285a.55.1781719460851; Wed, 17 Jun 2026 11:04:20 -0700 (PDT) Received: from ziepe.ca (crbknf0213w-47-54-130-67.pppoe-dynamic.high-speed.nl.bellaliant.net. [47.54.130.67]) by smtp.gmail.com with ESMTPSA id af79cd13be357-91619f06835sm1837505185a.14.2026.06.17.11.04.20 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 17 Jun 2026 11:04:20 -0700 (PDT) Received: from jgg by wakko with local (Exim 4.97) (envelope-from ) id 1wZucd-000000013fl-2f0g; Wed, 17 Jun 2026 15:04:19 -0300 Date: Wed, 17 Jun 2026 15:04:19 -0300 From: Jason Gunthorpe To: "Liam R. Howlett" Cc: Rik van Riel , linux-kernel@vger.kernel.org, kernel-team@meta.com, robin.murphy@arm.com, joro@8bytes.org, will@kernel.org, iommu@lists.linux.dev, kyle@mcmartin.ca, Rik van Riel , maple-tree@lists.infradead.org Subject: Re: [PATCH v3 3/3] iova: defer maple tree erase on GFP_ATOMIC failure Message-ID: <20260617180419.GA231643@ziepe.ca> References: <20260603033653.4144138-1-riel@surriel.com> <20260603033653.4144138-4-riel@surriel.com> <20260609130418.GI2764304@ziepe.ca> <61d51d4b5779d80145ceb38e9632a7cc8a79dbec.camel@surriel.com> <20260612164852.GL1066031@ziepe.ca> <20260612180303.GO1066031@ziepe.ca> <20260615115633.GQ1066031@ziepe.ca> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: On Wed, Jun 17, 2026 at 01:45:27PM -0400, Liam R. Howlett wrote: > On 26/06/15 08:56AM, Jason Gunthorpe wrote: > > On Fri, Jun 12, 2026 at 02:44:06PM -0400, Liam R. Howlett wrote: > > > > Currently it never returns a failure to the caller. Look at mas_erase(): > > > > > > > > entry = mas_state_walk(mas); > > > > if (!entry) > > > > return NULL; > > > > [..] > > > > if (mas_is_err(mas)) > > > > goto out; > > > > [..] > > > > out: > > > > mas_destroy(mas); > > > > return entry; > > > > > > > > There is no propogation of ENOMEM, it returns success. No caller > > > > checks for any error here either. > > > > > > At one point this was considered to be impossible to fail, and it is > > > documented to return the entry or null. > > > > I think that is the right API design.. > > Callers can check mas_is_err() and check for xa_err(mas) == -ENOMEM. > I'm going to add a note about it to the documentation of the erase > function. That's something for mas_erase, but doesn't help mtree_erase() .. > Why is the retry GFP_ATOMIC on a timer of 10ms? > > Also, why is this patch set using an external spinlock only created to > manage the tree? Why isn't it using an internal lock? Is it just to > avoid the possibility of the unlock? IDK, seem like good questions > > > > > I think, in your case, hitting an XA_ZERO_ENTRY would be necessary to > > > indicate that we cannot reuse this particular location until it is > > > correctly dealt with? Or is the maple tree the only reason it is > > > considered unusable? > > > > Yeah, it would be be a maple tree issue only. Defered rebalancing > > leave space unavailable. > > > > This is a case where there is no sane way to handle destroy > > failure. You can't return an error code from dma_unmap() for > > example. So the only reason for this complexity is because maple tree > > exposes a failable erase to it's caller.. > > I think this is even more complicated by the contexts it is called in - > that is, we cannot preallocate prior to going into this state either? Yes, in this case at least the context is GFP_ATOMIC and there is no way to pre-allocate.. But that seems like another issue since the mas_erase does not support GFP_ATOMIC anyhow.. > > Eg _iommufd_destroy_mmap() is in trouble too, it cargo culted the > > no-check mt_erase. > > Are you sure that's not okay? The mt_mmap tree is allocated with an > internal spinlock. In this case, the lock will be dropped, the > allocation will be satisfied and the erase operation will retry. Is this what I was asking before? Under some conditions the allocation can not fail because in the modern kernel we don't allow small GFP_KERNEL allocations to fail? Otherwise this: if (gfpflags_allow_blocking(gfp) && !mt_external_lock(mas->tree)) { mtree_unlock(mas->tree); mas_alloc_nodes(mas, gfp); mtree_lock(mas->tree); } else { mas_alloc_nodes(mas, gfp); } Is always called with GFP_KERNEL for erase. It doesn't seem like external lock has any impact if mas_alloc_nodes can fail or not? It looks like if you have an external lock then the hard wired GFP_KERNEL in mtree_erase/mas_erase mean the lock has to be a sleeping kind to use those functions. If that's the case it should be documented like this too :) Jason