mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Usama Arif <usamaarif642@gmail.com>
To: Andrew Morton <akpm@linux-foundation.org>
Cc: yuzhao@google.com, david@redhat.com, leitao@debian.org,
	huangzhaoyang@gmail.com, bharata@amd.com, willy@infradead.org,
	vbabka@suse.cz, linux-kernel@vger.kernel.org,
	kernel-team@meta.com, Johannes Weiner <hannes@cmpxchg.org>,
	zhaoyang.huang@unisoc.com, Rik van Riel <riel@surriel.com>
Subject: Re: [PATCH RESEND] mm: drop lruvec->lru_lock if contended when skipping folio
Date: Tue, 20 Aug 2024 11:45:11 -0400	[thread overview]
Message-ID: <29e481af-b5e1-4320-a672-8251f5099595@gmail.com> (raw)
In-Reply-To: <20240819181743.926f37da3b155215c088c809@linux-foundation.org>



On 20/08/2024 02:17, Andrew Morton wrote:
> On Mon, 19 Aug 2024 19:46:48 +0100 Usama Arif <usamaarif642@gmail.com> wrote:
> 
>> lruvec->lru_lock is highly contended and is held when calling
>> isolate_lru_folios. If the lru has a large number of CMA folios
>> consecutively, while the allocation type requested is not MIGRATE_MOVABLE,
>> isolate_lru_folios can hold the lock for a very long time while it
>> skips those. vmscan_lru_isolate tracepoint showed that skipped can go
>> above 70k in production and lockstat shows that waittime-max is x1000
>> higher without this patch.
>> This can cause lockups [1] and high memory pressure for extended periods of
>> time [2]. Hence release the lock if its contended when skipping a folio to
>> give other tasks a chance to acquire it and not stall.
>>
>> ...
>>
>> --- a/mm/vmscan.c
>> +++ b/mm/vmscan.c
>> @@ -1695,8 +1695,14 @@ static unsigned long isolate_lru_folios(unsigned long nr_to_scan,
>>  		if (folio_zonenum(folio) > sc->reclaim_idx ||
>>  				skip_cma(folio, sc)) {
>>  			nr_skipped[folio_zonenum(folio)] += nr_pages;
>> -			move_to = &folios_skipped;
>> -			goto move;
>> +			list_move(&folio->lru, &folios_skipped);
>> +			if (!spin_is_contended(&lruvec->lru_lock))
>> +				continue;
>> +			if (!list_empty(dst))
>> +				break;
>> +			spin_unlock_irq(&lruvec->lru_lock);
>> +			cond_resched();
>> +			spin_lock_irq(&lruvec->lru_lock);
>>  		}
> 
> Oh geeze ugly thing.  Must we do this?
> 
> The games that function plays with src, dst and move_to are a bit hard
> to follow.  Some tasteful comments explaining what's going on would
> help.
> 
> Also that test of !list_empty(dst).  It would be helpful to comment the
> dynamics which are happening in this case - why we're testing dst here.
> 
> 

So Johannes pointed out to me that this is not going to properly fix the problem of holding the lru_lock for a long time introduced in [1] because of 2 reasons:
- the task that is doing lock break is hoarding folios on folios_skipped and making the lru shorter, I didn't see it in the usecase I was trying, but it could be that yielding the lock to the other task is not of much use as it is going to go through a much shorter lru list or even an empty lru list and would OOM, while the folio it is looking for is on folios_skipped. We would be substituting one OOM problem for another with this patch.
- Compaction code goes through pages by pfn and not using the list, as this patch does not clear lru flag, compaction could claim this folio.

The patch in [1] is severely breaking production at Meta and its not a proper fix to the problem that the commit was trying to be solved. It results in holding the lru_lock for a very significant amount of time, stalling all other processes trying to claim memory, creating very high memory pressure for large periods of time and causing OOM.

The way forward would be to revert it and try to come up with a longer term solution that the original commit tried to solve. If no one is opposed to it, I will wait a couple of days for comments and send a revert patch.

[1] https://lore.kernel.org/all/1685501461-19290-1-git-send-email-zhaoyang.huang@unisoc.com/

Thanks,
Usama

  reply	other threads:[~2024-08-20 15:45 UTC|newest]

Thread overview: 7+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2024-08-19 18:46 Usama Arif
2024-08-20  1:17 ` Andrew Morton
2024-08-20 15:45   ` Usama Arif [this message]
2024-08-20 17:13     ` Breno Leitao
2024-08-21  0:51     ` Zhaoyang Huang
2024-08-21 18:33       ` Usama Arif
2024-08-20  4:38 ` Bharata B Rao

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=29e481af-b5e1-4320-a672-8251f5099595@gmail.com \
    --to=usamaarif642@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=bharata@amd.com \
    --cc=david@redhat.com \
    --cc=hannes@cmpxchg.org \
    --cc=huangzhaoyang@gmail.com \
    --cc=kernel-team@meta.com \
    --cc=leitao@debian.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=riel@surriel.com \
    --cc=vbabka@suse.cz \
    --cc=willy@infradead.org \
    --cc=yuzhao@google.com \
    --cc=zhaoyang.huang@unisoc.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®