From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from canpmsgout07.his.huawei.com (canpmsgout07.his.huawei.com [113.46.200.222]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1B63531987B for ; Fri, 30 Jan 2026 09:03:55 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=113.46.200.222 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1769763838; cv=none; b=MUJxPJPhSMOemtXVoeh6l73AKZnrmJZvnimNUTBbBGDQSICJc6Meo/lpbWuxZ8OusEQp2XGXPTLUUNe3L8kxw/INKfoKd5gq3rcWd23R+4G8b8l9/CtYKFAPM2yrrr8YJodh6WbjTam5h2WzfALQQlgLPldxcqDnk5aegF/4qGI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1769763838; c=relaxed/simple; bh=rTJiKCocAjxqV79JJxEsffsTKnnPxbyaPwpDTcn5s6s=; h=Message-ID:Date:MIME-Version:Subject:To:CC:References:From: In-Reply-To:Content-Type; b=c73agjDrR3Y5ME+xRgoQ5fYIdXnu4dBbgll27FbYCkpz8XMlvRO9z47MEEfcSfDFXShN3alKPpPYQMJODLhoi1TjizM7R5t+WpVavsP2PkmAESE1nbP3WmxZYzik08NotIm0bykIOCc2fAqbgnuAfzpiGdHdIedR5l+77hHRPLM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com; spf=pass smtp.mailfrom=huawei.com; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b=RP0LNUFB; arc=none smtp.client-ip=113.46.200.222 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=huawei.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=huawei.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=huawei.com header.i=@huawei.com header.b="RP0LNUFB" dkim-signature: v=1; a=rsa-sha256; d=huawei.com; s=dkim; c=relaxed/relaxed; q=dns/txt; h=From; bh=3u9we7d0FFmdmbQZ9gAmDCrJy57gOtA2Jv54WAcLIm8=; b=RP0LNUFBI5kS9e0LLpwT4f0RY3gFOJ+ToAGOtxz2CFnxVRDUPq74GeYHumB1kU4nvbuGdRxpC fy7oSuSzj7D6USfczYoC/31sdtl++FAhJgLPk4yrczHosDQFWI6RnzU26TNbHeCSyruPtVSLqMC NLZsrFbya7UvBdOr6StMgm4= Received: from mail.maildlp.com (unknown [172.19.163.163]) by canpmsgout07.his.huawei.com (SkyGuard) with ESMTPS id 4f2VMr75SdzLlwd; Fri, 30 Jan 2026 17:00:24 +0800 (CST) Received: from dggemv705-chm.china.huawei.com (unknown [10.3.19.32]) by mail.maildlp.com (Postfix) with ESMTPS id 83DFA4048B; Fri, 30 Jan 2026 17:03:53 +0800 (CST) Received: from kwepemq100012.china.huawei.com (7.202.195.195) by dggemv705-chm.china.huawei.com (10.3.19.32) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.11; Fri, 30 Jan 2026 17:03:51 +0800 Received: from [10.67.111.196] (10.67.111.196) by kwepemq100012.china.huawei.com (7.202.195.195) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.1544.11; Fri, 30 Jan 2026 17:03:50 +0800 Message-ID: <1594f461-549c-4db9-b80e-63c48818fc5b@huawei.com> Date: Fri, 30 Jan 2026 17:03:49 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH] sched: Re-evaluate scheduling when migrating queued tasks out of throttled cgroups To: CC: K Prateek Nayak , , , , , , , , , , , , , , , References: <20260120032549.186733-1-quzicheng@huawei.com> <20260130083438.1122457-1-quzicheng@huawei.com> From: Zicheng Qu In-Reply-To: <20260130083438.1122457-1-quzicheng@huawei.com> Content-Type: text/plain; charset="UTF-8"; format=flowed Content-Transfer-Encoding: 8bit X-ClientProxiedBy: kwepems100001.china.huawei.com (7.221.188.238) To kwepemq100012.china.huawei.com (7.202.195.195) On 1/30/2026 4:34 PM, Zicheng Qu wrote: > 4) For kernel <= 5.10: Later, cgroup A is unthrottled. However, the task > P has already been migrated out of cgroup A, so unthrottle_cfs_rq() > may observe load_weight == 0 and return early without resched_curr() > called. For kernel >= 6.6: The unthrottling path normally triggers > `resched_curr()` almost cases even when no runnable tasks remain in the > unthrottled cgroup, preventing the idle stall described above. However, > if cgroup A is removed before it gets unthrottled, the unthrottling path > for cgroup A is never executed. In a result, no `resched_curr()` can be > called. Hi Aaron, Apologies for the confusion in my earlier description — the original failure model was identified and analyzed on kernels based on LTS 5.10. Later I realized that on v6.6 and mainline, the issue becomes much harder to reproduce due to additional conditions introduced in the condition (cfs_rq->on_list) in unthrottle_cfs_rq(), which effectively mask the original reproduction path. As a result, I adjusted the reproducer accordingly. With the updated reproducer, the issue can still be triggered on mainline by explicitly bypassing the unthrottling reschedule path, as described in the commit message. The reproducer can be run directly via: ./make.sh My local /proc/cmdline is: systemd.unified_cgroup_hierarchy=0 nohz_full=2-15 rcu_nocbs=2-15 With this setup, the issue is reproducible on current mainline. make.sh ```sh #!/bin/bash gcc -O2 heartbeat.c -o heartbeat chmod +x ./run_test.sh && ./run_test.sh ``` heartbeat.c ```c #define _GNU_SOURCE #include #include #include #include static inline long long now_ns(void) {     struct timespec ts;     clock_gettime(CLOCK_MONOTONIC, &ts);     return ts.tv_sec * 1000000000LL + ts.tv_nsec; } int main(void) {     cpu_set_t set;     CPU_ZERO(&set);     CPU_SET(12, &set);  // CPU 12 is nohz_full     sched_setaffinity(0, sizeof(set), &set);     long long last = now_ns();     unsigned long long iter = 0;     while (1) {         iter++;         long long now = now_ns();         if (now - last > 1000 * 1000 * 1000) { // 1000ms             printf("[HB] sec=%lld pid=%d cpu=%d iter=%llu\n", now / 1000000000LL, getpid(), sched_getcpu(), iter);             fflush(stdout);             last = now_ns();         }     } } ``` ```sh #!/bin/bash # # run_test.sh # # Reproducer for a scheduling stall on nohz_full CPUs when migrating # queued tasks out of throttled cgroups. # # Test outline: #   1. Start a CPU-bound workload (heartbeat) that prints a heartbeat (HB) #      once per second. #   2. Migrate the task into a heavily throttled child cgroup. #   3. Migrate the task back to the root cgroup (potential trigger point). #   4. Immediately remove (destroy) the throttled cgroup before it gets #      unthrottled. #   5. Observe whether the heartbeat continues to advance. #      - If HB advances: no stall, continue to next round. #      - If HB stops advancing: scheduling stall detected, freeze the setup #        for debugging. # set -e ######################## # Basic configuration ######################## ROOT_CG=/sys/fs/cgroup/cpu THROTTLED_CG=$ROOT_CG/child_cgroup mkdir -p "$ROOT_CG" HB_LOG=heartbeat.log # Throttle settings: 1ms runtime per 1s period CFS_QUOTA_US=1000 CFS_PERIOD_US=1000000 # Timeout (in seconds) to consider the workload "stuck" STUCK_TIMEOUT=10 CHECK_INTERVAL=0.2 ######################## # Cleanup logic ######################## PID= cleanup() {     echo     echo "[!] cleanup: stopping workload"     if [[ -n "$PID" ]] && kill -0 "$PID" 2>/dev/null; then         echo "[!] killing pid $PID"         kill -TERM "$PID"         wait "$PID" 2>/dev/null || true     fi     echo "[!] cleanup done" } trap cleanup INT TERM EXIT ######################## # Start workload ######################## echo "[+] starting heartbeat workload" ./heartbeat | tee "$HB_LOG" & PID=$(($! - 1)) # temporary hack PID echo "[+] workload pid = $PID" echo ######################## # Helper functions ######################## # Extract the last printed heartbeat second from the log last_hb_sec() {     tail -n 1 "$HB_LOG" 2>/dev/null | awk '{         for (i = 1; i <= NF; i++) {             if ($i ~ /^sec=/) {                 split($i, a, "=");                 print a[2];                 exit;             }         }     }' } verify_cgroup_location() {     echo "  root cgroup:"     cat "$ROOT_CG/tasks" | grep "$PID" || true     echo "  throttled cgroup:"     cat "$THROTTLED_CG/tasks" | grep "$PID" || true } ######################## # Main test loop ######################## round=0 while true; do     # Recreate the throttled cgroup for the next iteration     mkdir -p "$THROTTLED_CG"     echo $CFS_QUOTA_US  > "$THROTTLED_CG/cpu.cfs_quota_us"     echo $CFS_PERIOD_US > "$THROTTLED_CG/cpu.cfs_period_us"     round=$((round + 1))     echo "========== ROUND $round =========="     echo "[1] move task into throttled cgroup"     echo "$PID" > "$THROTTLED_CG/tasks"     echo "[1.1] verify cgroup placement"     verify_cgroup_location     # Give the task some time to consume its quota and become throttled     sleep 0.2     echo "[2] migrate task back to root cgroup (potential trigger)"     echo "$PID" > "$ROOT_CG/tasks"     echo "[2.1] verify cgroup placement"     verify_cgroup_location     #     # IMPORTANT:     # For kernels >= 6.6, unthrottling normally triggers resched_curr().     # Removing the throttled cgroup before it gets unthrottled bypasses     # the unthrottle path and is required to reproduce the stall.     #     echo "[2.2] remove throttled cgroup before unthrottling"     rmdir "$THROTTLED_CG"     # Observe heartbeat after migration back to root     base_hb=$(last_hb_sec)     [[ -z "$base_hb" ]] && base_hb=0     echo "[3] observing heartbeat (base_hb_sec=$base_hb)"     start_ts=$(date +%s)     while true; do         cur_hb=$(last_hb_sec)         [[ -z "$cur_hb" ]] && cur_hb=0         if (( cur_hb > base_hb )); then             echo "[OK] heartbeat advanced: $base_hb -> $cur_hb"             break         fi         now_ts=$(date +%s)         if (( now_ts - start_ts >= STUCK_TIMEOUT )); then             echo             echo "[!!!] SCHEDULING STALL DETECTED AFTER MIGRATION !!!"             echo "[!!!] base_hb_sec=$base_hb cur_hb_sec=$cur_hb"             echo "[!!!] freezing setup for debugging for 20s"             echo             # Give some time to attach debuggers / tracing             sleep 20             echo "[!!!] workload still stuck, entering infinite sleep, and will continue to run now"             taskset -c 12 sleep 1 # more than 1 tasks, will break the nohz_full state             while true; do                 sleep 3600             done         fi         sleep "$CHECK_INTERVAL"     done     echo "[4] wait before next round"     sleep 1 done ``` Best regards, Zicheng