From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-lj1-f179.google.com (mail-lj1-f179.google.com [209.85.208.179]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 616762475CB for ; Sun, 14 Jun 2026 22:43:28 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.208.179 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781477010; cv=none; b=Hr7/wXcNbDmy/u9OorpgoMTXMGtDhyTq45+fpVrePNuHA8qH1ynFLq9aLq6Od9sBJLHkEL2Y0X1SqYgh3sbbUhMnObLQ4G8VDaQpWOjjLH2ikaawjMYOF38aDzM48BbW6iMmTKfDZrbDGLl7o7g2BQ9JEC5T/Dr7q6LQDrhqz98= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781477010; c=relaxed/simple; bh=wovvwwH8XeaJJNVmtu1Z5lMy311o0CJDJSRqrCqiLn0=; h=Message-ID:Date:MIME-Version:From:Subject:To:References:Cc: In-Reply-To:Content-Type; b=INRLvUzTHa43x6CH7Q9fm+BHRPDxjywMXebcmSJwZ40twlTKYwbqdQrqL4szOey5eTYK5H13h5Z5ZPPYhYRu4BQJUWpd4gZq9VkytmYT7XG64LQuV9wViEGkxeO2NZKVj1jZxc4CZ6B4+fJpeyrQdAwYCjEcBopOjthF2DImhJo= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=aIDseN9a; arc=none smtp.client-ip=209.85.208.179 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="aIDseN9a" Received: by mail-lj1-f179.google.com with SMTP id 38308e7fff4ca-399389dae7fso11955591fa.0 for ; Sun, 14 Jun 2026 15:43:28 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1781477006; x=1782081806; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:cc:content-language :references:to:subject:from:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=7/Jyft7mn323QdV6sssTyw1tRgw8bubOXcejq7qlLs4=; b=aIDseN9aEP3Kb+Qj91dJyviESP+VwcsrcxIXqCabpQurIjxY0ai5L+WfSfVduC00xO 4QasJI6wR7YU9WnmaVCmvLQ5z1lMnQmg0Ori2Ai01gA5x2wT+aMKqrRqgfEWNQZkUGCU jIzxBZd6nluINldBTmBHWsdsVEwDQMFNlByeNTcXMR99vV1+JhEsXM50p8FDUGkx/tVd Cj36qgiI9QByMgVQu6XJkFItr+i+UMxr9yiidoNfleWYLxOCW5WMJJJfheGdiR/P6ICH z4fR14DWAIaNrOlTo1JqN6ulvXHjFaF+bvrPpnPjFPY2rUH0ghdaRx6SOaeQl+L3NZ/l +GcQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781477006; x=1782081806; h=content-transfer-encoding:in-reply-to:cc:content-language :references:to:subject:from:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=7/Jyft7mn323QdV6sssTyw1tRgw8bubOXcejq7qlLs4=; b=GTSNvCayBV84nKRyIn3JoIQCOkdgiyozJfW7vH3bjP2VMVn0TgbCwumB5CHJJ5cFws cMvAYnAB1Iay5rwS6T7jOo4KND2cZicDYy8+YI7Rrv6wzz7xMY5s53vviKRU6EwOVpJp OUd85qFJgLaFHb/MFaeneiHKaz0O94RqZakkY5FNjZ/9zcK0UQA46QuZvP92lk8e8hEq 5m86VIG+CBEk2WPbUSoazYixn15RpcvjpdKGjocKvetxDqye/2iyfPeiCvuwRuuP4qQh 6HKdaugxv/imuIBCrrwMeKAipFg07BUudZ+dlWbfzlCw7/0ORa+70muVdLFVGK43Gkix IZTA== X-Forwarded-Encrypted: i=1; AFNElJ+Iupr3bPO97ZLvVUxxNBZ7oZl+jbmMSRYDAlBwdcxWYrrnnChsdSK1cCwctLOMvw0vA9S4bwgp2Ql0XmM=@vger.kernel.org X-Gm-Message-State: AOJu0Yx3wLDa2nwda8hN6oLQuVJCYFeId5LkerSO/86Toexg3yVUyVvv U/RBcVOjmfeZBKFOLd/JXzx/4iWCiIL6l9ESZ3Wg3U2KyFMUqqmmOW6AT4875DKz X-Gm-Gg: Acq92OH2yI16FxTM6ejgQYcLTnFk76dwif5DmqRcFKMXpeO3Rakm6G5UlnKmfg+Nz/B vUaJHtaMDziBuRAmZ8x8Zj6dDsel0CKlmKMcJrxo2vh8ECxHveRnqTQ1ID4im6FdjHuAWUbqsd/ Gfr9cmCqqRz/uTHDklZKO6dL90ydyo2xWOoaC9tmeXw7bFfLnS5fhC+IFkF8RZPtJvkozlVSqIh Lwh6TOq0AvkTMxZUz9aoI3VY07u7TySVlbZrv4l49BYJHtqaVqZUIu2eczQGC4lbjMZ+20fAXqU QajPM5CbLf1THcGIqVn0+WmmBpnjFLhcE0GUcS3IfoA6xNm2j12W5E6J3R4H6L4pR3KIKrZk8ws xoJclKDTn8RYImNd2Xm6Enc53CFhNV0Pf4xxrPD65FD/N4Mvq7LVZrQbI027oZFRTToGx2qgDXs BxGfGBq2oSA6yvGbkm+BJTzUNWqJfKSkoUOk/PXAHue135cdtyf1UXsvXHIZ1dYkWuiqfsjpYIg DHY X-Received: by 2002:a2e:a993:0:b0:38b:f0f0:e39b with SMTP id 38308e7fff4ca-3992bdb81bemr27273371fa.10.1781477006288; Sun, 14 Jun 2026 15:43:26 -0700 (PDT) Received: from [192.168.0.135] (95-24-171-3.broadband.corbina.ru. [95.24.171.3]) by smtp.gmail.com with ESMTPSA id 38308e7fff4ca-399313e9385sm17740011fa.7.2026.06.14.15.43.24 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Sun, 14 Jun 2026 15:43:24 -0700 (PDT) Message-ID: <4f9b8e2e-8c1e-469b-94e8-3125fbe62d33@gmail.com> Date: Mon, 15 Jun 2026 01:40:00 +0300 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird From: Konstantin Mikhaylov Subject: Re: perf AUX: race causes poll() hang To: Peter Zijlstra References: <20260518094121.GT3102624@noisy.programming.kicks-ass.net> Content-Language: en-US Cc: adrian.hunter@intel.com, mark.rutland@arm.com, linux-kernel@vger.kernel.org In-Reply-To: <20260518094121.GT3102624@noisy.programming.kicks-ass.net> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit On 5/18/26 12:41 PM, Peter Zijlstra wrote: > On Fri, May 15, 2026 at 01:35:48AM +0300, Константин Михайлов wrote: >> Hello Peter and Adrian, >> >> I'd like to report a potential race condition in perf AUX buffer handling. >> >> AUX tracing is designed to allow the tracee continue running when the AUX >> buffer fills. The PMU driver must disable tracing when AUX buffer is full. >> Typically, it schedules IRQ work to disable the event later. Meanwhile, a >> typical tracer's workflow looks like: poll() on perf FDs, consume the data, >> re-enable the event via PERF_EVENT_IOC_ENABLE ioctl(), then poll() again. >> >> Given this, the following race is possible: >> ------------------- >> | CPU #0 | CPU #1 | >> | tracee | tracer | >> ------------------- >> | ** | | tracee fills the AUX buffer completely with some data >> ------------------- >> | ** | | PMU driver updates aux_head accordingly and schedules IRQ works >> | ** | | to disable the event and wake up the tracer (setting rb->poll in >> | ** | | perf_output_wakeup() along the way) >> ------------------- >> | | ** | tracer consumes all the data from AUX buffer, >> | | ** | thus clears rb->poll in perf_poll() >> ------------------- >> | | ** | tracer re-enables the tracing (the event is still active, >> | | ** | so ioctl(...) returns immediately) >> ------------------- >> | | ** | tracer starts poll()'ing the AUX buffer again >> ------------------- >> | ** | | IRQ work handler finally disables the event and >> | ** | | wakes up tracer >> ------------------- >> | | ** | tracer obtains zero rb->poll and continues polling >> ------------------- >> As a result, tracee runs without PMU tracing, and tracer's poll() will >> never be woken up unless it has some timeout. >> >> I reproduced this on an x86 machine with intel_pt and kernel v6.17. >> Reproducing this race on the vanilla kernel is timing-sensitive, so I added >> 30 ms delay in the error path in intel_pt_interrupt() when >> pt_buffer_reset_markers() returns an error - this delay widens the window >> between aux_head update and actual event disable in IRQ work handler. I'm >> not sure that pt_buffer_reset_markers()'s error means that buffer >> overflowed, but this error branch is taken sometimes and all needed IRQ >> works are scheduled during a call to perf_aux_output_end(). I also added 3 >> ms delay in perf in __auxtrace_mmap__read() before itr->read_finish(), >> ensuring the ioctl() falls into that window. With these changes, some perf >> runs collected smaller traces than usual. I added traceprints to intel_pt >> driver and enabled tracing for sys_poll and sys_ioctl, which confirmed the >> exact sequence described above. The problem was sometimes mitigated by >> tracee migration to another cpu (as perf creates an event for every cpu, >> the event is re-enabled by kernel when it is set on a new cpu). Otherwise, >> tracee stayed on the same cpu and tracer hung on poll() until tracee exited. >> >> Could you please confirm if this analysis is correct? Should we move >> setting of rb->poll *after* the event is disabled in IRQ work handler? > > I *think* (its been a minute since I looked at this code), that you're > right. > > Does something like the below cure things? > > --- > kernel/events/core.c | 54 +++++++++++++++++++++++++++++----------------------- > 1 file changed, 30 insertions(+), 24 deletions(-) > > diff --git a/kernel/events/core.c b/kernel/events/core.c > index 7935d5663944..490407618f36 100644 > --- a/kernel/events/core.c > +++ b/kernel/events/core.c > @@ -2677,6 +2677,9 @@ static void __perf_event_disable(struct perf_event *event, > struct perf_event_context *ctx, > void *info) > { > + if (event->pending_disable) > + event->pending_disable = 0; > + > if (event->state < PERF_EVENT_STATE_INACTIVE) > return; > > @@ -3278,32 +3281,37 @@ static void _perf_event_enable(struct perf_event *event) > { > struct perf_event_context *ctx = event->ctx; > > - raw_spin_lock_irq(&ctx->lock); > - if (event->state >= PERF_EVENT_STATE_INACTIVE || > - event->state < PERF_EVENT_STATE_ERROR) { > -out: > - raw_spin_unlock_irq(&ctx->lock); > - return; > - } > + scoped_guard (raw_spinlock_irq, &ctx->lock) { > + if (event->state < PERF_EVENT_STATE_ERROR) > + return; > > - /* > - * If the event is in error state, clear that first. > - * > - * That way, if we see the event in error state below, we know that it > - * has gone back into error state, as distinct from the task having > - * been scheduled away before the cross-call arrived. > - */ > - if (event->state == PERF_EVENT_STATE_ERROR) { > /* > - * Detached SIBLING events cannot leave ERROR state. > + * If the event is in error state, clear that first. > + * > + * That way, if we see the event in error state below, we know that it > + * has gone back into error state, as distinct from the task having > + * been scheduled away before the cross-call arrived. > */ > - if (event->event_caps & PERF_EV_CAP_SIBLING && > - event->group_leader == event) > - goto out; > + if (event->state == PERF_EVENT_STATE_ERROR) { > + /* > + * Detached SIBLING events cannot leave ERROR state. > + */ > + if (event->event_caps & PERF_EV_CAP_SIBLING && > + event->group_leader == event) > + return; > > - event->state = PERF_EVENT_STATE_OFF; > + event->state = PERF_EVENT_STATE_OFF; > + } > + > + if (event->pending_disable) > + event->pending_disable = 0; > + > + /* > + * Already running, nothing to do. > + */ > + if (event->state >= PERF_EVENT_STATE_INACTIVE) > + return; > } > - raw_spin_unlock_irq(&ctx->lock); > > event_function_call(event, __perf_event_enable, NULL); > } > @@ -7612,10 +7620,8 @@ static void __perf_pending_disable(struct perf_event *event) > * Yay, we hit home and are in the context of the event. > */ > if (cpu == smp_processor_id()) { > - if (event->pending_disable) { > - event->pending_disable = 0; > + if (event->pending_disable) > perf_event_disable_local(event); > - } > return; > } > Sorry, I'm re-sending this response because my previous email (sent 3+ weeks ago) was not properly formatted as plain text and was not accepted by lkml.org. To ensure this reaches the thread, I'm re-sending it correctly formatted now. I tested this patch with the same delay on the error path in intel_pt driver. With this patch, the hardware tracing sometimes remains stopped until the next reschedule. The tracer does not re-enable the event when event->state >= PERF_EVENT_STATE_INACTIVE, since this condition is true while the driver is disabling hardware tracing. As a result, hardware tracing can stay disabled until the next reschedule even when the tracer has drained all data from the AUX buffer.