* ERESTARTSYS escaping from sem_wait with RTLinux patch
@ 2009-10-10 9:09 Blaise Gassend
2009-10-10 16:40 ` ERESTARTSYS escaping from sem_wait with Preempt-RT Blaise Gassend
2009-10-10 17:59 ` ERESTARTSYS escaping from sem_wait with RTLinux patch Thomas Gleixner
0 siblings, 2 replies; 11+ messages in thread
From: Blaise Gassend @ 2009-10-10 9:09 UTC (permalink / raw)
To: linux-kernel; +Cc: Jeremy Leibs
[-- Attachment #1: Type: text/plain, Size: 2285 bytes --]
The attached python program, in which 500 threads spin with microsecond
sleeps, crashes with a "sem_wait: Unknown error 512" (conditions
described below). This appears to be due to an ERESTARTSYS generated
from futex_wait escaping to user space (libc). My understanding is that
this should never happen and I am trying to track down what is going on.
Questions that would help me make progress:
-------------------------------------------
1) Where is the ERESTARTSYS being prevented from getting to user space?
The only likely place I see for preventing ERESTARTSYS from escaping to
user space is in arch/*/kernel/signal*.c. However, I don't see how the
code there is being called if there no signal pending. Is that a path
for ERESTARTSYS to escape from the kernel?
The following comment in kernel/futex.h in futex_wait makes me wonder if
two threads are getting marked as ERESTARTSYS. The first one to leave
the kernel processes the signal and restarts. The second one doesn't
have a signal to handle, so it returns to user space without getting
into signal*.c and wreaks havoc.
(...)
/*
* We expect signal_pending(current), but another thread may
* have handled it for us already.
*/
if (!abs_time)
return -ERESTARTSYS;
(...)
2) Why would this be happening only with RT kernels?
3) Any suggestions on the best place to patch/workaround this?
My understanding is that if I was to treat ERESTARTSYS as an EAGAIN,
most applications would be perfectly happy. Would bad things happen if I
replaced the ERESTARTSYS in futex_wait with an EAGAIN?
Crash conditions:
-----------------
- RTLinux only.
- More cores seems to make things worse. Lots of crashes on a dual-quad
core machine. None observed yet on dual core. At least one crash on a
dual-quad core when run with "taskset -c 1"
- Various versions, including 2.6.29.6-rt23, and whatever the latest was
earlier today.
- Seen on both ia64 and x86
- Ubuntu hardy and jaunty
- Sometimes hapens within 2 seconds on a dual quad-core machine, other
times will go for up to 30 minutes to an hour without crashing. I
suspect a dependence on system activity, but haven't noticed an obvious
pattern.
- Time to crash appears to drop fast with more CPU cores.
[-- Attachment #2: threadprocs8.py --]
[-- Type: text/x-python, Size: 222 bytes --]
import threading
import time
exiting = False
def spin():
while not exiting:
time.sleep(0.000001)
for i in range(0,500):
threading.Thread(target=spin).start()
try:
spin()
finally:
exiting = True
^ permalink raw reply [flat|nested] 11+ messages in thread* ERESTARTSYS escaping from sem_wait with Preempt-RT
2009-10-10 9:09 ERESTARTSYS escaping from sem_wait with RTLinux patch Blaise Gassend
@ 2009-10-10 16:40 ` Blaise Gassend
2009-10-10 17:59 ` ERESTARTSYS escaping from sem_wait with RTLinux patch Thomas Gleixner
1 sibling, 0 replies; 11+ messages in thread
From: Blaise Gassend @ 2009-10-10 16:40 UTC (permalink / raw)
To: linux-kernel
Apologies, the message below is for Preempt-RT, not RTLinux. The
original message with full details can be found here:
http://lkml.org/lkml/2009/10/10/36
On Sat, 2009-10-10 at 02:09 -0700, Blaise Gassend wrote:
> The attached python program, in which 500 threads spin with microsecond
> sleeps, crashes with a "sem_wait: Unknown error 512" (conditions
> described below). This appears to be due to an ERESTARTSYS generated
> from futex_wait escaping to user space (libc). My understanding is that
> this should never happen and I am trying to track down what is going on.
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: ERESTARTSYS escaping from sem_wait with RTLinux patch
2009-10-10 9:09 ERESTARTSYS escaping from sem_wait with RTLinux patch Blaise Gassend
2009-10-10 16:40 ` ERESTARTSYS escaping from sem_wait with Preempt-RT Blaise Gassend
@ 2009-10-10 17:59 ` Thomas Gleixner
2009-10-10 19:08 ` Jeremy Leibs
2009-10-11 5:48 ` Jeremy Leibs
1 sibling, 2 replies; 11+ messages in thread
From: Thomas Gleixner @ 2009-10-10 17:59 UTC (permalink / raw)
To: Blaise Gassend; +Cc: LKML, Jeremy Leibs, Darren Hart, Peter Zijlstra
Blaise,
On Sat, 10 Oct 2009, Blaise Gassend wrote:
> 1) Where is the ERESTARTSYS being prevented from getting to user space?
>
> The only likely place I see for preventing ERESTARTSYS from escaping to
> user space is in arch/*/kernel/signal*.c. However, I don't see how the
> code there is being called if there no signal pending. Is that a path
> for ERESTARTSYS to escape from the kernel?
>
> The following comment in kernel/futex.h in futex_wait makes me wonder if
> two threads are getting marked as ERESTARTSYS. The first one to leave
> the kernel processes the signal and restarts. The second one doesn't
> have a signal to handle, so it returns to user space without getting
> into signal*.c and wreaks havoc.
>
> (...)
> /*
> * We expect signal_pending(current), but another thread may
> * have handled it for us already.
> */
> if (!abs_time)
> return -ERESTARTSYS;
> (...)
If the task is woken by a signal, then the task private flag
TIF_SIGPENDING is set, but in case of a process wide signal the signal
might have been handled by another thread of the same process before
that thread reaches the signal handling code, but then ERESTARTSYS is
handled gracefully. So you seem to trigger a code path which does not
go through do_signal.
> 2) Why would this be happening only with RT kernels?
Slightly different timing and locking semantics.
> 3) Any suggestions on the best place to patch/workaround this?
>
> My understanding is that if I was to treat ERESTARTSYS as an EAGAIN,
> most applications would be perfectly happy. Would bad things happen if I
> replaced the ERESTARTSYS in futex_wait with an EAGAIN?
No workarounds please. We really want to know what's wrong.
Two things to look at:
1) Does that happen with 2.6.31.2-rt13 as well ?
2) Add a check to the code path where ERESTARTSYS is returned:
if (!signal_pending(current))
printk(KERN_ERR ".....");
If you can see that message then we'll look further. I'll give your
script a test ride on my systems as well.
Thanks,
tglx
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: ERESTARTSYS escaping from sem_wait with RTLinux patch
2009-10-10 17:59 ` ERESTARTSYS escaping from sem_wait with RTLinux patch Thomas Gleixner
@ 2009-10-10 19:08 ` Jeremy Leibs
2009-10-11 2:07 ` Jeremy Leibs
2009-10-11 5:48 ` Jeremy Leibs
1 sibling, 1 reply; 11+ messages in thread
From: Jeremy Leibs @ 2009-10-10 19:08 UTC (permalink / raw)
To: Thomas Gleixner; +Cc: Blaise Gassend, LKML, Darren Hart, Peter Zijlstra
Thomas, thanks for the quick reply.
On Sat, Oct 10, 2009 at 10:59 AM, Thomas Gleixner <tglx@linutronix.de> wrote:
> Blaise,
>
> On Sat, 10 Oct 2009, Blaise Gassend wrote:
>> 1) Where is the ERESTARTSYS being prevented from getting to user space?
>>
>> The only likely place I see for preventing ERESTARTSYS from escaping to
>> user space is in arch/*/kernel/signal*.c. However, I don't see how the
>> code there is being called if there no signal pending. Is that a path
>> for ERESTARTSYS to escape from the kernel?
>>
>> The following comment in kernel/futex.h in futex_wait makes me wonder if
>> two threads are getting marked as ERESTARTSYS. The first one to leave
>> the kernel processes the signal and restarts. The second one doesn't
>> have a signal to handle, so it returns to user space without getting
>> into signal*.c and wreaks havoc.
>>
>> (...)
>> /*
>> * We expect signal_pending(current), but another thread may
>> * have handled it for us already.
>> */
>> if (!abs_time)
>> return -ERESTARTSYS;
>> (...)
>
> If the task is woken by a signal, then the task private flag
> TIF_SIGPENDING is set, but in case of a process wide signal the signal
> might have been handled by another thread of the same process before
> that thread reaches the signal handling code, but then ERESTARTSYS is
> handled gracefully. So you seem to trigger a code path which does not
> go through do_signal.
>
>> 2) Why would this be happening only with RT kernels?
>
> Slightly different timing and locking semantics.
>
>> 3) Any suggestions on the best place to patch/workaround this?
>>
>> My understanding is that if I was to treat ERESTARTSYS as an EAGAIN,
>> most applications would be perfectly happy. Would bad things happen if I
>> replaced the ERESTARTSYS in futex_wait with an EAGAIN?
>
> No workarounds please. We really want to know what's wrong.
>
> Two things to look at:
>
> 1) Does that happen with 2.6.31.2-rt13 as well ?
I am nearly certain we saw the problems with the newer kernel as well,
although that was back with a much less concise test and I've since
reinstalled over that machine in the process of trying a number of
different 32/64 hardy/jaunty configurations on different hardware.
I'll do a fresh install of that particular kernel with default
configuration options on our hardware and let you know a little later
today.
> 2) Add a check to the code path where ERESTARTSYS is returned:
>
> if (!signal_pending(current))
> printk(KERN_ERR ".....");
>
> If you can see that message then we'll look further. I'll give your
> script a test ride on my systems as well.
>
> Thanks,
>
> tglx
>
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: ERESTARTSYS escaping from sem_wait with RTLinux patch
2009-10-10 19:08 ` Jeremy Leibs
@ 2009-10-11 2:07 ` Jeremy Leibs
0 siblings, 0 replies; 11+ messages in thread
From: Jeremy Leibs @ 2009-10-11 2:07 UTC (permalink / raw)
To: Thomas Gleixner; +Cc: Blaise Gassend, LKML, Darren Hart, Peter Zijlstra
On Sat, Oct 10, 2009 at 12:08 PM, Jeremy Leibs <leibs@willowgarage.com> wrote:
> Thomas, thanks for the quick reply.
>
> On Sat, Oct 10, 2009 at 10:59 AM, Thomas Gleixner <tglx@linutronix.de> wrote:
-----snip-----
>>
>> 1) Does that happen with 2.6.31.2-rt13 as well ?
>
> I am nearly certain we saw the problems with the newer kernel as well,
> although that was back with a much less concise test and I've since
> reinstalled over that machine in the process of trying a number of
> different 32/64 hardy/jaunty configurations on different hardware.
> I'll do a fresh install of that particular kernel with default
> configuration options on our hardware and let you know a little later
> today.
Running on fresh install of 64 bit jaunty:
leibs@c1:~$ uname -rv
2.6.31.2-rt13 #1 SMP PREEMPT RT Sat Oct 10 15:34:13 PDT 2009
leibs@c1:~$ python threadprocs8.py
sem_wait: Unknown error 512
Fatal Python error: ceval: orphan tstate
Aborted
^ permalink raw reply [flat|nested] 11+ messages in thread
* Re: ERESTARTSYS escaping from sem_wait with RTLinux patch
2009-10-10 17:59 ` ERESTARTSYS escaping from sem_wait with RTLinux patch Thomas Gleixner
2009-10-10 19:08 ` Jeremy Leibs
@ 2009-10-11 5:48 ` Jeremy Leibs
2009-10-12 14:16 ` Darren Hart
1 sibling, 1 reply; 11+ messages in thread
From: Jeremy Leibs @ 2009-10-11 5:48 UTC (permalink / raw)
To: Thomas Gleixner; +Cc: Blaise Gassend, LKML, Darren Hart, Peter Zijlstra
On Sat, Oct 10, 2009 at 10:59 AM, Thomas Gleixner <tglx@linutronix.de> wrote:
> Blaise,
>
> On Sat, 10 Oct 2009, Blaise Gassend wrote:
>> 1) Where is the ERESTARTSYS being prevented from getting to user space?
>>
>> The only likely place I see for preventing ERESTARTSYS from escaping to
>> user space is in arch/*/kernel/signal*.c. However, I don't see how the
>> code there is being called if there no signal pending. Is that a path
>> for ERESTARTSYS to escape from the kernel?
>>
>> The following comment in kernel/futex.h in futex_wait makes me wonder if
>> two threads are getting marked as ERESTARTSYS. The first one to leave
>> the kernel processes the signal and restarts. The second one doesn't
>> have a signal to handle, so it returns to user space without getting
>> into signal*.c and wreaks havoc.
>>
>> (...)
>> /*
>> * We expect signal_pending(current), but another thread may
>> * have handled it for us already.
>> */
>> if (!abs_time)
>> return -ERESTARTSYS;
>> (...)
>
> If the task is woken by a signal, then the task private flag
> TIF_SIGPENDING is set, but in case of a process wide signal the signal
> might have been handled by another thread of the same process before
> that thread reaches the signal handling code, but then ERESTARTSYS is
> handled gracefully. So you seem to trigger a code path which does not
> go through do_signal.
>
>> 2) Why would this be happening only with RT kernels?
>
> Slightly different timing and locking semantics.
>
>> 3) Any suggestions on the best place to patch/workaround this?
>>
>> My understanding is that if I was to treat ERESTARTSYS as an EAGAIN,
>> most applications would be perfectly happy. Would bad things happen if I
>> replaced the ERESTARTSYS in futex_wait with an EAGAIN?
>
> No workarounds please. We really want to know what's wrong.
>
> Two things to look at:
>
> 1) Does that happen with 2.6.31.2-rt13 as well ?
>
> 2) Add a check to the code path where ERESTARTSYS is returned:
>
> if (!signal_pending(current))
> printk(KERN_ERR ".....");
>
Ok, in 2.6.31.2-rt13, I modified futex.c as:
-----
/*
* We expect signal_pending(current), but another thread may
* have handled it for us already.
*/
ret = -ERESTARTSYS;
if (!abs_time)
{
if (!signal_pending(current))
printk(KERN_ERR ".....");
goto out_put_key;
}
-----
Then when I cause the crash:
leibs@c1:~$ python threadprocs8.py
sem_wait: Unknown error 512
Segmentation fault
dmesg shows me the corresponding:
[ 82.232999] .....
[ 82.233177] python[2834]: segfault at 48 ip 00000000004b0177 sp
00007f9429788ad8 error 4 in python2.6[400000+216000]
--Jeremy
^ permalink raw reply [flat|nested] 11+ messages in thread* Re: ERESTARTSYS escaping from sem_wait with RTLinux patch
2009-10-11 5:48 ` Jeremy Leibs
@ 2009-10-12 14:16 ` Darren Hart
[not found] ` <1255384010.10236.123.camel@lts.willowgarage.com>
0 siblings, 1 reply; 11+ messages in thread
From: Darren Hart @ 2009-10-12 14:16 UTC (permalink / raw)
To: Jeremy Leibs; +Cc: Thomas Gleixner, Blaise Gassend, LKML, Peter Zijlstra
Jeremy Leibs wrote:
> On Sat, Oct 10, 2009 at 10:59 AM, Thomas Gleixner <tglx@linutronix.de> wrote:
>> Blaise,
>>
>> On Sat, 10 Oct 2009, Blaise Gassend wrote:
>>> 1) Where is the ERESTARTSYS being prevented from getting to user space?
>>>
>>> The only likely place I see for preventing ERESTARTSYS from escaping to
>>> user space is in arch/*/kernel/signal*.c. However, I don't see how the
>>> code there is being called if there no signal pending. Is that a path
>>> for ERESTARTSYS to escape from the kernel?
>>>
>>> The following comment in kernel/futex.h in futex_wait makes me wonder if
>>> two threads are getting marked as ERESTARTSYS. The first one to leave
>>> the kernel processes the signal and restarts. The second one doesn't
>>> have a signal to handle, so it returns to user space without getting
>>> into signal*.c and wreaks havoc.
>>>
>>> (...)
>>> /*
>>> * We expect signal_pending(current), but another thread may
>>> * have handled it for us already.
>>> */
>>> if (!abs_time)
>>> return -ERESTARTSYS;
>>> (...)
>> If the task is woken by a signal, then the task private flag
>> TIF_SIGPENDING is set, but in case of a process wide signal the signal
>> might have been handled by another thread of the same process before
>> that thread reaches the signal handling code, but then ERESTARTSYS is
>> handled gracefully. So you seem to trigger a code path which does not
>> go through do_signal.
>>
>>> 2) Why would this be happening only with RT kernels?
>> Slightly different timing and locking semantics.
>>
>>> 3) Any suggestions on the best place to patch/workaround this?
>>>
>>> My understanding is that if I was to treat ERESTARTSYS as an EAGAIN,
>>> most applications would be perfectly happy. Would bad things happen if I
>>> replaced the ERESTARTSYS in futex_wait with an EAGAIN?
>> No workarounds please. We really want to know what's wrong.
>>
>> Two things to look at:
>>
>> 1) Does that happen with 2.6.31.2-rt13 as well ?
>>
>> 2) Add a check to the code path where ERESTARTSYS is returned:
>>
>> if (!signal_pending(current))
>> printk(KERN_ERR ".....");
>>
>
> Ok, in 2.6.31.2-rt13, I modified futex.c as:
> -----
> /*
> * We expect signal_pending(current), but another thread may
> * have handled it for us already.
> */
> ret = -ERESTARTSYS;
> if (!abs_time)
> {
> if (!signal_pending(current))
> printk(KERN_ERR ".....");
> goto out_put_key;
> }
> -----
>
> Then when I cause the crash:
>
> leibs@c1:~$ python threadprocs8.py
> sem_wait: Unknown error 512
> Segmentation fault
>
> dmesg shows me the corresponding:
> [ 82.232999] .....
> [ 82.233177] python[2834]: segfault at 48 ip 00000000004b0177 sp
> 00007f9429788ad8 error 4 in python2.6[400000+216000]
OK, so I suspect one of two things.
1) Recent changes to futex.c have somehow created a wakeup race and
unqueue_me() doesn't detect it was woken with FUTEX_WAKE, then falls
out through the ERESTARTSYS path.
2) Recent changes have exposed an existing race in unqueue_me().
I'll do some runs on my 8-way systems and see if I can:
o Identify the guilty patch
o Identify the race in question
Thanks for the test case! Now... why is sem_wait() being used in a timer
call....
--
Darren Hart
IBM Linux Technology Center
Real-Time Linux Team
^ permalink raw reply [flat|nested] 11+ messages in thread
end of thread, other threads:[~2009-10-13 18:46 UTC | newest]
Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2009-10-10 9:09 ERESTARTSYS escaping from sem_wait with RTLinux patch Blaise Gassend
2009-10-10 16:40 ` ERESTARTSYS escaping from sem_wait with Preempt-RT Blaise Gassend
2009-10-10 17:59 ` ERESTARTSYS escaping from sem_wait with RTLinux patch Thomas Gleixner
2009-10-10 19:08 ` Jeremy Leibs
2009-10-11 2:07 ` Jeremy Leibs
2009-10-11 5:48 ` Jeremy Leibs
2009-10-12 14:16 ` Darren Hart
[not found] ` <1255384010.10236.123.camel@lts.willowgarage.com>
[not found] ` <4AD3BD57.6080703@us.ibm.com>
[not found] ` <4AD3D6AE.2050609@us.ibm.com>
[not found] ` <4AD3FFB0.5030405@us.ibm.com>
2009-10-13 4:54 ` Darren Hart
2009-10-13 8:56 ` Blaise Gassend
2009-10-13 15:13 ` Darren Hart
2009-10-13 18:45 ` Thomas Gleixner
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®