* [PATCH 1/3] vsyscall_64: add missing ifdef CONFIG_SECCOMP
@ 2012-07-14 15:32 Will Drewry
2012-07-14 15:32 ` [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip Will Drewry
2012-07-14 15:32 ` [PATCH 3/3] Documentation: add a caveat for seccomp filter and vsyscall emulation Will Drewry
0 siblings, 2 replies; 10+ messages in thread
From: Will Drewry @ 2012-07-14 15:32 UTC (permalink / raw)
To: linux-kernel, torvalds
Cc: fengxj325, eparis, keescook, james.l.morris, hpa, cevans, luto,
rob, linux-doc, Will Drewry
vsyscall_seccomp introduced a dependency on __secure_computing. On
configurations with CONFIG_SECCOMP disabled, compilation will fail.
Reported-by: feng xiangjun <fengxj325@gmail.com>
Signed-off-by: Will Drewry <wad@chromium.org>
---
arch/x86/kernel/vsyscall_64.c | 4 ++++
1 file changed, 4 insertions(+)
diff --git a/arch/x86/kernel/vsyscall_64.c b/arch/x86/kernel/vsyscall_64.c
index 08a18d0..5db36ca 100644
--- a/arch/x86/kernel/vsyscall_64.c
+++ b/arch/x86/kernel/vsyscall_64.c
@@ -139,6 +139,7 @@ static int addr_to_vsyscall_nr(unsigned long addr)
return nr;
}
+#ifdef CONFIG_SECCOMP
static int vsyscall_seccomp(struct task_struct *tsk, int syscall_nr)
{
if (!seccomp_mode(&tsk->seccomp))
@@ -147,6 +148,9 @@ static int vsyscall_seccomp(struct task_struct *tsk, int syscall_nr)
task_pt_regs(tsk)->ax = syscall_nr;
return __secure_computing(syscall_nr);
}
+#else
+#define vsyscall_seccomp(_tsk, _nr) 0
+#endif
static bool write_ok_or_segv(unsigned long ptr, size_t size)
{
--
1.7.9.5
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip
2012-07-14 15:32 [PATCH 1/3] vsyscall_64: add missing ifdef CONFIG_SECCOMP Will Drewry
@ 2012-07-14 15:32 ` Will Drewry
[not found] ` <CAObL_7GX0q_qywY4g1S2iRZWPS95ar01wtvv9wDW=zswwaZ6fQ@mail.gmail.com>
2012-07-14 15:32 ` [PATCH 3/3] Documentation: add a caveat for seccomp filter and vsyscall emulation Will Drewry
1 sibling, 1 reply; 10+ messages in thread
From: Will Drewry @ 2012-07-14 15:32 UTC (permalink / raw)
To: linux-kernel, torvalds
Cc: fengxj325, eparis, keescook, james.l.morris, hpa, cevans, luto,
rob, linux-doc, Will Drewry
Current quirky ptrace behavior with vsyscall and seccomp
does not allow tracers to bypass the call. This change
provides that ability by checking if orig_ax changed.
Signed-off-by: Will Drewry <wad@chromium.org>
---
arch/x86/kernel/vsyscall_64.c | 10 +++++++---
1 file changed, 7 insertions(+), 3 deletions(-)
diff --git a/arch/x86/kernel/vsyscall_64.c b/arch/x86/kernel/vsyscall_64.c
index 5db36ca..5f9640c 100644
--- a/arch/x86/kernel/vsyscall_64.c
+++ b/arch/x86/kernel/vsyscall_64.c
@@ -142,11 +142,15 @@ static int addr_to_vsyscall_nr(unsigned long addr)
#ifdef CONFIG_SECCOMP
static int vsyscall_seccomp(struct task_struct *tsk, int syscall_nr)
{
+ int ret;
if (!seccomp_mode(&tsk->seccomp))
return 0;
task_pt_regs(tsk)->orig_ax = syscall_nr;
task_pt_regs(tsk)->ax = syscall_nr;
- return __secure_computing(syscall_nr);
+ ret = __secure_computing(syscall_nr);
+ if (task_pt_regs(tsk)->orig_ax != syscall_nr)
+ return 1; /* ptrace syscall skip */
+ return ret;
}
#else
#define vsyscall_seccomp(_tsk, _nr) 0
@@ -278,9 +282,9 @@ bool emulate_vsyscall(struct pt_regs *regs, unsigned long address)
current_thread_info()->sig_on_uaccess_error = prev_sig_on_uaccess_error;
if (skip) {
- if ((long)regs->ax <= 0L) /* seccomp errno emulation */
+ if ((long)regs->ax <= 0L || skip == 1) /* seccomp errno/trace */
goto do_ret;
- goto done; /* seccomp trace/trap */
+ goto done; /* seccomp trap */
}
if (ret == -EFAULT) {
--
1.7.9.5
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 3/3] Documentation: add a caveat for seccomp filter and vsyscall emulation
2012-07-14 15:32 [PATCH 1/3] vsyscall_64: add missing ifdef CONFIG_SECCOMP Will Drewry
2012-07-14 15:32 ` [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip Will Drewry
@ 2012-07-14 15:32 ` Will Drewry
1 sibling, 0 replies; 10+ messages in thread
From: Will Drewry @ 2012-07-14 15:32 UTC (permalink / raw)
To: linux-kernel, torvalds
Cc: fengxj325, eparis, keescook, james.l.morris, hpa, cevans, luto,
rob, linux-doc, Will Drewry
With the addition of seccomp support to vsyscall emulation:
http://permalink.gmane.org/gmane.linux.kernel/1327732
and the prior patch in this series.
Update the documentation to indicate quirky behaviors when the 'ip' is
in the vsyscall page and vsyscall emulation is in effect.
Signed-off-by: Will Drewry <wad@chromium.org>
---
Documentation/prctl/seccomp_filter.txt | 22 ++++++++++++++++++++++
1 file changed, 22 insertions(+)
diff --git a/Documentation/prctl/seccomp_filter.txt b/Documentation/prctl/seccomp_filter.txt
index 597c3c5..67ed88b 100644
--- a/Documentation/prctl/seccomp_filter.txt
+++ b/Documentation/prctl/seccomp_filter.txt
@@ -161,3 +161,25 @@ architecture supports both ptrace_event and seccomp, it will be able to
support seccomp filter with minor fixup: SIGSYS support and seccomp return
value checking. Then it must just add CONFIG_HAVE_ARCH_SECCOMP_FILTER
to its arch-specific Kconfig.
+
+
+Caveats
+-------
+
+On x86-64 with vsyscall emulation enabled and while servicing a
+vsyscall-emulated system call:
+- A return value of SECCOMP_RET_TRAP will set a si_call_addr pointing to
+ the vsyscall entry for the given call and not the address after the
+ 'syscall' instruction. Any code which wants to restart the call
+ should return to that address and code wishing to return simulating
+ completion may either sigreturn normally or simulate a ret instruction
+ and use the return address from the stack.
+- A return value of SECCOMP_RET_TRACE will signal the tracer as usual,
+ but the syscall may not be changed to another system call using the
+ orig_rax register. It may only be changed to a different value in
+ order to skip the currently emulated call and any change will result
+ in that behavior. The remainder of the registers may be altered as
+ usual.
+- Detection of this quirky behavior may be done by checking for getcpu,
+ time, or gettimeofday and if the si_call_addr or rip is in the
+ vsyscall page, specifically at the start of the specific entry call.
--
1.7.9.5
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip
[not found] ` <CAObL_7GX0q_qywY4g1S2iRZWPS95ar01wtvv9wDW=zswwaZ6fQ@mail.gmail.com>
@ 2012-07-14 15:50 ` Will Drewry
2012-07-14 15:57 ` Will Drewry
0 siblings, 1 reply; 10+ messages in thread
From: Will Drewry @ 2012-07-14 15:50 UTC (permalink / raw)
To: Andrew Lutomirski
Cc: Kees Cook, james.l.morris, rob, linux-doc, cevans, hpa,
linux-kernel, torvalds, eparis, fengxj325
On Sat, Jul 14, 2012 at 10:44 AM, Andrew Lutomirski <luto@mit.edu> wrote:
> I think I'd prefer if changing to something other than whatever value is
> used to cancel the syscall resulted in a crash rather than just being
> ignored.
I was trying to keep as much seccomp-ptrace behavior intact rather
than making it terminal in this special case. Is there a reason why
it'd make more sense to crash?
> How hard is it for a page fault to return into the syscall entry path? It
> should be possible to do this for rel, although it could be messy and not
> worth it.
Not sure, tbh. I think given vsyscall's status and the fact that
ptrace+seccomp+vsyscall=emulate isn't horrible, I think it's fine to
either ignore (what is in tree now) or to allow ptrace to skip,
without providing full functionality. But obviously, my view my be
biased!
thanks!
will
>
> On Jul 14, 2012 10:35 AM, "Will Drewry" <wad@chromium.org> wrote:
>>
>> Current quirky ptrace behavior with vsyscall and seccomp
>> does not allow tracers to bypass the call. This change
>> provides that ability by checking if orig_ax changed.
>>
>> Signed-off-by: Will Drewry <wad@chromium.org>
>> ---
>> arch/x86/kernel/vsyscall_64.c | 10 +++++++---
>> 1 file changed, 7 insertions(+), 3 deletions(-)
>>
>> diff --git a/arch/x86/kernel/vsyscall_64.c b/arch/x86/kernel/vsyscall_64.c
>> index 5db36ca..5f9640c 100644
>> --- a/arch/x86/kernel/vsyscall_64.c
>> +++ b/arch/x86/kernel/vsyscall_64.c
>> @@ -142,11 +142,15 @@ static int addr_to_vsyscall_nr(unsigned long addr)
>> #ifdef CONFIG_SECCOMP
>> static int vsyscall_seccomp(struct task_struct *tsk, int syscall_nr)
>> {
>> + int ret;
>> if (!seccomp_mode(&tsk->seccomp))
>> return 0;
>> task_pt_regs(tsk)->orig_ax = syscall_nr;
>> task_pt_regs(tsk)->ax = syscall_nr;
>> - return __secure_computing(syscall_nr);
>> + ret = __secure_computing(syscall_nr);
>> + if (task_pt_regs(tsk)->orig_ax != syscall_nr)
>> + return 1; /* ptrace syscall skip */
>> + return ret;
>> }
>> #else
>> #define vsyscall_seccomp(_tsk, _nr) 0
>> @@ -278,9 +282,9 @@ bool emulate_vsyscall(struct pt_regs *regs, unsigned
>> long address)
>> current_thread_info()->sig_on_uaccess_error =
>> prev_sig_on_uaccess_error;
>>
>> if (skip) {
>> - if ((long)regs->ax <= 0L) /* seccomp errno emulation */
>> + if ((long)regs->ax <= 0L || skip == 1) /* seccomp
>> errno/trace */
>> goto do_ret;
>> - goto done; /* seccomp trace/trap */
>> + goto done; /* seccomp trap */
>> }
>>
>> if (ret == -EFAULT) {
>> --
>> 1.7.9.5
>>
>
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip
2012-07-14 15:50 ` Will Drewry
@ 2012-07-14 15:57 ` Will Drewry
2012-07-14 16:15 ` Andrew Lutomirski
0 siblings, 1 reply; 10+ messages in thread
From: Will Drewry @ 2012-07-14 15:57 UTC (permalink / raw)
To: Andrew Lutomirski
Cc: Kees Cook, james.l.morris, rob, linux-doc, cevans, hpa,
linux-kernel, torvalds, eparis, fengxj325
On Sat, Jul 14, 2012 at 10:50 AM, Will Drewry <wad@chromium.org> wrote:
> On Sat, Jul 14, 2012 at 10:44 AM, Andrew Lutomirski <luto@mit.edu> wrote:
>> I think I'd prefer if changing to something other than whatever value is
>> used to cancel the syscall resulted in a crash rather than just being
>> ignored.
>
> I was trying to keep as much seccomp-ptrace behavior intact rather
> than making it terminal in this special case. Is there a reason why
> it'd make more sense to crash?
Unless you meant something the tracer could catch? That may make
sense, but they could also use singlestep or whatever else to get
similar behavior. But maybe I'm missing the bigger picture!
thanks!
will
>> How hard is it for a page fault to return into the syscall entry path? It
>> should be possible to do this for rel, although it could be messy and not
>> worth it.
>
> Not sure, tbh. I think given vsyscall's status and the fact that
> ptrace+seccomp+vsyscall=emulate isn't horrible, I think it's fine to
> either ignore (what is in tree now) or to allow ptrace to skip,
> without providing full functionality. But obviously, my view my be
> biased!
>
> thanks!
> will
>
>>
>> On Jul 14, 2012 10:35 AM, "Will Drewry" <wad@chromium.org> wrote:
>>>
>>> Current quirky ptrace behavior with vsyscall and seccomp
>>> does not allow tracers to bypass the call. This change
>>> provides that ability by checking if orig_ax changed.
>>>
>>> Signed-off-by: Will Drewry <wad@chromium.org>
>>> ---
>>> arch/x86/kernel/vsyscall_64.c | 10 +++++++---
>>> 1 file changed, 7 insertions(+), 3 deletions(-)
>>>
>>> diff --git a/arch/x86/kernel/vsyscall_64.c b/arch/x86/kernel/vsyscall_64.c
>>> index 5db36ca..5f9640c 100644
>>> --- a/arch/x86/kernel/vsyscall_64.c
>>> +++ b/arch/x86/kernel/vsyscall_64.c
>>> @@ -142,11 +142,15 @@ static int addr_to_vsyscall_nr(unsigned long addr)
>>> #ifdef CONFIG_SECCOMP
>>> static int vsyscall_seccomp(struct task_struct *tsk, int syscall_nr)
>>> {
>>> + int ret;
>>> if (!seccomp_mode(&tsk->seccomp))
>>> return 0;
>>> task_pt_regs(tsk)->orig_ax = syscall_nr;
>>> task_pt_regs(tsk)->ax = syscall_nr;
>>> - return __secure_computing(syscall_nr);
>>> + ret = __secure_computing(syscall_nr);
>>> + if (task_pt_regs(tsk)->orig_ax != syscall_nr)
>>> + return 1; /* ptrace syscall skip */
>>> + return ret;
>>> }
>>> #else
>>> #define vsyscall_seccomp(_tsk, _nr) 0
>>> @@ -278,9 +282,9 @@ bool emulate_vsyscall(struct pt_regs *regs, unsigned
>>> long address)
>>> current_thread_info()->sig_on_uaccess_error =
>>> prev_sig_on_uaccess_error;
>>>
>>> if (skip) {
>>> - if ((long)regs->ax <= 0L) /* seccomp errno emulation */
>>> + if ((long)regs->ax <= 0L || skip == 1) /* seccomp
>>> errno/trace */
>>> goto do_ret;
>>> - goto done; /* seccomp trace/trap */
>>> + goto done; /* seccomp trap */
>>> }
>>>
>>> if (ret == -EFAULT) {
>>> --
>>> 1.7.9.5
>>>
>>
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip
2012-07-14 15:57 ` Will Drewry
@ 2012-07-14 16:15 ` Andrew Lutomirski
2012-07-14 16:21 ` Andrew Lutomirski
0 siblings, 1 reply; 10+ messages in thread
From: Andrew Lutomirski @ 2012-07-14 16:15 UTC (permalink / raw)
To: Will Drewry
Cc: Kees Cook, james.l.morris, rob, linux-doc, cevans, hpa,
linux-kernel, torvalds, eparis, fengxj325
On Sat, Jul 14, 2012 at 8:57 AM, Will Drewry <wad@chromium.org> wrote:
> On Sat, Jul 14, 2012 at 10:50 AM, Will Drewry <wad@chromium.org> wrote:
>> On Sat, Jul 14, 2012 at 10:44 AM, Andrew Lutomirski <luto@mit.edu> wrote:
>>> I think I'd prefer if changing to something other than whatever value is
>>> used to cancel the syscall resulted in a crash rather than just being
>>> ignored.
>>
>> I was trying to keep as much seccomp-ptrace behavior intact rather
>> than making it terminal in this special case. Is there a reason why
>> it'd make more sense to crash?
>
> Unless you meant something the tracer could catch? That may make
> sense, but they could also use singlestep or whatever else to get
> similar behavior. But maybe I'm missing the bigger picture!
I think it would be nice to not introduce any special behavior that
things might rely on if we do this better in the future. Similarly,
for almost all purposes, a tracer could change gettimeofday to write,
but there would be a silent behavior change if gettimeofday were
entered via vsyscall.
What's the standard way of skipping a syscall? sigreturn? (I don't
know off the top of my head what sigreturn does.) sys_ni_syscall?
I'd be all for making those continue to work but making anything that
can't be emulated correctly do something sufficiently unpleasant that
people won't do it. Is there a syscall that does nothing at all?
I wish we could just increment rip by 7 and set a flag to allow the
vsyscall page instructions to be fully emulated until one of the ret
instructions happens, but I don't know how to do that without
monkeying with the entry assembly -- the do_page_fault path doesn't
look enough like system_call to pull it off easily.
In any case, can you change the docs to indicate that the special
behavior only happens iff rip & ~0x0c00 == 0xffffffffff600000? That
way a hypothetical future emulator could add better emulation
(incrementing rip by 7, for example) and everything would still work.
(There's no need to check the syscall number.)
FWIW, this crap is why I'm sort of tempted to say that seccomp should
just ignore vsyscalls entirely (except in mode 1). None of them
actually do anything other than querying things that can be read out
of the vsyscall page (for the most part) without any kernel entry.
(Making *that* part of the ABI would be bad, though.) Better ideas
are welcome.
--Andy
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip
2012-07-14 16:15 ` Andrew Lutomirski
@ 2012-07-14 16:21 ` Andrew Lutomirski
2012-07-14 16:55 ` Will Drewry
0 siblings, 1 reply; 10+ messages in thread
From: Andrew Lutomirski @ 2012-07-14 16:21 UTC (permalink / raw)
To: Will Drewry
Cc: Kees Cook, james.l.morris, rob, linux-doc, cevans, hpa,
linux-kernel, torvalds, eparis, fengxj325
On Sat, Jul 14, 2012 at 9:15 AM, Andrew Lutomirski <luto@mit.edu> wrote:
> On Sat, Jul 14, 2012 at 8:57 AM, Will Drewry <wad@chromium.org> wrote:
>> On Sat, Jul 14, 2012 at 10:50 AM, Will Drewry <wad@chromium.org> wrote:
>>> On Sat, Jul 14, 2012 at 10:44 AM, Andrew Lutomirski <luto@mit.edu> wrote:
>>>> I think I'd prefer if changing to something other than whatever value is
>>>> used to cancel the syscall resulted in a crash rather than just being
>>>> ignored.
>>>
>>> I was trying to keep as much seccomp-ptrace behavior intact rather
>>> than making it terminal in this special case. Is there a reason why
>>> it'd make more sense to crash?
>>
>> Unless you meant something the tracer could catch? That may make
>> sense, but they could also use singlestep or whatever else to get
>> similar behavior. But maybe I'm missing the bigger picture!
>
> I think it would be nice to not introduce any special behavior that
> things might rely on if we do this better in the future. Similarly,
> for almost all purposes, a tracer could change gettimeofday to write,
> but there would be a silent behavior change if gettimeofday were
> entered via vsyscall.
>
> What's the standard way of skipping a syscall? sigreturn? (I don't
> know off the top of my head what sigreturn does.) sys_ni_syscall?
> I'd be all for making those continue to work but making anything that
> can't be emulated correctly do something sufficiently unpleasant that
> people won't do it. Is there a syscall that does nothing at all?
Here's my suggestion for now: any attempt by the seccomp filter to do
anything other than executing the syscall unchanged, skipping it
entirely, or killing results in a trap or sigsys with the syscall
number unchanged.
In the RET_TRACE case, the SIGSYS restarts at the mov nr,%rax
instruction. The tracer can deal with it accordingly (via checking
the rip) -- if the tracer changes rax, it'll get changed right back by
the emulated mov nr,%rax. If this results in an infinite loop, so be
it.
In the RET_TRAP case, do the same thing but via the special ptrace
event instead of SIGSYS.
This way, the special rip & ~0x0c00 == 0xffffffffff600000 handling
becomes part of the ABI, but it looks like the mov instruction
trapped, which it more or less did. If we want RET_TRAP to be able to
skip the syscall, let it happen the normal way.
This stuff probably barely matters. For example, ptrace currently
doesn't work right when vsyscalls are being emulated. No one appears
to care.
--Andy
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip
2012-07-14 16:21 ` Andrew Lutomirski
@ 2012-07-14 16:55 ` Will Drewry
2012-08-10 9:14 ` James Morris
0 siblings, 1 reply; 10+ messages in thread
From: Will Drewry @ 2012-07-14 16:55 UTC (permalink / raw)
To: Andrew Lutomirski
Cc: Kees Cook, james.l.morris, rob, linux-doc, cevans, hpa,
linux-kernel, torvalds, eparis, fengxj325
On Sat, Jul 14, 2012 at 11:21 AM, Andrew Lutomirski <luto@mit.edu> wrote:
> On Sat, Jul 14, 2012 at 9:15 AM, Andrew Lutomirski <luto@mit.edu> wrote:
>> On Sat, Jul 14, 2012 at 8:57 AM, Will Drewry <wad@chromium.org> wrote:
>>> On Sat, Jul 14, 2012 at 10:50 AM, Will Drewry <wad@chromium.org> wrote:
>>>> On Sat, Jul 14, 2012 at 10:44 AM, Andrew Lutomirski <luto@mit.edu> wrote:
>>>>> I think I'd prefer if changing to something other than whatever value is
>>>>> used to cancel the syscall resulted in a crash rather than just being
>>>>> ignored.
>>>>
>>>> I was trying to keep as much seccomp-ptrace behavior intact rather
>>>> than making it terminal in this special case. Is there a reason why
>>>> it'd make more sense to crash?
>>>
>>> Unless you meant something the tracer could catch? That may make
>>> sense, but they could also use singlestep or whatever else to get
>>> similar behavior. But maybe I'm missing the bigger picture!
>>
>> I think it would be nice to not introduce any special behavior that
>> things might rely on if we do this better in the future. Similarly,
>> for almost all purposes, a tracer could change gettimeofday to write,
>> but there would be a silent behavior change if gettimeofday were
>> entered via vsyscall.
>>
>> What's the standard way of skipping a syscall? sigreturn? (I don't
>> know off the top of my head what sigreturn does.) sys_ni_syscall?
Most implementations I've seen just set orig_rax to -1, but in
practice any system call that is invalid will result in the system
call being skipped in the normal syscall path. So I could tweak it to
check <0 > __NR_syscalls and add a sigsys generator if it is valid.
Do you think that'd be the most correct move? (As you say below, the
impact of _any_ changes here are very minor :)
>> I'd be all for making those continue to work but making anything that
>> can't be emulated correctly do something sufficiently unpleasant that
>> people won't do it. Is there a syscall that does nothing at all?
>
> Here's my suggestion for now: any attempt by the seccomp filter to do
> anything other than executing the syscall unchanged, skipping it
> entirely, or killing results in a trap or sigsys with the syscall
> number unchanged.
>
> In the RET_TRACE case, the SIGSYS restarts at the mov nr,%rax
> instruction. The tracer can deal with it accordingly (via checking
> the rip) -- if the tracer changes rax, it'll get changed right back by
> the emulated mov nr,%rax. If this results in an infinite loop, so be
> it.
>
> In the RET_TRAP case, do the same thing but via the special ptrace
> event instead of SIGSYS.
You've lost me here. The flow of a RET_TRAP is:
- enter vsyscall at 0xffffffffff600[]00
- run seccomp
- get ret_trap
- queue up a synchronous trap
- skip the vsyscall
- return
Even if the trap were to interrupt the current vsyscall trap handler,
the skip value would still be -1 and the syscall would still get
skipped without touching sp or ip.
Traps cannot provide kernel-side arbitration of allowed syscalls, they
can just emulate them in userspace before "returning" to the original
callsite. In this case, vsyscall makes returning there harder and the
quirk is to return to the prior address if the ip is in vsyscall's
range.
> This way, the special rip & ~0x0c00 == 0xffffffffff600000 handling
> becomes part of the ABI, but it looks like the mov instruction
> trapped, which it more or less did. If we want RET_TRAP to be able to
> skip the syscall, let it happen the normal way.
I *think* trap is ok?
> This stuff probably barely matters. For example, ptrace currently
> doesn't work right when vsyscalls are being emulated. No one appears
> to care.
Agreed :) I don't mind making tweaks to get it right, but this only
matters to users that want to:
- use seccomp filter
- with ptrace (or trap with resumption and not sigreturn)
- of time, gettimeofday, and getcpu
since they will then have to include quirk management _just_ in case
their code is linked against something using vsyscall and
vsyscall=emulate is in effect!
thanks!
will
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip
2012-07-14 16:55 ` Will Drewry
@ 2012-08-10 9:14 ` James Morris
2012-08-13 18:17 ` Andrew Lutomirski
0 siblings, 1 reply; 10+ messages in thread
From: James Morris @ 2012-08-10 9:14 UTC (permalink / raw)
To: Will Drewry
Cc: Andrew Lutomirski, Kees Cook, james.l.morris, rob, linux-doc,
cevans, hpa, linux-kernel, torvalds, eparis, fengxj325
On Sat, 14 Jul 2012, Will Drewry wrote:
> Agreed :) I don't mind making tweaks to get it right, but this only
> matters to users that want to:
> - use seccomp filter
> - with ptrace (or trap with resumption and not sigreturn)
> - of time, gettimeofday, and getcpu
> since they will then have to include quirk management _just_ in case
> their code is linked against something using vsyscall and
> vsyscall=emulate is in effect!
I think these patches came out during the merge window -- is there any
further discussion on them?
- James
--
James Morris
<jmorris@namei.org>
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip
2012-08-10 9:14 ` James Morris
@ 2012-08-13 18:17 ` Andrew Lutomirski
0 siblings, 0 replies; 10+ messages in thread
From: Andrew Lutomirski @ 2012-08-13 18:17 UTC (permalink / raw)
To: James Morris
Cc: Will Drewry, Kees Cook, james.l.morris, rob, linux-doc, cevans,
hpa, linux-kernel, torvalds, eparis, fengxj325
On Fri, Aug 10, 2012 at 2:14 AM, James Morris <jmorris@namei.org> wrote:
> On Sat, 14 Jul 2012, Will Drewry wrote:
>
>> Agreed :) I don't mind making tweaks to get it right, but this only
>> matters to users that want to:
>> - use seccomp filter
>> - with ptrace (or trap with resumption and not sigreturn)
>> - of time, gettimeofday, and getcpu
>> since they will then have to include quirk management _just_ in case
>> their code is linked against something using vsyscall and
>> vsyscall=emulate is in effect!
>
> I think these patches came out during the merge window -- is there any
> further discussion on them?
The patch I sent on Aug 1 is the latest version. I don't think
there's been any further discussion -- everyone seems reasonably happy
with that version.
--Andy
>
>
> - James
> --
> James Morris
> <jmorris@namei.org>
^ permalink raw reply [flat|nested] 10+ messages in thread
end of thread, other threads:[~2012-08-13 18:17 UTC | newest]
Thread overview: 10+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2012-07-14 15:32 [PATCH 1/3] vsyscall_64: add missing ifdef CONFIG_SECCOMP Will Drewry
2012-07-14 15:32 ` [PATCH 2/3] vsyscall_64: allow SECCOMP_RET_TRACErs to skip Will Drewry
[not found] ` <CAObL_7GX0q_qywY4g1S2iRZWPS95ar01wtvv9wDW=zswwaZ6fQ@mail.gmail.com>
2012-07-14 15:50 ` Will Drewry
2012-07-14 15:57 ` Will Drewry
2012-07-14 16:15 ` Andrew Lutomirski
2012-07-14 16:21 ` Andrew Lutomirski
2012-07-14 16:55 ` Will Drewry
2012-08-10 9:14 ` James Morris
2012-08-13 18:17 ` Andrew Lutomirski
2012-07-14 15:32 ` [PATCH 3/3] Documentation: add a caveat for seccomp filter and vsyscall emulation Will Drewry
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®