* Re: sigopen() vs. /dev/sigtimedwait
@ 2001-08-07 14:20 Erich Nahum
0 siblings, 0 replies; 4+ messages in thread
From: Erich Nahum @ 2001-08-07 14:20 UTC (permalink / raw)
To: linux-kernel
Abhishek Chandra and I are benchmarking /dev/epoll vs. RT signals
with signal-per-FD, and we wanted to chip in some thoughts along
these lines.
First of all, Davide Libenzi's /dev/epoll does not have the same
semantics as the original /dev/poll that Sun did. Select, poll,
and the original /dev/poll are all state-based mechanisms, whereas
/dev/epoll and RT signals are event-based mechanisms. In the state-based
approach, the application can ask the kernel which file descriptors
are ready to read or write to. In the event-based approach, the
kernel notifies the application when something changes. This has
serious implications for how one develops the server; in the
event-based case, the server has to keep track of the state of the
connections more carefully. For more discussion of event-based vs.
state-based, see the original Banga/Druschel/Mogul work via Dan
Kegel's c10k page (http://www.kegel.com/c10k.html)
Both /dev/epoll and RT signals with sig-per-fd have the property that
the event queue never overflows, since events are coalesced on a per-fd
basis, assuming the user-space server isn't broken. If the server
underestimages the max number of file descriptors it can use, the
event queue can overflow in either scenario.
Event-based interfaces have some conditions that the server developer
has to be aware of. For example, when a server using writes to a
socket for the first time, /dev/poll will tell you the socket is ready,
whereas no event will show up on /dev/epoll, since the socket write
state hasn't changed. If you naively wait for a write event to happen
(as we did before we realized this), you'll wait a long time.
Some race conditions can also occur. One is when the data arrives
on the socket after the accept but before the kernel is notifyied via
/dev/epoll, thus never generating an event. Another involves getting
stray events after the fd is closed (soon to be fixed according to
Davide Libenzi). A third is when you have simultaneous reads and writes
going on a socket, as happens with HTTP 1.1.
As far as we can tell, the /dev/epoll has the same semantics as the
RT signals with signal-per-fd. The differences are in the interfaces,
which may have some performance implications. For example,
/dev/epoll can get batches of events through the read to /dev/epoll,
whereas RT signals get one signal at a time through sigtimedwait().
On the other hand, RT signals don't have to explicitly notify the kernel
with each new or closed connection the way /dev/epoll does. Instead,
it's done implicitly through the setsockopt/fcntl call to make the socket
asynchronous/non-blocking, which the server has to do anyway.
As I mentioned earlier, we're benchmarking these to see what the
performance difference is, if any. Davide Libenzi is also pursuing
this comparison.
So far Abhishek and I haven't looked at Ben LeHaises async I/O interface,
but it's on our schedule.
-Erich
--
Erich M. Nahum IBM T.J. Watson Research Center
Networking Research P.O. Box 704
nahum@watson.ibm.com Yorktown Heights NY 10598
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: sigopen() vs. /dev/sigtimedwait
2001-08-04 1:38 ` Petru Paler
@ 2001-08-04 2:10 ` Dan Kegel
0 siblings, 0 replies; 4+ messages in thread
From: Dan Kegel @ 2001-08-04 2:10 UTC (permalink / raw)
To: Petru Paler; +Cc: Christopher Smith, linux-kernel, Zach Brown
Petru Paler wrote:
>
> On Fri, Aug 03, 2001 at 06:32:52PM -0700, Dan Kegel wrote:
> > So I'm proposing the following user story:
> >
> > // open a fd linked to signal mysignum
> > int fd = open("/dev/sigtimedwait", O_RDWR);
> > int sigs[1]; sigs[0] = mysignum;
> > write(fd, sigs, sizeof(sigs[0]));
> >
> > // memory map a result buffer
> > struct siginfo_t *map = mmap(NULL, mapsize, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);
> >
> > for (;;) {
> > // grab recent siginfo_t's
> > struct devsiginfo dsi;
> > dsi.dsi_nsis = 1000;
> > dsi.dsi_sis = NULL; // NULL means "use map instead of buffer"
> > dsi.dsi_timeout = 1;
> > int nsis = ioctl(fd, DS_SIGTIMEDWAIT, &dvp);
> >
> > // use 'em. Some might be completion notifications; some might be readiness notifications.
> > for (i=0; i<nsis; i++)
> > handle_siginfo(map+i);
> > }
>
> And the advantage of this over /dev/epoll would be that you don't have to
> explicitly add/remove fd's?
The advantage is that it can be used to collect
completion notifications for aio. (It can also be
used to collect readiness notification via either
linux's traditional rtsig stuff, or the signal-per-fd stuff,
so this unifies readiness notification and completion notification,
in case you happen to want to use both in the same thread.)
> I ask because yesterday I used /dev/epoll in a project and it behaves *very*
> well, so I'm wondering what advantages your interface would bring.
I am a huge fan of /dev/epoll and would like to see it integrated
into the ac series. /dev/epoll doesn't address the needs of those
who are doing aio, though.
> How do you handle signal queue overflow? signal-per-fd helps, but you still
> have to have the queue as big as the maximum number of fds is...
I am not addressing that issue. However, when doing aio, the
application can simply avoid issuing more than N I/O operations,
where N is comfortably lower than the current size of the signal queue.
When I get around to reading the kernel source finally, maybe I'll
have a look at what the costs of large signal queues are.
- Dan
--
"I have seen the future, and it licks itself clean." -- Bucky Katt
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: sigopen() vs. /dev/sigtimedwait
2001-08-04 1:32 Dan Kegel
@ 2001-08-04 1:38 ` Petru Paler
2001-08-04 2:10 ` Dan Kegel
0 siblings, 1 reply; 4+ messages in thread
From: Petru Paler @ 2001-08-04 1:38 UTC (permalink / raw)
To: Dan Kegel; +Cc: Christopher Smith, linux-kernel, Michael Elkins, Zach Brown
On Fri, Aug 03, 2001 at 06:32:52PM -0700, Dan Kegel wrote:
> So I'm proposing the following user story:
>
> // open a fd linked to signal mysignum
> int fd = open("/dev/sigtimedwait", O_RDWR);
> int sigs[1]; sigs[0] = mysignum;
> write(fd, sigs, sizeof(sigs[0]));
>
> // memory map a result buffer
> struct siginfo_t *map = mmap(NULL, mapsize, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);
>
> for (;;) {
> // grab recent siginfo_t's
> struct devsiginfo dsi;
> dsi.dsi_nsis = 1000;
> dsi.dsi_sis = NULL; // NULL means "use map instead of buffer"
> dsi.dsi_timeout = 1;
> int nsis = ioctl(fd, DS_SIGTIMEDWAIT, &dvp);
>
> // use 'em. Some might be completion notifications; some might be readiness notifications.
> for (i=0; i<nsis; i++)
> handle_siginfo(map+i);
> }
And the advantage of this over /dev/epoll would be that you don't have to
explicitly add/remove fd's?
I ask because yesterday I used /dev/epoll in a project and it behaves *very*
well, so I'm wondering what advantages your interface would bring.
> Comments?
How do you handle signal queue overflow? signal-per-fd helps, but you still
have to have the queue as big as the maximum number of fds is...
Petru
^ permalink raw reply [flat|nested] 4+ messages in thread
* sigopen() vs. /dev/sigtimedwait
@ 2001-08-04 1:32 Dan Kegel
2001-08-04 1:38 ` Petru Paler
0 siblings, 1 reply; 4+ messages in thread
From: Dan Kegel @ 2001-08-04 1:32 UTC (permalink / raw)
To: Christopher Smith, linux-kernel; +Cc: Michael Elkins, Zach Brown
So I've been thinking about the sigopen() system call I proposed.
(To recap: sigopen() would let you use read() instead of sigwaitinfo()
to retrieve lots of realtime signals at one go, AND would
protect your signal from being swiped by hostile code elsewhere
in the application, a la Sun's JDK.)
Upon further consideration, maybe I should model it after
/dev/epoll. That would get rid of nagging questions like
"but read() can't leave holes like sigtimedwait could",
and would be even higher performance than read()
(see graphs at http://www.xmailserver.org/linux-patches/nio-improve.html )
So I'm proposing the following user story:
// open a fd linked to signal mysignum
int fd = open("/dev/sigtimedwait", O_RDWR);
int sigs[1]; sigs[0] = mysignum;
write(fd, sigs, sizeof(sigs[0]));
// memory map a result buffer
struct siginfo_t *map = mmap(NULL, mapsize, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);
for (;;) {
// grab recent siginfo_t's
struct devsiginfo dsi;
dsi.dsi_nsis = 1000;
dsi.dsi_sis = NULL; // NULL means "use map instead of buffer"
dsi.dsi_timeout = 1;
int nsis = ioctl(fd, DS_SIGTIMEDWAIT, &dvp);
// use 'em. Some might be completion notifications; some might be readiness notifications.
for (i=0; i<nsis; i++)
handle_siginfo(map+i);
}
Sure, the interface is crap, but it's fast, and at least it doesn't
add any syscalls (the sigopen() proposal required two new syscalls: sigopen()
and timedread()).
Comments?
BTW I'm halfway thru "Understanding the Linux Kernel" and it's
a very good read (modulo some strange lingo, e.g. "cycle" for "loop"
and "table" for "record" or "struct").
So since I only halfway understand the linux kernel, the above proposal
may be half baked.
- Dan
--
"I have seen the future, and it licks itself clean." -- Bucky Katt
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2001-08-07 14:20 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2001-08-07 14:20 sigopen() vs. /dev/sigtimedwait Erich Nahum
-- strict thread matches above, loose matches on Subject: below --
2001-08-04 1:32 Dan Kegel
2001-08-04 1:38 ` Petru Paler
2001-08-04 2:10 ` Dan Kegel
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®