* [PATCH] anycast support for IPv6, linux-2.5.31
@ 2002-08-28 21:44 David Stevens
2002-08-28 22:22 ` Christoph Hellwig
2002-08-30 6:37 ` Pekka Savola
0 siblings, 2 replies; 4+ messages in thread
From: David Stevens @ 2002-08-28 21:44 UTC (permalink / raw)
To: linux-kernel, linux-net
[-- Attachment #1: Type: text/plain, Size: 8642 bytes --]
Below is a patch relative to the mainline 2.5.31 code for an
implementation of anycast support for IPv6. This code was submitted and accepted
in the USAGI tree last Fall. Below is a high-level description of the
implementation:
1) The API
Although the RFC's liken anycasting to ordinary unicasting, I think
it's more appropriate to tie it closely to particular applications, so I've
chosen an API similar to multicasting. So, rather than having a permanent
anycast address associated with the machine, particular applications
that use anycasting can join or leave "anycast groups," and the machine will
recognize the anycast addresses as its own when one or more applications have
joined the group.
So, for example, someone using anycasting for DNS high availability
can add a join to the anycast group in the server and as long as the DNS server
is running, the machine will answer to that anycast address. But the machine
will not respond to anycasts when the service that's using it isn't available,
so a broken server application that has exited won't deny that service if
there are other working members of the anycast group on other hosts.
I don't know if that's controversial or not-- the RFC's are written
more from the external context, but seem to imply a model along the lines of
using "ifconfig" to add anycast addresses. I think that model doesn't fit the
best uses of anycasting, but I'd like to hear your thoughts on it.
The application interface for joining and leaving anycast groups is 2
new setsockopt() calls: IPV6_JOIN_ANYCAST and IPV6_LEAVE_ANYCAST. The arguments
are the same as the corresponding multicast operations. The kernel keeps a
reference count of members; when that goes to zero, the anycast address is not
recognized as a local address. While nonzero, the host listens on the solicited
node for that address, sends advertisements in response to solicitations (with
override=0) and delivers packets sent to the anycast address to upper layers.
There's also an in-kernel interface described below, which is used by
IPv6 mobility, for example.
2) Security Model
RFC 2373 states:
"
o An anycast address must not be assigned to an IPv6 host, that is, it may be
assigned to an IPv6 router only."
This patch violates this in 1 special case, and I'll explain why.
a) The restriction on host use of anycast is to avoid carrying individual host
routes for anycast addresses spread out among multiple physical
networks. I think the initial application sets are exactly things that
won't be on off-the-shelf routers (high availabily servers (DNS, http,
etc) and mobile IPv6) and the particular cases don't have the problem of
requiring host routes or participation in the routing system. They use
anycast addresses with a prefix common to a unicast address on the
system, so ordinary routing gets you to the right network, anyway, and
there's no external penalty on the routing system for using those types
of anycast addresses. For that reason, I allow anycast addresses that
match an existing unicast prefix even on hosts.
Finally (for security considerations), I had to choose whether anycast
should require root privilege or not. Multicasting does not, but it'd obviously
be a spoofing issue if an application joined an "anycast" that was actually the
unicast address of another machine on that network. On the other hand, it's
handy for non-root users to be able to make use of anycasting where that use
doesn't pose any security risks.
The code below allows non-root users to join anycast groups that have
matching prefixes (don't require special-route propagation) with existing
unicast addresses, and require root (really "CAP_NET_ADMIN") and a router for
off-link anycasts (disallowed completely on hosts). I think that should be
extended to require CAP_NET_ADMIN for any anycasts (even on-link ones) that are
not well-known anycasts (to avoid the spoofing of on-link unicast addresses).
4) The Implementation
The code maintains a list of anycast addresses that are in use for
a given interface. The code is a modifed version of the existing multicast
code, with some things cleaned up, and operations on the anycast list instead
of the multicast list. Because the anycast address list is separate from the
ordinary address list, anycast addresses in general won't be selected as a
source address, or available for inappropriate uses. Protocols (like ICMP ECHO)
that respond by swapping the source and destination address have a separate
check for anycasts and set the source to zero in that case-- allows IPv6 to
choose the outbound source address.
The code has the setsockopt() interface for joining and leaving anycast
groups, but does not yet have changes needed for UDP and TCP to work with them.
TCP is problematic, because the PCB lookup mechanism relies on the destination
address which must change-- it should be disallowed initially. UDP may work
with an INADDR_ANY-bound listener, but I haven't made changes to support it
yet. It will probably use the anycast address as the source, so it'll need a
modification similar to what I've done with ICMP, but should be straightforward.
Ultimately, I think we want to allow binding to anycast addresses as well.
Our immediate application is mobile IPv6, so this patch doesn't include
any of the upper-layer changes that may be needed for general application
support.
For in-kernel use, applications (like mobile IPv6) can call join and
drop functions for anycast addresses, and a function that checks if a device
is in an anycast group (if dev == 0, checks if any device is in that group).
They are (similar to multicast functions):
int ipv6_dev_ac_inc(struct net_device *dev, struct in6_addr *addr)
- add "addr" as an anycast address on "dev"
int ipv6_dev_ac_dec(struct net_device *dev, struct in6_addr *addr)
- remove "addr" as an anycast address on "dev"
these use reference counts, so only the first call to "inc" for a particular
address will add a new address, and only when all references are removed via
"dec" will the address be removed as a local address.
The function:
int ipv6_chk_acast_addr(struct net_device *dev, struct in6_addr *addr)
returns true if "addr" is an anycast address on "dev", false otherwise. If
"dev" is 0, it searches all devices for "addr".
Those 3 functions provide the in-kernel interface.
4) Things of Note
I think we want the ip6_addr_type() to check *only* the well-known
anycasts, since it seems inappropriate to me that that function should be
searching linked lists of anycast addresses. It would also need a "dev"
argument it doesn't have now, since anycast addresses, like unicast and
multicast addresses, in this implementation are associated with particular
devices. Use of those address on other devices should not return type ANYCAST,
but should for the device that has the anycast address. So, in most cases,
ipv6_chk_acast_addr() and not ipv6_addr_type() will be more appropriate.
ipv6_addr_type(), with modifications included for reserved anycast
addresses, will still be useful for cases where the address is known to
*always* be an anycast (for example, disallowing reserved anycasts through
"ifconfig" being set as an ordinary address), but for the lower-level code,
it'll usually need a per-device check. So, I recommend we keep both, and use
ipv6_chk_acast_addr() to answer if it is a configured anycast address, use
ipv6_addr_type() to answer if the address is reserved for anycast (whether
configured or not).
That's what this code does.
5) Testing
I wrote programs to join and leave anycast groups and I checked through
the /proc/net interface (file "anycast6") the presence of the groups. I've
used network sniffers to watch the neighbor discovery sequence and verify the
override bit is cleared, and I've tested with multiple hosts in the anycast
group talking to an unmodifed host that pings the anycast address. I also
verified that the existing code handles "override=0" correctly (it does).
In addition, our mobile IPv6 team has used the code to test the use of
anycasting for Dynamic Home Agent address discovery, with several different
topologies and configurations.
We've done tests with uniprocessor and SMP kernels on multiprocessor
machines.
6) TODO
I think the next steps are to flesh out the UDP part so ordinary
user-level applications can make full use of anycasting.
+-DLS
(See attached file: anycast-2.5.31.patch)
[-- Attachment #2: anycast-2.5.31.patch --]
[-- Type: application/octet-stream, Size: 33492 bytes --]
diff -urN linux-2.5.31/include/linux/in6.h linux-2.5.31AC/include/linux/in6.h
--- linux-2.5.31/include/linux/in6.h Sat Aug 10 18:41:42 2002
+++ linux-2.5.31AC/include/linux/in6.h Tue Aug 20 12:39:50 2002
@@ -56,6 +56,8 @@
int ipv6mr_ifindex;
};
+#define ipv6mr_acaddr ipv6mr_multiaddr
+
struct in6_flowlabel_req
{
struct in6_addr flr_dst;
@@ -156,6 +158,9 @@
#define IPV6_MTU_DISCOVER 23
#define IPV6_MTU 24
#define IPV6_RECVERR 25
+/* 26 is IPV6_V6ONLY in USAGI code */
+#define IPV6_JOIN_ANYCAST 27
+#define IPV6_LEAVE_ANYCAST 28
/* IPV6_MTU_DISCOVER values */
#define IPV6_PMTUDISC_DONT 0
diff -urN linux-2.5.31/include/linux/ipv6.h linux-2.5.31AC/include/linux/ipv6.h
--- linux-2.5.31/include/linux/ipv6.h Sat Aug 10 18:41:45 2002
+++ linux-2.5.31AC/include/linux/ipv6.h Tue Aug 20 12:45:16 2002
@@ -155,6 +155,7 @@
pmtudisc:2;
struct ipv6_mc_socklist *ipv6_mc_list;
+ struct ipv6_ac_socklist *ipv6_ac_list;
struct ipv6_fl_socklist *ipv6_fl_list;
__u32 dst_cookie;
diff -urN linux-2.5.31/include/linux/netdevice.h linux-2.5.31AC/include/linux/netdevice.h
--- linux-2.5.31/include/linux/netdevice.h Sat Aug 10 18:41:55 2002
+++ linux-2.5.31AC/include/linux/netdevice.h Tue Aug 20 12:35:50 2002
@@ -462,6 +462,10 @@
extern void dev_add_pack(struct packet_type *pt);
extern void dev_remove_pack(struct packet_type *pt);
extern int dev_get(const char *name);
+extern struct net_device *dev_getany(unsigned short flags,
+ unsigned short mask);
+extern struct net_device *__dev_getany(unsigned short flags,
+ unsigned short mask);
extern struct net_device *dev_get_by_name(const char *name);
extern struct net_device *__dev_get_by_name(const char *name);
extern struct net_device *dev_alloc(const char *name, int *err);
diff -urN linux-2.5.31/include/net/addrconf.h linux-2.5.31AC/include/net/addrconf.h
--- linux-2.5.31/include/net/addrconf.h Sat Aug 10 18:41:24 2002
+++ linux-2.5.31AC/include/net/addrconf.h Wed Aug 21 15:12:35 2002
@@ -58,7 +58,15 @@
extern int ipv6_get_saddr(struct dst_entry *dst,
struct in6_addr *daddr,
struct in6_addr *saddr);
+extern int ipv6_dev_get_saddr(struct net_device *dev,
+ struct in6_addr *daddr,
+ struct in6_addr *saddr,
+ int onlink);
extern int ipv6_get_lladdr(struct net_device *dev, struct in6_addr *);
+extern void addrconf_join_solict(struct net_device *dev,
+ struct in6_addr *addr);
+extern void addrconf_leave_solict(struct net_device *dev,
+ struct in6_addr *addr);
/*
* multicast prototypes (mcast.c)
@@ -88,6 +96,26 @@
extern void addrconf_prefix_rcv(struct net_device *dev,
u8 *opt, int len);
+/*
+ * anycast prototypes (anycast.c)
+ */
+extern int ipv6_sock_ac_join(struct sock *sk,
+ int ifindex,
+ struct in6_addr *addr);
+extern int ipv6_sock_ac_drop(struct sock *sk,
+ int ifindex,
+ struct in6_addr *addr);
+extern void ipv6_sock_ac_close(struct sock *sk);
+extern int inet6_ac_check(struct sock *sk, struct in6_addr *addr, int ifindex);
+
+extern int ipv6_dev_ac_inc(struct net_device *dev,
+ struct in6_addr *addr);
+extern int ipv6_dev_ac_dec(struct net_device *dev,
+ struct in6_addr *addr);
+extern int ipv6_chk_acast_addr(struct net_device *dev,
+ struct in6_addr *addr);
+
+
/* Device notifier */
extern int register_inet6addr_notifier(struct notifier_block *nb);
extern int unregister_inet6addr_notifier(struct notifier_block *nb);
diff -urN linux-2.5.31/include/net/if_inet6.h linux-2.5.31AC/include/net/if_inet6.h
--- linux-2.5.31/include/net/if_inet6.h Sat Aug 10 18:41:56 2002
+++ linux-2.5.31AC/include/net/if_inet6.h Tue Aug 20 12:35:50 2002
@@ -69,6 +69,25 @@
spinlock_t mca_lock;
};
+/* Anycast stuff */
+
+struct ipv6_ac_socklist
+{
+ struct in6_addr acl_addr;
+ int acl_ifindex;
+ struct ipv6_ac_socklist *acl_next;
+};
+
+struct ifacaddr6
+{
+ struct in6_addr aca_addr;
+ struct inet6_dev *aca_idev;
+ struct ifacaddr6 *aca_next;
+ int aca_users;
+ atomic_t aca_refcnt;
+ spinlock_t aca_lock;
+};
+
#define IFA_HOST IPV6_ADDR_LOOPBACK
#define IFA_LINK IPV6_ADDR_LINKLOCAL
#define IFA_SITE IPV6_ADDR_SITELOCAL
@@ -96,6 +115,7 @@
struct inet6_ifaddr *addr_list;
struct ifmcaddr6 *mc_list;
+ struct ifacaddr6 *ac_list;
rwlock_t lock;
atomic_t refcnt;
__u32 if_flags;
diff -urN linux-2.5.31/net/core/dev.c linux-2.5.31AC/net/core/dev.c
--- linux-2.5.31/net/core/dev.c Sat Aug 10 18:41:30 2002
+++ linux-2.5.31AC/net/core/dev.c Tue Aug 20 12:35:50 2002
@@ -541,6 +541,50 @@
}
/**
+ * dev_getany - find any device with given flags
+ * @if_flags: IFF_* values
+ * @mask: bitmask of bits in if_flags to check
+ *
+ * Search for any interface with the given flags. Returns NULL if a device
+ * is not found or a pointer to the device. The device returned has
+ * had a reference added and the pointer is safe until the user calls
+ * dev_put to indicate they have finished with it.
+ */
+
+struct net_device * dev_getany(unsigned short if_flags, unsigned short mask)
+{
+ struct net_device *dev;
+
+ read_lock(&dev_base_lock);
+ dev = __dev_getany(if_flags, mask);
+ if (dev)
+ dev_hold(dev);
+ read_unlock(&dev_base_lock);
+ return dev;
+}
+
+/**
+ * __dev_getany - find any device with given flags
+ * @if_flags: IFF_* values
+ * @mask: bitmask of bits in if_flags to check
+ *
+ * Search for any interface with the given flags. Returns NULL if a device
+ * is not found or a pointer to the device. The caller must hold either
+ * the RTNL semaphore or @dev_base_lock.
+ */
+
+struct net_device *__dev_getany(unsigned short if_flags, unsigned short mask)
+{
+ struct net_device *dev;
+
+ for (dev = dev_base; dev != NULL; dev = dev->next) {
+ if (((dev->flags ^ if_flags) & mask) == 0)
+ return dev;
+ }
+ return NULL;
+}
+
+/**
* dev_alloc_name - allocate a name for a device
* @dev: device
* @name: name format string
diff -urN linux-2.5.31/net/ipv6/Makefile linux-2.5.31AC/net/ipv6/Makefile
--- linux-2.5.31/net/ipv6/Makefile Sat Aug 10 18:41:48 2002
+++ linux-2.5.31AC/net/ipv6/Makefile Tue Aug 20 12:46:35 2002
@@ -6,7 +6,7 @@
obj-$(CONFIG_IPV6) += ipv6.o
-ipv6-objs := af_inet6.o ip6_output.o ip6_input.o addrconf.o sit.o \
+ipv6-objs := af_inet6.o anycast.o ip6_output.o ip6_input.o addrconf.o sit.o \
route.o ip6_fib.o ipv6_sockglue.o ndisc.o udp.o raw.o \
protocol.o icmp.o mcast.o reassembly.o tcp_ipv6.o \
exthdrs.o sysctl_net_ipv6.o datagram.o proc.o \
diff -urN linux-2.5.31/net/ipv6/addrconf.c linux-2.5.31AC/net/ipv6/addrconf.c
--- linux-2.5.31/net/ipv6/addrconf.c Sat Aug 10 18:41:55 2002
+++ linux-2.5.31AC/net/ipv6/addrconf.c Wed Aug 21 15:13:38 2002
@@ -134,19 +134,13 @@
int ipv6_addr_type(struct in6_addr *addr)
{
+ int type;
u32 st;
st = addr->s6_addr32[0];
- /* Consider all addresses with the first three bits different of
- 000 and 111 as unicasts.
- */
- if ((st & __constant_htonl(0xE0000000)) != __constant_htonl(0x00000000) &&
- (st & __constant_htonl(0xE0000000)) != __constant_htonl(0xE0000000))
- return IPV6_ADDR_UNICAST;
-
if ((st & __constant_htonl(0xFF000000)) == __constant_htonl(0xFF000000)) {
- int type = IPV6_ADDR_MULTICAST;
+ type = IPV6_ADDR_MULTICAST;
switch((st & __constant_htonl(0x00FF0000))) {
case __constant_htonl(0x00010000):
@@ -163,12 +157,28 @@
};
return type;
}
+ /* check for reserved anycast addresses */
+
+ if ((st & __constant_htonl(0xE0000000)) &&
+ ((addr->s6_addr32[2] == __constant_htonl(0xFDFFFFFF) &&
+ (addr->s6_addr32[3] | __constant_htonl(0x7F)) == (u32)~0) ||
+ (addr->s6_addr32[2] == 0 && addr->s6_addr32[3] == 0)))
+ type = IPV6_ADDR_ANYCAST;
+ else
+ type = IPV6_ADDR_UNICAST;
+
+ /* Consider all addresses with the first three bits different of
+ 000 and 111 as finished.
+ */
+ if ((st & __constant_htonl(0xE0000000)) != __constant_htonl(0x00000000) &&
+ (st & __constant_htonl(0xE0000000)) != __constant_htonl(0xE0000000))
+ return type;
if ((st & __constant_htonl(0xFFC00000)) == __constant_htonl(0xFE800000))
- return (IPV6_ADDR_LINKLOCAL | IPV6_ADDR_UNICAST);
+ return (IPV6_ADDR_LINKLOCAL | type);
if ((st & __constant_htonl(0xFFC00000)) == __constant_htonl(0xFEC00000))
- return (IPV6_ADDR_SITELOCAL | IPV6_ADDR_UNICAST);
+ return (IPV6_ADDR_SITELOCAL | type);
if ((addr->s6_addr32[0] | addr->s6_addr32[1]) == 0) {
if (addr->s6_addr32[2] == 0) {
@@ -176,16 +186,24 @@
return IPV6_ADDR_ANY;
if (addr->s6_addr32[3] == __constant_htonl(0x00000001))
- return (IPV6_ADDR_LOOPBACK | IPV6_ADDR_UNICAST);
+ return (IPV6_ADDR_LOOPBACK | type);
- return (IPV6_ADDR_COMPATv4 | IPV6_ADDR_UNICAST);
+ return (IPV6_ADDR_COMPATv4 | type);
}
if (addr->s6_addr32[2] == __constant_htonl(0x0000ffff))
return IPV6_ADDR_MAPPED;
}
- return IPV6_ADDR_RESERVED;
+ st &= __constant_htonl(0xFF000000);
+ if (st == 0)
+ return IPV6_ADDR_RESERVED;
+ st &= __constant_htonl(0xFE000000);
+ if (st == __constant_htonl(0x02000000))
+ return IPV6_ADDR_RESERVED; /* for NSAP */
+ if (st == __constant_htonl(0x04000000))
+ return IPV6_ADDR_RESERVED; /* for IPX */
+ return type;
}
static void addrconf_del_timer(struct inet6_ifaddr *ifp)
@@ -221,7 +239,6 @@
add_timer(&ifp->timer);
}
-
/* Nobody refers to this device, we may destroy it. */
void in6_dev_finish_destroy(struct inet6_dev *idev)
@@ -300,24 +317,91 @@
return idev;
}
+void ipv6_addr_prefix(struct in6_addr *prefix,
+ struct in6_addr *addr, int prefix_len)
+{
+ unsigned long mask;
+ int ncopy, nbits;
+
+ memset(prefix, 0, sizeof(*prefix));
+
+ if (prefix_len <= 0)
+ return;
+ if (prefix_len > 128)
+ prefix_len = 128;
+
+ ncopy = prefix_len / 32;
+ switch (ncopy) {
+ case 4: prefix->s6_addr32[3] = addr->s6_addr32[3];
+ case 3: prefix->s6_addr32[2] = addr->s6_addr32[2];
+ case 2: prefix->s6_addr32[1] = addr->s6_addr32[1];
+ case 1: prefix->s6_addr32[0] = addr->s6_addr32[0];
+ case 0: break;
+ }
+ nbits = prefix_len % 32;
+ if (nbits == 0)
+ return;
+
+ mask = ~((1 << (32 - nbits)) - 1);
+ mask = htonl(mask);
+
+ prefix->s6_addr32[ncopy] = addr->s6_addr32[ncopy] & mask;
+}
+
+
+static void dev_forward_change(struct inet6_dev *idev)
+{
+ struct net_device *dev;
+ struct inet6_ifaddr *ifa;
+ struct in6_addr addr;
+
+ if (!idev)
+ return;
+ dev = idev->dev;
+ if (dev && (dev->flags & IFF_MULTICAST)) {
+ ipv6_addr_all_routers(&addr);
+
+ if (idev->cnf.forwarding)
+ ipv6_dev_mc_inc(dev, &addr);
+ else
+ ipv6_dev_mc_dec(dev, &addr);
+ }
+ for (ifa=idev->addr_list; ifa; ifa=ifa->if_next) {
+ ipv6_addr_prefix(&addr, &ifa->addr, ifa->prefix_len);
+ if (addr.s6_addr32[0] == 0 && addr.s6_addr32[1] == 0 &&
+ addr.s6_addr32[2] == 0 && addr.s6_addr32[3] == 0)
+ continue;
+ if (idev->cnf.forwarding)
+ ipv6_dev_ac_inc(idev->dev, &addr);
+ else
+ ipv6_dev_ac_dec(idev->dev, &addr);
+ }
+}
+
+
static void addrconf_forward_change(struct inet6_dev *idev)
{
struct net_device *dev;
- if (idev)
+ if (idev) {
+ dev_forward_change(idev);
return;
+ }
read_lock(&dev_base_lock);
for (dev=dev_base; dev; dev=dev->next) {
read_lock(&addrconf_lock);
idev = __in6_dev_get(dev);
- if (idev)
+ if (idev) {
idev->cnf.forwarding = ipv6_devconf.forwarding;
+ dev_forward_change(idev);
+ }
read_unlock(&addrconf_lock);
}
read_unlock(&dev_base_lock);
}
+
/* Nobody refers to this ifaddr, destroy it */
void inet6_ifa_finish_destroy(struct inet6_ifaddr *ifp)
@@ -455,29 +539,20 @@
* an address of the attached interface
* iii) don't use deprecated addresses
*/
-int ipv6_get_saddr(struct dst_entry *dst,
- struct in6_addr *daddr, struct in6_addr *saddr)
+int ipv6_dev_get_saddr(struct net_device *dev,
+ struct in6_addr *daddr, struct in6_addr *saddr, int onlink)
{
- int scope;
struct inet6_ifaddr *ifp = NULL;
struct inet6_ifaddr *match = NULL;
- struct net_device *dev = NULL;
struct inet6_dev *idev;
- struct rt6_info *rt;
+ int scope;
int err;
- rt = (struct rt6_info *) dst;
- if (rt)
- dev = rt->rt6i_dev;
- scope = ipv6_addr_scope(daddr);
- if (rt && (rt->rt6i_flags & RTF_ALLONLINK)) {
- /*
- * route for the "all destinations on link" rule
- * when no routers are present
- */
+ if (!onlink)
+ scope = ipv6_addr_scope(daddr);
+ else
scope = IFA_LINK;
- }
/*
* known dev
@@ -565,6 +640,24 @@
return err;
}
+
+int ipv6_get_saddr(struct dst_entry *dst,
+ struct in6_addr *daddr, struct in6_addr *saddr)
+{
+ struct rt6_info *rt;
+ struct net_device *dev = NULL;
+ int onlink;
+
+ rt = (struct rt6_info *) dst;
+ if (rt)
+ dev = rt->rt6i_dev;
+
+ onlink = (rt && (rt->rt6i_flags & RTF_ALLONLINK));
+
+ return ipv6_dev_get_saddr(dev, daddr, saddr, onlink);
+}
+
+
int ipv6_get_lladdr(struct net_device *dev, struct in6_addr *addr)
{
struct inet6_dev *idev;
@@ -657,7 +750,7 @@
/* Join to solicited addr multicast group. */
-static void addrconf_join_solict(struct net_device *dev, struct in6_addr *addr)
+void addrconf_join_solict(struct net_device *dev, struct in6_addr *addr)
{
struct in6_addr maddr;
@@ -668,7 +761,7 @@
ipv6_dev_mc_inc(dev, &maddr);
}
-static void addrconf_leave_solict(struct net_device *dev, struct in6_addr *addr)
+void addrconf_leave_solict(struct net_device *dev, struct in6_addr *addr)
{
struct in6_addr maddr;
@@ -1553,6 +1646,15 @@
addrconf_mod_timer(ifp, AC_RS, ifp->idev->cnf.rtr_solicit_interval);
spin_unlock_bh(&ifp->lock);
}
+
+ if (ifp->idev->cnf.forwarding) {
+ struct in6_addr addr;
+
+ ipv6_addr_prefix(&addr, &ifp->addr, ifp->prefix_len);
+ if (addr.s6_addr32[0] || addr.s6_addr32[1] ||
+ addr.s6_addr32[2] || addr.s6_addr32[3])
+ ipv6_dev_ac_inc(ifp->idev->dev, &addr);
+ }
}
#ifdef CONFIG_PROC_FS
@@ -1833,6 +1935,14 @@
break;
case RTM_DELADDR:
addrconf_leave_solict(ifp->idev->dev, &ifp->addr);
+ if (ifp->idev->cnf.forwarding) {
+ struct in6_addr addr;
+
+ ipv6_addr_prefix(&addr, &ifp->addr, ifp->prefix_len);
+ if (addr.s6_addr32[0] || addr.s6_addr32[1] ||
+ addr.s6_addr32[2] || addr.s6_addr32[3])
+ ipv6_dev_ac_dec(ifp->idev->dev, &addr);
+ }
if (!ipv6_chk_addr(&ifp->addr, NULL))
ip6_rt_addr_del(&ifp->addr, ifp->idev->dev);
break;
@@ -1855,11 +1965,7 @@
struct inet6_dev *idev = NULL;
if (valp != &ipv6_devconf.forwarding) {
- struct net_device *dev = dev_get_by_index(ctl->ctl_name);
- if (dev) {
- idev = in6_dev_get(dev);
- dev_put(dev);
- }
+ idev = (struct inet6_dev *)ctl->extra1;
if (idev == NULL)
return ret;
} else
@@ -1869,8 +1975,6 @@
if (*valp)
rt6_purge_dflt_routers(0);
- if (idev)
- in6_dev_put(idev);
}
return ret;
@@ -1947,6 +2051,7 @@
for (i=0; i<sizeof(t->addrconf_vars)/sizeof(t->addrconf_vars[0])-1; i++) {
t->addrconf_vars[i].data += (char*)p - (char*)&ipv6_devconf;
t->addrconf_vars[i].de = NULL;
+ t->addrconf_vars[i].extra1 = idev; /* embedded; no ref */
}
if (dev) {
t->addrconf_dev[0].procname = dev->name;
diff -urN linux-2.5.31/net/ipv6/af_inet6.c linux-2.5.31AC/net/ipv6/af_inet6.c
--- linux-2.5.31/net/ipv6/af_inet6.c Sat Aug 10 18:41:36 2002
+++ linux-2.5.31AC/net/ipv6/af_inet6.c Tue Aug 20 13:14:22 2002
@@ -76,6 +76,7 @@
/* IPv6 procfs goodies... */
#ifdef CONFIG_PROC_FS
+extern int anycast6_get_info(char *, char **, off_t, int);
extern int raw6_get_info(char *, char **, off_t, int);
extern int tcp6_get_info(char *, char **, off_t, int);
extern int udp6_get_info(char *, char **, off_t, int);
@@ -380,6 +381,9 @@
/* Free mc lists */
ipv6_sock_mc_close(sk);
+ /* Free ac lists */
+ ipv6_sock_ac_close(sk);
+
return inet_release(sock);
}
@@ -712,6 +716,8 @@
goto proc_sockstat6_fail;
if (!proc_net_create("snmp6", 0, afinet6_get_snmp))
goto proc_snmp6_fail;
+ if (!proc_net_create("anycast6", 0, anycast6_get_info))
+ goto proc_anycast6_fail;
#endif
ipv6_netdev_notif_init();
ipv6_packet_init();
@@ -727,6 +733,8 @@
return 0;
#ifdef CONFIG_PROC_FS
+proc_anycast6_fail:
+ proc_net_remove("anycast6");
proc_snmp6_fail:
proc_net_remove("sockstat6");
proc_sockstat6_fail:
@@ -762,6 +770,7 @@
proc_net_remove("udp6");
proc_net_remove("sockstat6");
proc_net_remove("snmp6");
+ proc_net_remove("anycast6");
#endif
/* Cleanup code parts. */
sit_cleanup();
diff -urN linux-2.5.31/net/ipv6/anycast.c linux-2.5.31AC/net/ipv6/anycast.c
--- linux-2.5.31/net/ipv6/anycast.c Wed Dec 31 16:00:00 1969
+++ linux-2.5.31AC/net/ipv6/anycast.c Wed Aug 21 14:24:41 2002
@@ -0,0 +1,508 @@
+/* $Header$ */
+
+/*
+ * Anycast support for IPv6
+ * Linux INET6 implementation
+ *
+ * Authors:
+ * David L Stevens (dlsteven@us.ibm.com)
+ *
+ * $Id$
+ *
+ * based heavily on net/ipv6/mcast.c
+ *
+ * This program is free software; you can redistribute it and/or
+ * modify it under the terms of the GNU General Public License
+ * as published by the Free Software Foundation; either version
+ * 2 of the License, or (at your option) any later version.
+ */
+
+/* Changes:
+ *
+ */
+
+#define __NO_VERSION__
+#include <linux/config.h>
+#include <linux/module.h>
+#include <linux/errno.h>
+#include <linux/types.h>
+#include <linux/random.h>
+#include <linux/string.h>
+#include <linux/socket.h>
+#include <linux/sockios.h>
+#include <linux/sched.h>
+#include <linux/net.h>
+#include <linux/in6.h>
+#include <linux/netdevice.h>
+#include <linux/if_arp.h>
+#include <linux/route.h>
+#include <linux/init.h>
+#include <linux/proc_fs.h>
+
+#include <net/sock.h>
+#include <net/snmp.h>
+
+#include <net/ipv6.h>
+#include <net/protocol.h>
+#include <net/if_inet6.h>
+#include <net/ndisc.h>
+#include <net/addrconf.h>
+#include <net/ip6_route.h>
+
+#include <net/checksum.h>
+
+#ifdef CONFIG_IPV6_MLD6_DEBUG
+#include <linux/inet.h>
+#endif
+
+/* Big ac list lock for all the sockets */
+static rwlock_t ipv6_sk_ac_lock = RW_LOCK_UNLOCKED;
+
+/* XXX ip6_addr_match() and ip6_onlink() really belong in net/core.c */
+
+static int
+ip6_addr_match(struct in6_addr *addr1, struct in6_addr *addr2, int prefix)
+{
+ __u32 mask;
+ int i;
+
+ if (prefix > 128 || prefix < 0)
+ return 0;
+ if (prefix == 0)
+ return 1;
+ for (i=0; i<4; ++i) {
+ if (prefix >= 32)
+ mask = ~0;
+ else
+ mask = htonl(~0 << (32 - prefix));
+ if ((addr1->s6_addr32[i] ^ addr2->s6_addr32[i]) & mask)
+ return 0;
+ prefix -= 32;
+ if (prefix <= 0)
+ break;
+ }
+ return 1;
+}
+
+static int
+ip6_onlink(struct in6_addr *addr, struct net_device *dev)
+{
+ struct inet6_dev *idev;
+ struct inet6_ifaddr *ifa;
+ int onlink;
+
+ onlink = 0;
+ read_lock(&addrconf_lock);
+ idev = __in6_dev_get(dev);
+ if (idev) {
+ read_lock_bh(&idev->lock);
+ for (ifa=idev->addr_list; ifa; ifa=ifa->if_next) {
+ onlink = ip6_addr_match(addr, &ifa->addr,
+ ifa->prefix_len);
+ if (onlink)
+ break;
+ }
+ read_unlock_bh(&idev->lock);
+ }
+ read_unlock(&addrconf_lock);
+ return onlink;
+}
+
+
+/*
+ * socket join an anycast group
+ */
+
+int ipv6_sock_ac_join(struct sock *sk, int ifindex, struct in6_addr *addr)
+{
+ struct ipv6_pinfo *np = inet6_sk(sk);
+ struct net_device *dev = NULL;
+ struct inet6_dev *idev;
+ struct ipv6_ac_socklist *pac;
+ int ishost = !ipv6_devconf.forwarding;
+ int err = 0;
+
+ if (ipv6_addr_type(addr) & IPV6_ADDR_MULTICAST)
+ return -EINVAL;
+
+ pac = sock_kmalloc(sk, sizeof(struct ipv6_ac_socklist), GFP_KERNEL);
+ if (pac == NULL)
+ return -ENOMEM;
+ pac->acl_next = NULL;
+ ipv6_addr_copy(&pac->acl_addr, addr);
+
+ if (ifindex == 0) {
+ struct rt6_info *rt;
+
+ rt = rt6_lookup(addr, NULL, 0, 0);
+ if (rt) {
+ dev = rt->rt6i_dev;
+ dev_hold(dev);
+ dst_release(&rt->u.dst);
+ } else if (ishost) {
+ sock_kfree_s(sk, pac, sizeof(*pac));
+ return -EADDRNOTAVAIL;
+ } else {
+ /* router, no matching interface: just pick one */
+
+ dev = dev_getany(IFF_UP, IFF_UP|IFF_LOOPBACK);
+ }
+ } else
+ dev = dev_get_by_index(ifindex);
+
+ if (dev == NULL) {
+ sock_kfree_s(sk, pac, sizeof(*pac));
+ return -ENODEV;
+ }
+
+ idev = in6_dev_get(dev);
+ if (!idev) {
+ sock_kfree_s(sk, pac, sizeof(*pac));
+ dev_put(dev);
+ if (ifindex)
+ return -ENODEV;
+ else
+ return -EADDRNOTAVAIL;
+ }
+ /* reset ishost, now that we have a specific device */
+ ishost = !idev->cnf.forwarding;
+ in6_dev_put(idev);
+
+ pac->acl_ifindex = dev->ifindex;
+
+ /* XXX
+ * For hosts, allow link-local or matching prefix anycasts.
+ * This obviates the need for propagating anycast routes while
+ * still allowing some non-router anycast participation.
+ *
+ * allow anyone to join anycasts that don't require a special route
+ * and can't be spoofs of unicast addresses (reserved anycast only)
+ */
+ if (!ip6_onlink(addr, dev)) {
+ if (ishost)
+ err = -EADDRNOTAVAIL;
+ else if (!capable(CAP_NET_ADMIN))
+ err = -EPERM;
+ if (err) {
+ sock_kfree_s(sk, pac, sizeof(*pac));
+ dev_put(dev);
+ return err;
+ }
+ } else if (!(ipv6_addr_type(addr) & IPV6_ADDR_ANYCAST) &&
+ !capable(CAP_NET_ADMIN))
+ return -EPERM;
+
+ err = ipv6_dev_ac_inc(dev, addr);
+ if (err) {
+ sock_kfree_s(sk, pac, sizeof(*pac));
+ dev_put(dev);
+ return err;
+ }
+
+ write_lock_bh(&ipv6_sk_ac_lock);
+ pac->acl_next = np->ipv6_ac_list;
+ np->ipv6_ac_list = pac;
+ write_unlock_bh(&ipv6_sk_ac_lock);
+
+ dev_put(dev);
+
+ return 0;
+}
+
+/*
+ * socket leave an anycast group
+ */
+int ipv6_sock_ac_drop(struct sock *sk, int ifindex, struct in6_addr *addr)
+{
+ struct ipv6_pinfo *np = inet6_sk(sk);
+ struct net_device *dev;
+ struct ipv6_ac_socklist *pac, *prev_pac;
+
+ write_lock_bh(&ipv6_sk_ac_lock);
+ prev_pac = 0;
+ for (pac = np->ipv6_ac_list; pac; pac = pac->acl_next) {
+ if ((ifindex == 0 || pac->acl_ifindex == ifindex) &&
+ ipv6_addr_cmp(&pac->acl_addr, addr) == 0)
+ break;
+ prev_pac = pac;
+ }
+ if (!pac) {
+ write_unlock_bh(&ipv6_sk_ac_lock);
+ return -ENOENT;
+ }
+ if (prev_pac)
+ prev_pac->acl_next = pac->acl_next;
+ else
+ np->ipv6_ac_list = pac->acl_next;
+
+ write_unlock_bh(&ipv6_sk_ac_lock);
+
+ dev = dev_get_by_index(pac->acl_ifindex);
+ if (dev) {
+ ipv6_dev_ac_dec(dev, &pac->acl_addr);
+ dev_put(dev);
+ }
+ sock_kfree_s(sk, pac, sizeof(*pac));
+ return 0;
+}
+
+void ipv6_sock_ac_close(struct sock *sk)
+{
+ struct ipv6_pinfo *np = inet6_sk(sk);
+ struct net_device *dev = 0;
+ struct ipv6_ac_socklist *pac;
+ int prev_index;
+
+ write_lock_bh(&ipv6_sk_ac_lock);
+ pac = np->ipv6_ac_list;
+ np->ipv6_ac_list = 0;
+ write_unlock_bh(&ipv6_sk_ac_lock);
+
+ prev_index = 0;
+ while (pac) {
+ struct ipv6_ac_socklist *next = pac->acl_next;
+
+ if (pac->acl_ifindex != prev_index) {
+ if (dev)
+ dev_put(dev);
+ dev = dev_get_by_index(pac->acl_ifindex);
+ prev_index = pac->acl_ifindex;
+ }
+ if (dev)
+ ipv6_dev_ac_dec(dev, &pac->acl_addr);
+ sock_kfree_s(sk, pac, sizeof(*pac));
+ pac = next;
+ }
+ if (dev)
+ dev_put(dev);
+}
+
+int inet6_ac_check(struct sock *sk, struct in6_addr *addr, int ifindex)
+{
+ struct ipv6_ac_socklist *pac;
+ struct ipv6_pinfo *np = inet6_sk(sk);
+ int found;
+
+ found = 0;
+ read_lock(&ipv6_sk_ac_lock);
+ for (pac=np->ipv6_ac_list; pac; pac=pac->acl_next) {
+ if (ifindex && pac->acl_ifindex != ifindex)
+ continue;
+ found = ipv6_addr_cmp(&pac->acl_addr, addr) == 0;
+ if (found)
+ break;
+ }
+ read_unlock(&ipv6_sk_ac_lock);
+
+ return found;
+}
+
+static void aca_put(struct ifacaddr6 *ac)
+{
+ if (atomic_dec_and_test(&ac->aca_refcnt)) {
+ in6_dev_put(ac->aca_idev);
+ kfree(ac);
+ }
+}
+
+/*
+ * device anycast group inc (add if not found)
+ */
+int ipv6_dev_ac_inc(struct net_device *dev, struct in6_addr *addr)
+{
+ struct ifacaddr6 *aca;
+ struct inet6_dev *idev;
+
+ idev = in6_dev_get(dev);
+
+ if (idev == NULL)
+ return -EINVAL;
+
+ write_lock_bh(&idev->lock);
+ if (idev->dead) {
+ write_unlock_bh(&idev->lock);
+ in6_dev_put(idev);
+ return -ENODEV;
+ }
+
+ for (aca = idev->ac_list; aca; aca = aca->aca_next) {
+ if (ipv6_addr_cmp(&aca->aca_addr, addr) == 0) {
+ aca->aca_users++;
+ write_unlock_bh(&idev->lock);
+ in6_dev_put(idev);
+ return 0;
+ }
+ }
+
+ /*
+ * not found: create a new one.
+ */
+
+ aca = kmalloc(sizeof(struct ifacaddr6), GFP_ATOMIC);
+
+ if (aca == NULL) {
+ write_unlock_bh(&idev->lock);
+ in6_dev_put(idev);
+ return -ENOMEM;
+ }
+
+ memset(aca, 0, sizeof(struct ifacaddr6));
+
+ ipv6_addr_copy(&aca->aca_addr, addr);
+ aca->aca_idev = idev;
+ aca->aca_users = 1;
+ atomic_set(&aca->aca_refcnt, 2);
+ aca->aca_lock = SPIN_LOCK_UNLOCKED;
+
+ aca->aca_next = idev->ac_list;
+ idev->ac_list = aca;
+ write_unlock_bh(&idev->lock);
+
+ ip6_rt_addr_add(&aca->aca_addr, dev);
+
+ addrconf_join_solict(dev, &aca->aca_addr);
+
+ aca_put(aca);
+ return 0;
+}
+
+/*
+ * device anycast group decrement
+ */
+int ipv6_dev_ac_dec(struct net_device *dev, struct in6_addr *addr)
+{
+ struct inet6_dev *idev;
+ struct ifacaddr6 *aca, *prev_aca;
+
+ idev = in6_dev_get(dev);
+ if (idev == NULL)
+ return -ENODEV;
+
+ write_lock_bh(&idev->lock);
+ prev_aca = 0;
+ for (aca = idev->ac_list; aca; aca = aca->aca_next) {
+ if (ipv6_addr_cmp(&aca->aca_addr, addr) == 0)
+ break;
+ prev_aca = aca;
+ }
+ if (!aca) {
+ write_unlock_bh(&idev->lock);
+ in6_dev_put(idev);
+ return -ENOENT;
+ }
+ if (--aca->aca_users > 0) {
+ write_unlock_bh(&idev->lock);
+ in6_dev_put(idev);
+ return 0;
+ }
+ if (prev_aca)
+ prev_aca->aca_next = aca->aca_next;
+ else
+ idev->ac_list = aca->aca_next;
+ write_unlock_bh(&idev->lock);
+ addrconf_leave_solict(dev, &aca->aca_addr);
+
+ ip6_rt_addr_del(&aca->aca_addr, dev);
+
+ aca_put(aca);
+ in6_dev_put(idev);
+ return 0;
+}
+
+/*
+ * check if the interface has this anycast address
+ */
+static int ipv6_chk_acast_dev(struct net_device *dev, struct in6_addr *addr)
+{
+ struct inet6_dev *idev;
+ struct ifacaddr6 *aca;
+
+ idev = in6_dev_get(dev);
+ if (idev) {
+ read_lock_bh(&idev->lock);
+ for (aca = idev->ac_list; aca; aca = aca->aca_next)
+ if (ipv6_addr_cmp(&aca->aca_addr, addr) == 0)
+ break;
+ read_unlock_bh(&idev->lock);
+ in6_dev_put(idev);
+ return aca != 0;
+ }
+ return 0;
+}
+
+/*
+ * check if given interface (or any, if dev==0) has this anycast address
+ */
+int ipv6_chk_acast_addr(struct net_device *dev, struct in6_addr *addr)
+{
+ if (dev)
+ return ipv6_chk_acast_dev(dev, addr);
+ read_lock(&dev_base_lock);
+ for (dev=dev_base; dev; dev=dev->next)
+ if (ipv6_chk_acast_dev(dev, addr))
+ break;
+ read_unlock(&dev_base_lock);
+ return dev != 0;
+}
+
+
+/* IPv6 device initialization. */
+
+void ipv6_ac_init_dev(struct inet6_dev *idev)
+{
+}
+
+#ifdef CONFIG_PROC_FS
+int anycast6_get_info(char *buffer, char **start, off_t offset, int length)
+{
+ off_t pos=0, begin=0;
+ struct ifacaddr6 *im;
+ int len=0;
+ struct net_device *dev;
+
+ read_lock(&dev_base_lock);
+ for (dev = dev_base; dev; dev = dev->next) {
+ struct inet6_dev *idev;
+
+ if ((idev = in6_dev_get(dev)) == NULL)
+ continue;
+
+ read_lock_bh(&idev->lock);
+ for (im = idev->ac_list; im; im = im->aca_next) {
+ int i;
+
+ len += sprintf(buffer+len,"%-4d %-15s ", dev->ifindex, dev->name);
+
+ for (i=0; i<16; i++)
+ len += sprintf(buffer+len, "%02x", im->aca_addr.s6_addr[i]);
+
+ len += sprintf(buffer+len, " %5d\n", im->aca_users);
+
+ pos=begin+len;
+ if (pos < offset) {
+ len=0;
+ begin=pos;
+ }
+ if (pos > offset+length) {
+ read_unlock_bh(&idev->lock);
+ in6_dev_put(idev);
+ goto done;
+ }
+ }
+ read_unlock_bh(&idev->lock);
+ in6_dev_put(idev);
+ }
+
+done:
+ read_unlock(&dev_base_lock);
+
+ *start=buffer+(offset-begin);
+ len-=(offset-begin);
+ if(len>length)
+ len=length;
+ if (len<0)
+ len=0;
+ return len;
+}
+
+#endif
diff -urN linux-2.5.31/net/ipv6/icmp.c linux-2.5.31AC/net/ipv6/icmp.c
--- linux-2.5.31/net/ipv6/icmp.c Sat Aug 10 18:41:25 2002
+++ linux-2.5.31AC/net/ipv6/icmp.c Tue Aug 20 12:35:50 2002
@@ -395,7 +395,8 @@
saddr = &skb->nh.ipv6h->daddr;
- if (ipv6_addr_type(saddr) & IPV6_ADDR_MULTICAST)
+ if (ipv6_addr_type(saddr) & IPV6_ADDR_MULTICAST ||
+ ipv6_chk_acast_addr(0, saddr))
saddr = NULL;
msg.icmph.icmp6_type = ICMPV6_ECHO_REPLY;
diff -urN linux-2.5.31/net/ipv6/ipv6_sockglue.c linux-2.5.31AC/net/ipv6/ipv6_sockglue.c
--- linux-2.5.31/net/ipv6/ipv6_sockglue.c Sat Aug 10 18:41:19 2002
+++ linux-2.5.31AC/net/ipv6/ipv6_sockglue.c Tue Aug 20 12:35:50 2002
@@ -351,6 +351,24 @@
retv = ipv6_sock_mc_drop(sk, mreq.ipv6mr_ifindex, &mreq.ipv6mr_multiaddr);
break;
}
+ case IPV6_JOIN_ANYCAST:
+ case IPV6_LEAVE_ANYCAST:
+ {
+ struct ipv6_mreq mreq;
+
+ if (optlen != sizeof(struct ipv6_mreq))
+ goto e_inval;
+
+ retv = -EFAULT;
+ if (copy_from_user(&mreq, optval, sizeof(struct ipv6_mreq)))
+ break;
+
+ if (optname == IPV6_JOIN_ANYCAST)
+ retv = ipv6_sock_ac_join(sk, mreq.ipv6mr_ifindex, &mreq.ipv6mr_acaddr);
+ else
+ retv = ipv6_sock_ac_drop(sk, mreq.ipv6mr_ifindex, &mreq.ipv6mr_acaddr);
+ break;
+ }
case IPV6_ROUTER_ALERT:
retv = ip6_ra_control(sk, val, NULL);
break;
diff -urN linux-2.5.31/net/ipv6/ndisc.c linux-2.5.31AC/net/ipv6/ndisc.c
--- linux-2.5.31/net/ipv6/ndisc.c Sat Aug 10 18:41:41 2002
+++ linux-2.5.31AC/net/ipv6/ndisc.c Wed Aug 21 16:33:20 2002
@@ -316,7 +316,10 @@
struct in6_addr *daddr, struct in6_addr *solicited_addr,
int router, int solicited, int override, int inc_opt)
{
+ static struct in6_addr tmpaddr;
+ struct inet6_ifaddr *ifp;
struct sock *sk = ndisc_socket->sk;
+ struct in6_addr *src_addr;
struct nd_msg *msg;
int len;
struct sk_buff *skb;
@@ -338,13 +341,23 @@
ND_PRINTK1("send_na: alloc skb failed\n");
return;
}
+ /* for anycast or proxy, solicited_addr != src_addr */
+ ifp = ipv6_get_ifaddr(solicited_addr, dev);
+ if (ifp) {
+ src_addr = solicited_addr;
+ in6_ifa_put(ifp);
+ } else {
+ if (ipv6_dev_get_saddr(dev, daddr, &tmpaddr, 0))
+ return;
+ src_addr = &tmpaddr;
+ }
if (ndisc_build_ll_hdr(skb, dev, daddr, neigh, len) == 0) {
kfree_skb(skb);
return;
}
- ip6_nd_hdr(sk, skb, dev, solicited_addr, daddr, IPPROTO_ICMPV6, len);
+ ip6_nd_hdr(sk, skb, dev, src_addr, daddr, IPPROTO_ICMPV6, len);
msg = (struct nd_msg *) skb_put(skb, len);
@@ -364,7 +377,7 @@
ndisc_fill_option(msg->opt, ND_OPT_TARGET_LL_ADDR, dev->dev_addr, dev->addr_len);
/* checksum */
- msg->icmph.icmp6_cksum = csum_ipv6_magic(solicited_addr, daddr, len,
+ msg->icmph.icmp6_cksum = csum_ipv6_magic(src_addr, daddr, len,
IPPROTO_ICMPV6,
csum_partial((__u8 *) msg,
len, 0));
@@ -1080,6 +1093,51 @@
}
}
in6_ifa_put(ifp);
+ } else if (ipv6_chk_acast_addr(dev, &msg->target)) {
+ struct inet6_dev *idev = in6_dev_get(dev);
+ int addr_type = ipv6_addr_type(saddr);
+
+ /* anycast */
+
+ if (!idev) {
+ /* XXX: count this drop? */
+ return 0;
+ }
+
+ if (addr_type == IPV6_ADDR_ANY) {
+ struct in6_addr maddr;
+
+ ipv6_addr_all_nodes(&maddr);
+ ndisc_send_na(dev, NULL, &maddr, &msg->target,
+ idev->cnf.forwarding, 0, 0, 1);
+ in6_dev_put(idev);
+ return 0;
+ }
+
+ if (addr_type & IPV6_ADDR_UNICAST) {
+ int inc = ipv6_addr_type(daddr)&IPV6_ADDR_MULTICAST;
+ if (inc)
+ nd_tbl.stats.rcv_probes_mcast++;
+ else
+ nd_tbl.stats.rcv_probes_ucast++;
+
+ /*
+ * update / create cache entry
+ * for the source adddress
+ */
+
+ neigh = ndisc_recv_ns(saddr, skb);
+
+ if (neigh || !dev->hard_header) {
+ ndisc_send_na(dev, neigh, saddr,
+ &msg->target,
+ idev->cnf.forwarding, 1, 0, inc);
+ if (neigh)
+ neigh_release(neigh);
+ }
+ }
+ in6_dev_put(idev);
+
} else {
struct inet6_dev *in6_dev = in6_dev_get(dev);
int addr_type = ipv6_addr_type(saddr);
diff -urN linux-2.5.31/net/netsyms.c linux-2.5.31AC/net/netsyms.c
--- linux-2.5.31/net/netsyms.c Sat Aug 10 18:41:27 2002
+++ linux-2.5.31AC/net/netsyms.c Tue Aug 20 12:35:50 2002
@@ -462,6 +462,8 @@
EXPORT_SYMBOL(unregister_netdevice);
EXPORT_SYMBOL(netdev_state_change);
EXPORT_SYMBOL(dev_new_index);
+EXPORT_SYMBOL(dev_getany);
+EXPORT_SYMBOL(__dev_getany);
EXPORT_SYMBOL(dev_get_by_index);
EXPORT_SYMBOL(__dev_get_by_index);
EXPORT_SYMBOL(dev_get_by_name);
^ permalink raw reply [flat|nested] 4+ messages in thread* Re: [PATCH] anycast support for IPv6, linux-2.5.31
2002-08-28 21:44 [PATCH] anycast support for IPv6, linux-2.5.31 David Stevens
@ 2002-08-28 22:22 ` Christoph Hellwig
2002-08-30 6:37 ` Pekka Savola
1 sibling, 0 replies; 4+ messages in thread
From: Christoph Hellwig @ 2002-08-28 22:22 UTC (permalink / raw)
To: David Stevens; +Cc: linux-kernel, linux-net
On Wed, Aug 28, 2002 at 03:44:57PM -0600, David Stevens wrote:
>
> Below is a patch relative to the mainline 2.5.31 code for an
I think it would make sense to Cc the netdev list at oss.sgi.com..
(and please inline the patch, makes it much easier to respond..)
diff -urN linux-2.5.31/net/ipv6/anycast.c linux-2.5.31AC/net/ipv6/anycast.c
--- linux-2.5.31/net/ipv6/anycast.c Wed Dec 31 16:00:00 1969
+++ linux-2.5.31AC/net/ipv6/anycast.c Wed Aug 21 14:24:41 2002
@@ -0,0 +1,508 @@
+/* $Header$ */
+
+/*
+ * Anycast support for IPv6
+ * Linux INET6 implementation
+ *
+ * Authors:
+ * David L Stevens (dlsteven@us.ibm.com)
+ *
+ * $Id$
+ *
+ * based heavily on net/ipv6/mcast.c
+ *
+ * This program is free software; you can redistribute it and/or
+ * modify it under the terms of the GNU General Public License
+ * as published by the Free Software Foundation; either version
+ * 2 of the License, or (at your option) any later version.
+ */
+
+/* Changes:
+ *
+ */
Umm, in the kernel tree $Header$ and $Id$ will never be filled out.
Also empty changes comments are pretty useless.
+
+#define __NO_VERSION__
Not needed in 2.4/2.5
+#ifdef CONFIG_IPV6_MLD6_DEBUG
+#include <linux/inet.h>
+#endif
This patch doesn't reference CONFIG_IPV6_MLD6_DEBUG anywhere else..
+
+void ipv6_ac_init_dev(struct inet6_dev *idev)
+{
+}
I can't see this actually beeing used anywhere..
+#ifdef CONFIG_PROC_FS
+int anycast6_get_info(char *buffer, char **start, off_t offset, int length)
+{
+ off_t pos=0, begin=0;
+ struct ifacaddr6 *im;
+ int len=0;
+ struct net_device *dev;
+
+ read_lock(&dev_base_lock);
+ for (dev = dev_base; dev; dev = dev->next) {
+ struct inet6_dev *idev;
+
+ if ((idev = in6_dev_get(dev)) == NULL)
+ continue;
+
+ read_lock_bh(&idev->lock);
+ for (im = idev->ac_list; im; im = im->aca_next) {
+ int i;
This function would really benefit from use of the seq_file API..
^ permalink raw reply [flat|nested] 4+ messages in thread* Re: [PATCH] anycast support for IPv6, linux-2.5.31
2002-08-28 21:44 [PATCH] anycast support for IPv6, linux-2.5.31 David Stevens
2002-08-28 22:22 ` Christoph Hellwig
@ 2002-08-30 6:37 ` Pekka Savola
1 sibling, 0 replies; 4+ messages in thread
From: Pekka Savola @ 2002-08-30 6:37 UTC (permalink / raw)
To: David Stevens; +Cc: linux-kernel, linux-net, netdev
On Wed, 28 Aug 2002, David Stevens wrote:
> 1) The API
> Although the RFC's liken anycasting to ordinary unicasting, I think
> it's more appropriate to tie it closely to particular applications, so I've
> chosen an API similar to multicasting. So, rather than having a permanent
> anycast address associated with the machine, particular applications
> that use anycasting can join or leave "anycast groups," and the machine will
> recognize the anycast addresses as its own when one or more applications have
> joined the group.
> So, for example, someone using anycasting for DNS high availability
> can add a join to the anycast group in the server and as long as the DNS server
> is running, the machine will answer to that anycast address. But the machine
> will not respond to anycasts when the service that's using it isn't available,
> so a broken server application that has exited won't deny that service if
> there are other working members of the anycast group on other hosts.
> I don't know if that's controversial or not-- the RFC's are written
> more from the external context, but seem to imply a model along the lines of
> using "ifconfig" to add anycast addresses. I think that model doesn't fit the
> best uses of anycasting, but I'd like to hear your thoughts on it.
> The application interface for joining and leaving anycast groups is 2
> new setsockopt() calls: IPV6_JOIN_ANYCAST and IPV6_LEAVE_ANYCAST. The arguments
> are the same as the corresponding multicast operations. The kernel keeps a
> reference count of members; when that goes to zero, the anycast address is not
> recognized as a local address. While nonzero, the host listens on the solicited
> node for that address, sends advertisements in response to solicitations (with
> override=0) and delivers packets sent to the anycast address to upper layers.
> There's also an in-kernel interface described below, which is used by
> IPv6 mobility, for example.
Before going too much down this path, I think one should write an Internet
Draft about the proposed API (should be quite short & simple) and see what
kind of response it has in the relevant working groups.
--
Pekka Savola "Tell me of difficulties surmounted,
Netcore Oy not those you stumble over and fall"
Systems. Networks. Security. -- Robert Jordan: A Crown of Swords
^ permalink raw reply [flat|nested] 4+ messages in thread
* Re: [PATCH] anycast support for IPv6, linux-2.5.31
@ 2002-08-30 8:16 David Stevens
0 siblings, 0 replies; 4+ messages in thread
From: David Stevens @ 2002-08-30 8:16 UTC (permalink / raw)
To: Pekka Savola; +Cc: linux-kernel, linux-net, netdev
Pekka,
You wrote:
>Before going too much down this path, I think one should write an Internet
>Draft about the proposed API (should be quite short & simple) and see what
>kind of response it has in the relevant working groups.
I don't disagree with that, for informational purposes, but it doesn't
conflict
with the RFC's, which of course don't cover API's, and don't specify any
interface
for anycasting.
However, my primary goal is to get anycasting support with an in-kernel
interface
in 2.5 before the freeze. :-) I used the setsockopt() API for testing, and
left it
in the patch for others to do the same. Though I think it's the right
approach, for
the reasons I mentioned, I'd rather see that portion pulled from the patch
if it's
controversial, than have the in-kernel interface and anycasting proper
delayed over
that.
The one use of anycast I'm aware of right now is for IPv6 mobility, which
needs the in-kernel interface. The
user-level interface is important for future applications, and a
reference-counted setsockopt() interface doesn't
mean we can't also have an ip/ifconfig interface for permanent anycast
addresses, too (the required anycast
addresses in this patch are permanent, for example). So I don't see it as
committing to one choice, but having
in-kernel anycast support (soon) I think is the more important first step.
+-DLS
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2002-08-30 8:12 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2002-08-28 21:44 [PATCH] anycast support for IPv6, linux-2.5.31 David Stevens
2002-08-28 22:22 ` Christoph Hellwig
2002-08-30 6:37 ` Pekka Savola
2002-08-30 8:16 David Stevens
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®