mirror of
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git
synced 2026-08-31 03:35:32 -04:00
Merge branch 'net-nexthop-per-nexthop-udp-dst-port-for-fdb-vxlan-nexthops'
Jack Ma says:
====================
net: nexthop: per-nexthop UDP dst port for fdb (VXLAN) nexthops
FDB nexthops let a VXLAN fdb entry point at a group of remote VTEPs, with the
kernel flow-hashing across the group (commit 1274e1cc42 ("vxlan: ecmp support
for mac fdb entries")). Each leg carries its own remote IP, but the UDP
destination port is always taken from the VXLAN device (vxlan->cfg.dst_port)
and cannot be set per leg.
This series adds an optional per-nexthop UDP destination port for fdb nexthops,
so a group's legs can share a remote IP and differ only in UDP port.
Motivation
The deployment runs an overlay in which each tenant's traffic is terminated by
a "forwarder": a pod that hosts the VXLAN VTEP, decapsulates the tenant's
overlay, and relays it to and from that tenant's workload. Forwarders for many
different tenants are packed onto the same receiver node behind one
mesh-routable underlay IP, and are demultiplexed purely by UDP destination
port. The host does a stateless outer-UDP demux by port; it never terminates
the tunnel:
receiver node -- one mesh-routable underlay IP (NodeIP_A)
+----------------------------------------------------+
| host netns: stateless outer-UDP demux by dst port |
| (host does NOT terminate the tunnel) |
| |
| dst :40000 dst :40001 dst :40002 |
| | | | |
| +-----v----+ +-----v----+ +-----v----+ |
| | pod0 ns | | pod1 ns | | pod2 ns | |
| | vxlan | | vxlan | | vxlan | |
| | VTEP | | VTEP | | VTEP | |
| | decap | | decap | | decap | |
| +----------+ +----------+ +----------+ |
+----------------------------------------------------+
(up to ~10 forwarder pods packed per node)
The packed pods are unrelated: each belongs to a different tenant on its own
VXLAN VNI, so the per-pod UDP port is node-level demux, not an HA construct.
The host, which only demuxes outer UDP, never has to reason about tenancy.
A single forwarder is made highly available by running replicas. The replicas
of one forwarder share a single anycast overlay identity: one inner MAC and IP.
Clients address that one identity, and a sender spreads flows across the live
replicas with an fdb nexthop group. Failover is transparent: a dead replica is
just dropped from the group, with no client re-resolution or route change. The
single identity is deliberate; the endpoint is consumed one layer up as a
single stable address, so giving each replica its own address would push
multi-address handling and health-checking up into that consumer.
Anti-affinity keeps the two replicas of one HA set on different nodes, so a
group's legs land on distinct node IPs. But each leg is still reachable only
at (node IP, that pod's UDP port), so within one group the legs differ in IP
*and* port. A group can already carry a distinct IP per leg, but it takes the
UDP port from the device (a single value), so it cannot send each leg to its
own port. That is the gap this series closes.
Zooming into one forwarder pod, there is nothing for the host to load-balance:
the tunnel terminates on a vxlan device inside the pod's own netns, and the pod
reaches its tenant through a separate NIC:
one forwarder pod -- its own netns, tenant VNI X
+-------------------------------------------------+
| |
| on/off-ramp NIC <--- customer data plane |
| | on-ramp (ingress) / off-ramp (egress) |
| | inner packet |
| vxlan (VTEP) encap / decap for VNI X, |
| | listens on this pod's UDP port |
| | outer VXLAN UDP |
| eth0 (underlay) NodeIP:port |
| | to peer VTEPs over the |
| v mesh underlay |
| |
+-------------------------------------------------+
Existing mechanisms do not fit this shape:
- L3 multipath in the overlay needs each leg to be a distinct routable
nexthop with its own address. Since an HA set is a single anycast address
by design, there are no distinct per-leg addresses to route over; the fdb
nexthop group bridging to that shared MAC is what load-balances.
- Host-side fan-out (XDP / TC / SO_REUSEPORT) assumes a shared host datapath
that is not there. SO_REUSEPORT balances sockets within one netns, but the
receivers are in different netns (in fact different tenants). An XDP/TC
fan-out would require the host to terminate the tunnel and re-dispatch
inner traffic across netns and VNI boundaries, i.e. become a VTEP, which
puts the host into the tenant datapath and largely duplicates what an fdb
nexthop group already does.
- Demuxing on VNI instead of port (one shared 4789 socket, multiple vxlan
devices differing only in VNI, moved into each pod's netns) works when the
co-located pods have different VNIs. It does not help two same-VNI HA sets
on one node: their outer headers are identical, so the host would again
have to terminate the tunnel to tell them apart. It also costs packing
density: with one shared underlay IP, VNI demux allows at most one VTEP per
(VNI, node), so N same-VNI HA sets of two replicas need 2N nodes, whereas a
per-pod port fits them on two nodes with anti-affinity preserved.
This series adds the attribute:
- Patch 1 adds a netlink attribute NHA_DST_PORT (__be16, mirroring NDA_PORT),
stored in struct nh_info and echoed back on dump. It is only accepted
together with NHA_FDB and NHA_GATEWAY. Control-plane only; datapath
behaviour is unchanged.
- Patch 2 wires it into the VXLAN datapath: vxlan_fdb_nh_path_select() sets
rdst->remote_port to the selected leg's port. vxlan_xmit_one() already
prefers rdst->remote_port when non-zero and otherwise falls back to the
device port, so nexthops without a port are unaffected (backward
compatible).
- Patch 3 extends the fdb nexthop selftests.
On the uAPI: this does not add a new datapath concept. A single fdb entry
already carries a per-destination UDP port (NDA_PORT), and vxlan_xmit_one()
already prefers rdst->remote_port when set. NHA_DST_PORT is the nexthop analog
of that existing attribute: control-plane only, no datapath change, and
backward compatible (a leg with no port falls back to the device port as
today). It sits at the nexthop level rather than under NHA_ENCAP because fdb
nexthops do not use the NHA_ENCAP / LWT infrastructure.
Example:
ip nexthop add id 1 via 192.0.2.10 fdb dst_port 4789
ip nexthop add id 2 via 192.0.2.10 fdb dst_port 5789
ip nexthop add id 10 group 1/2 fdb
bridge fdb add 00:11:22:33:44:55 dev vxlan0 nhid 10
Both legs share gateway 192.0.2.10 and differ only in UDP port; the kernel
hashes flows across them.
====================
Link: https://patch.msgid.link/20260724-b4-vxlan-fdb-port-v5-0-cd1c6aeee058@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
This commit is contained in:
@@ -28,6 +28,7 @@ struct nh_config {
|
||||
u8 nh_protocol;
|
||||
u8 nh_blackhole;
|
||||
u8 nh_fdb;
|
||||
__be16 nh_dst_port;
|
||||
u32 nh_flags;
|
||||
|
||||
int nh_ifindex;
|
||||
@@ -63,6 +64,7 @@ struct nh_info {
|
||||
u8 family;
|
||||
bool reject_nh;
|
||||
bool fdb_nh;
|
||||
__be16 dst_port;
|
||||
|
||||
union {
|
||||
struct fib_nh_common fib_nhc;
|
||||
@@ -574,7 +576,8 @@ struct fib_nh_common *nexthop_fdb_nhc(struct nexthop *nh)
|
||||
}
|
||||
|
||||
static inline struct fib_nh_common *nexthop_path_fdb_result(struct nexthop *nh,
|
||||
int hash)
|
||||
int hash,
|
||||
__be16 *dst_port)
|
||||
{
|
||||
struct nh_info *nhi;
|
||||
struct nexthop *nhp;
|
||||
@@ -583,6 +586,7 @@ static inline struct fib_nh_common *nexthop_path_fdb_result(struct nexthop *nh,
|
||||
if (unlikely(!nhp))
|
||||
return NULL;
|
||||
nhi = rcu_dereference(nhp->nh_info);
|
||||
*dst_port = nhi->dst_port;
|
||||
return &nhi->fib_nhc;
|
||||
}
|
||||
#endif
|
||||
|
||||
@@ -567,8 +567,9 @@ static inline bool vxlan_fdb_nh_path_select(struct nexthop *nh,
|
||||
struct vxlan_rdst *rdst)
|
||||
{
|
||||
struct fib_nh_common *nhc;
|
||||
__be16 dst_port = 0;
|
||||
|
||||
nhc = nexthop_path_fdb_result(nh, hash >> 1);
|
||||
nhc = nexthop_path_fdb_result(nh, hash >> 1, &dst_port);
|
||||
if (unlikely(!nhc))
|
||||
return false;
|
||||
|
||||
@@ -583,6 +584,8 @@ static inline bool vxlan_fdb_nh_path_select(struct nexthop *nh,
|
||||
break;
|
||||
}
|
||||
|
||||
rdst->remote_port = dst_port;
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
|
||||
@@ -83,6 +83,9 @@ enum {
|
||||
/* u32; read-only; whether any driver collects HW stats */
|
||||
NHA_HW_STATS_USED,
|
||||
|
||||
/* be16; UDP destination port for an fdb nexthop (e.g. VXLAN) */
|
||||
NHA_DST_PORT,
|
||||
|
||||
__NHA_MAX,
|
||||
};
|
||||
|
||||
|
||||
@@ -39,6 +39,7 @@ static const struct nla_policy rtm_nh_policy_new[] = {
|
||||
[NHA_ENCAP_TYPE] = { .type = NLA_U16 },
|
||||
[NHA_ENCAP] = { .type = NLA_NESTED },
|
||||
[NHA_FDB] = { .type = NLA_FLAG },
|
||||
[NHA_DST_PORT] = NLA_POLICY_MIN(NLA_BE16, 1),
|
||||
[NHA_RES_GROUP] = { .type = NLA_NESTED },
|
||||
[NHA_HW_STATS_ENABLE] = NLA_POLICY_MAX(NLA_U32, true),
|
||||
};
|
||||
@@ -956,6 +957,9 @@ static int nh_fill_node(struct sk_buff *skb, struct nexthop *nh,
|
||||
} else if (nhi->fdb_nh) {
|
||||
if (nla_put_flag(skb, NHA_FDB))
|
||||
goto nla_put_failure;
|
||||
if (nhi->dst_port &&
|
||||
nla_put_be16(skb, NHA_DST_PORT, nhi->dst_port))
|
||||
goto nla_put_failure;
|
||||
} else {
|
||||
const struct net_device *dev;
|
||||
|
||||
@@ -1055,6 +1059,9 @@ static size_t nh_nlmsg_size_single(struct nexthop *nh)
|
||||
break;
|
||||
}
|
||||
|
||||
if (nhi->dst_port)
|
||||
sz += nla_total_size(2); /* NHA_DST_PORT */
|
||||
|
||||
if (nhi->fib_nhc.nhc_lwtstate) {
|
||||
sz += lwtunnel_get_encap_size(nhi->fib_nhc.nhc_lwtstate);
|
||||
sz += nla_total_size(2); /* NHA_ENCAP_TYPE */
|
||||
@@ -2965,8 +2972,10 @@ static struct nexthop *nexthop_create(struct net *net, struct nh_config *cfg,
|
||||
nhi->family = cfg->nh_family;
|
||||
nhi->fib_nhc.nhc_scope = RT_SCOPE_LINK;
|
||||
|
||||
if (cfg->nh_fdb)
|
||||
if (cfg->nh_fdb) {
|
||||
nhi->fdb_nh = 1;
|
||||
nhi->dst_port = cfg->nh_dst_port;
|
||||
}
|
||||
|
||||
if (cfg->nh_blackhole) {
|
||||
nhi->reject_nh = 1;
|
||||
@@ -3156,6 +3165,15 @@ static int rtm_to_nh_config(struct net *net, struct sk_buff *skb,
|
||||
cfg->nh_fdb = nla_get_flag(tb[NHA_FDB]);
|
||||
}
|
||||
|
||||
if (tb[NHA_DST_PORT]) {
|
||||
if (!tb[NHA_FDB] || !tb[NHA_GATEWAY]) {
|
||||
NL_SET_ERR_MSG(extack,
|
||||
"Destination port can only be set on fdb nexthops that have a gateway");
|
||||
goto out;
|
||||
}
|
||||
cfg->nh_dst_port = nla_get_be16(tb[NHA_DST_PORT]);
|
||||
}
|
||||
|
||||
if (tb[NHA_GROUP]) {
|
||||
if (nhm->nh_family != AF_UNSPEC) {
|
||||
NL_SET_ERR_MSG(extack, "Invalid family for group");
|
||||
|
||||
@@ -30,6 +30,7 @@ IPV4_TESTS="
|
||||
ipv4_large_res_grp
|
||||
ipv4_compat_mode
|
||||
ipv4_fdb_grp_fcnal
|
||||
ipv4_fdb_port_fcnal
|
||||
ipv4_mpath_select
|
||||
ipv4_torture
|
||||
ipv4_res_torture
|
||||
@@ -44,6 +45,7 @@ IPV6_TESTS="
|
||||
ipv6_large_res_grp
|
||||
ipv6_compat_mode
|
||||
ipv6_fdb_grp_fcnal
|
||||
ipv6_fdb_port_fcnal
|
||||
ipv6_mpath_select
|
||||
ipv6_torture
|
||||
ipv6_res_torture
|
||||
@@ -432,6 +434,15 @@ check_nexthop_fdb_support()
|
||||
fi
|
||||
}
|
||||
|
||||
check_nexthop_fdb_port_support()
|
||||
{
|
||||
$IP nexthop help 2>&1 | grep -q "dst_port"
|
||||
if [ $? -ne 0 ]; then
|
||||
echo "SKIP: iproute2 too old, missing nexthop dst_port support"
|
||||
return $ksft_skip
|
||||
fi
|
||||
}
|
||||
|
||||
check_nexthop_res_support()
|
||||
{
|
||||
$IP nexthop help 2>&1 | grep -q resilient
|
||||
@@ -541,6 +552,42 @@ ipv6_fdb_grp_fcnal()
|
||||
$IP link del dev vx10
|
||||
}
|
||||
|
||||
ipv6_fdb_port_fcnal()
|
||||
{
|
||||
echo
|
||||
echo "IPv6 fdb nexthop dst_port functional"
|
||||
echo "------------------------------------"
|
||||
|
||||
check_nexthop_fdb_port_support
|
||||
if [ $? -eq $ksft_skip ]; then
|
||||
return $ksft_skip
|
||||
fi
|
||||
|
||||
# NHA_DST_PORT: optional per-nexthop VXLAN destination UDP port,
|
||||
# letting an fdb nexthop group balance a flow across legs that share
|
||||
# an underlay IP but listen on different UDP ports.
|
||||
run_cmd "$IP nexthop add id 80 via 2001:db8:91::2 fdb dst_port 4790"
|
||||
check_nexthop "id 80" \
|
||||
"id 80 via 2001:db8:91::2 scope link fdb dst_port 4790"
|
||||
log_test $? 0 "Fdb nexthop with dst_port"
|
||||
|
||||
run_cmd "$IP nexthop add id 81 fdb dst_port 4790"
|
||||
log_test $? 2 "Fdb nexthop with dst_port but no gateway"
|
||||
|
||||
run_cmd "$IP nexthop add id 81 via 2001:db8:91::2 fdb dst_port 0"
|
||||
log_test $? 2 "Fdb nexthop with dst_port 0"
|
||||
|
||||
run_cmd "$IP nexthop add id 82 via 2001:db8:91::2 fdb dst_port 4789"
|
||||
run_cmd "$IP nexthop add id 83 via 2001:db8:91::3 fdb dst_port 5789"
|
||||
run_cmd "$IP nexthop add id 106 group 82/83 fdb"
|
||||
check_nexthop "id 106" "id 106 group 82/83 fdb"
|
||||
log_test $? 0 "Fdb nexthop group with legs differing in dst_port"
|
||||
|
||||
run_cmd "$IP nexthop add id 84 via 2001:db8:91::2 fdb"
|
||||
check_nexthop "id 84" "id 84 via 2001:db8:91::2 scope link fdb"
|
||||
log_test $? 0 "Fdb nexthop without dst_port omits dst_port"
|
||||
}
|
||||
|
||||
ipv4_fdb_grp_fcnal()
|
||||
{
|
||||
local rc
|
||||
@@ -641,6 +688,42 @@ ipv4_fdb_grp_fcnal()
|
||||
$IP link del dev vx10
|
||||
}
|
||||
|
||||
ipv4_fdb_port_fcnal()
|
||||
{
|
||||
echo
|
||||
echo "IPv4 fdb nexthop dst_port functional"
|
||||
echo "------------------------------------"
|
||||
|
||||
check_nexthop_fdb_port_support
|
||||
if [ $? -eq $ksft_skip ]; then
|
||||
return $ksft_skip
|
||||
fi
|
||||
|
||||
# NHA_DST_PORT: optional per-nexthop VXLAN destination UDP port,
|
||||
# letting an fdb nexthop group balance a flow across legs that share
|
||||
# an underlay IP but listen on different UDP ports.
|
||||
run_cmd "$IP nexthop add id 30 via 172.16.1.2 fdb dst_port 4790"
|
||||
check_nexthop "id 30" \
|
||||
"id 30 via 172.16.1.2 scope link fdb dst_port 4790"
|
||||
log_test $? 0 "Fdb nexthop with dst_port"
|
||||
|
||||
run_cmd "$IP nexthop add id 31 fdb dst_port 4790"
|
||||
log_test $? 2 "Fdb nexthop with dst_port but no gateway"
|
||||
|
||||
run_cmd "$IP nexthop add id 31 via 172.16.1.2 fdb dst_port 0"
|
||||
log_test $? 2 "Fdb nexthop with dst_port 0"
|
||||
|
||||
run_cmd "$IP nexthop add id 32 via 172.16.1.2 fdb dst_port 4789"
|
||||
run_cmd "$IP nexthop add id 33 via 172.16.1.3 fdb dst_port 5789"
|
||||
run_cmd "$IP nexthop add id 105 group 32/33 fdb"
|
||||
check_nexthop "id 105" "id 105 group 32/33 fdb"
|
||||
log_test $? 0 "Fdb nexthop group with legs differing in dst_port"
|
||||
|
||||
run_cmd "$IP nexthop add id 34 via 172.16.1.2 fdb"
|
||||
check_nexthop "id 34" "id 34 via 172.16.1.2 scope link fdb"
|
||||
log_test $? 0 "Fdb nexthop without dst_port omits dst_port"
|
||||
}
|
||||
|
||||
ipv4_mpath_select()
|
||||
{
|
||||
local rc dev match h addr
|
||||
|
||||
@@ -56,6 +56,17 @@ tc_stats_get()
|
||||
tc_rule_handle_stats_get "dev dummy1 egress" 101 ".packets" "-n $ns1"
|
||||
}
|
||||
|
||||
nh_stats_get_port()
|
||||
{
|
||||
ip -n "$ns1" -s -j nexthop show id 20 | \
|
||||
jq ".[][\"group_stats\"][][\"packets\"]"
|
||||
}
|
||||
|
||||
tc_stats_get_port()
|
||||
{
|
||||
tc_rule_handle_stats_get "dev dummy1 egress" 102 ".packets" "-n $ns1"
|
||||
}
|
||||
|
||||
basic_tx_common()
|
||||
{
|
||||
local af_str=$1; shift
|
||||
@@ -90,6 +101,31 @@ basic_tx_common()
|
||||
busywait "$BUSYWAIT_TIMEOUT" until_counter_is "== 1" tc_stats_get > /dev/null
|
||||
check_err $? "tc filter stats did not increase"
|
||||
|
||||
# Add a second FDB nexthop group whose nexthop carries a per-nexthop
|
||||
# destination port (NHA_DST_PORT) that differs from the VXLAN device
|
||||
# default. Matching outer traffic must egress with that port, so a
|
||||
# separate flower filter keyed on the new port catches it.
|
||||
run_cmd "tc -n $ns1 filter add dev dummy1 egress proto $proto \
|
||||
pref 1 handle 102 flower ip_proto udp dst_ip $remote_addr \
|
||||
dst_port 4790 action pass"
|
||||
|
||||
run_cmd "ip -n $ns1 nexthop add id 2 via $remote_addr fdb dst_port 4790"
|
||||
run_cmd "ip -n $ns1 nexthop add id 20 group 2 fdb"
|
||||
|
||||
run_cmd "bridge -n $ns1 fdb add 00:11:22:33:44:66 dev vx0 \
|
||||
self static nhid 20"
|
||||
|
||||
run_cmd "ip netns exec $ns1 mausezahn vx0 -a own \
|
||||
-b 00:11:22:33:44:66 -c 1 -q"
|
||||
|
||||
busywait "$BUSYWAIT_TIMEOUT" until_counter_is "== 1" \
|
||||
nh_stats_get_port > /dev/null
|
||||
check_err $? "FDB nexthop group stats did not increase (with port)"
|
||||
|
||||
busywait "$BUSYWAIT_TIMEOUT" until_counter_is "== 1" \
|
||||
tc_stats_get_port > /dev/null
|
||||
check_err $? "tc filter stats did not increase (with port)"
|
||||
|
||||
log_test "VXLAN FDB nexthop: $af_str basic Tx"
|
||||
}
|
||||
|
||||
@@ -210,8 +246,8 @@ require_command arping
|
||||
require_command ndisc6
|
||||
require_command jq
|
||||
|
||||
if ! ip nexthop help 2>&1 | grep -q "stats"; then
|
||||
echo "SKIP: iproute2 ip too old, missing nexthop stats support"
|
||||
if ! ip nexthop help 2>&1 | grep -q "dst_port"; then
|
||||
echo "SKIP: iproute2 ip too old, missing nexthop dst_port support"
|
||||
exit "$ksft_skip"
|
||||
fi
|
||||
|
||||
|
||||
Reference in New Issue
Block a user