Spine-Leaf EVPN-VXLAN Fabric Design
EVPN-VXLAN has become the de facto standard for modern data centers with spine-leaf architectures. It combines BGP EVPN (control plane) and VXLAN (data plane) to overcome the limitations of STP/HSRP. This comprehensive design guide covers topology, symmetric vs asymmetric IRB selection, multi-homing, and multi-site deployments.
Spine-leaf architecture
- Spine: high-density 100G/400G switches (e.g., Cisco Nexus 9300-GX, Juniper QFX5130)
- Leaf: switches with 25G server ports and 100G uplinks (e.g., Nexus 93180YC-FX, QFX5120-48Y)
- Spine-leaf links: full-mesh ECMP
- No spine-to-spine links and no leaf-to-leaf links
- Scale: N spines × M leaves = 1:N oversubscription (e.g., 4 spines × 32 leaves = 8:1 oversubscription)
Underlay routing
eBGP underlay
The most common choice. Each leaf has a unique private AS, while the spines share an AS. eBGP multihop peering runs between loopbacks. It is simple, scales well, and is easy to automate.
OSPF/IS-IS underlay
The traditional alternative. OSPF area 0 spans the entire fabric, with aggressive BFD timers (50/150ms). It is less scalable than eBGP beyond 50 switches.
BGP EVPN overlay
- iBGP EVPN: spines = route reflectors, leaves = RR clients
- eBGP EVPN: possible, but more complex
- The l2vpn evpn address family is enabled
- Route targets (RTs) provide multi-tenant isolation
Symmetric vs Asymmetric IRB
Symmetric IRB (recommended)
Each leaf performs local routing for all subnets. Inter-subnet traffic is forwarded through an L3 transit VNI. It scales better and requires fewer routes per leaf.
Asymmetric IRB
The source leaf routes, while the destination leaf bridges. This requires every VNI on every leaf and does not scale beyond 10-15 VNIs.
Multi-homing (ESI LAG)
- EVPN Ethernet Segment Identifier (ESI) for multi-homed servers
- Type-4 routes for auto-discovery
- Type-1 routes for per-ES auto-discovery
- Per-flow load balancing (hashing) or active/active MAC forwarding
- A more flexible alternative to vPC (Cisco) or MLAG
Multi-site EVPN
For DC1 + DC2 interconnection:
- Border leaves (or border gateways) at each DC
- EVPN multi-site: inter-DC RT import/export
- DCI (Data Center Interconnect): dark fiber, MPLS, Internet IPsec
- Global anycast gateway: out of scope without VPC-LACP (Cisco) or MH-EVPN (Juniper)
Fabric sizing
Small (50 racks)
- 2× QFX5120-32C spine switches
- 10× QFX5120-48Y leaf switches (48×25G server ports)
- Budget: ~€280,000 excluding tax
- Capacity: 480 25G servers
Medium (200 racks)
- 4× QFX5200-32C spine switches
- 40× leaf switches
- Budget: ~€1.2M excluding tax
Large (>500 racks, super-spine)
- 5-stage architecture: super-spine + spine + leaf
- Pod-based design
- Budget: €5-15M excluding tax
Alternatives
- Cisco ACI: intent-based networking + GUI; APIC controllers required
- VMware NSX: hypervisor overlay, not a hardware fabric
- Juniper Apstra: multi-vendor automation for EVPN-VXLAN fabrics
Order from OPTINOC
Turnkey spine-leaf EVPN-VXLAN fabric. Cisco Nexus, Juniper QFX, and Arista. Compatible 100G/400G transceivers offering 50% savings. Quotes for data centers with 20-500 racks within 48 hours.
