Lab 01 β OSPF underlay + iBGP-EVPN (full mesh)¶
Complete, self-contained guide. Build a working VXLAN-EVPN fabric from bare vJunos switches, one layer at a time. Read the Study track first for the theory; this lab is the hands-on part.
β Validated end-to-end on vJunos-switch 23.2R1.14.
This is the foundational lab. It uses the simplest overlay β a full mesh between the two leaves β so you can see EVPN in its clearest form. (The production version swaps that for spine route-reflectors.)
What you'll build¶
| Layer | Choice |
|---|---|
| Underlay | OSPF, single area 0 |
| Overlay | iBGP-EVPN, AS 65000, leaf-to-leaf full mesh (spines carry no EVPN) |
| Services | one L2VNI: VLAN 100 β VNI 10100, two hosts in one subnet |
graph TB
S1["spine1<br/>lo0 10.0.0.11"]
S2["spine2<br/>lo0 10.0.0.12"]
L1["leaf1 Β· VTEP<br/>lo0 10.0.0.21"]
L2["leaf2 Β· VTEP<br/>lo0 10.0.0.22"]
H1["host1<br/>10.100.10.10"]
H2["host2<br/>10.100.10.11"]
S1 ---|"10.10.1.0/31"| L1
S1 ---|"10.10.2.0/31"| L2
S2 ---|"10.10.3.0/31"| L1
S2 ---|"10.10.4.0/31"| L2
L1 ---|"VLAN 100"| H1
L2 ---|"VLAN 100"| H2
classDef spine fill:#1565c0,stroke:#90caf9,color:#ffffff,stroke-width:2px,font-weight:bold;
classDef leaf fill:#2e7d32,stroke:#a5d6a7,color:#ffffff,stroke-width:2px,font-weight:bold;
classDef host fill:#e65100,stroke:#ffb74d,color:#ffffff,stroke-width:2px,font-weight:bold;
class S1,S2 spine; class L1,L2 leaf; class H1,H2 host;
Addresses (full plan in common/ipplan.md):
| Device | lo0 (router-id / VTEP) | to spine1 | to spine2 |
|---|---|---|---|
| spine1 | 10.0.0.11 | β | β |
| spine2 | 10.0.0.12 | β | β |
| leaf1 | 10.0.0.21 | 10.10.1.1/31 | 10.10.3.1/31 |
| leaf2 | 10.0.0.22 | 10.10.2.1/31 | 10.10.4.1/31 |
Interfaces: vJunos-switch uses
ge-0/0/N. containerlabeth1βge-0/0/0,eth2βge-0/0/1,eth3βge-0/0/2(a +1 offset). Login:admin/admin@123.
Before you start¶
- Host set up (GCP + containerlab + vJunos image) β see Host Setup.
- This lab runs on its own fabric (
clab-evpn-lab-*).
β οΈ Pre-flight β only ONE lab at a time. A 2Γ2 vJunos fabric needs ~16 GB RAM;
two at once starve the host and boot unhealthy. Before deploying, check
nothing else is running:
docker ps --format '{{.Names}}' | grep '^clab-' || echo "clean β nothing running"
sudo docker rm -f $(docker ps -aq --filter name=clab-) # force-remove all clab containers
deploy.sh and reset.sh also refuse to start if another fabric is up, so
you can't hit this by accident β but checking first is good habit.)
How to run it¶
./scripts/deploy.sh 01-ospf-ibgp # boot the fabric (~5-8 min/node)
# then EITHER build it all at once:
./scripts/apply.sh 01-ospf-ibgp all
# OR learn by hand β type each step below yourself, or one step at a time:
./scripts/apply.sh 01-ospf-ibgp 02 # e.g. just Step 2
Check the fabric is ready (vJunos takes ~5β8 min/node to boot):
docker ps --filter "name=clab-evpn-lab" --format "table {{.Names}}\t{{.Status}}"
Up β¦ (health: starting) | still booting β wait |
| Up β¦ (healthy) | β
ready β safe to apply.sh |
Wait until all four switches read (healthy) (the two hosts just show Up).
Watch it live: watch -n 5 'docker ps --filter "name=clab-evpn-lab" --format "table {{.Names}}\t{{.Status}}"'.
Or just log in to confirm: ssh admin@clab-evpn-lab-leaf1 (password admin@123).
apply.shwaits for each node's CLI on its own, so you can run it right after deploy β it holds until nodes are ready (up to ~2 min/node).
Wipe or redo:
./scripts/destroy.sh 01-ospf-ibgp # wipe (no redeploy)
./scripts/reset.sh 01-ospf-ibgp # wipe + redeploy clean
To do it by hand: ssh admin@clab-evpn-lab-leaf1 (password admin@123),
configure, paste the step's block, commit.
π Command cheat-sheet (copy-paste)¶
Everything runs from the clab host (~/netforge-labs) unless marked Junos CLI.
Reusable helper β paste once into your shell, then check any node without logging in:
jrun() { sshpass -p 'admin@123' ssh -o StrictHostKeyChecking=no \
-o UserKnownHostsFile=/dev/null -o LogLevel=ERROR \
admin@clab-evpn-lab-"$1" "${@:2}"; }
# usage: jrun leaf1 "show bgp summary"
Deploy & check the fabric
./scripts/deploy.sh 01-ospf-ibgp # boot the fabric (~5-8 min/node)
docker ps --filter name=clab-evpn-lab \
--format "table {{.Names}}\t{{.Status}}" # all 4 switches must read (healthy)
Build the config
./scripts/apply.sh 01-ospf-ibgp all # β build every step 01β05 in order (recommended)
./scripts/apply.sh 01-ospf-ibgp 01 # a single step (only if 01..N-1 are already applied)
./scripts/apply.sh 01-ospf-ibgp 01-03 # a range of steps, in order
Iterate without a reboot β wipe config to baseline and rebuild (seconds, not minutes):
./scripts/clean.sh 01-ospf-ibgp # wipe lab config, keep mgmt/SSH β NO reboot
./scripts/apply.sh 01-ospf-ibgp all # then rebuild from scratch
Switch to another design WITHOUT a reboot β labs 01β04 share this 2Γ2 topology
and the same container name (clab-evpn-lab-*), so one booted fabric runs any
of them. Just clean and apply the other lab's config β no override, no new deploy:
./scripts/clean.sh 02-ospf-ibgp-rr # wipe config (~30s)
./scripts/apply.sh 02-ospf-ibgp-rr all # apply lab 02's RR design in place
Verify from the host (no login needed, via the jrun helper):
jrun leaf1 "show ospf neighbor" # underlay: both spines Full
jrun leaf1 "show bgp summary" # overlay: peer 10.0.0.22 Establ
jrun leaf1 "show route table bgp.evpn.0" # EVPN routes (Type-3 then Type-2)
jrun leaf1 "show ethernet-switching vxlan-tunnel-end-point remote" # tunnel up
Hosts + final proof (host commands run on the clab host, not Junos):
docker exec clab-evpn-lab-host1 sh -c "ip addr add 10.100.10.10/24 dev eth1; ip link set eth1 up"
docker exec clab-evpn-lab-host2 sh -c "ip addr add 10.100.10.11/24 dev eth1; ip link set eth1 up"
docker exec clab-evpn-lab-host1 ping -c3 10.100.10.11 # π 0% loss = done
Tear down
./scripts/destroy.sh 01-ospf-ibgp # wipe containers (no redeploy)
./scripts/reset.sh 01-ospf-ibgp # destroy + redeploy clean (slow β last resort)
β οΈ Steps are cumulative β build bottom-up. Each step depends on the one below it (
01 fabric β 02 OSPF β 03 iBGP β 04 L2VNI β 05 access). On a freshly-cleaned fabric you must apply 01 through N in order β jumping straight to a later step commits config onto an empty box that silently can't work. When unsure, just runapply.sh 01-ospf-ibgp all.
The build¶
Do the steps in order. Each one follows the same rhythm β Apply β Verify β β DONE β and you must pass the check before moving to the next.
lo0 reachable (ping) β BGP Establ β Type-3 + tunnel β Type-2 β host ping
Step 2 Step 3 Step 4/5 Step 5 Step 5
Step 1 β Fabric: interfaces & loopbacks¶
Why: every switch needs its fabric-link /31 IPs and a /32 loopback. On a
leaf, lo0 is the router-id, the BGP peering address, and the VXLAN tunnel
source β the single most important address on the box.
spine1
set interfaces ge-0/0/0 unit 0 family inet address 10.10.1.0/31
set interfaces ge-0/0/1 unit 0 family inet address 10.10.2.0/31
set interfaces lo0 unit 0 family inet address 10.0.0.11/32
set routing-options router-id 10.0.0.11
set interfaces ge-0/0/0 unit 0 family inet address 10.10.3.0/31
set interfaces ge-0/0/1 unit 0 family inet address 10.10.4.0/31
set interfaces lo0 unit 0 family inet address 10.0.0.12/32
set routing-options router-id 10.0.0.12
set interfaces ge-0/0/0 unit 0 family inet address 10.10.1.1/31
set interfaces ge-0/0/1 unit 0 family inet address 10.10.3.1/31
set interfaces lo0 unit 0 family inet address 10.0.0.21/32
set routing-options router-id 10.0.0.21
set interfaces ge-0/0/0 unit 0 family inet address 10.10.2.1/31
set interfaces ge-0/0/1 unit 0 family inet address 10.10.4.1/31
set interfaces lo0 unit 0 family inet address 10.0.0.22/32
set routing-options router-id 10.0.0.22
./scripts/apply.sh 01-ospf-ibgp 01 Β· or paste the four blocks above by hand.
Verify (from the host):
jrun leaf1 "show interfaces terse | match ge-" # fabric links admin/link up
jrun leaf1 "ping 10.10.1.0 count 3" # directly-connected /31 replies
Step 1 β DONE β
Fabric links up, loopbacks present, /31 ping replies. β Step 2
Step 2 β Underlay: OSPF¶
Why: the underlay's one job is to make every loopback reachable from every other, over both spines (ECMP). Loopbacks are advertised passive (announced, but no neighbour to form there); fabric links are point-to-point.
Identical on all four switches:
set protocols ospf area 0 interface lo0.0 passive
set protocols ospf area 0 interface ge-0/0/0.0 interface-type p2p
set protocols ospf area 0 interface ge-0/0/1.0 interface-type p2p
./scripts/apply.sh 01-ospf-ibgp 02 Β· the same three lines on all four switches.
Verify (from the host):
jrun leaf1 "show ospf neighbor" # both spines in state Full
jrun leaf1 "ping 10.0.0.22 source 10.0.0.21 count 3" # leaf-to-leaf loopback, ttl=63 (one spine hop)
Step 2 β DONE β
Loopback-to-loopback ping works. β Step 3. If it fails, stop β nothing above works without the underlay.
Step 3 β Overlay: iBGP-EVPN (full mesh)¶
Why: the overlay is a BGP session carrying the evpn family, so leaves learn
each other's hosts without flooding. In full mesh the leaves peer directly with
each other (one session for two leaves); the spines run no EVPN.
leaf1
set routing-options autonomous-system 65000
set protocols bgp group overlay type internal
set protocols bgp group overlay local-address 10.0.0.21
set protocols bgp group overlay family evpn signaling
set protocols bgp group overlay neighbor 10.0.0.22
local-address 10.0.0.22, neighbor 10.0.0.21.
Apply: ./scripts/apply.sh 01-ospf-ibgp 03 Β· leaf1 above; leaf2 mirrors it.
Verify (from the host):
jrun leaf1 "show bgp summary" # peer 10.0.0.22 = Establ; bgp.evpn.0 present (0 routes is correct β no VXLAN yet)
The
License key missing; requires 'bgp'warning is a benign vJunos-eval message.
Step 3 β DONE β
The EVPN session to the other leaf is Establ. β Step 4
Step 4 β EVPN + VXLAN glue¶
Why: this turns on the VTEP. protocols evpn picks VXLAN + which VNIs;
switch-options sets the tunnel source (lo0.0), the RD (unique per leaf) and RT
(shared per VNI); and a VLANβVNI mapping bridges VLAN 100 onto VNI 10100.
Leaves only β spines are not VTEPs.
leaf1 (leaf2 mirrors, RD 10.0.0.22:1)
set protocols evpn encapsulation vxlan
set protocols evpn extended-vni-list all
set switch-options vtep-source-interface lo0.0
set switch-options route-distinguisher 10.0.0.21:1
set switch-options vrf-target target:65000:1
set vlans v100 vlan-id 100
set vlans v100 vxlan vni 10100
./scripts/apply.sh 01-ospf-ibgp 04 Β· leaf1 above; leaf2 mirrors (RD 10.0.0.22:1).
β οΈ Expect NO routes yet. show route table bgp.evpn.0 is still empty β this is
correct. Junos only advertises a VNI once its VLAN has an up member interface
(unlike Cisco). It lights up in Step 5.
Verify by CONFIG PRESENCE β the EVPN route table does not appear yet:
jrun leaf1 "show configuration protocols evpn" # encapsulation vxlan + extended-vni-list
jrun leaf1 "show configuration vlans" # v100 β vni 10100
β οΈ No
bgp.evpn.0/default-switch.evpn.0table yet β this is correct.show route table ?still lists onlyinet.0/mgmt_junos.*. The EVPN table only materialises in Step 5, once the access port brings VLAN 100 up β the same reason Type-3 waits for an up member. Verify Step 4 by config, not routes.
Step 4 β DONE β
protocols evpn + switch-options + VLAN 100 β VNI 10100 are committed.
(No EVPN route table yet β that's expected; it appears in Step 5.) β Step 5
Step 5 β Services: attach hosts & prove it¶
Why: put the host ports into VLAN 100. The moment the port is up, the leaf advertises its Type-3 (IMET) route, the VXLAN tunnel forms, and once hosts talk, Type-2 (MAC/IP) routes teach both leaves where each host is.
5a β access ports, leaf1 and leaf2 (same):
set interfaces ge-0/0/2 unit 0 family ethernet-switching interface-mode access
set interfaces ge-0/0/2 unit 0 family ethernet-switching vlan members v100
./scripts/apply.sh 01-ospf-ibgp 05 Β· or paste the two lines above on both leaves.
Verify 5a (from the host):
jrun leaf1 "show route table bgp.evpn.0" # two Type-3 (3:) routes appear
jrun leaf1 "show ethernet-switching vxlan-tunnel-end-point remote" # tunnel to the other leaf
5b β give the hosts their IPs (clab host shell, not Junos):
docker exec clab-evpn-lab-host1 sh -c "ip addr add 10.100.10.10/24 dev eth1; ip link set eth1 up"
docker exec clab-evpn-lab-host2 sh -c "ip addr add 10.100.10.11/24 dev eth1; ip link set eth1 up"
docker exec clab-evpn-lab-host1 ping -c3 10.100.10.11
Step 5 β DONE β Β· the finish line π
host1 β host2 ping returns 0% packet loss across the VXLAN fabric.
Once traffic flows, jrun leaf1 "show route table bgp.evpn.0" also shows the Type-2 (2:) MAC/IP routes.
β Full checklist β deploy to ping¶
Work top to bottom. Each item lists the command that proves it β don't move on
until it passes. (jrun is the helper from the cheat-sheet above.)
Fabric up
- [ ] All 4 switches (healthy) β docker ps --filter name=clab-evpn-lab --format "table {{.Names}}\t{{.Status}}"
Step 1 Β· interfaces & loopbacks
- [ ] Fabric links up/up β jrun leaf1 "show interfaces terse | match ge-"
- [ ] Directly-connected /31 replies β jrun leaf1 "ping 10.10.1.0 count 3"
Step 2 Β· OSPF underlay
- [ ] Both spines Full β jrun leaf1 "show ospf neighbor"
- [ ] Leaf-to-leaf loopback ping, ttl=63 β jrun leaf1 "ping 10.0.0.22 source 10.0.0.21 count 3"
Step 3 Β· iBGP-EVPN overlay
- [ ] Peer 10.0.0.22 Establ β jrun leaf1 "show bgp summary" (0 routes here is correct)
Step 4 Β· EVPN + VXLAN glue (verify by config β no EVPN route table until Step 5)
- [ ] protocols evpn present β jrun leaf1 "show configuration protocols evpn"
- [ ] VLAN β VNI present β jrun leaf1 "show configuration vlans" (no bgp.evpn.0 yet β appears in Step 5)
Step 5 Β· services + proof
- [ ] Two Type-3 (3:) routes β jrun leaf1 "show route table bgp.evpn.0"
- [ ] Tunnel to other leaf β jrun leaf1 "show ethernet-switching vxlan-tunnel-end-point remote"
- [ ] Remote host MAC via vtep.xxxx β jrun leaf1 "show ethernet-switching table"
- [ ] host1 β host2 ping, 0% loss β docker exec clab-evpn-lab-host1 ping -c3 10.100.10.11 π
Break-it exercises¶
Predict the symptom, break it, find the show that exposes it, then fix it.
- Underlay link:
deactivate interfaces ge-0/0/0on leaf1 β loopback stays reachable via the other spine (show route 10.0.0.22). Reactivate. - BGP source: point
local-addressat the wrong IP β session never Establishes (show bgp summary). Restore. - VNI mismatch: set leaf2's VLAN 100 to
vni 10199β tunnel/host ping breaks (show evpn database). Restore to 10100. - VTEP source:
delete switch-options vtep-source-interfaceon leaf1 β Type-3 withdrawn, tunnel drops. Restorelo0.0.
Lessons from the live build¶
- Interfaces are
ge-0/0/N; clabethNβge-0/0/(N-1). family inetcommits clean onge-ports (noethernet-switchingto delete).- β Junos originates Type-3 only when the VLAN has an up member β biggest difference vs Cisco; the tunnel appears at Step 5, not Step 4.
- Management (
fxp0) is on10.0.0.0/24, overlapping the loopbacks but isolated in themgmt_junosinstance β harmless. - "OSPF instance is not running" right after commit is just timing β wait ~30 s.
Troubleshooting¶
| Symptom | Cause | Fix |
|---|---|---|
apply.sh says container not found on every node |
Wrong lab folder β the running fabric is evpn-lab (lab 01), you ran a different lab's name |
Use 01-ospf-ibgp (it matches clab-evpn-lab-*) |
A step committed but nothing works |
Steps applied out of order β config landed on an empty fabric | Rebuild bottom-up: apply.sh 01-ospf-ibgp all |
clean.sh shows syntax error, expecting <identifier> on one node |
A command got garbled over that node's slow pty (harmless if the node was already clean) | Re-run clean.sh, or confirm the node is baseline: jrun spine1 "show configuration \| display set \| match routing-instances" β only mgmt_junos lines = already clean |
Commits take 30β90 s or apply reports FAILED |
Contention while all 4 nodes boot/converge at once | Wait until the fabric is idle (top β low load, 0.0 st steal), then re-run; iterate with clean.sh, not reset.sh |
gcloud: SERVFAIL / command not found |
You typed gcloud inside the VM or a device CLI |
Run gcloud only from your laptop or Cloud Shell, never inside the VM |
| Lost SSH after a config wipe | An old wholesale delete removed the mgmt instance |
Already fixed in clean.sh (it deletes only named lab hierarchies); if stuck, recover via console/telnet |
show bgp summary warns License key missing |
Benign vJunos-eval message | Ignore |
A node exits with FileNotFoundError: init.conf |
It was docker restarted β clab nodes lose init.conf on restart |
Never docker restart a clab node. Recover with sudo containerlab deploy --reconfigure -t <topo> (rebuilds the fabric β clab 0.77 can't recreate a single node) |
Golden rule for iterating: boot the fabric once with deploy.sh, then loop
clean.sh β apply.sh. Never docker restart a node (it wipes init.conf and
the node exits), and you don't need a fresh fabric per lab β labs 01β04 share this
topology, so run any design on the same nodes with FABRIC=<prefix> (see the
cheat-sheet). Reach for reset.sh / --reconfigure only when a node is truly dead.