DHCP lease churn in a Proxmox homelab: when Authelia down x87 was a lost IP
On 2 September my self-healing monitor logged DOWN: Authelia 87 times in a row. Authelia, my SSO portal, was never down. Its own address answered 200 the whole time.
What had gone was an IP address. This post covers that incident and the others that turned out to be the same bug: a smart bulb that took a server's address, a cloned container sharing a MAC, and a re-IP that switched off my tracing.
"It's always been on that address" is not "it's pinned"
Both Proxmox nodes rebooted at about 14:16 that afternoon. The container that runs my reverse proxy came back and DHCP gave it a different address.
Everything pointed at its usual address: AdGuard rewrites, Cloudflare records, proxy configs, monitors. But there was no DHCP reservation for it on the router. The address was only ever a sticky lease: dnsmasq tends to hand a returning client its old address, until the day it doesn't. The pool had 150 addresses on 12-hour leases.
That container runs my only Caddy instance, and six hostnames resolve to it, including the SSO portal. One lost lease took out the front door and all SSO while every backend stayed healthy. My monitors probe by hostname, so they blamed the wrong service.
The lesson I wrote down: when a monitor says a service on another box is down, probe that service's own address and port before touching it.
Making it static, and the lease that wouldn't leave
I set a static address in the container's Proxmox network config with pct set ... -net0 ...,ip=<addr>/24,gw=<gw>. That rewrites the container's systemd-networkd file to a fixed address with DHCP off, and hot-adds the static address. But the old dynamic address and its proto dhcp default route survived a systemctl restart systemd-networkd. I had to delete both by hand with ip addr del and ip route del. Until then the container held both addresses.
After that, SSO answered again and the healer went quiet.
The bulb that took a server's address
That evening I swept every service to check it was really back. My password manager's hostname returned 502, while its container reported itself healthy. Caddy's log showed its upstream dial failing with connect: connection refused.
The fault was at layer 2. The password manager's container was set to a static address, but that address sat inside the DHCP pool. So the router had leased it to something else, and the MAC answering for it began with an OUI registered to WiZ, the smart bulb maker. Ping times to it were around 15 ms, which is wireless, not a container on the same host.
The bulb won the ARP race across the whole network, even on the Proxmox host. A static IP inside a DHCP pool is not a static IP. The router doesn't know it's taken and will hand it to the next device that asks.
The fix was to move the container to a static address below the start of the pool, where the router can never lease it. Only one live reference needed changing: the reverse_proxy line in the Caddyfile, then systemctl reload caddy.
The third one, and a false green
Same night, a third service: my 3D printer monitoring app returned 502 while the app was healthy. My NPMplus reverse proxy on another box still had its upstream set to the container's pre-reboot address. I repointed it and reloaded nginx inside the container.
The monitoring was worse. Uptime Kuma had two monitors for that service, both on the old address:
- the HTTP check was red with
ECONNREFUSED: the right alarm for the wrong reason; - the ping check was green, because an unrelated household device had inherited the old lease and was answering pings.
A stale IP in a ping monitor doesn't go red. It goes green on whatever takes the address. Meanwhile my healer, which only reads Kuma, had started escalating to rebooting a container that was never unhealthy. Only its limit of three reboots an hour held it back.
Two more lessons came out of this one:
-
A raw-IP probe could never have worked. The app is Django using the sites framework, which picks the site by the request's
Hostheader. Probed by IP and port, the login page returned 500. With the rightHostheader it returned 200. The old monitor had only ever passed because a site record existed for the old IP and port. I pointed the monitor at the public URL instead, the only probe that would have caught the real user-facing fault in the proxy. -
A 3xx is not proof of health. I called that service "restored" off a 302. The 302 was the redirect to the login page, which was returning 500. Kuma followed the redirect, which is how it caught what my
curlmissed.
Kuma's history showed the old monitor green until 08:49 that morning, before the 14:16 reboot. That container's lease had turned over mid-morning, not at the reboot at all.
The tally for one unreserved address, over that night and the next day: two container IPs, a Caddy upstream, two proxy upstream fixes, five references in source files and two monitors, all changed by hand. One of those source files would have quietly recreated the broken raw-IP monitor on its next run.
Reservations and static config protect different things
On 3 September I added DHCP reservations on the router for all ten containers, then checked each still held its expected address and every front-door hostname answered.
This is the part I had misunderstood:
- Static config in the container stops the container asking for a new address.
- A reservation on the router stops the router giving that address to someone else.
The bulb incident was the missing second half. The rule I now follow: static hosts live below the pool, DHCP keeps the pool, and anything that other config names by address gets a reservation.
An older version of the same bug: two containers, one MAC
In August I'd already met the extreme case. On 5 August I migrated my agent container to a newer Proxmox node for more RAM, and never stopped the original. Both were set to start on boot. They were byte-identical clones: same MAC, so the same DHCP lease and the same IP, the same hostname, and the same NetBird identity.
The switch learns a MAC wherever it last saw a frame, so replies went to whichever twin had spoken last. That explained the "flaky networking" I had been chasing:
- SSH banner stalls in roughly 20–30% of connections;
- Uptime Kuma's login acknowledgement never arriving;
- NetBird reconnecting 11,898 times in 24 hours on that one peer, against 2–3 for every other peer, because two clients held one identity and kept evicting each other.
ARP looked clean, because both twins answered with the same MAC. I first called it a layer-2 loop, which was wrong. A paired tcpdump on both ends finally showed a reset from our address that our container had never sent. That was the twin.
I stopped the old container on 9 August. NetBird's reconnect log for that peer dropped from about 24 per three minutes to 1.
The re-IP that quietly switched off tracing
The latest instance was this week. My Langfuse container came back from a node reboot on 22 September on a static address that someone or something had set in its config. It was inside the DHCP pool too, and different from its reservation. Uptime Kuma went red with EHOSTUNREACH for 2,382 heartbeats. Langfuse itself was healthy throughout, but my agents' tracing, pointed at the reserved address, went dark.
On 23 September I moved the container back to its reserved address rather than repointing everything else, because the reservation is the source of truth and a static address inside the pool is the bulb trap again.
That wasn't the end. On 22 September every specialist agent profile's config had been repointed to the wrong address, and none of those files went back when the container did. So every specialist agent stayed untraced until I found it on 24 September, repointed them, restarted the gateway and confirmed new traces arriving.
An IP change has to update every config that names the address, including the per-profile copies, and the monitors.
What I check now
- Is this address reserved on the router, or just sticky? "It's always been there" proves nothing.
- Is any static address inside the DHCP pool? Move it below the pool.
- After any IP change: grep configs on every host, including proxy configs inside containers on other boxes, per-profile config files and every monitor.
- Prefer hostname probes that follow redirects over pings to an IP.
- After cloning or migrating a container, make sure the original is actually stopped.
The longer version of this, with the DNS, DHCP and tunnel traps side by side and a checklist, is a short field report: Home Network Traps.
🤖 Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.