Latest Posts

  • Where Guest Network Pro Runs Out

    The topology post ended on a promise. I said Guest Network Pro on the RT-BE88U would let me finally segment the IoT and media gear off the main network, and that having hardware which supported it natively removed my last excuse for not doing it.

    It removed about seventy percent of the excuse. This is where the other thirty percent lives.

    network diagram of vlans in my home network

    From looking at the SNB Forums, it was obvious that it wasn’t the easiest solution to use, especially when compared to the Ubiquiti love fest on r/homelab and r/HomeNetworking. While I look on enviously at the single pane management interface across the entire network stack, I can’t bring myself to abandon the off-the-shelf prosumer segment I call home. Making shit that isn’t supposed to work together integrate is always more fun!

    What Guest Network Pro actually does

    Credit first. Guest Network Pro genuinely works. You create a profile, give it a VLAN ID, and the router hands you an SSID that drops its clients into that VLAN with its own subnet and DHCP scope. Four profiles, four SSIDs, four isolated networks. For a wireless-only setup it is the whole job done in a web UI.

    It also emits properly tagged 802.1Q frames on ordinary LAN ports, which I did not expect and spent an embarrassing amount of time failing to prove. More on that later.

    Limitation 1: three port modes, and none of them is the one you need

    The router has a VLAN Switch Control page giving each LAN port one of three modes:

    • All, the default
    • Access, which binds the port to a single profile, untagged
    • Trunk, which carries multiple tagged VLANs

    Trunk sounds like what you want until you read the fine print: “non-tagged will be dropped.” My house is structured cabling. One link runs from the router to the head-end switch, and every room hangs off it. Setting that port to Trunk cuts the untagged path for every device in the house at once. That is not a configuration change, it is an outage, and that was one lesson I had the misfortune of learning the hard way. It’s no fun when your wife’s K-drama is hanging and she starts screaming at you!

    The port I actually needed was one carrying untagged VLAN 1 plus several tagged VLANs simultaneously. That is a hybrid port, and the BE88U cannot express it. Not a bug, just not a feature it has.

    The device that needs it most is the AiMesh node in the master bedroom. It serves every SSID the main router serves, including the guest profiles, so its wired backhaul has to carry ordinary untagged traffic and the guest VLANs at the same time. Trunk mode kills its backhaul. Access mode kills its guest SSIDs. There is no third option on the router.

    Limitation 2: a DHCP setting you will not find by looking

    Guest Network Pro will not hand a DHCP lease to a client that arrives with a VLAN tag, unless you have found and enabled a per-profile toggle called “Enable the DHCP Server”.

    I recorded this as a hard platform limitation and designed around it. I planned an external DHCP server. I wrote it into my notes as a constraint on the whole architecture. It was a checkbox.

    What makes it nasty is that the failure is completely silent. The client sends DISCOVERs, the router ignores them, and nothing anywhere reports a problem. Ping works. ARP works. The device just never gets an address, and you go looking at your switch config.

    Limitation 3: no ACLs, so isolation is all-or-nothing

    Each profile has an “Access Intranet” toggle. Off means the guest network reaches the internet and nothing else. On means it reaches everything.

    There is no middle. You cannot say “the AV VLAN may reach the media server on port 8096 and nothing else.” So the moment you need one service to cross a boundary, the router cannot help you, and the crossing has to happen somewhere else.

    In my case that somewhere is the Proxmox layer. Anything that must live in two VLANs gets a second tagged virtual NIC:

    • Pi-hole, twice over, so every VLAN gets filtered DNS instead of falling back to the router
    • Jellyfin, which keeps its read-only NAS mounts on the trusted side and serves media into the AV VLAN. The NAS itself never appears there
    • Home Assistant, which needs to reach IoT devices on one VLAN and media devices on another

    Each one is a single command and a static address with no gateway. The pattern is the useful part: the router does isolation, the hypervisor does the deliberate holes.

    The surprise: All mode was already doing what I wanted

    Here is the part that cost me the most time and turned out to be the least complicated.

    I spent two evenings convinced that a port left on All emits no guest-VLAN traffic, and that Trunk was mandatory. I had test results backing it. I built an entire isolated test rig to prove it, carried a laptop across the house, and cabled it directly to a spare router port so the unmanaged switch could not muddy the result.

    All mode carries tagged and untagged traffic in both directions, including frames the router originates. It always did. My earlier tests failed because at the time no profile with that VLAN ID existed, so there was nothing to emit. I had invented a mechanism to explain a null result, written it into my notes next to genuine vendor documentation, and then believed it.

    The practical upshot is good: the router-to-switch uplink stays on All, so the whole-house-outage hazard of Trunk mode (i.e. the K-drama blow up) never has to be touched at all.

    What the managed switch actually bought me

    A D-Link DXS-F108T, eight 10GBase-T ports, replacing the unmanaged 10G switch at the head end. It does three things the router cannot:

    1. Real hybrid ports. The study run carries untagged VLAN 1 plus four tagged VLANs, so the hypervisors can tag their own virtual NICs. The mesh node gets untagged VLAN 1 plus the guest VLANs. Neither is expressible on the router.
    2. Untagging onto access ports. The media console switch hangs off a port set to untagged VLAN 30. Every TV and AV receiver behind it sees plain ethernet, has no idea a VLAN exists, and gets DHCP normally. Dumb devices stay dumb.
    3. Loop protection that suits the topology. RSTP only catches loops where both ends land on the managed switch. The loop I actually caused once, months ago, was a bridge misconfiguration on a hypervisor behind four cascaded unmanaged switches, invisible to spanning tree. Loopback detection catches that, because unmanaged switches obligingly flood the probe frame straight back.

    Does AiMesh tag, or tunnel?

    This was the last unknown, and it decided whether the mesh node needed that hybrid port or would have been fine on a plain access port.

    The test: with tags present, connect an iPad to the AV SSID through that node specifically, and load Jellyfin. Then strip the tagged VLANs from the port and try again. Then put them back.

    Loads, fails, loads. AiMesh tags guest SSID traffic over the wired backhaul. It does not tunnel it. The managed switch was not an upgrade, it was a prerequisite.

    The third step matters more than it looks. A single failure could have been the iPad dropping off for its own reasons. Watching the outcome track the variable in both directions is what makes it a result instead of an anecdote.

    What broke, because something always does

    Moving the media console to its own VLAN broke things quietly:

    • Kodi vanished from Home Assistant. The integration stores a hard-coded IP. The device changed subnet, the integration kept looking at the old address, and the entity simply went unavailable. Nothing raised an error
    • Kodi also lost its media. Its NAS shares live on the trusted side and are now unreachable by design. It gets repointed at Jellyfin, which is exactly what the media gateway is for
    • Plex is stranded. It runs on the NAS, which is deliberately not in the AV VLAN. Anything that wants Plex needs to be on the trusted side, or Plex needs to move. Fortunately, i Have Remote Access enabled, so its still my primary media service
    • Casting is gone. mDNS does not cross VLANs and the router will not reflect it. I do not cast from my phone, so I have accepted it. If you do, this is the single biggest cost of segmenting AV gear

    If you are planning something similar, audit your automation platform for hard-coded device IPs before you move anything. That is the failure mode that will find you weeks later.

    What I would tell myself

    Four times in this project I recorded a confident conclusion that turned out to be wrong, and not one of them was a hardware fault. Every single one was a test that measured something other than what I thought it measured.

    A DHCP probe that returned a firm failure had never actually transmitted a packet. A negative result got explained with an invented mechanism that then hardened into a documented fact. A script I wrote reported PASS while its own packet capture showed zero packets, because I (ok ok, it was the LLM) had grepped for a string that also appears inside the failure message.

    The fix was not better tools. It was putting a control in every test. Prove the link works before blaming the service running over it. Include something that must fail, so a false positive is visible. I intentionally ran a test to an known good VLAN and an uncreated VLAN during every test. Reach for the cheapest probe that isolates the layer you are actually asking about. And when a negative result seems to confirm what you already believed, that is exactly the moment to check whether your probe ran at all.

    The network is segmented now. The router does the VLANs, the switch decides which ports get them tagged and which get them untagged, and the hypervisor punches the deliberate holes. Guest Network Pro got me most of the way there. It just could not do the one port that mattered most. Now comes the arduous task of moving moving the different devices to their forever-home VLANs.

  • The Layer the Cable Map Doesn’t Show

    I’ve already drawn the cables. This is where the packets actually go.

    Back in May I published the full topology of the house: every switch, every 10G run, every zone. That was the easy half. Cables tell you what’s physically possible. They tell you nothing about where a DNS query goes, which door a request came through, or why one machine’s traffic leaves the country and another’s doesn’t. If you go back far enough in this blog, you would have noticed that I’ve been using the Asus RT routers since the original N56U (in 2011!). While they can do most of the basic functions (and some of the advanced ones with the Merlin firmware), I’ve taken advantage of my homelab to add on a few more bells n whistles.

    its always DNS

    The Router Does Three Separate Jobs

    Everything starts at an Asus RT-BE88U running Merlin firmware. Stock firmware would route fine. I run Merlin for three features that turned out to be the backbone of everything downstream.

    DHCP with reservations. Boring, until it isn’t. My reverse proxy references backend IPs directly, so a container that gets a new lease silently breaks a service while the proxy keeps happily forwarding to an address nobody’s at. I once had four services sitting proxied-but-unreserved for four days without noticing. Reserve the address before you point anything at it.

    DNS Director. This force-redirects port 53 for most client on the LAN to my Pi-holes, so a device that hardcodes 8.8.8.8 gets intercepted and filtered anyway. It works completely, which is the problem. My Pi-holes are LAN clients too. For a while their own upstream queries were being redirected back at the router, so the recursive resolver behind them never reached the root servers at all. Recursion was an illusion. Any host running its own resolver needs to be set to “No Redirection” — that one exemption is the whole trick.

    VPN Director. Policy-based routing: named clients leave through a commercial VPN, everything else takes the native connection. My desktop is on the list; the Proxmox nodes aren’t. The catch nobody mentions is that it’s IPv4 only. Checked from my desktop while writing this: IPv4 exits at a CDN77 address, IPv6 exits on my actual ISP prefix. Anything that resolves AAAA first walks straight past the VPN with my real address attached. I spent a while treating that as a bug before accepting it as the design.


    DNS: Pi-hole in Front, Unbound Behind

    Two Pi-holes, both containers on the Proxmox cluster. One is the source of truth, the other a replica, and nebula-sync copies one to the other every fifteen minutes. Two of them because DNS failing takes the whole house with it.

    Behind each Pi-hole sits unbound, doing full recursion out to the root servers. No Cloudflare, no Quad9, nobody upstream keeping a log of what I look up. Pi-hole blocks and forwards; unbound resolves.

    Getting that right cost me about 3,975 SERVFAILs. Unbound validates DNSSEC itself. I had also ticked DNSSEC on in Pi-hole, so every signed lookup got validated twice and the second validator kept rejecting the first one’s work. The fix was one checkbox. Leave Pi-hole’s DNSSEC off when unbound is upstream — unbound is the validator, and it still rejects bogus signatures without any help.


    Caddy: One Certificate, Thirty-Three Names

    Everything internal answers on a single wildcard domain, served by a Caddy container on one of the cluster nodes. One Let’s Encrypt certificate, issued over DNS-01 against Cloudflare, which means nothing has to be exposed to the internet to get it. Renewal happens entirely through a DNS record.

    Thirty-three hostnames as of this week. Each is a Pi-hole record pointing at Caddy’s address plus a five-line block in one config file. Adding a service takes two minutes and no new certificate. I tried to get NGINX set up 3 times and failed miserably, so I got my LLM to set up and manage Caddy for me. It worked at the first try!

    The one I got wrong: for eleven days my router’s admin interface was proxied on that domain. The highest-value target on the network, protected by nothing but the router’s own login, reachable from anything that could resolve LAN DNS. I had added it, never written it down, and only found it during an audit. I deleted the entry rather than putting an authentication layer in front of it. Removing the exposure was cheaper than defending it.


    Tailscale: The Private Overlay

    Tailscale is how I get in from outside — laptop, phone, ipad, wherever. The interesting part is DNS, because MagicDNS and Pi-hole both want to be the resolver.

    The tailnet uses split DNS: queries for my domain go to both Pi-holes, everything else resolves normally. Those records answer with LAN addresses, not Tailscale addresses, so devices already on the LAN keep working unchanged. Remote devices reach the same addresses through a subnet router advertising the whole /24.

    That subnet router is currently one node. If it goes down, remote access to the LAN goes with it. I know. A second one is on the list.

    The trap here is subtle and cost me an afternoon. 100.100.100.100 is a forwarder, not a resolver. Tailscale answers tailnet names and hands everything else to whatever the system’s real nameservers were at the moment it took over. One node had its interfaces bounced during a NIC migration, tailscaled restarted while the config was momentarily empty, and it captured nothing as its upstream. Every public name on that host failed silently for eleven hours. All three nodes had byte-identical resolver config. The state that mattered lived in a backup file that none of the obvious commands show you.


    Cloudflare Tunnel: The Public Door

    External access runs on a different domain entirely, through a Cloudflare tunnel. No inbound ports open anywhere. Most hostnames sit behind Cloudflare Access; Plex is deliberately bare, because it is built to face the internet and enforces its own login (with 2FA, naturally).

    The two domains differ by a single hyphen. That is a genuinely bad idea I am now stuck with, and I have written down which is which in three separate places. (The one with the hypen is external!)

    The tunnel originally ran one connector, on the NAS, which was also the origin for one of the services it fronted. A DSM update took down the door and one of the rooms behind it, orphaning three perfectly healthy services that had nothing to do with the NAS. There are two connectors now, on separate boxes, serving the same tunnel. I tested it by stopping the first: twelve of twelve requests still served.

    Getting that wrong is easy. A second connector has to use the existing tunnel’s token. Clicking “create a tunnel” gives you two tunnels and zero redundancy.


    The Pattern

    Five systems, and most of the faults above were the same shape: two layers both trying to do one job. Two DNSSEC validators. A router intercepting the resolver it was pointing at. A door mounted on the room it opened onto. Draw the logical layer, not just the cables. You cannot spot a duplicate when you are only looking at one diagram.

    None of them announced themselves. They presented as “the internet is slow”, “flatpak is broken”, “apt can’t reach the repos” — never as the thing that was actually wrong. Each fix was a single setting. Each one took days to find. Just remember the wise advice of the interwebs: if you have a problem, its always DNS.

    Next up for me is figuring out VLANS. My new 10G Managed switch (DXS-F108T) just arrived today, so stay tuned!

  • I Built a Dashboard for the Things I Actually Check

    Five separate logins became one page. The interesting part was how often I turned out to be wrong.

    I check five things regularly: whether the homelab is healthy, my portfolio, my health trend, whatever trip is coming up, and my to-do list. Each one lived somewhere different. Each had its own login. None of them talked to each other. I have to go looking for information by opening a browser, going to my bookmarks page and finding the right tile to click:

    My Heimdall Dashboard of links to my self-hosted services.

    Information you have to go and fetch is information you check less often than you should.


    What It Looks Like Now

    One page. Five modules, each running in its own container, pulled together by a dashboard shell called Glance. The shell is configuration, not a product — it reads each service’s API and renders the numbers. That’s the whole trick (admittedly greatly helped by LLMs to design):

    The architecture for my data dashboard

    The health module was the biggest piece by some distance. I exported four years of Fitbit history and imported it alongside live Garmin data, so the trend line runs continuously from March 2022 to today instead of restarting when I changed watches. About 7 million data points. There’s a real gap of six days in January where neither device recorded anything, which is the week I switched. I’ve left it in. It’s honest.

    The finance module pulls positions from my broker and reconciles them against manually entered accounts for the things a broker doesn’t know about — Provident Fund, cash, the rest. Everything now sits behind a single login.


    Three of My Four API Guesses Were Wrong

    This is the part worth writing down.

    I planned the whole thing on paper first (with the help of AI, of course, including this post), including which endpoint each module would call and which fields it would read. When I came to build it, three of the four were wrong.

    The monitoring tool returned a flat structure where I’d assumed a nested one (I needed the LLM to explain this to me too!). The finance app had moved its performance endpoint to a new API version, and the old path returned a 404. The travel app used different field names to what I was expecting.

    None of these would have crashed anything. Each would have quietly rendered an error tile on a dashboard I’d then have trusted. The fix was boring and it worked: read a live response from every endpoint before writing a single line of template code.


    The Service I Decided Not to Build

    The last gap was authentication. Several services were reachable on my network with no login at all, and the obvious answer was to put single sign-on in front of everything.

    So I drafted the whole build. Container, secrets, rollout order, backups. Then I actually tested my assumptions and found that most of the services I’d listed as unprotected were, in fact, protected. The real list was less than half of what I’d written down (i.e. in my issue backlog in my Obsidian Vault). A lot of the services were things I wanted open so I could use them on the fly on an ipad or my phone.

    I closed the gap with tools already running instead. I stopped proxying my router’s admin interface entirely rather than putting a login in front of it, and turned off open registration on one app that had it. That took an afternoon and added nothing new to maintain.

    A new service is a permanent cost. Backups, updates, secrets to store, another thing that breaks at 2am. Since I already stayed up late last night to figure out why Jellyfin was transcoding when no one was using it, I figured “Simplex sigillum veri“.


    What I’d Tell Someone Starting

    Verify before you build. Every assumption I wrote in my .md vault was cheap to check and expensive to get wrong.

    Back up the databases properly. A snapshot of a running database is a snapshot of a database mid-write. Every module here dumps itself fifteen minutes before the nightly backup, and each dump refuses to complete unless it passes an integrity check. A backup you haven’t verified is a hope.

    And scope the problem before you scope the solution. I nearly built an identity provider to solve a problem that turned out to be four services and a router.