154 lines
24 KiB
Markdown
154 lines
24 KiB
Markdown
# Homelab Network Overhaul — Status & Decisions
|
||
|
||
Living document. Update checkboxes and tables as things change instead of re-deriving this from scratch each time.
|
||
|
||
Location: Linz, Austria. ISP: LIWEST (cable). Modem: Technicolor CGA4233-EU.
|
||
|
||
## Current TODO
|
||
|
||
See [todo.md](todo.md) for the live checklist. Quick status (as of 2026-09-26): bridge mode, controller adoption, firmware update, VLANs, ACLs, WireGuard VPN + INWX DynDNS, the SG2016P switch, all three SSIDs, NAS on Servers with QuickConnect disabled, and NPM/Pi-hole local naming are all done. In progress: the first Hyper Backup of the NAS. Next: second AP + living-room switch (purchase decided), office AP ceiling mount + Wi-Fi re-measure, partner's Synology app logins, hardware shopping list. Everything else is waiting on the NUC.
|
||
|
||
## Hardware decisions
|
||
|
||
### Network core
|
||
|
||
| Item | Model | Status |
|
||
|---|---|---|
|
||
| Router | TP-Link Omada ER605 v2.30, firmware 2.4.5 | Received, adopted, updated |
|
||
| Main switch | TP-Link Omada SG2016P (16-port, 8×PoE+, 120W budget) | Received, adopted, wired in; powers the office AP via PoE (make sure the AP is in an actual **PoE** port — only 8 of the 16 are) |
|
||
| Access point #1 | TP-Link Omada EAP650 | Received — mounting in the **office** first, powered via its own DC adapter until the switch's PoE+ is available |
|
||
| Access point #2 | TP-Link Omada EAP650 (same model as AP #1) | **Decided to buy** — chosen over the EAP610 (AX1800, no 160 MHz on 5 GHz) so both APs are identical; the peak-speed difference is irrelevant to the actual couch problem, which is signal quality. Living-room/dining-area coverage, mounted above the TV (see Wi-Fi section) |
|
||
| Living-room switch | TP-Link Omada **ES205GP** (5-port, 4× PoE+ 802.3at/af, 65W budget) | **Decided to buy** — note the **P**: the plain ES205G in earlier notes has no PoE. Easy Managed tier, but supports 802.1Q VLANs (up to 32 groups) via Omada Controller, adopted like the other devices. 3 devices need ports (AP, TV, streaming Pi), so 5 ports is plenty — no need for an 8-port unit |
|
||
|
||
### Homeserver
|
||
|
||
- **Used Intel NUC** (8th-gen Core i3, 16GB DDR4, 128GB SSD), bought via Willhaben for ~€155 — replaced the original new-Beelink-EQ14 plan.
|
||
- OS: plain Debian/Ubuntu + Docker Compose. **No Proxmox/hypervisor.**
|
||
- Home Assistant via the official **"Home Assistant Container"** image (not Home Assistant OS) — matches running everything else as plain Docker.
|
||
- Config management: **Ansible**, run from David's own machine over SSH (no footprint on the server itself).
|
||
- **Kubernetes explicitly rejected for home services** (no benefit on a single node). Will build a separate **k3s learning cluster** later on already-owned Raspberry Pis (RPi4 and older), kept fully separate from production services — motivated by wanting K8s reps for work, not by home-service needs.
|
||
|
||
### Living room
|
||
|
||
- Single existing Ethernet cable feeds this area from the wall. **Plan: run that connection to a spot above the TV**, and put the new ES205GP switch + second AP (EAP650) there, so the TV and streaming Pi are each only a short patch cable away. (Alternative considered and rejected: switch + AP at the wall outlet above the cabinet, which would need two ~5 m runs to the TV and Pi instead of one longer run.)
|
||
- Connected there: the second AP (EAP650, **PoE-powered** from the ES205GP — no separate adapter), the TV (wired, on **IoT**), and the Raspberry Pi for game streaming (on **Trusted**).
|
||
- VLAN plumbing this needs: the ES205GP uplink port and the AP's port are **trunks** carrying Trusted (10) / IoT (20) / Guest (40) tagged; TV port = access port on IoT; Pi port = access port on Trusted. The SG2016P port feeding this cable must be a matching trunk (remember the lesson from the first AP: every VLAN needs to be tagged on **every** hop between the AP and the ER605, or clients associate but never get an IP).
|
||
- Streaming Steam games to the TV: try the already-owned **RPi4 with Moonlight** first (well-proven, hardware-decode, zero cost) before buying anything. Fallback if insufficient: new RPi5 (~€150) or a cheap N100 mini PC (similar price, more flexible, slightly higher latency/idle power).
|
||
- Streaming Pi should sit on the **Trusted VLAN** (same as the gaming PC) to avoid an unnecessary routing hop.
|
||
|
||
### Full wired device list (needs a switch port)
|
||
|
||
Linux PC · Windows gaming PC · partner's PC · printer (optional) · Synology NAS · homeserver (NUC) · OctoPrint Pi · AP (via living-room switch) · streaming Pi (via living-room switch)
|
||
|
||
## Network architecture
|
||
|
||
### VLAN plan
|
||
|
||
| VLAN | Subnet | Purpose |
|
||
|---|---|---|
|
||
| 10 — Trusted | `192.168.10.0/24` | Personal devices, PCs, streaming Pi (avoids an extra routing hop for game streaming) |
|
||
| 20 — IoT / Home Automation | `192.168.20.0/24` | OctoPrint + webcam, sensors, light bulbs, camera, robot vacuum, TV — internet allowed (needed for TV streaming apps + vacuum's cloud control), LAN access denied by ACL |
|
||
| 30 — Servers / Homelab | `192.168.30.0/24` | NAS, homeserver (NUC) |
|
||
| 40 — Guest | `192.168.40.0/24` | Isolated (Omada "Isolate Network" flag), internet-only |
|
||
| 99 — Management | `192.168.99.0/24` | Switch/AP/controller control plane |
|
||
|
||
- Window/door/temperature sensors are **Zigbee via a USB coordinator on the homeserver**, not IP devices — they never touch VLAN 20's network; Home Assistant talks to them over the local Zigbee radio directly.
|
||
- **Default-deny between VLANs is not automatic on Omada** — inter-VLAN routing is allow-all by default; ACL rules must be built manually (e.g., Home Assistant → IoT devices).
|
||
- ACL direction for alerts/control is **Servers → IoT** (Home Assistant reaching in to poll/control devices), not the reverse — IoT devices never need to initiate connections out to notify anyone.
|
||
- TV casting only needs **Trusted → IoT** allowed (the casting source connects out to the TV); return traffic flows back automatically as part of that established connection, no reverse rule needed. Discovery is handled by the ER605's mDNS repeater.
|
||
- **ACL rules built** (Omada ACL types: `Network` = exact match, `!Network` = any network except the one picked; Direction `LAN-LAN` vs `LAN-WAN` vs `WAN-IN` are separate scopes — a `LAN-LAN` rule has zero effect on internet access):
|
||
- `Network: IoT` → `!Network: IoT`, Direction `LAN-LAN`, **Deny** — one rule blocks IoT from reaching every other VLAN (Trusted/Servers/Management/Guest) at once; internet access for IoT is untouched since WAN isn't in scope for a `LAN-LAN` rule.
|
||
- `Network: Servers` → `Network: Management`, Direction `LAN-LAN`, **Deny**.
|
||
- `Network: Guest` → `Network: Management`, Direction `LAN-LAN`, **Deny** (defense-in-depth alongside Guest's Isolate Network flag).
|
||
- Everything else (Trusted↔Servers, Trusted→Management, Servers→IoT, Trusted→IoT, all internet access) relies on Omada's allow-all-by-default behavior — no rule needed.
|
||
- **WAN-IN rules generally aren't needed** — NAT already blocks all unsolicited inbound traffic by default; WAN-IN only matters for scoping something you've already port-forwarded (e.g. the VPN port).
|
||
- ER605 supports full per-port 802.1Q tagging/untagging/PVID; no practical VLAN-count limit for this setup.
|
||
- ER605's built-in **mDNS repeater** bridges Bonjour/mDNS discovery across VLANs (relevant for reaching OctoPrint/HA by local hostname).
|
||
- ER605's **LAN DNS** feature (firmware 2.3.0+) allows manual local hostname → IP entries — but like all DNS, it cannot encode a port number. Non-standard-port services (OctoPrint:5000, HA:8123, Omada Controller:8043, Spoolman:7912, Portainer:9000, etc.) still need a reverse proxy for clean no-port URLs.
|
||
|
||
### Wi-Fi
|
||
|
||
- Router is headless by design (lives in a cupboard — its own Wi-Fi radio would be useless there anyway).
|
||
- First AP (EAP650) is in the **office**. Measured signal (interim position, AP on top of a closet — not yet at its final ceiling mount): office desk **-42 dBm**, office at the 3D printer **-39**, master bedroom **-42**, kitchen **-57**, storage room (door closed) **-59**, living-room couch **-70** (matches a measured drop from ~50 to 10-20 Mbit there). Coverage is good everywhere except the couch nook, which sits behind the bathroom and hallway walls relative to the office. Rough thresholds: better than -60 good, to about -67 reliable for streaming/calls, -70 marginal, worse than -75 weak.
|
||
- **Decision: add a second AP (EAP650) in the living/dining area, mounted above the TV** — a single weak spot alone wouldn't have justified it (phone use at -70 is fine, and the TV/streaming Pi are wired), but David wants to keep working from the couch or dining table on a laptop, where -70 dBm is genuinely marginal. Wired backhaul, so no mesh needed.
|
||
- A ping to an idle phone is **not** a valid coverage test — phones in Wi-Fi power-save routinely show 100-500 ms round trips with a strong signal. Use the AP's own reading (Controller → Clients → the device's signal/link rate), a screen-on ping, or a throughput test instead.
|
||
- Worth re-measuring the couch once the office AP is at its real ceiling position and again after the second AP is up, for a before/after comparison.
|
||
|
||
### VPN
|
||
|
||
- **WireGuard runs on the ER605 itself**, not on the homeserver — avoids extra port-forwarding + static routing that a NUC-hosted VPN would need. **Working as of firmware 2.4.5.**
|
||
- **ER605 firmware must be at least 2.4.5 for WireGuard VPN Server to appear as an option in the Controller at all** — on the pre-update firmware (2.3.3), WireGuard wasn't listed, and the OpenVPN server that *was* available silently never actually started (config looked correct in the Controller — enabled, user linked, no errors in Audit Logs — but the daemon never bound to its port, confirmed via `ECONNREFUSED` on local, WAN-direct, and external tests). Updating firmware fixed both.
|
||
- **Local Networks scoped to Servers (30) + IoT (20)** — remote access to the NAS/self-hosted services and things like OctoPrint, not to personal devices on Trusted. (Was briefly set to Trusted-only, then discovered to actually have *every* network selected — same "field silently left on all networks" class of bug hit earlier with OpenVPN's Local Networks. Corrected to the minimal actual-need scope: Servers + IoT, explicitly excluding Management for now — no current need for it, easy to widen later.) Tunnel Mode **Split** (only the selected networks' traffic goes through the tunnel, not general browsing).
|
||
- VPN client IP pool: `192.168.66.0/24` — deliberately outside all VLAN subnets, no fixed convention requirement (doesn't have to be a `10.x` range).
|
||
- Needs its own port-forward rule (Settings → Transmission → NAT → Port Forwarding) separate from any VLAN/ACL config: WAN1, external+internal port `51820`, internal IP `192.168.0.1` (the gateway's own Default-network address, since the VPN terminates on the gateway itself), protocol UDP or All — **not TCP only** (a real bug hit during OpenVPN testing: the port-forward was accidentally set to TCP-only while OpenVPN used UDP, causing a silent timeout that looked identical to "nothing is running" until diagnosed via matching local vs. external `ECONNREFUSED` signals).
|
||
- No additional WAN-IN ACL rule was needed on top of the port-forward — confirmed via the external test returning an active `ECONNREFUSED` rather than a silent timeout, which proves traffic was already reaching the gateway.
|
||
- Android: WireGuard/OpenVPN aren't part of Android's built-in VPN framework (that only covers IKEv2/IPsec and L2TP natively) — use the official **WireGuard** app (supports QR-code import — zoom the browser in if the phone camera won't scan it off a monitor) or **OpenVPN Connect** app.
|
||
- **DynDNS instead of paying LIWEST for a static IP** — real annual cost difference (~€200+/year for static IP vs. free DynDNS), and cable IPs don't change often in practice anyway. **Decided against DuckDNS** — David already owns a domain via **INWX**, which has its own DynDNS2-compatible service, so used that instead (no shared `duckdns.org` domain, and INWX's DomRobot API has a ready-made `acme.sh` plugin for future Let's Encrypt DNS-01 certs on NPM).
|
||
- Set up an INWX **DynDNS account** (their control panel's dedicated DynDNS section, not general domain/DNS management) — this single step both creates the DNS record for the chosen hostname and generates a **separate DynDNS-scoped username/password** (not the main INWX account login — narrower blast radius if it ever leaked).
|
||
- Configured on the ER605 via **Device Config → Gateway → DNS → Dynamic DNS**, Service Provider **Custom**:
|
||
- Update-URL: `https://[USERNAME]:[PASSWORD]@dyndns.inwx.com/nic/update?hostname=[DOMAIN]&myip=[IP]` (Omada's Custom DDNS placeholders — substituted from the Username/Password/Domain Name fields, not typed literally into the URL)
|
||
- The actual filled-in URL/credentials are **not** written anywhere in this repo — keep them out of anything committed.
|
||
- **Working** — verified via `dig +short <hostname>` returning the ER605's current public IP.
|
||
- WireGuard client `Endpoint` doesn't auto-update to a hostname from the Controller UI (it always embeds the raw WAN IP in exported configs) — WireGuard's config format supports hostnames in `Endpoint` natively though, so just manually edit the line (in the downloaded `.conf` or directly in the WireGuard app's peer settings) to `Endpoint = <hostname>:51820` after generating a client config.
|
||
- **General WireGuard gotcha, same shape as the Endpoint one above**: fields set on the *server* config (e.g. Primary DNS Server) do **not** retroactively apply to peers whose config was already generated/imported before the change — WireGuard bakes settings into the client's config at generation time, it isn't pushed live the way DHCP options are. After changing a server-side field, either regenerate and re-import the client config, or edit the existing peer directly in the WireGuard app (same trick as the Endpoint fix).
|
||
- Bridge mode confirmed active on the modem — ER605 WAN gets the public IP directly (see TODO above).
|
||
|
||
### Local DNS & naming
|
||
|
||
- Naming scheme settled on: **`<name>.home.staffenberger.at`** — a subdomain of David's own already-owned INWX domain, not `.local` (reserved for mDNS/Bonjour, RFC 6762 — using it for manual DNS records risks conflicting with automatic mDNS resolution) and not `.home` or `.home.arpa` (the latter is the IETF-correct reserved suffix per RFC 8375, but rejected as too unfriendly for the partner to read/type). These records exist only in Pi-hole's local DNS — never published to INWX's real public DNS, so nothing is internet-exposed by using a real owned domain for the naming.
|
||
- Records added so far as individual **Local DNS Records** in Pi-hole (not yet the wildcard approach — see TODO): `pihole.home.staffenberger.at`, `nas.home.staffenberger.at`, both pointing at the NAS's Servers-VLAN IP (`192.168.30.10`).
|
||
- **Gotcha: a secondary/fallback DNS server on a VLAN's DHCP config causes intermittent resolution failures for internal-only names.** Trusted's DHCP had Pi-hole as Primary DNS but `8.8.8.8` (Google public DNS) as Secondary — client OS resolvers don't reliably always-prefer-primary (can race/round-robin), so any query that happened to hit `8.8.8.8` for an internal `.home.staffenberger.at` name failed immediately, since a public resolver has no knowledge of it. Symptom looked like "works for me, not for my partner" / inconsistent across devices. Fix: don't set a secondary DNS pointing outside Pi-hole at all — Pi-hole already forwards normal internet lookups upstream itself.
|
||
- Same **stale-DHCP-lease pattern hit all week applies to DNS server changes too** — a device won't pick up a newly-configured DNS server until it renews its lease (Wi-Fi toggle, `ipconfig /release`+`/renew`, etc.), same as VLAN reassignment.
|
||
|
||
### Omada Controller
|
||
|
||
- Self-hosted **Omada Software Controller** (Docker) instead of TP-Link's cloud controller — avoids a vendor cloud relay, in line with the general "no cloud relay" networking principle.
|
||
- Plan: run it temporarily on David's own Linux desktop now, adopt the ER605 immediately, then migrate to the NUC later via Omada's **Controller Migration** (backup/restore) feature once the NUC is set up.
|
||
- Watch for controller version mismatches between export and import — TP-Link/Omada requires matching major.minor(.patch) versions for restore; fallback is installing a specific matching version first, restoring, then upgrading.
|
||
|
||
### Reverse proxy
|
||
|
||
- **Chosen: Nginx Proxy Manager (NPM)** — single GUI, single source of truth for all proxy configs, works uniformly whether the target is a Docker container on the NUC or a bare-metal service elsewhere (e.g. OctoPrint on the Pi).
|
||
- Traefik considered and rejected for home use (label-based config fights against wanting one central place to look); will get real Traefik exposure anyway via the k3s learning cluster, where it's the default ingress controller.
|
||
|
||
### Switch port configuration (SG2016P)
|
||
|
||
- **Access port** = one VLAN, untagged — for every VLAN-unaware end device (PCs, NAS, printer, Pis, TV). Set the port's Native/Untagged network to the device's VLAN and leave Tagged empty; don't leave Default involved. **Trunk** = several VLANs tagged over one cable — only for infrastructure links (switch↔router, switch↔AP).
|
||
- **Every VLAN must be tagged on every hop** between an AP and the ER605, or clients associate to the SSID but never get an IP ("IP configuration failure"). Hit this twice: the first AP's port, then the switch's uplink to the ER605 (missing VLAN 10). The per-Network "Select Device Port" step and the per-port VLAN config are two views of the same thing; editing the **port** directly avoids the "cannot deselect the interface which selects the LAN as PVID" error.
|
||
- The AP's port native/untagged VLAN and the switch/AP management traffic deliberately stay on **Default** for now (same L2 segment as the Controller on David's desktop) — migrate management to VLAN 99 together with the Controller move to the NUC, at which point Default can be retired.
|
||
- Extra safety net: Guest VLAN also enabled on the currently-used ports so anything unexpectedly plugged in lands isolated.
|
||
- After any port/VLAN or DHCP-scope change, a device keeps its old lease until it renews: toggle the port off/on in the Controller (equivalent to unplugging), `ipconfig /release`+`/renew` on Windows, or Wi-Fi off/on on phones.
|
||
- The `Network` column in the Controller's Clients list is blank for wireless clients — cosmetic; the client's IP range is the real check of which VLAN it landed in.
|
||
- Switch PoE is only on 8 of the 16 ports; a "dead" AP was just plugged into a non-PoE port (and then needed a minute to boot).
|
||
|
||
## Backups
|
||
|
||
- **Omada Controller**: auto-backup enabled, **daily** for now (config is changing a lot; switch to weekly later). Local only on the Controller host for now — backups live in the container's data volume, so `docker compose down -v` deletes them. Off-machine rsync copy to a dedicated, quota-limited Synology user is deferred until the NUC migration (see todo.md). Backups contain secrets (Device Account, WireGuard keys, DDNS settings) — never commit them.
|
||
- **Synology NAS → 6TB USB drive via Hyper Backup**, deliberately **manually connected** (offline drive also guards against ransomware/accidental deletion), cadence **every two weeks**, recurring calendar reminder. Drive arrived as **exFAT** (DSM 7.3 mounts exFAT natively — the "exFAT Access" package no longer appears in Package Center — but exFAT has no journal, so an unclean unplug leaves a dirty flag and DSM reports "not ejected safely"); reformatted to **ext4** via DSM (Control Panel → External Devices → Format — select the device row and Eject first if Format is greyed out). Hyper Backup has no plug-in trigger on DSM 7, so it's "plug in → Back up now → **eject in DSM** → unplug".
|
||
- Wizard order is **destination first, then sources**. Backup type **Multiple versions**, **Folders and Packages** (not LUNs).
|
||
- Sources: all shared folders (incl. `photo`, Drive team folders, `docker`) + `homes` (~200 GB) + all applications (Synology Drive Server and Photos matter most — they carry the version history / albums / metadata that plain file copies lose), config backup on, **client-side encryption on** (password in Enpass), no schedule, Smart Recycle retention.
|
||
- `homes` is a reserved system share: its permissions can't be edited ("homes is reserved for the system"), and only Administrators-group accounts see it in Hyper Backup — neither is a problem.
|
||
- Entire-system backup (restorable NAS image) isn't available to a local USB destination, only to a remote NAS / C2 — restoring means reinstalling DSM and packages and restoring data + config from the backup.
|
||
- **Still to confirm after the first run finishes: a test restore** of a personal Drive file and a personal photo (both users) to prove the application backup really covers personal data. Later: occasional off-site copy (3-2-1).
|
||
- **Planned (post-NUC): power the drive via a Home Assistant smart plug** for a weekly window instead of plugging it in by hand — goal is protection against hardware failure; the slightly weaker ransomware isolation is accepted (off-site copy if that ever matters). Power-off only after a scripted clean eject. Details in todo.md.
|
||
- Hyper Backup vs Hyper Backup Vault: Vault is only the *receiving* side on a second Synology (relevant for a future off-site copy), not needed for a local USB drive.
|
||
|
||
## Cabling & power
|
||
|
||
- **Cupboard patch cables**: flat Cat6 is fine for short (2-3 m or less) **data-only** links — buy pure copper (CCA and very thin conductors are the risk), from a reputable maker (deleyCON flat U/UTP is AWG 27 stranded copper, unshielded, 1.5 mm thick, sold as 5-packs in white and individually in white/black, down to 25 cm; Goobay, InLine, Digitus/Delock are comparable). Unshielded is fine — twisted-pair geometry does the noise rejection. Flat cables are fragile where the cable meets the plug: no pull or sharp bend there. **PoE runs (to the APs) and longer runs (e.g. to above the TV) should be round pure-copper cable, not flat.** Verify each link shows 1000 Mbps in the Controller after swapping; a link that negotiates 100 Mbps or shows errors = swap the cable. (The TV port showed 100 Mbps earlier — probably the TV's own NIC, but unconfirmed.) Measure the real cable path with string and add 20-30 cm before choosing 2 m vs 3 m. Color-by-length coding isn't available in flat cables (only white/black) — use labels/Velcro, and a colored round multi-length set for loose spares.
|
||
- **Socket strips** (Austria: Schuko/Type F): a strip adds sockets, not capacity — plug two strips into *different* wall outlets, no daisy-chaining. Look for VDE/GS/ÖVE mark, 16 A, 3×1.5 mm² cable, surge protection for the PC/desk strip, and **no switch (or a guarded one) for the router/switch/NAS strip**. Mount with **screws** (keyhole slots/mounting plate), sockets facing sideways/down, never adhesive. The workshop-style metal strip David likes is bracket-mounted (Brennenstuhl Premium-Alu-Line is the closest wall-mountable match found; the AJ Produkte MOTION and Manutan ones are for their own workbench frames and only list CE, not VDE/GS).
|
||
|
||
## Home automation plans
|
||
|
||
- Sensors: temperature/humidity, window-open monitoring, light bulb monitoring — via Zigbee-style sensors + a USB coordinator dongle (e.g. Sonoff Zigbee 3.0, ~€15–20) plugged into the homeserver.
|
||
- Camera to check on the cat: casual live viewing is computationally free. If smart detection (Frigate) is wanted later, add a Coral USB accelerator (~€70) rather than upgrading the CPU.
|
||
|
||
## Planned services on the homeserver (NUC)
|
||
|
||
- Home Assistant (Container image)
|
||
- Pi-hole
|
||
- Filament spool manager (Docker)
|
||
- Omada Software Controller (Docker — migrated from the temporary desktop instance)
|
||
- Nginx Proxy Manager
|
||
- Self-written Docker tools (future)
|
||
- All managed via Ansible playbooks
|