Infrastructure/docs/todo.md

103 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Homelab Network Overhaul — TODO
Living checklist. See [status.md](status.md) for the full decisions/context behind each item.
## Switch wiring — in progress
- [x] SG2016P wired in and adopted; AP moved over to it (PoE-powered).
- [x] SSID plan implemented: Trusted (10), Guest (40), IoT (20, hidden) all working on their correct VLAN subnets. Root cause of the initial "Trusted Wi-Fi won't get an IP" bug: the switch's uplink port to the ER605 was missing VLAN 10 from its tagged list — same bug independently blocked the partner's PC from getting onto Trusted directly. Fixed by adding all VLANs to that uplink port's tagged list.
- [x] OctoPrint Pi wired to IoT, confirmed reachable.
- [x] Printer moved to IoT (after fixing its port — was left as untagged-Default with IoT only tagged, which a VLAN-unaware printer could never actually reach; fixed to native/untagged = IoT).
- [x] DHCP reservations done for devices with stable MACs.
- [x] Guest VLAN enabled as a safety net on currently-in-use switch ports, so anything unexpectedly plugged in lands isolated by default.
- [x] **Synology NAS** → Servers (30), DHCP-reserved.
- [x] **Windows gaming PC** and **partner's PC** both on Trusted (the uplink-port VLAN bug above was what had blocked them).
- [x] Homeserver (OptiPlex) on Servers at `192.168.30.20`; David's Linux desktop moved to Trusted (fixed `192.168.10.10`) after the controller left it; switches/APs moved to Management VLAN 99; Default is now empty (2026-10-04).
## Homeserver (OptiPlex) — follow-ups
Set up 2026-10-04, see [homeserver.md](homeserver.md).
**Network**
- [ ] **Switch DHCP DNS** in Omada from `192.168.30.10` (NAS Pi-hole) to `192.168.30.20` on every VLAN that uses Pi-hole, renew leases, then **turn off the NAS Pi-hole**. Keep no secondary DNS outside Pi-hole (see status.md, Local DNS).
- [ ] **Narrow ACL exceptions for DNS** to Pi-hole (`192.168.30.20`, port 53) from IoT and Guest, placed above rules 1 and 2. Replaces the old "should IoT use Pi-hole" question below.
- [ ] Add **Trusted and Default** to the destinations of rule 3 (Server Outwards).
- [ ] **Lock down the Default network**: allow only the controller ports 29810–29817 to the OptiPlex, block internet.
- [ ] Optional: restrict Management to controller, DNS and internet only.
- [ ] Check the **DDNS and LAN DNS settings** in the controller after the migration.
- [ ] Optional: NPM proxy host for NPM's own admin UI.
**Server**
- [ ] **Off-box backups** of `~/omada`, `~/pihole`, `~/npm`, `~/homeassistant` — encrypted, since NPM's data contains the INWX password. Supersedes the Omada-only rsync plan under "Later".
- [ ] **Bring the OptiPlex under Ansible**: SSH hardening, `resolved` stub-listener drop-in, GRUB `pcie_aspm=off`, Docker install, compose stacks from [`../docker/`](../docker/). Everything was done by hand so far.
- [ ] Cupboard cooling (fan and vent) — the CPU hits 88 °C under stress and the NVMe needs its thermal pad.
- [ ] Optional: replace the NVMe (~€30–40, 500 GB) and drop `pcie_aspm=off`.
**Services**
- [ ] Portainer (behind HTTPS; for viewing and restarts only, compose files stay the source of truth).
- [ ] Vaultwarden, with backups to the NAS.
- [ ] Own Docker tools.
**Home Assistant**
- [ ] Fix login through `ha.home.staffenberger.at` — first check Websockets Support on the NPM proxy host.
- [ ] Zigbee dongle and IKEA lights.
- [ ] ACL exceptions so Home Assistant can reach IoT devices (Servers → IoT is allowed today; check the new rules don't break that).
- [ ] Synology DSM integration (separate DSM user).
- [ ] Cat camera.
## Next up
- [ ] **Hardware shopping list**: ~~**EAP650** (2nd AP) + **ES205GP** (living-room PoE switch)~~ (bought and installed); **flat Cat6 patch cables** for the cupboard (deleyCON flat U/UTP, pure copper — 25 cm 5-pack for short links, plus 0.5 m / 1 m / 2 m, and a **3 m** for the PC runs *if* a string measurement of the real path is over ~1.7 m; verify on each listing that it says copper and is the flat U/UTP version); optionally a colored multi-length **round** Cat6 set for grab-and-go spares (flat only matters for the permanent cupboard cables); **socket strips** — two, each plugged into a *different* wall outlet, no daisy-chaining, wall-mountable metal-housing (screws, no adhesive), VDE/GS/ÖVE mark, 3×1.5 mm² cable, surge-protected for the desk/PC strip, and no switch (or a guarded one) for the strip carrying router/switch/NAS. **PoE-powered cables (to the APs) should be round pure-copper cable, not flat/thin ones.**
- [x] **Living-room second AP + switch** — EAP650 + ES205GP bought, mounted and configured (as of 2026-09-29). Still open: **re-measure the couch** (baseline **-70 dBm**). Original plan (see status.md Wi-Fi and Living room sections): buy an **EAP650** (same as the first AP) and an **ES205GP** (the PoE+ variant, not the plain ES205G). Mount both above the TV, one cable run from the wall to there. Then: adopt both, set the ES205GP uplink + AP ports as trunks (Trusted/IoT/Guest tagged) and the SG2016P port feeding it likewise, TV port → access on IoT, Pi port → access on Trusted. Re-measure the couch afterward (baseline was **-70 dBm**).
- [ ] **Mount the office AP at its final ceiling position**, then re-measure the couch and the other rooms (baseline readings recorded in status.md) — cheap to do first, and useful as the before/after reference for the second AP.
- [x] **Omada Controller auto-backup enabled, set to daily** (config is changing a lot right now). Weekly switch tracked separately in Backlog. Original notes: top priority, cheap insurance for all the VLAN/ACL/VPN config already built. **Scope for now: local only** — backups stay on the machine running the Controller (accepted for now); also download a manual copy occasionally. The off-machine copy to the Synology is deferred, see "Later". Note auto-backups live in the container's data volume, so `docker compose down -v` deletes them along with everything else.
- [x] **Robot vacuum** and **cat feeder** connected to the IoT Wi-Fi SSID, named in the Controller for easy recognition. No DHCP reservation — not needed for devices only reached via their own cloud apps.
- [x] **Synology NAS access**: QuickConnect **disabled** — NAS is now only reachable via the WireGuard VPN, confirmed working. Synology Drive/Photos apps switched from QuickConnect ID to manual local-address login (David's phone and the partner's accounts both done).
- [ ] **Decide whether IoT devices should use Pi-hole for DNS** (now tracked as the DNS ACL exception under Homeserver follow-ups — tick both together): Pi-hole sits on Servers (30), same as always planned. The existing `IoT → !IoT deny` ACL rule currently blocks IoT from reaching Pi-hole too — meaning any IoT device pointed at it for DNS would fail to resolve. **Left as-is for now** (IoT just uses ER605/ISP DNS directly, no ad-blocking there). If wanted later: add a narrow exception, `Network: IoT → Network: Servers`, port 53 only, Direction `LAN-LAN`, Allow — evaluated before the broader IoT deny-all rule.
- [ ] **▶ Synology NAS backup — first run completed (started 2026-09-26 overnight).** Task is created (Multiple versions, Folders and Packages, all shared folders + `homes` + all applications, config backup + encryption on, key in Enpass, no schedule). **Next: (1)** ~~check the run finished~~ done, **(2) test-restore** a personal Drive file and a personal photo for both users to confirm the Drive Server / Photos application backups really cover personal data, **(3) eject in DSM and unplug**, **(4) set the every-two-weeks reminder**, then tick this item. Full setup notes are in status.md's Backups section. Original notes: (currently no backup existed at all — top data-safety item): 6TB USB drive on hand, but the **USB 3.0 Micro-B cable is misplaced** (find it, or buy a replacement). Chosen workflow: **manually connected, not permanently attached** — the offline drive also protects against ransomware/accidental deletion. Steps: format the drive as **ext4** in DSM (Control Panel → External Devices → Format; erases it — check for existing data first), install **Hyper Backup**, create a Data backup task ("Local folder & USB") covering the important shared folders + Applications, **encryption on** (password in Enpass), version retention (Smart Recycle) + periodic integrity check, DSM notifications on failure, and **test-restore one folder**. To run: plug in, "Back up now", **eject in DSM** (Control Panel → External Devices → Eject) before unplugging, store the drive away from the NAS. DSM 7 has no plug-in-triggered backup, so the habit needs a **recurring calendar reminder** — chosen cadence: **every two weeks** (see how it goes; also run one before big photo imports or major DSM updates). Drive was exFAT out of the box (DSM couldn't cleanly mount it and reported "not ejected safely" — exFAT has no journal); reformatted to ext4 in DSM. The first full backup will be slow on the DS223j's 1 GB RAM — run it overnight. Later: an occasional off-site copy (3-2-1 rule) since one drive at home doesn't cover fire/theft.
- [ ] **Real (non-self-signed) certificate for the NAS's own DSM access** — not crucial, DSM's cert warning is just cosmetic for local access. **Easiest now:** a `nas.home.staffenberger.at` proxy host in NPM (Pi-hole record → `192.168.30.20`) reuses the existing wildcard cert, no extra credentials. Older notes: Leaning toward **`mkcert`** (local CA, install its root cert as trusted on your own devices once, zero external credentials/services involved) over Let's Encrypt DNS-01 through INWX — the latter would need either a scoped INWX API key (check if INWX offers one) or `acme.sh`'s manual mode (no credentials, but manual renewal every ~90 days); full INWX account credentials handed to a script was correctly ruled out as too broad a permission grant for this.
## Ansible — learning, started 2026-09-27
Setup lives in [`../ansible/`](../ansible/README.md). OctoPrint Pi is the guinea pig; the OptiPlex homeserver is the real goal (status.md: Ansible from David's PC, no footprint on the server).
- [x] SSH key auth to the Pi: one key per *client device* (`id_ed25519_homelab`), not per server; `~/.ssh/config` entry with `IdentitiesOnly yes`.
- [x] `ansible octoprint -m ping` works; `update.yml` run for real on 2026-09-27 (248 packages, kernel → 6.12.109). Gotchas: David's sudo needs a password → run with `-K`; Raspberry Pi OS does **not** create `/var/run/reboot-required` after kernel updates, so the playbook compares the running kernel with the newest installed one of the same flavour.
- [ ] **Audit the Pi's hand-made changes** so a rebuild can reproduce them: `apt-mark showmanual`, enabled services, crontabs, `/boot/firmware/config.txt`, installed OctoPrint plugins. **Removed by hand 2026-09-27** (one-off cleanups don't belong in the config playbook — it describes what *should* be there, and an "ensure absent" task would fight future experiments): **Docker** (installed Sep 2026 for a Spoolman test, Spoolman already gone — packages, `docker.list` repo + `docker.asc` key, `/var/lib/docker`, `/var/lib/containerd`, `docker` group) and **nginx** (only the default site, never listening — `haproxy` holds 80/443 and fronts OctoPrint on `127.0.0.1:5000`). Rule going forward: experiment by hand, then either add it to the playbook or remove it; spool-manager-type services belong on the homeserver anyway. The Pi 4 (**1 GB RAM**) boots the 32-bit `rpi-v7` kernel instead of the Bookworm default (64-bit `v8` kernel, 32-bit userland) — `config.txt` has an explicit `arm_64bit=0`, which appears to come with the OctoPi image (apt history shows the Pi is essentially a stock OctoPi flash, not years of in-place upgrades). **Decided: leave it.** With 1 GB there's no RAM to gain, and 32-bit userland uses slightly less memory; switching only the kernel isn't worth it before the rebuild. For the rebuild: a **32-bit** OctoPi image is fine on this Pi, using whatever kernel it boots by default. Still wanted from the audit: `config.txt`, enabled services, apt history.
- [x] First *configuration* playbook `playbooks/octoprint.yml` (run 2026-09-27; verified: password SSH refused, timezone set, no Bluetooth device): timezone, ModemManager masked (probes the printer's serial port), Bluetooth off (services + `dtoverlay=disable-bt`), SSH password login off. Wi-Fi stays on — the Pi is on the IoT **Wi-Fi** now, no longer on the cable. nginx removal checked: webcam uses MJPEG (`webcamd`), not the HLS stream (`ffmpeg_hls`, which needed nginx).
- [ ] **Rebuild procedure** (document once tested): flash current 32-bit OctoPi with Raspberry Pi Imager's OS customisation — hostname, user `david`, IoT Wi-Fi, SSH **public-key only** with `id_ed25519_homelab.pub` (Wi-Fi password stays out of the repo) → `ansible-playbook playbooks/octoprint.yml -K` → restore the OctoPrint backup. Test on a spare SD card.
- [ ] **Automated OctoPrint backups** (deferred — basics first). Rebuild plan: flash card → run playbook → restore OctoPrint backup; OctoPrint's own state (settings, plugins, profiles) comes from its Backup & Restore zip, not Ansible (OctoPrint rewrites its own `config.yaml`, so templating it would fight the UI). Plan:
1. Ansible deploys a nightly timer on the Pi running the backup CLI (`octoprint plugins backup:backup` — verify syntax/exclude flags first), excluding uploads/timelapses, keeping the last few. Alternative: the "Backup Scheduler" plugin, but that's UI config outside the repo.
2. Off-device copy must be **pulled** — the `IoT → !IoT` ACL means the Pi can't push to the NAS (keep it that way). Interim: `update.yml` takes a backup and fetches it to David's PC before upgrading. Later: nightly pull from the OptiPlex (Servers → IoT allowed) with a **dedicated key restricted via `rrsync`** to read-only access to the backup folder.
3. Backup zips contain user hashes + API keys — store outside this repo (decide location, e.g. NAS share).
- [ ] Split into roles (`common`, `updates`) + `group_vars`; `ansible-vault` for any secrets (e.g. OctoPrint API key for a "skip while printing" check).
- [ ] Add the two personal Linux machines (`community.general.pacman` for Arch-based ones — not fully unattended, Arch updates occasionally need manual steps).
- [ ] Provision the OptiPlex with it (common role + Docker + Compose stacks) — see the Homeserver follow-ups.
## Backlog / lower priority
- [ ] **Change the Omada Controller auto-backup from daily to weekly** — daily is deliberate for now because so much config is changing; revisit once the switch/AP/VLAN work has settled and changes become infrequent.
- [ ] UPS/power protection for the always-on gear (router/switch/homeserver/NAS) — optional cost/complexity tradeoff, not urgent.
- [ ] Basic uptime/service monitoring (e.g. Uptime Kuma) on the OptiPlex — nice-to-have, not essential.
## Later
- [x] **Local DNS + reverse proxy plan**: first ran temporarily on the Synology; **moved to the OptiPlex on 2026-10-04** with clean 80/443/81 ports (see [`../docker/npm/`](../docker/npm/README.md)). Original notes: NPM deployed (SQLite-based), remapped to ports **9080/9443/9081** since DSM's own internal nginx already held 80/443/81 — proxied HTTPS URLs need an explicit `:9443` until this moves to the NUC, where a clean 80/443/81 mapping will work with no conflict (revert `docker/npm/docker-compose.yml`'s `ports:` at that point). Pi-hole Local DNS Records in use (e.g. `pihole.home.staffenberger.at`, `nas.home.staffenberger.at`) rather than the wildcard approach so far — worth switching to the wildcard dnsmasq config later to stop needing a new record per service.
- [x] **TLS certs for NPM**: decided and done 2026-10-04 — **Let's Encrypt wildcard** `*.home.staffenberger.at` via NPM's INWX DNS challenge, using a dedicated minimal-rights INWX sub-user (that resolved the credential-scope concern that had pointed toward `mkcert`). Details in [`../docker/npm/`](../docker/npm/README.md).
- [ ] **Off-machine Omada Controller backups to the Synology** — superseded by the OptiPlex off-box backup item under Homeserver follow-ups (covers all four stacks); the design notes below still apply. Autobackups are at `~/omada/data/autobackup` on the host (bind mount). Original notes: Decided approach: a scheduled **rsync** job on the Controller host copying the Controller's autobackup folder (believed to be `/opt/tplink/EAPController/data/autobackup` inside the container — verify with `docker exec omada-controller ls` once a backup has run) to the NAS. **Not** a NAS folder mounted into the container (avoids the container's startup depending on the NAS, and keeps a failed copy from affecting the Controller). Use a **dedicated Synology user with a small storage quota** and key-based auth, so the credential on the Controller host can only touch that one backup folder — check that the DS223j's volume filesystem supports user quotas, and how DSM allows a non-admin user to do rsync/SFTP (SSH is admin-only by default). Backups contain secrets (Device Account, WireGuard keys, DDNS settings), so restrict the folder to that user and keep the files out of this Git repo. Once the Synology's own Hyper Backup to the 6TB drive exists, these are covered there too.
- [x] Migrate Omada Controller from the desktop to the homeserver — done 2026-10-04 (OptiPlex), see [omada-controller-migration.md](omada-controller-migration.md).
## Future ideas
- [ ] **Parents' network overhaul**: plan refined after discussion — not a site-to-site link, not adopting their hardware into this Omada Controller (TP-Link **Deco is a separate, incompatible ecosystem** — can't be adopted into Omada Controller under any configuration; a "workaround" of switching Deco to AP mode just makes it a dumber Wi-Fi extender still managed by the Deco app, doesn't add controller compatibility).
- Buy a **second ER605** to replace their current limited ISP-provided modem/router — get their modem into bridge mode (same pattern as home), ER605 becomes their real router.
- Switch their existing **Deco units into Access Point Mode** (a supported Deco app toggle) so they stop doing their own NAT/DHCP behind the new ER605 — otherwise double-NAT. Note: even in AP mode, Deco can't map SSIDs to VLANs like the EAP650 can — it just bridges onto whatever single network the port it's plugged into carries. Since they don't need segmentation (confirmed — no IoT/multi-trust-tier need there), this doesn't matter: **one flat subnet** is enough, no VLANs needed on their side.
- Pick a subnet clearly distinct from home's (`192.168.10/20/30/40/66/99.0/24`) — e.g. `10.50.0.0/24`, using a different address family entirely so it's visually obvious which network you're on.
- Set up an **independent WireGuard VPN server** on that second ER605 (own DynDNS entry via INWX, same pattern as home) with its own client IP pool distinct from home's `192.168.66.0/24` — **not** a site-to-site link. Connect to it only on-demand from your own devices when help is actually needed; no persistent link between the two networks, no changes/risk to your own network at all.
- Doesn't remove the Deco mesh's own reliance on TP-Link's cloud for its own management (that's inherent to Deco) — separate, optional concern from the remote-access goal, not solved by any of the above.
- Next actual step: investigate their current ISP/modem situation (does it support bridge mode? which ISP?) before buying anything.