Introduces Optiplex to the home setup. Install the first software. Introduce management vlan 99

This commit is contained in:
David Staffenberger 2026-10-04 21:09:00 +02:00
parent e8ffeb0856
commit f7beda4c59
17 changed files with 454 additions and 206 deletions

View file

@ -5,6 +5,8 @@ and Docker/Ansible-managed services.
- [docs/status.md](docs/status.md) — decisions and status - [docs/status.md](docs/status.md) — decisions and status
- [docs/todo.md](docs/todo.md) — live TODO checklist - [docs/todo.md](docs/todo.md) — live TODO checklist
- [docs/homeserver.md](docs/homeserver.md) — Dell OptiPlex homeserver: hardware, SSD fix, OS setup, service overview
- [docs/er605-firmware-update.md](docs/er605-firmware-update.md) — ER605 firmware update steps - [docs/er605-firmware-update.md](docs/er605-firmware-update.md) — ER605 firmware update steps
- [docker/omada-controller/](docker/omada-controller/) — self-hosted Omada Software Controller (Docker) - [docs/omada-controller-migration.md](docs/omada-controller-migration.md) — moving the controller between hosts (done once: desktop → OptiPlex)
- [docs/omada-controller-migration.md](docs/omada-controller-migration.md) — moving the controller from desktop to the NUC later - [docker/](docker/) — compose files running on the OptiPlex: [omada](docker/omada/), [pihole](docker/pihole/), [npm](docker/npm/), [homeassistant](docker/homeassistant/)
- [ansible/](ansible/README.md) — Ansible setup (OctoPrint Pi so far)

View file

@ -0,0 +1,24 @@
# Home Assistant
Runs on the OptiPlex homeserver (`192.168.30.20`) in `~/homeassistant`, as the
official container image (not Home Assistant OS). UI:
`https://ha.home.staffenberger.at` (via NPM) or `http://192.168.30.20:8123`.
## Reverse proxy
Configured in the UI, **not** in `configuration.yaml`: Settings → System →
Network → HTTP server → Reverse proxy:
- **Trust X-Forwarded-For** on
- Trusted proxies: `172.16.0.0/12` (Docker's internal range; NPM connects from `172.19.0.2`)
Gotcha: this version migrated the YAML `http:` block into the UI once and
ignores it afterwards. A typo during migration (`use_x_forward_for`) left the
switch off, causing `400: Bad Request`. Keep **no** `http:` block in the YAML.
## Notes
- Bluetooth errors in the log are harmless; installing `bluez` on the host or
removing the Bluetooth integration stops them.
- Login through `ha.home.staffenberger.at` was unreliable — first check that
Websockets Support is on for the NPM proxy host (see [../../docs/todo.md](../../docs/todo.md)).

View file

@ -0,0 +1,13 @@
services:
homeassistant:
image: ghcr.io/home-assistant/home-assistant:stable
container_name: homeassistant
restart: unless-stopped
# Host networking for device discovery (mDNS, SSDP, ...).
network_mode: host
environment:
TZ: Europe/Vienna
volumes:
- ./config:/config
- /etc/localtime:/etc/localtime:ro
- /run/dbus:/run/dbus:ro # Bluetooth via the host's BlueZ

View file

@ -1,33 +1,59 @@
# Nginx Proxy Manager (temporary, on the Synology) # Nginx Proxy Manager
Running here until the NUC homeserver exists — see Runs on the OptiPlex homeserver (`192.168.30.20`) in `~/npm`. Admin UI:
[../../docs/todo.md](../../docs/todo.md) for the full local-DNS + reverse-proxy `http://192.168.30.20:81` (not proxied yet).
plan (Pi-hole/LAN DNS points internal hostnames at this box's IP; NPM routes
each hostname to the right backend service/port).
Deploy via Container Manager's **Project** feature (paste/import this Every service gets a name through two entries: a Pi-hole **Local DNS record**
compose file — same as however Pi-hole was set up there) or over SSH with pointing `<name>.home.staffenberger.at` at `192.168.30.20`, and an NPM **proxy
`docker compose up -d` from this directory. host** forwarding to the real address and port.
## Port note ## Proxy hosts
Ports 80/443 are already taken on this Synology by DSM's own internal nginx | Name | Scheme | Forward to | Extras |
(`netstat -tulpn | grep :443` showed `nginx: worker` — likely Web Station or |---|---|---|---|
the Login Portal's HTTP→HTTPS redirect feature); 8080/8443/81 turned out to | pihole.home.staffenberger.at | http | 192.168.30.20:8081 | Advanced: redirect `/` to `/admin/` |
also already be in use by other running packages. Rather than chase down and | omada.home.staffenberger.at | https | 192.168.30.20:8043 | Websockets on |
touch existing DSM config for a temporary deployment, NPM is remapped to | ha.home.staffenberger.at | http | 192.168.30.20:8123 | Websockets on |
host ports **9080/9443/9081** instead — meaning proxied HTTPS URLs need an | octoprint.home.staffenberger.at | http | OctoPrint Pi, port 80 | Websockets on |
explicit `:9443` and the admin UI is at `http://<synology-ip>:9081` until
this moves to the NUC, where a clean `80:80`/`443:443`/`81:81` mapping will
work with no conflict. Revert the compose file's `ports:` section at that
point.
## Notes All hosts use the wildcard certificate with Force SSL and HTTP/2 on; HSTS is off for now.
- No MariaDB container — NPM's built-in SQLite database is enough for this ## Wildcard certificate
scale and keeps the footprint light on the DS223j's 1GB RAM.
- Admin UI: `http://<synology-ip>:81` — default login is `admin@example.com` Let's Encrypt certificate for `*.home.staffenberger.at`, issued and renewed by
/ `changeme` on first run; change both immediately. NPM via a **DNS challenge at INWX** (the names point at internal addresses, so
- TLS certs: plan is Let's Encrypt via DNS-01 using INWX's API (an `acme.sh` an HTTP challenge can't work, and wildcards need DNS validation anyway).
plugin exists for this) once that part of the plan is implemented — see
[../../docs/todo.md](../../docs/todo.md). Add Certificate → Let's Encrypt via DNS → provider INWX, key type ECDSA 256,
propagation 120 seconds:
```
dns_inwx_url = https://api.domrobot.com/xmlrpc/
dns_inwx_username = <sub-user>
dns_inwx_password = """<password>"""
```
- Uses a **dedicated INWX sub-user** with minimal rights and no 2FA. NPM stores
the password in plain text in its database and credentials file, so `./data`
and `./letsencrypt` are secrets — never commit them, and encrypt any backup.
- Covers exactly one level below `home.staffenberger.at`; nothing outside `*.home` is affected.
- If it leaks: change the sub-user's password at INWX and update it in NPM.
- Alternative considered: acme-dns via a CNAME on `_acme-challenge.home`,
limiting the credential to a single TXT record. Possible later, ideally
self-hosted on the git server.
## Lessons learned
- The **Forward Hostname** field takes only an IP or host name, never a path.
A path there caused a proxy loop (`414 Request-URI Too Large`).
- The **Advanced** tab is the gear icon at the top right of the proxy host
dialog. Redirect used for Pi-hole:
```nginx
location = / {
return 301 /admin/;
}
```
- For large OctoPrint uploads, add `client_max_body_size 0;` in the Advanced tab if needed.
- No MariaDB container — NPM's built-in SQLite is enough at this scale.

14
docker/npm/compose.yaml Normal file
View file

@ -0,0 +1,14 @@
services:
npm:
# Built-in SQLite database — no DB_MYSQL_* / DB_SQLITE_* env vars needed.
image: jc21/nginx-proxy-manager:latest
container_name: npm
restart: unless-stopped
ports:
- "80:80" # HTTP, proxied traffic
- "443:443" # HTTPS, proxied traffic
- "81:81" # Admin UI
volumes:
# Both contain the INWX DNS-challenge password in plain text — never commit.
- ./data:/data
- ./letsencrypt:/etc/letsencrypt

View file

@ -1,24 +0,0 @@
services:
app:
container_name: nginx-proxy-manager
# Multi-arch tag — auto-selects the arm64 build on the Synology DS223j.
# Never set the DB_MYSQL_* / DB_SQLITE_* env vars unless you specifically
# want MariaDB — omitting them keeps NPM on its lightweight built-in
# SQLite database, no separate DB container needed.
image: 'jc21/nginx-proxy-manager:latest'
restart: unless-stopped
ports:
# Remapped off 80/443/81 — all three were already claimed on this
# Synology (DSM's own internal nginx plus other running packages).
# Temporary deployment only; revert to a clean 80:80/443:443/81:81
# once this moves to the NUC, which won't have that conflict.
- '9080:80' # HTTP, proxied traffic
- '9443:443' # HTTPS, proxied traffic
- '9081:81' # Admin UI
volumes:
- ./data:/data
- ./letsencrypt:/etc/letsencrypt
volumes:
data:
letsencrypt:

View file

@ -1,51 +0,0 @@
# Omada Software Controller (temporary, on David's desktop)
Runs the self-hosted Omada Software Controller via the community-maintained
[`mbentley/omada-controller`](https://github.com/mbentley/docker-omada-controller)
image, avoiding TP-Link's cloud controller.
This instance is meant to be **temporary**: it runs on David's Linux desktop just
long enough to adopt the ER605 (and later the switch/AP) instead of leaving them
in standalone mode, then gets migrated to the NUC once that's provisioned — see
[../../docs/omada-controller-migration.md](../../docs/omada-controller-migration.md).
The desktop does not need to run 24/7. The ER605 keeps forwarding traffic on its
last-known config even if the controller is offline — you only lose live
management/statistics visibility while it's down.
## Prerequisites
- Docker + Docker Compose installed on the desktop.
- ER605 firmware already updated — see [../../docs/er605-firmware-update.md](../../docs/er605-firmware-update.md). Do this first; it's harder once the device is controller-managed.
- Desktop and ER605 on the same L2 network segment (needed for adoption's broadcast discovery).
## Bring it up
```bash
docker compose up -d
```
Then open `https://<desktop-ip>:8043` and walk through the initial setup wizard
(create the controller admin account, name the site, etc).
## Adopt the ER605
1. In the controller UI, go to the site's **Devices** view — the ER605 should
appear as "Pending" once discovery finds it on the network.
2. Click **Adopt** and enter the router's current admin credentials when
prompted.
3. Wait for adoption to finish and the device to show **Connected**.
## Notes
- Image tag is pinned to a `major.minor` version (currently `6.3`) — check
[Docker Hub tags](https://hub.docker.com/r/mbentley/omada-controller/tags)
before bumping it, and never move backwards to an older version once the
Mongo database has been touched by a newer one.
- `network_mode: host` is used because Omada discovery/adoption depends on a
wide range of UDP/TCP ports and broadcast traffic; bridging those
individually is more fragile than just sharing the host network.
- Controller data lives in the `omada-data` / `omada-logs` named Docker
volumes, not bind mounts — back them up via the controller's own
**Settings → Maintenance → Backup** feature (see the migration guide),
not by copying the volume directly.

View file

@ -1,45 +0,0 @@
services:
omada-controller:
container_name: omada-controller
# Pin major.minor — never use `latest`, a version bump can break the Mongo schema mid-upgrade.
# Check https://hub.docker.com/r/mbentley/omada-controller/tags for the current recommended tag.
image: mbentley/omada-controller:6.3
restart: unless-stopped
ulimits:
nofile:
soft: 4096
hard: 8192
stop_grace_period: 60s
# Host networking: Omada device discovery/adoption relies on broadcast + a wide range of
# UDP/TCP ports below. Simplest to just share the desktop's network stack directly.
network_mode: host
environment:
- PUID=508
- PGID=508
- TZ=Europe/Vienna
- MANAGE_HTTP_PORT=8088
- MANAGE_HTTPS_PORT=8043
- PORTAL_HTTP_PORT=8088
- PORTAL_HTTPS_PORT=8843
- UPGRADE_HTTPS_PORT=8044
- PORT_APP_DISCOVERY=27001
- PORT_DISCOVERY=29810
- PORT_MANAGER_V1=29811
- PORT_ADOPT_V1=29812
- PORT_UPGRADE_V1=29813
- PORT_MANAGER_V2=29814
- PORT_TRANSFER_V2=29815
- PORT_RTTY=29816
- PORT_DEVICE_MONITOR=29817
- WEB_CONFIG_OVERRIDE=false
- SHOW_SERVER_LOGS=true
- SHOW_MONGODB_LOGS=false
- SSL_CERT_NAME=tls.crt
- SSL_KEY_NAME=tls.key
volumes:
- omada-data:/opt/tplink/EAPController/data
- omada-logs:/opt/tplink/EAPController/logs
volumes:
omada-data:
omada-logs:

37
docker/omada/README.md Normal file
View file

@ -0,0 +1,37 @@
# Omada Software Controller
Runs on the OptiPlex homeserver (`192.168.30.20`, Servers VLAN) in `~/omada`,
using the community-maintained
[`mbentley/omada-controller`](https://github.com/mbentley/docker-omada-controller)
image instead of TP-Link's cloud controller. UI: `https://omada.home.staffenberger.at`
(via NPM) or `https://192.168.30.20:8043`.
Migrated from the temporary instance on David's desktop on 2026-10-04 — see
[../../docs/omada-controller-migration.md](../../docs/omada-controller-migration.md).
## Operate
```bash
docker compose up -d
docker compose logs -f
```
Version check:
```bash
curl -sk https://localhost:8043/api/info | grep -o '"controllerVer":"[^"]*"'
```
## Notes
- Image tag is pinned to `major.minor` (`6.3`, controller 6.3.0.45). **Never use
`latest`** — it still points at v5. Check
[Docker Hub tags](https://hub.docker.com/r/mbentley/omada-controller/tags)
before bumping, and never move backwards once the Mongo database has been
touched by a newer version.
- `network_mode: host` because discovery/adoption depends on broadcast and a
wide range of UDP/TCP ports.
- Data is bind-mounted to `./data` and `./logs`. Auto-backups land in
`./data/autobackup`. Export a manual backup after every bigger change
(Settings → Maintenance → Backup & Restore). Backups contain secrets (device
account, WireGuard keys, DDNS settings) — keep them out of git.

21
docker/omada/compose.yaml Normal file
View file

@ -0,0 +1,21 @@
services:
omada-controller:
# Pin major.minor — never use `latest`, it still points at v5 and a version
# bump can break the Mongo schema mid-upgrade.
# Check https://hub.docker.com/r/mbentley/omada-controller/tags before bumping.
image: mbentley/omada-controller:6.3
container_name: omada-controller
restart: unless-stopped
# Host networking: device discovery/adoption relies on broadcast and a wide
# range of UDP/TCP ports (8043, 8088, 27001, 29810-29817, ...).
network_mode: host
stop_grace_period: 60s
ulimits:
nofile:
soft: 4096
hard: 8192
environment:
- TZ=Europe/Vienna
volumes:
- ./data:/opt/tplink/EAPController/data
- ./logs:/opt/tplink/EAPController/logs

25
docker/pihole/README.md Normal file
View file

@ -0,0 +1,25 @@
# Pi-hole
Runs on the OptiPlex homeserver (`192.168.30.20`) in `~/pihole`. Web UI on
port **8081** (`https://pihole.home.staffenberger.at` via NPM), so 80/443 stay
free for NPM. Replaces the Pi-hole on the Synology (`192.168.30.10`), which is
switched off once DHCP points at the new one.
## Setup notes
- Port 53 on the host is freed by disabling systemd-resolved's stub listener:
`DNSStubListener=no` in `/etc/systemd/resolved.conf.d/no-stub.conf`, and
`/etc/resolv.conf` linked to `/run/systemd/resolve/resolv.conf`.
- `FTLCONF_dns_listeningMode: "all"` is required — otherwise Pi-hole ignores
clients from other VLANs.
- Admin password: `docker exec -it pihole pihole setpassword` (stored in the
password manager, not in the compose file).
- Settings were imported from the NAS Pi-hole via **Teleporter**.
- **Local DNS records:** every `*.home.staffenberger.at` name points at
`192.168.30.20` (NPM) — including services on other devices such as OctoPrint.
## Test
```bash
nslookup google.com 192.168.30.20
```

View file

@ -0,0 +1,17 @@
services:
pihole:
image: pihole/pihole:latest
container_name: pihole
restart: unless-stopped
ports:
- "53:53/tcp"
- "53:53/udp"
- "8081:80/tcp" # Web UI — 80/443 stay free for NPM
environment:
TZ: Europe/Vienna
# Answer clients from other VLANs, not just the local subnet.
FTLCONF_dns_listeningMode: "all"
# Admin password is set with `docker exec -it pihole pihole setpassword`,
# not here.
volumes:
- ./etc-pihole:/etc/pihole

View file

@ -19,7 +19,7 @@ Do this even though the router is barely configured yet — it's a 30-second saf
1. Log into the router's web UI (default `https://192.168.0.1`, or whatever IP it's currently reachable at; default credentials are `admin`/`admin` unless already changed). 1. Log into the router's web UI (default `https://192.168.0.1`, or whatever IP it's currently reachable at; default credentials are `admin`/`admin` unless already changed).
2. Go to **System Tools / Settings → Backup & Restore** (exact path varies slightly by firmware version). 2. Go to **System Tools / Settings → Backup & Restore** (exact path varies slightly by firmware version).
3. Download the current config backup somewhere safe (e.g. `docker/omada-controller/` is *not* the place — keep router backups out of git, or in a private/untracked location). 3. Download the current config backup somewhere safe (e.g. `docker/omada/` is *not* the place — keep router backups out of git, or in a private/untracked location).
## 4. Upload the new firmware ## 4. Upload the new firmware
@ -35,7 +35,7 @@ Do this even though the router is barely configured yet — it's a 30-second saf
## Next step ## Next step
Once firmware is current, move on to standing up the Omada Software Controller: [../docker/omada-controller/](../docker/omada-controller/). Once firmware is current, move on to standing up the Omada Software Controller: [../docker/omada/](../docker/omada/).
## References ## References

112
docs/homeserver.md Normal file
View file

@ -0,0 +1,112 @@
# Homeserver — Dell OptiPlex 3000
The OptiPlex (`192.168.30.20`, Servers VLAN) is the homeserver. It runs the
Omada controller, Pi-hole, Nginx Proxy Manager and Home Assistant in Docker;
all web UIs are served over HTTPS under `*.home.staffenberger.at`. Set up on
2026-10-04.
| Service | Runs on | Internal address | Name (via NPM, HTTPS) |
|---|---|---|---|
| Omada controller | OptiPlex (Docker, host network) | https://192.168.30.20:8043 | omada.home.staffenberger.at |
| Pi-hole | OptiPlex (Docker) | http://192.168.30.20:8081/admin | pihole.home.staffenberger.at |
| Nginx Proxy Manager | OptiPlex (Docker) | http://192.168.30.20:81 | (not proxied yet) |
| Home Assistant | OptiPlex (Docker, host network) | http://192.168.30.20:8123 | ha.home.staffenberger.at |
| OctoPrint | Raspberry Pi | Pi IP, port 80 | octoprint.home.staffenberger.at |
| Old Pi-hole | Synology NAS | 192.168.30.10 | to be switched off |
Compose files live in `~/omada`, `~/pihole`, `~/npm` and `~/homeassistant` on
the OptiPlex; copies with notes are in [`../docker/`](../docker/). SSH from the
PC: `ssh optiplex`.
## Hardware check
Dell OptiPlex 3000: Core i3-12300T (12th gen, 35 W "T" part), 16 GB DDR4, 512 GB NVMe SSD, genuine Dell 65 W power supply — matched the listing.
| Check | Result |
|---|---|
| Memtest86+ | 2 passes, 0 errors (needs Secure Boot off to boot) |
| CPU stress test (stress-ng, 10 min) | Max 88 °C, no throttling |
| Dell ePSA diagnostics | Passed |
| Network link | 1000 Mb/s |
| USB ports | All working |
| SSD SMART | ~15,400 power-on hours, 0 media errors, extended self-test passed twice |
The Windows key was read with ShowKeyPlus before wiping. An OEM key in the
firmware stays there after installing Linux and only reactivates Windows on
this machine.
## SSD problem and fix
The system froze under disk load because the NVMe link threw fatal PCIe
errors. Stable since adding a thermal pad and the kernel option `pcie_aspm=off`.
**The drive:** Toshiba XG4 THNSN5512GPUK (512 GB NVMe), an HP OEM part (HP P/N
902944-002, made 2017), fitted by a refurbisher — not the original Dell drive.
**Symptoms:** load average around 37 with idle CPU, commands hanging, kernel
log entries `PCIe Bus Error: severity=Uncorrectable (Fatal)` followed by
controller resets.
**Findings:**
- The board's "M.2 SSD THERMAL PAD AREA" under the SSD was bare. The drive
reached 81 °C (warning 78 °C, critical 82 °C). With a thermal pad (two
stacked), idle dropped from 68 °C to ~46 °C, ~70 °C under sustained writes.
- The errors persisted at low temperature, so heat was not the root cause. The
SSD's AER registers show an FCP error (flow-control protocol, classified
fatal), most likely triggered by the L1.2 link power state the BIOS enables.
- The BIOS declares ASPM unsupported (FADT), so the kernel never controls ASPM.
`pcie_aspm=off` makes the kernel leave AER and LTR to the BIOS, which
tolerates the error. With kernel control (`pcie_aspm.policy=performance`) the
fatal errors returned every 30–60 seconds.
- The link trains at full speed (8 GT/s, x4). fio: ~1,380 MB/s read, ~500 MB/s write.
**Current setting:** `GRUB_CMDLINE_LINUX_DEFAULT="pcie_aspm=off"` in
`/etc/default/grub`. The APST option (`nvme_core.default_ps_max_latency_us=0`)
was removed again — it only made the drive run hotter.
**Health checks:**
```bash
sensors | grep -A1 Composite
sudo smartctl -a /dev/nvme0n1 | grep -iE "temp|media"
sudo lspci -vv -s 01:00.0 | grep -E "UESta|CESta"
```
Baseline: Warning Comp. Temperature Time = 5 minutes. A new 500 GB NVMe
(~€30–40) would remove the workaround entirely.
## Ubuntu Server and SSH
Ubuntu Server 26.04 LTS, headless, no disk encryption, SSH key logins only.
- **Install choices:** full disk via LVM, OpenSSH server, no featured snaps, no
encryption (so it boots unattended after power cuts).
- **BIOS:** Secure Boot on, AC Recovery set to Power On.
- **SSH key:** `~/.ssh/id_ed25519_homelab` (same per-client key as for the Pi),
copied with `ssh-copy-id`. Entry in `~/.ssh/config` on the PC:
```
Host optiplex
HostName 192.168.30.20
User david
IdentityFile ~/.ssh/id_ed25519_homelab
IdentitiesOnly yes
```
- **Password login disabled** in `/etc/ssh/sshd_config.d/00-no-password.conf`
(`PasswordAuthentication no`, `KbdInteractiveAuthentication no`). The `00-`
prefix makes it win over Ubuntu's `50-cloud-init.conf`.
- **Port 53 freed for Pi-hole:** `DNSStubListener=no` in
`/etc/systemd/resolved.conf.d/no-stub.conf`, and `/etc/resolv.conf` linked to
`/run/systemd/resolve/resolv.conf`.
## Docker
Installed from Docker's official apt repository (`docker-ce`,
`docker-compose-plugin`); user `david` is in the `docker` group. Service
details: [omada](../docker/omada/README.md), [pihole](../docker/pihole/README.md),
[npm](../docker/npm/README.md), [homeassistant](../docker/homeassistant/README.md).
All of the above was done by hand. Turning it into an Ansible playbook is
tracked in [todo.md](todo.md).

View file

@ -1,6 +1,15 @@
# Migrating the Omada Controller: desktop → NUC # Migrating the Omada Controller: desktop → homeserver
Once the NUC homeserver is provisioned and set up with Docker (via Ansible), > **Done 2026-10-04** — the controller now runs on the OptiPlex
> (`192.168.30.20`, see [homeserver.md](homeserver.md)). What actually
> happened: backup exported on the desktop controller (6.3.0.45), restored in
> the new controller's setup wizard (same version), then the controller's
> built-in **device migration** pointed all five devices at `192.168.30.20` —
> no device reboots needed. The desktop instance was stopped with
> `docker compose down`. The guide below is kept for reference (e.g. a future
> hardware swap).
Once a new homeserver is provisioned and set up with Docker,
move the Omada Software Controller there permanently so it doesn't depend on move the Omada Software Controller there permanently so it doesn't depend on
the desktop being on. This is a **backup/restore** move, not a live migration — the desktop being on. This is a **backup/restore** move, not a live migration —
plan for a few minutes of controller downtime (the ER605 itself keeps routing plan for a few minutes of controller downtime (the ER605 itself keeps routing
@ -9,25 +18,25 @@ traffic the whole time; only controller management is briefly unavailable).
## Before you start ## Before you start
- **Note the exact controller version** running on the desktop (`Settings → About` or the version shown in the container image tag, e.g. `mbentley/omada-controller:6.3.0.x`). - **Note the exact controller version** running on the desktop (`Settings → About` or the version shown in the container image tag, e.g. `mbentley/omada-controller:6.3.0.x`).
- The NUC's Compose setup **must run the same major.minor(.patch) version** as the desktop at restore time. TP-Link's controller cannot restore a backup taken by a newer version into an older one — it corrupts the database. If you want to upgrade, upgrade *after* the restore succeeds on matching versions, not before. - The new host's Compose setup **must run the same major.minor(.patch) version** as the desktop at restore time. TP-Link's controller cannot restore a backup taken by a newer version into an older one — it corrupts the database. If you want to upgrade, upgrade *after* the restore succeeds on matching versions, not before.
- The NUC needs to be reachable from the ER605 for adoption to keep working post-migration — either on the same L2 segment, or with the gateway's controller "inform URL" pointed at the NUC's address (Omada gateways support setting this manually under the standalone/local management settings if they're ever on different segments). - The new host needs to be reachable from the ER605 for adoption to keep working post-migration — either on the same L2 segment, or with the gateway's controller "inform URL" pointed at the new host's address (Omada gateways support setting this manually under the standalone/local management settings if they're ever on different segments).
## Steps ## Steps
1. **On the desktop controller:** `Settings → Maintenance → Backup` → create a manual backup (or use Auto Backup if one's already scheduled) and download the backup file. 1. **On the desktop controller:** `Settings → Maintenance → Backup` → create a manual backup (or use Auto Backup if one's already scheduled) and download the backup file.
2. **On the NUC:** copy [`../docker/omada-controller/docker-compose.yml`](../docker/omada-controller/docker-compose.yml) over via the Ansible-managed Docker setup, keeping the **image tag identical** to what the desktop was running. 2. **On the new host:** copy [`../docker/omada/compose.yaml`](../docker/omada/compose.yaml) over via the Ansible-managed Docker setup, keeping the **image tag identical** to what the desktop was running.
3. Bring the container up on the NUC (`docker compose up -d`), open its controller UI, and go through the setup wizard just far enough to reach the option to restore from a backup instead of creating a new site — this is usually offered right at first login (`Restore` alongside `Create New Controller`). 3. Bring the container up on the new host (`docker compose up -d`), open its controller UI, and go through the setup wizard just far enough to reach the option to restore from a backup instead of creating a new site — this is usually offered right at first login (`Restore` alongside `Create New Controller`).
4. Upload the backup file from step 1 and let it restore. The controller restarts once done. 4. Upload the backup file from step 1 and let it restore. The controller restarts once done.
5. Confirm the ER605 (and anything else already adopted) shows up and reconnects on the NUC controller. Give discovery a minute — devices need to re-inform to the new controller address. 5. Confirm the ER605 (and anything else already adopted) shows up and reconnects on the new controller. Give discovery a minute — devices need to re-inform to the new controller address.
6. Once confirmed working, stop and remove the desktop instance (`docker compose down`, and clean up the `omada-data`/`omada-logs` volumes there if you're done with them) — don't leave two controllers able to fight over the same devices. 6. Once confirmed working, stop and remove the desktop instance (`docker compose down`, and clean up the `omada-data`/`omada-logs` volumes there if you're done with them) — don't leave two controllers able to fight over the same devices.
7. If you *do* want to move to a newer controller version, do that now, as a separate step, against the NUC instance only — bump the image tag, `docker compose up -d`, and let it run its own internal DB migration. 7. If you *do* want to move to a newer controller version, do that now, as a separate step, against the new instance only — bump the image tag, `docker compose up -d`, and let it run its own internal DB migration.
## If versions don't match ## If versions don't match
If the desktop somehow ended up ahead of what you initially deploy on the NUC: If the desktop somehow ended up ahead of what you initially deploy on the new host:
1. Deploy the **exact matching version** on the NUC first (adjust the image tag). 1. Deploy the **exact matching version** on the new host first (adjust the image tag).
2. Restore the backup — this will succeed since versions match. 2. Restore the backup — this will succeed since versions match.
3. *Then* bump the NUC's image tag to the newer version and let it self-upgrade the database in place. 3. *Then* bump the new host's image tag to the newer version and let it self-upgrade the database in place.
Never try to restore a newer-version backup directly into an older controller — see the note in [status.md](status.md#omada-controller). Never try to restore a newer-version backup directly into an older controller — see the note in [status.md](status.md#omada-controller).

View file

@ -6,7 +6,7 @@ Location: Linz, Austria. ISP: LIWEST (cable). Modem: Technicolor CGA4233-EU.
## Current TODO ## Current TODO
See [todo.md](todo.md) for the live checklist. Quick status (as of 2026-09-26): bridge mode, controller adoption, firmware update, VLANs, ACLs, WireGuard VPN + INWX DynDNS, the SG2016P switch, all three SSIDs, NAS on Servers with QuickConnect disabled, and NPM/Pi-hole local naming are all done. In progress: the first Hyper Backup of the NAS. Next: second AP + living-room switch (purchase decided), office AP ceiling mount + Wi-Fi re-measure, partner's Synology app logins, hardware shopping list. Everything else is waiting on the NUC. See [todo.md](todo.md) for the live checklist. Quick status (as of 2026-10-04): bridge mode, controller adoption, firmware update, VLANs, ACLs, WireGuard VPN + INWX DynDNS, the SG2016P switch, all three SSIDs, NAS on Servers with QuickConnect disabled (both users' Synology apps on local login), the first Hyper Backup run, and the living-room EAP650 + ES205GP are all done. **The Dell OptiPlex homeserver is up** ([homeserver.md](homeserver.md)): Omada controller migrated there, Pi-hole + NPM + Home Assistant running, all web UIs on HTTPS via a Let's Encrypt wildcard for `*.home.staffenberger.at`; switches/APs moved to Management VLAN 99, PC to Trusted, Default emptied. Next: switch DHCP DNS to the new Pi-hole and turn off the NAS one, tighten the ACLs, off-box backups of the OptiPlex, Hyper Backup test restore, office AP ceiling mount + Wi-Fi re-measure.
## Hardware decisions ## Hardware decisions
@ -17,13 +17,13 @@ See [todo.md](todo.md) for the live checklist. Quick status (as of 2026-09-26):
| Router | TP-Link Omada ER605 v2.30, firmware 2.4.5 | Received, adopted, updated | | Router | TP-Link Omada ER605 v2.30, firmware 2.4.5 | Received, adopted, updated |
| Main switch | TP-Link Omada SG2016P (16-port, 8×PoE+, 120W budget) | Received, adopted, wired in; powers the office AP via PoE (make sure the AP is in an actual **PoE** port — only 8 of the 16 are) | | Main switch | TP-Link Omada SG2016P (16-port, 8×PoE+, 120W budget) | Received, adopted, wired in; powers the office AP via PoE (make sure the AP is in an actual **PoE** port — only 8 of the 16 are) |
| Access point #1 | TP-Link Omada EAP650 | Received — mounting in the **office** first, powered via its own DC adapter until the switch's PoE+ is available | | Access point #1 | TP-Link Omada EAP650 | Received — mounting in the **office** first, powered via its own DC adapter until the switch's PoE+ is available |
| Access point #2 | TP-Link Omada EAP650 (same model as AP #1) | **Decided to buy** — chosen over the EAP610 (AX1800, no 160 MHz on 5 GHz) so both APs are identical; the peak-speed difference is irrelevant to the actual couch problem, which is signal quality. Living-room/dining-area coverage, mounted above the TV (see Wi-Fi section) | | Access point #2 | TP-Link Omada EAP650 (same model as AP #1) | **Installed and configured** (living room) — chosen over the EAP610 (AX1800, no 160 MHz on 5 GHz) so both APs are identical; the peak-speed difference is irrelevant to the actual couch problem, which is signal quality. Living-room/dining-area coverage, mounted above the TV (see Wi-Fi section) |
| Living-room switch | TP-Link Omada **ES205GP** (5-port, 4× PoE+ 802.3at/af, 65W budget) | **Decided to buy** — note the **P**: the plain ES205G in earlier notes has no PoE. Easy Managed tier, but supports 802.1Q VLANs (up to 32 groups) via Omada Controller, adopted like the other devices. 3 devices need ports (AP, TV, streaming Pi), so 5 ports is plenty — no need for an 8-port unit | | Living-room switch | TP-Link Omada **ES205GP** (5-port, 4× PoE+ 802.3at/af, 65W budget) | **Installed and configured** — note the **P**: the plain ES205G in earlier notes has no PoE. Easy Managed tier, but supports 802.1Q VLANs (up to 32 groups) via Omada Controller, adopted like the other devices. 3 devices need ports (AP, TV, streaming Pi), so 5 ports is plenty — no need for an 8-port unit |
### Homeserver ### Homeserver
- **Used Intel NUC** (8th-gen Core i3, 16GB DDR4, 128GB SSD), bought via Willhaben for ~€155 — replaced the original new-Beelink-EQ14 plan. - **Dell OptiPlex 3000** (Core i3-12300T, 12th gen, 16GB DDR4, 512GB NVMe), `192.168.30.20` on Servers — the actual homeserver, set up 2026-10-04. Full setup log, hardware checks and the NVMe `pcie_aspm=off` workaround: [homeserver.md](homeserver.md). (Earlier plans: a new Beelink EQ14, then a used Intel NUC from Willhaben, 8th-gen i3/16GB/128GB, ~€155.)
- OS: plain Debian/Ubuntu + Docker Compose. **No Proxmox/hypervisor.** - OS: **Ubuntu Server 26.04 LTS** + Docker Compose, headless, SSH keys only, no disk encryption (boots unattended after power cuts). **No Proxmox/hypervisor.**
- Home Assistant via the official **"Home Assistant Container"** image (not Home Assistant OS) — matches running everything else as plain Docker. - Home Assistant via the official **"Home Assistant Container"** image (not Home Assistant OS) — matches running everything else as plain Docker.
- Config management: **Ansible**, run from David's own machine over SSH (no footprint on the server itself). - Config management: **Ansible**, run from David's own machine over SSH (no footprint on the server itself).
- **Kubernetes explicitly rejected for home services** (no benefit on a single node). Will build a separate **k3s learning cluster** later on already-owned Raspberry Pis (RPi4 and older), kept fully separate from production services — motivated by wanting K8s reps for work, not by home-service needs. - **Kubernetes explicitly rejected for home services** (no benefit on a single node). Will build a separate **k3s learning cluster** later on already-owned Raspberry Pis (RPi4 and older), kept fully separate from production services — motivated by wanting K8s reps for work, not by home-service needs.
@ -38,7 +38,7 @@ See [todo.md](todo.md) for the live checklist. Quick status (as of 2026-09-26):
### Full wired device list (needs a switch port) ### Full wired device list (needs a switch port)
Linux PC · Windows gaming PC · partner's PC · printer (optional) · Synology NAS · homeserver (NUC) · OctoPrint Pi · AP (via living-room switch) · streaming Pi (via living-room switch) Linux PC · Windows gaming PC · partner's PC · printer (optional) · Synology NAS · homeserver (OptiPlex) · OctoPrint Pi · AP (via living-room switch) · streaming Pi (via living-room switch)
## Network architecture ## Network architecture
@ -46,21 +46,42 @@ Linux PC · Windows gaming PC · partner's PC · printer (optional) · Synology
| VLAN | Subnet | Purpose | | VLAN | Subnet | Purpose |
|---|---|---| |---|---|---|
| 10 — Trusted | `192.168.10.0/24` | Personal devices, PCs, streaming Pi (avoids an extra routing hop for game streaming) | | 1 — Default | `192.168.0.0/24` | **Empty** since 2026-10-04 — kept only as the recovery/adoption network for factory-reset devices |
| 20 — IoT / Home Automation | `192.168.20.0/24` | OctoPrint + webcam, sensors, light bulbs, camera, robot vacuum, TV — internet allowed (needed for TV streaming apps + vacuum's cloud control), LAN access denied by ACL | | 10 — Trusted | `192.168.10.0/24` | Personal devices, PCs (David's PC fixed at `.10`), streaming Pi (avoids an extra routing hop for game streaming) |
| 30 — Servers / Homelab | `192.168.30.0/24` | NAS, homeserver (NUC) | | 20 — IoT / Home Automation | `192.168.20.0/24` | OctoPrint + webcam, sensors, light bulbs, camera, robot vacuum, TV, printer — internet allowed (needed for TV streaming apps + vacuum's cloud control), LAN access denied by ACL |
| 30 — Servers / Homelab | `192.168.30.0/24` | NAS (`.10`), homeserver OptiPlex (`.20`) |
| 40 — Guest | `192.168.40.0/24` | Isolated (Omada "Isolate Network" flag), internet-only | | 40 — Guest | `192.168.40.0/24` | Isolated (Omada "Isolate Network" flag), internet-only |
| 99 — Management | `192.168.99.0/24` | Switch/AP/controller control plane | | 99 — Management | `192.168.99.0/24` | Switches and APs (in use since 2026-10-04) |
**Infrastructure addresses:**
| Device | Model | Address | Fallback IP |
|---|---|---|---|
| Office – VPN Router | ER605 | Gateway in every VLAN (`.1`) | – |
| Office – 16 Port Switch | SG2016P | `192.168.99.10` (static) | – |
| Living Room – 5 Port Switch | ES205GP | `192.168.99.11` | `192.168.99.211` |
| Office – WiFi AP | EAP650 | `192.168.99.20` | `192.168.99.220` |
| Living Room – WiFi AP | EAP650 | `192.168.99.21` | `192.168.99.221` |
- Window/door/temperature sensors are **Zigbee via a USB coordinator on the homeserver**, not IP devices — they never touch VLAN 20's network; Home Assistant talks to them over the local Zigbee radio directly. - Window/door/temperature sensors are **Zigbee via a USB coordinator on the homeserver**, not IP devices — they never touch VLAN 20's network; Home Assistant talks to them over the local Zigbee radio directly.
- **Default-deny between VLANs is not automatic on Omada** — inter-VLAN routing is allow-all by default; ACL rules must be built manually (e.g., Home Assistant → IoT devices). - **Default-deny between VLANs is not automatic on Omada** — inter-VLAN routing is allow-all by default; ACL rules must be built manually (e.g., Home Assistant → IoT devices).
- ACL direction for alerts/control is **Servers → IoT** (Home Assistant reaching in to poll/control devices), not the reverse — IoT devices never need to initiate connections out to notify anyone. - ACL direction for alerts/control is **Servers → IoT** (Home Assistant reaching in to poll/control devices), not the reverse — IoT devices never need to initiate connections out to notify anyone.
- TV casting only needs **Trusted → IoT** allowed (the casting source connects out to the TV); return traffic flows back automatically as part of that established connection, no reverse rule needed. Discovery is handled by the ER605's mDNS repeater. - TV casting only needs **Trusted → IoT** allowed (the casting source connects out to the TV); return traffic flows back automatically as part of that established connection, no reverse rule needed. Discovery is handled by the ER605's mDNS repeater.
- **ACL rules built** (Omada ACL types: `Network` = exact match, `!Network` = any network except the one picked; Direction `LAN-LAN` vs `LAN-WAN` vs `WAN-IN` are separate scopes — a `LAN-LAN` rule has zero effect on internet access): - **ACL rules built** (Omada ACL types: `Network` = exact match, `!Network` = any network except the one picked; Direction `LAN-LAN` vs `LAN-WAN` vs `WAN-IN` are separate scopes — a `LAN-LAN` rule has zero effect on internet access):
- `Network: IoT` → `!Network: IoT`, Direction `LAN-LAN`, **Deny** — one rule blocks IoT from reaching every other VLAN (Trusted/Servers/Management/Guest) at once; internet access for IoT is untouched since WAN isn't in scope for a `LAN-LAN` rule. Gateway ACLs as of 2026-10-04 (LAN→LAN, evaluated top to bottom):
- `Network: Servers` → `Network: Management`, Direction `LAN-LAN`, **Deny**.
- `Network: Guest` → `Network: Management`, Direction `LAN-LAN`, **Deny** (defense-in-depth alongside Guest's Isolate Network flag). | # | Name | Policy | Source | Destination |
- Everything else (Trusted↔Servers, Trusted→Management, Servers→IoT, Trusted→IoT, all internet access) relies on Omada's allow-all-by-default behavior — no rule needed. |---|---|---|---|---|
| 1 | IoT Outwards | Deny | IoT | `!IoT` (everything except IoT) |
| 2 | Guest Outwards | Deny | Guest | `!Guest` (everything except Guest) |
| 3 | Server Outwards | Deny | Servers | Management, Guest |
| 4 | Trusted Outwards | Deny | Trusted | Guest |
- Internet access is untouched by all of these, since WAN isn't in scope for a `LAN-LAN` rule.
- Everything else (Trusted↔Servers, Trusted→Management, Servers→IoT, Trusted→IoT) relies on Omada's allow-all-by-default behavior.
- The ACLs are **stateful**: replies to allowed connections pass, so the controller on Servers reaches VLAN 99 devices without an extra rule even though rule 3 denies Servers → Management (the devices open the connection to the controller).
- Printer discovery from Trusted uses an Omada **Bonjour (mDNS) rule**, IoT → Trusted.
- Planned tightening (rule 3 also covering Trusted/Default, DNS exceptions, locking down Default/Management): see [todo.md](todo.md).
- **WAN-IN rules generally aren't needed** — NAT already blocks all unsolicited inbound traffic by default; WAN-IN only matters for scoping something you've already port-forwarded (e.g. the VPN port). - **WAN-IN rules generally aren't needed** — NAT already blocks all unsolicited inbound traffic by default; WAN-IN only matters for scoping something you've already port-forwarded (e.g. the VPN port).
- ER605 supports full per-port 802.1Q tagging/untagging/PVID; no practical VLAN-count limit for this setup. - ER605 supports full per-port 802.1Q tagging/untagging/PVID; no practical VLAN-count limit for this setup.
- ER605's built-in **mDNS repeater** bridges Bonjour/mDNS discovery across VLANs (relevant for reaching OctoPrint/HA by local hostname). - ER605's built-in **mDNS repeater** bridges Bonjour/mDNS discovery across VLANs (relevant for reaching OctoPrint/HA by local hostname).
@ -76,7 +97,7 @@ Linux PC · Windows gaming PC · partner's PC · printer (optional) · Synology
### VPN ### VPN
- **WireGuard runs on the ER605 itself**, not on the homeserver — avoids extra port-forwarding + static routing that a NUC-hosted VPN would need. **Working as of firmware 2.4.5.** - **WireGuard runs on the ER605 itself**, not on the homeserver — avoids extra port-forwarding + static routing that a homeserver-hosted VPN would need. **Working as of firmware 2.4.5.**
- **ER605 firmware must be at least 2.4.5 for WireGuard VPN Server to appear as an option in the Controller at all** — on the pre-update firmware (2.3.3), WireGuard wasn't listed, and the OpenVPN server that *was* available silently never actually started (config looked correct in the Controller — enabled, user linked, no errors in Audit Logs — but the daemon never bound to its port, confirmed via `ECONNREFUSED` on local, WAN-direct, and external tests). Updating firmware fixed both. - **ER605 firmware must be at least 2.4.5 for WireGuard VPN Server to appear as an option in the Controller at all** — on the pre-update firmware (2.3.3), WireGuard wasn't listed, and the OpenVPN server that *was* available silently never actually started (config looked correct in the Controller — enabled, user linked, no errors in Audit Logs — but the daemon never bound to its port, confirmed via `ECONNREFUSED` on local, WAN-direct, and external tests). Updating firmware fixed both.
- **Local Networks scoped to Servers (30) + IoT (20)** — remote access to the NAS/self-hosted services and things like OctoPrint, not to personal devices on Trusted. (Was briefly set to Trusted-only, then discovered to actually have *every* network selected — same "field silently left on all networks" class of bug hit earlier with OpenVPN's Local Networks. Corrected to the minimal actual-need scope: Servers + IoT, explicitly excluding Management for now — no current need for it, easy to widen later.) Tunnel Mode **Split** (only the selected networks' traffic goes through the tunnel, not general browsing). - **Local Networks scoped to Servers (30) + IoT (20)** — remote access to the NAS/self-hosted services and things like OctoPrint, not to personal devices on Trusted. (Was briefly set to Trusted-only, then discovered to actually have *every* network selected — same "field silently left on all networks" class of bug hit earlier with OpenVPN's Local Networks. Corrected to the minimal actual-need scope: Servers + IoT, explicitly excluding Management for now — no current need for it, easy to widen later.) Tunnel Mode **Split** (only the selected networks' traffic goes through the tunnel, not general browsing).
- VPN client IP pool: `192.168.66.0/24` — deliberately outside all VLAN subnets, no fixed convention requirement (doesn't have to be a `10.x` range). - VPN client IP pool: `192.168.66.0/24` — deliberately outside all VLAN subnets, no fixed convention requirement (doesn't have to be a `10.x` range).
@ -96,26 +117,31 @@ Linux PC · Windows gaming PC · partner's PC · printer (optional) · Synology
### Local DNS & naming ### Local DNS & naming
- Naming scheme settled on: **`<name>.home.staffenberger.at`** — a subdomain of David's own already-owned INWX domain, not `.local` (reserved for mDNS/Bonjour, RFC 6762 — using it for manual DNS records risks conflicting with automatic mDNS resolution) and not `.home` or `.home.arpa` (the latter is the IETF-correct reserved suffix per RFC 8375, but rejected as too unfriendly for the partner to read/type). These records exist only in Pi-hole's local DNS — never published to INWX's real public DNS, so nothing is internet-exposed by using a real owned domain for the naming. - Naming scheme settled on: **`<name>.home.staffenberger.at`** — a subdomain of David's own already-owned INWX domain, not `.local` (reserved for mDNS/Bonjour, RFC 6762 — using it for manual DNS records risks conflicting with automatic mDNS resolution) and not `.home` or `.home.arpa` (the latter is the IETF-correct reserved suffix per RFC 8375, but rejected as too unfriendly for the partner to read/type). These records exist only in Pi-hole's local DNS — never published to INWX's real public DNS, so nothing is internet-exposed by using a real owned domain for the naming.
- Records added so far as individual **Local DNS Records** in Pi-hole (not yet the wildcard approach — see TODO): `pihole.home.staffenberger.at`, `nas.home.staffenberger.at`, both pointing at the NAS's Servers-VLAN IP (`192.168.30.10`). - Pi-hole now runs on the OptiPlex (`192.168.30.20`, settings imported from the NAS instance via Teleporter). Every service name is an individual **Local DNS Record** pointing at `192.168.30.20` (NPM), which forwards to the real address and port — including services on other devices such as OctoPrint. Current names: `omada`, `pihole`, `ha`, `octoprint`, plus `nas` (moved over from the NAS Pi-hole) (`.home.staffenberger.at`). DHCP still hands out the NAS Pi-hole (`192.168.30.10`) until the switch-over in todo.md.
- **Gotcha: a secondary/fallback DNS server on a VLAN's DHCP config causes intermittent resolution failures for internal-only names.** Trusted's DHCP had Pi-hole as Primary DNS but `8.8.8.8` (Google public DNS) as Secondary — client OS resolvers don't reliably always-prefer-primary (can race/round-robin), so any query that happened to hit `8.8.8.8` for an internal `.home.staffenberger.at` name failed immediately, since a public resolver has no knowledge of it. Symptom looked like "works for me, not for my partner" / inconsistent across devices. Fix: don't set a secondary DNS pointing outside Pi-hole at all — Pi-hole already forwards normal internet lookups upstream itself. - **Gotcha: a secondary/fallback DNS server on a VLAN's DHCP config causes intermittent resolution failures for internal-only names.** Trusted's DHCP had Pi-hole as Primary DNS but `8.8.8.8` (Google public DNS) as Secondary — client OS resolvers don't reliably always-prefer-primary (can race/round-robin), so any query that happened to hit `8.8.8.8` for an internal `.home.staffenberger.at` name failed immediately, since a public resolver has no knowledge of it. Symptom looked like "works for me, not for my partner" / inconsistent across devices. Fix: don't set a secondary DNS pointing outside Pi-hole at all — Pi-hole already forwards normal internet lookups upstream itself.
- Same **stale-DHCP-lease pattern hit all week applies to DNS server changes too** — a device won't pick up a newly-configured DNS server until it renews its lease (Wi-Fi toggle, `ipconfig /release`+`/renew`, etc.), same as VLAN reassignment. - Same **stale-DHCP-lease pattern hit all week applies to DNS server changes too** — a device won't pick up a newly-configured DNS server until it renews its lease (Wi-Fi toggle, `ipconfig /release`+`/renew`, etc.), same as VLAN reassignment.
### Omada Controller ### Omada Controller
- Self-hosted **Omada Software Controller** (Docker) instead of TP-Link's cloud controller — avoids a vendor cloud relay, in line with the general "no cloud relay" networking principle. - Self-hosted **Omada Software Controller** (Docker) instead of TP-Link's cloud controller — avoids a vendor cloud relay, in line with the general "no cloud relay" networking principle.
- Plan: run it temporarily on David's own Linux desktop now, adopt the ER605 immediately, then migrate to the NUC later via Omada's **Controller Migration** (backup/restore) feature once the NUC is set up. - Ran temporarily on David's desktop, **migrated to the OptiPlex on 2026-10-04** (backup → restore in the setup wizard → built-in device migration; all five devices reconnected without rebooting). Image `mbentley/omada-controller:6.3` (6.3.0.45) — never the `latest` tag, which still points at v5. See [omada-controller-migration.md](omada-controller-migration.md) and [`../docker/omada/`](../docker/omada/).
- Watch for controller version mismatches between export and import — TP-Link/Omada requires matching major.minor(.patch) versions for restore; fallback is installing a specific matching version first, restoring, then upgrading. - Watch for controller version mismatches between export and import — TP-Link/Omada requires matching major.minor(.patch) versions for restore; fallback is installing a specific matching version first, restoring, then upgrading.
### Reverse proxy ### Reverse proxy
- **Chosen: Nginx Proxy Manager (NPM)** — single GUI, single source of truth for all proxy configs, works uniformly whether the target is a Docker container on the NUC or a bare-metal service elsewhere (e.g. OctoPrint on the Pi). - **Chosen: Nginx Proxy Manager (NPM)** — single GUI, single source of truth for all proxy configs, works uniformly whether the target is a Docker container on the homeserver or a bare-metal service elsewhere (e.g. OctoPrint on the Pi). Runs on the OptiPlex on clean 80/443/81; proxy hosts and lessons learned in [`../docker/npm/`](../docker/npm/README.md).
- **TLS: Let's Encrypt wildcard for `*.home.staffenberger.at`** via NPM's DNS challenge at INWX (decided 2026-10-04, replacing the earlier `mkcert` lean). The names point at internal IPs, so only DNS validation works, and a wildcard needs it anyway. The credential is a **dedicated INWX sub-user** with minimal rights, so the "no full INWX account credentials in a script" concern is met; NPM still stores that password in plain text, so its data directory counts as a secret. Alternative kept in mind: acme-dns with a CNAME on `_acme-challenge.home`, limiting the credential to a single TXT record.
- Traefik considered and rejected for home use (label-based config fights against wanting one central place to look); will get real Traefik exposure anyway via the k3s learning cluster, where it's the default ingress controller. - Traefik considered and rejected for home use (label-based config fights against wanting one central place to look); will get real Traefik exposure anyway via the k3s learning cluster, where it's the default ingress controller.
### Switch port configuration (SG2016P) ### Switch port configuration (SG2016P)
- **Access port** = one VLAN, untagged — for every VLAN-unaware end device (PCs, NAS, printer, Pis, TV). Set the port's Native/Untagged network to the device's VLAN and leave Tagged empty; don't leave Default involved. **Trunk** = several VLANs tagged over one cable — only for infrastructure links (switch↔router, switch↔AP). - **Access port** = one VLAN, untagged — for every VLAN-unaware end device (PCs, NAS, printer, Pis, TV). Set the port's Native/Untagged network to the device's VLAN and leave Tagged empty; don't leave Default involved. **Trunk** = several VLANs tagged over one cable — only for infrastructure links (switch↔router, switch↔AP).
- **Every VLAN must be tagged on every hop** between an AP and the ER605, or clients associate to the SSID but never get an IP ("IP configuration failure"). Hit this twice: the first AP's port, then the switch's uplink to the ER605 (missing VLAN 10). The per-Network "Select Device Port" step and the per-port VLAN config are two views of the same thing; editing the **port** directly avoids the "cannot deselect the interface which selects the LAN as PVID" error. - **Every VLAN must be tagged on every hop** between an AP and the ER605, or clients associate to the SSID but never get an IP ("IP configuration failure"). Hit this twice: the first AP's port, then the switch's uplink to the ER605 (missing VLAN 10). The per-Network "Select Device Port" step and the per-port VLAN config are two views of the same thing; editing the **port** directly avoids the "cannot deselect the interface which selects the LAN as PVID" error.
- The AP's port native/untagged VLAN and the switch/AP management traffic deliberately stay on **Default** for now (same L2 segment as the Controller on David's desktop) — migrate management to VLAN 99 together with the Controller move to the NUC, at which point Default can be retired. - **Management moved to VLAN 99 on 2026-10-04** (together with the controller move), Default is now empty and kept only as the recovery/adoption network:
- Uplink ports carry Default untagged (native) and all other VLANs tagged. VLAN 99 stays **tagged**, because the devices tag their own management traffic.
- APs: Config → IP Settings (network Management, fixed IP, fallback IP and gateway) **plus** the separate **Management VLAN** setting (Custom → Management). The IP Settings network alone does not change the VLAN.
- Switches: Config → Interface → VLAN 99 interface with Management VLAN enabled. The old VLAN 1 interface was disabled only after the new address answered.
- The ER605 stays as it is — it has an address in every VLAN.
- Extra safety net: Guest VLAN also enabled on the currently-used ports so anything unexpectedly plugged in lands isolated. - Extra safety net: Guest VLAN also enabled on the currently-used ports so anything unexpectedly plugged in lands isolated.
- After any port/VLAN or DHCP-scope change, a device keeps its old lease until it renews: toggle the port off/on in the Controller (equivalent to unplugging), `ipconfig /release`+`/renew` on Windows, or Wi-Fi off/on on phones. - After any port/VLAN or DHCP-scope change, a device keeps its old lease until it renews: toggle the port off/on in the Controller (equivalent to unplugging), `ipconfig /release`+`/renew` on Windows, or Wi-Fi off/on on phones.
- The `Network` column in the Controller's Clients list is blank for wireless clients — cosmetic; the client's IP range is the real check of which VLAN it landed in. - The `Network` column in the Controller's Clients list is blank for wireless clients — cosmetic; the client's IP range is the real check of which VLAN it landed in.
@ -123,7 +149,8 @@ Linux PC · Windows gaming PC · partner's PC · printer (optional) · Synology
## Backups ## Backups
- **Omada Controller**: auto-backup enabled, **daily** for now (config is changing a lot; switch to weekly later). Local only on the Controller host for now — backups live in the container's data volume, so `docker compose down -v` deletes them. Off-machine rsync copy to a dedicated, quota-limited Synology user is deferred until the NUC migration (see todo.md). Backups contain secrets (Device Account, WireGuard keys, DDNS settings) — never commit them. - **Omada Controller**: auto-backup enabled, **daily** for now (config is changing a lot; switch to weekly later). Local only on the OptiPlex for now, in `~/omada/data/autobackup` (bind mount). Export a manual backup after every bigger change. Off-box copy is the next step — see the OptiPlex backup item in todo.md.
- **OptiPlex** (`~/omada`, `~/pihole`, `~/npm`, `~/homeassistant`): **no off-box backup yet**. Must be encrypted, because NPM's data contains the INWX sub-user password. Backups contain secrets (Device Account, WireGuard keys, DDNS settings) — never commit them.
- **Synology NAS → 6TB USB drive via Hyper Backup**, deliberately **manually connected** (offline drive also guards against ransomware/accidental deletion), cadence **every two weeks**, recurring calendar reminder. Drive arrived as **exFAT** (DSM 7.3 mounts exFAT natively — the "exFAT Access" package no longer appears in Package Center — but exFAT has no journal, so an unclean unplug leaves a dirty flag and DSM reports "not ejected safely"); reformatted to **ext4** via DSM (Control Panel → External Devices → Format — select the device row and Eject first if Format is greyed out). Hyper Backup has no plug-in trigger on DSM 7, so it's "plug in → Back up now → **eject in DSM** → unplug". - **Synology NAS → 6TB USB drive via Hyper Backup**, deliberately **manually connected** (offline drive also guards against ransomware/accidental deletion), cadence **every two weeks**, recurring calendar reminder. Drive arrived as **exFAT** (DSM 7.3 mounts exFAT natively — the "exFAT Access" package no longer appears in Package Center — but exFAT has no journal, so an unclean unplug leaves a dirty flag and DSM reports "not ejected safely"); reformatted to **ext4** via DSM (Control Panel → External Devices → Format — select the device row and Eject first if Format is greyed out). Hyper Backup has no plug-in trigger on DSM 7, so it's "plug in → Back up now → **eject in DSM** → unplug".
- Wizard order is **destination first, then sources**. Backup type **Multiple versions**, **Folders and Packages** (not LUNs). - Wizard order is **destination first, then sources**. Backup type **Multiple versions**, **Folders and Packages** (not LUNs).
- Sources: all shared folders (incl. `photo`, Drive team folders, `docker`) + `homes` (~200 GB) + all applications (Synology Drive Server and Photos matter most — they carry the version history / albums / metadata that plain file copies lose), config backup on, **client-side encryption on** (password in Enpass), no schedule, Smart Recycle retention. - Sources: all shared folders (incl. `photo`, Drive team folders, `docker`) + `homes` (~200 GB) + all applications (Synology Drive Server and Photos matter most — they carry the version history / albums / metadata that plain file copies lose), config backup on, **client-side encryption on** (password in Enpass), no schedule, Smart Recycle retention.
@ -142,12 +169,19 @@ Linux PC · Windows gaming PC · partner's PC · printer (optional) · Synology
- Sensors: temperature/humidity, window-open monitoring, light bulb monitoring — via Zigbee-style sensors + a USB coordinator dongle (e.g. Sonoff Zigbee 3.0, ~€15–20) plugged into the homeserver. - Sensors: temperature/humidity, window-open monitoring, light bulb monitoring — via Zigbee-style sensors + a USB coordinator dongle (e.g. Sonoff Zigbee 3.0, ~€15–20) plugged into the homeserver.
- Camera to check on the cat: casual live viewing is computationally free. If smart detection (Frigate) is wanted later, add a Coral USB accelerator (~€70) rather than upgrading the CPU. - Camera to check on the cat: casual live viewing is computationally free. If smart detection (Frigate) is wanted later, add a Coral USB accelerator (~€70) rather than upgrading the CPU.
## Planned services on the homeserver (NUC) ## Services on the homeserver (OptiPlex)
- Home Assistant (Container image) Running (see [homeserver.md](homeserver.md) and [`../docker/`](../docker/)):
- Pi-hole
- Filament spool manager (Docker) - Omada Software Controller (migrated from the desktop 2026-10-04)
- Omada Software Controller (Docker — migrated from the temporary desktop instance) - Pi-hole (replacing the NAS instance)
- Nginx Proxy Manager - Nginx Proxy Manager
- Self-written Docker tools (future) - Home Assistant (Container image)
- All managed via Ansible playbooks
Planned:
- Portainer — behind HTTPS, for viewing and restarts only; the compose files stay the source of truth
- Vaultwarden, with backups to the NAS
- Filament spool manager
- Self-written Docker tools
- Everything set up by hand so far; bring it under Ansible (see todo.md)

View file

@ -12,51 +12,85 @@ Living checklist. See [status.md](status.md) for the full decisions/context behi
- [x] Guest VLAN enabled as a safety net on currently-in-use switch ports, so anything unexpectedly plugged in lands isolated by default. - [x] Guest VLAN enabled as a safety net on currently-in-use switch ports, so anything unexpectedly plugged in lands isolated by default.
- [x] **Synology NAS** → Servers (30), DHCP-reserved. - [x] **Synology NAS** → Servers (30), DHCP-reserved.
- [x] **Windows gaming PC** and **partner's PC** both on Trusted (the uplink-port VLAN bug above was what had blocked them). - [x] **Windows gaming PC** and **partner's PC** both on Trusted (the uplink-port VLAN bug above was what had blocked them).
- [ ] Still to wire/move: homeserver (NUC), blocked until the NUC itself is provisioned. David's own Linux desktop deliberately stays on Default for now (it hosts the Omada Controller) — moves to Trusted once the Controller moves to the NUC, which is also when the Default network can finally be retired. - [x] Homeserver (OptiPlex) on Servers at `192.168.30.20`; David's Linux desktop moved to Trusted (fixed `192.168.10.10`) after the controller left it; switches/APs moved to Management VLAN 99; Default is now empty (2026-10-04).
## Homeserver (OptiPlex) — follow-ups
Set up 2026-10-04, see [homeserver.md](homeserver.md).
**Network**
- [ ] **Switch DHCP DNS** in Omada from `192.168.30.10` (NAS Pi-hole) to `192.168.30.20` on every VLAN that uses Pi-hole, renew leases, then **turn off the NAS Pi-hole**. Keep no secondary DNS outside Pi-hole (see status.md, Local DNS).
- [ ] **Narrow ACL exceptions for DNS** to Pi-hole (`192.168.30.20`, port 53) from IoT and Guest, placed above rules 1 and 2. Replaces the old "should IoT use Pi-hole" question below.
- [ ] Add **Trusted and Default** to the destinations of rule 3 (Server Outwards).
- [ ] **Lock down the Default network**: allow only the controller ports 29810–29817 to the OptiPlex, block internet.
- [ ] Optional: restrict Management to controller, DNS and internet only.
- [ ] Check the **DDNS and LAN DNS settings** in the controller after the migration.
- [ ] Optional: NPM proxy host for NPM's own admin UI.
**Server**
- [ ] **Off-box backups** of `~/omada`, `~/pihole`, `~/npm`, `~/homeassistant` — encrypted, since NPM's data contains the INWX password. Supersedes the Omada-only rsync plan under "Later".
- [ ] **Bring the OptiPlex under Ansible**: SSH hardening, `resolved` stub-listener drop-in, GRUB `pcie_aspm=off`, Docker install, compose stacks from [`../docker/`](../docker/). Everything was done by hand so far.
- [ ] Cupboard cooling (fan and vent) — the CPU hits 88 °C under stress and the NVMe needs its thermal pad.
- [ ] Optional: replace the NVMe (~€30–40, 500 GB) and drop `pcie_aspm=off`.
**Services**
- [ ] Portainer (behind HTTPS; for viewing and restarts only, compose files stay the source of truth).
- [ ] Vaultwarden, with backups to the NAS.
- [ ] Own Docker tools.
**Home Assistant**
- [ ] Fix login through `ha.home.staffenberger.at` — first check Websockets Support on the NPM proxy host.
- [ ] Zigbee dongle and IKEA lights.
- [ ] ACL exceptions so Home Assistant can reach IoT devices (Servers → IoT is allowed today; check the new rules don't break that).
- [ ] Synology DSM integration (separate DSM user).
- [ ] Cat camera.
## Next up ## Next up
- [ ] **Hardware shopping list**: **EAP650** (2nd AP) + **ES205GP** (living-room PoE switch); **flat Cat6 patch cables** for the cupboard (deleyCON flat U/UTP, pure copper — 25 cm 5-pack for short links, plus 0.5 m / 1 m / 2 m, and a **3 m** for the PC runs *if* a string measurement of the real path is over ~1.7 m; verify on each listing that it says copper and is the flat U/UTP version); optionally a colored multi-length **round** Cat6 set for grab-and-go spares (flat only matters for the permanent cupboard cables); **socket strips** — two, each plugged into a *different* wall outlet, no daisy-chaining, wall-mountable metal-housing (screws, no adhesive), VDE/GS/ÖVE mark, 3×1.5 mm² cable, surge-protected for the desk/PC strip, and no switch (or a guarded one) for the strip carrying router/switch/NAS. **PoE-powered cables (to the APs) should be round pure-copper cable, not flat/thin ones.** - [ ] **Hardware shopping list**: ~~**EAP650** (2nd AP) + **ES205GP** (living-room PoE switch)~~ (bought and installed); **flat Cat6 patch cables** for the cupboard (deleyCON flat U/UTP, pure copper — 25 cm 5-pack for short links, plus 0.5 m / 1 m / 2 m, and a **3 m** for the PC runs *if* a string measurement of the real path is over ~1.7 m; verify on each listing that it says copper and is the flat U/UTP version); optionally a colored multi-length **round** Cat6 set for grab-and-go spares (flat only matters for the permanent cupboard cables); **socket strips** — two, each plugged into a *different* wall outlet, no daisy-chaining, wall-mountable metal-housing (screws, no adhesive), VDE/GS/ÖVE mark, 3×1.5 mm² cable, surge-protected for the desk/PC strip, and no switch (or a guarded one) for the strip carrying router/switch/NAS. **PoE-powered cables (to the APs) should be round pure-copper cable, not flat/thin ones.**
- [ ] **Living-room second AP + switch** (decided — see status.md Wi-Fi and Living room sections): buy an **EAP650** (same as the first AP) and an **ES205GP** (the PoE+ variant, not the plain ES205G). Mount both above the TV, one cable run from the wall to there. Then: adopt both, set the ES205GP uplink + AP ports as trunks (Trusted/IoT/Guest tagged) and the SG2016P port feeding it likewise, TV port → access on IoT, Pi port → access on Trusted. Re-measure the couch afterward (baseline was **-70 dBm**). - [x] **Living-room second AP + switch** — EAP650 + ES205GP bought, mounted and configured (as of 2026-09-29). Still open: **re-measure the couch** (baseline **-70 dBm**). Original plan (see status.md Wi-Fi and Living room sections): buy an **EAP650** (same as the first AP) and an **ES205GP** (the PoE+ variant, not the plain ES205G). Mount both above the TV, one cable run from the wall to there. Then: adopt both, set the ES205GP uplink + AP ports as trunks (Trusted/IoT/Guest tagged) and the SG2016P port feeding it likewise, TV port → access on IoT, Pi port → access on Trusted. Re-measure the couch afterward (baseline was **-70 dBm**).
- [ ] **Mount the office AP at its final ceiling position**, then re-measure the couch and the other rooms (baseline readings recorded in status.md) — cheap to do first, and useful as the before/after reference for the second AP. - [ ] **Mount the office AP at its final ceiling position**, then re-measure the couch and the other rooms (baseline readings recorded in status.md) — cheap to do first, and useful as the before/after reference for the second AP.
- [x] **Omada Controller auto-backup enabled, set to daily** (config is changing a lot right now). Weekly switch tracked separately in Backlog. Original notes: top priority, cheap insurance for all the VLAN/ACL/VPN config already built. **Scope for now: local only** — backups stay on the machine running the Controller (accepted for now); also download a manual copy occasionally. The off-machine copy to the Synology is deferred, see "Later". Note auto-backups live in the container's data volume, so `docker compose down -v` deletes them along with everything else. - [x] **Omada Controller auto-backup enabled, set to daily** (config is changing a lot right now). Weekly switch tracked separately in Backlog. Original notes: top priority, cheap insurance for all the VLAN/ACL/VPN config already built. **Scope for now: local only** — backups stay on the machine running the Controller (accepted for now); also download a manual copy occasionally. The off-machine copy to the Synology is deferred, see "Later". Note auto-backups live in the container's data volume, so `docker compose down -v` deletes them along with everything else.
- [x] **Robot vacuum** and **cat feeder** connected to the IoT Wi-Fi SSID, named in the Controller for easy recognition. No DHCP reservation — not needed for devices only reached via their own cloud apps. - [x] **Robot vacuum** and **cat feeder** connected to the IoT Wi-Fi SSID, named in the Controller for easy recognition. No DHCP reservation — not needed for devices only reached via their own cloud apps.
- [x] **Synology NAS access**: QuickConnect **disabled** — NAS is now only reachable via the WireGuard VPN, confirmed working. Synology Drive/Photos apps switched from QuickConnect ID to manual local-address login (David's phone done; **partner's accounts still to do — planned for tomorrow**). - [x] **Synology NAS access**: QuickConnect **disabled** — NAS is now only reachable via the WireGuard VPN, confirmed working. Synology Drive/Photos apps switched from QuickConnect ID to manual local-address login (David's phone and the partner's accounts both done).
- [ ] **Decide whether IoT devices should use Pi-hole for DNS**: the NAS (running Pi-hole as a Docker container) sits on Servers (30), same as always planned. The existing `IoT → !IoT deny` ACL rule currently blocks IoT from reaching Pi-hole too — meaning any IoT device pointed at it for DNS would fail to resolve. **Left as-is for now** (IoT just uses ER605/ISP DNS directly, no ad-blocking there). If wanted later: add a narrow exception, `Network: IoT → Network: Servers`, port 53 only, Direction `LAN-LAN`, Allow — evaluated before the broader IoT deny-all rule. - [ ] **Decide whether IoT devices should use Pi-hole for DNS** (now tracked as the DNS ACL exception under Homeserver follow-ups — tick both together): Pi-hole sits on Servers (30), same as always planned. The existing `IoT → !IoT deny` ACL rule currently blocks IoT from reaching Pi-hole too — meaning any IoT device pointed at it for DNS would fail to resolve. **Left as-is for now** (IoT just uses ER605/ISP DNS directly, no ad-blocking there). If wanted later: add a narrow exception, `Network: IoT → Network: Servers`, port 53 only, Direction `LAN-LAN`, Allow — evaluated before the broader IoT deny-all rule.
- [ ] **Backup strategy for the NUC/homeserver's own data** — Home Assistant config, Docker volumes, Omada Controller data, etc. Not addressed anywhere yet; the Synology's own backups don't cover this. - [ ] **▶ Synology NAS backup — first run completed (started 2026-09-26 overnight).** Task is created (Multiple versions, Folders and Packages, all shared folders + `homes` + all applications, config backup + encryption on, key in Enpass, no schedule). **Next: (1)** ~~check the run finished~~ done, **(2) test-restore** a personal Drive file and a personal photo for both users to confirm the Drive Server / Photos application backups really cover personal data, **(3) eject in DSM and unplug**, **(4) set the every-two-weeks reminder**, then tick this item. Full setup notes are in status.md's Backups section. Original notes: (currently no backup existed at all — top data-safety item): 6TB USB drive on hand, but the **USB 3.0 Micro-B cable is misplaced** (find it, or buy a replacement). Chosen workflow: **manually connected, not permanently attached** — the offline drive also protects against ransomware/accidental deletion. Steps: format the drive as **ext4** in DSM (Control Panel → External Devices → Format; erases it — check for existing data first), install **Hyper Backup**, create a Data backup task ("Local folder & USB") covering the important shared folders + Applications, **encryption on** (password in Enpass), version retention (Smart Recycle) + periodic integrity check, DSM notifications on failure, and **test-restore one folder**. To run: plug in, "Back up now", **eject in DSM** (Control Panel → External Devices → Eject) before unplugging, store the drive away from the NAS. DSM 7 has no plug-in-triggered backup, so the habit needs a **recurring calendar reminder** — chosen cadence: **every two weeks** (see how it goes; also run one before big photo imports or major DSM updates). Drive was exFAT out of the box (DSM couldn't cleanly mount it and reported "not ejected safely" — exFAT has no journal); reformatted to ext4 in DSM. The first full backup will be slow on the DS223j's 1 GB RAM — run it overnight. Later: an occasional off-site copy (3-2-1 rule) since one drive at home doesn't cover fire/theft.
- [ ] **▶ Synology NAS backup — first run in progress (2026-09-26, overnight).** Task is created (Multiple versions, Folders and Packages, all shared folders + `homes` + all applications, config backup + encryption on, key in Enpass, no schedule). **Next: (1)** check the run finished without per-folder/per-app errors, **(2) test-restore** a personal Drive file and a personal photo for both users to confirm the Drive Server / Photos application backups really cover personal data, **(3) eject in DSM and unplug**, **(4) set the every-two-weeks reminder**, then tick this item. Full setup notes are in status.md's Backups section. Original notes: (currently no backup existed at all — top data-safety item): 6TB USB drive on hand, but the **USB 3.0 Micro-B cable is misplaced** (find it, or buy a replacement). Chosen workflow: **manually connected, not permanently attached** — the offline drive also protects against ransomware/accidental deletion. Steps: format the drive as **ext4** in DSM (Control Panel → External Devices → Format; erases it — check for existing data first), install **Hyper Backup**, create a Data backup task ("Local folder & USB") covering the important shared folders + Applications, **encryption on** (password in Enpass), version retention (Smart Recycle) + periodic integrity check, DSM notifications on failure, and **test-restore one folder**. To run: plug in, "Back up now", **eject in DSM** (Control Panel → External Devices → Eject) before unplugging, store the drive away from the NAS. DSM 7 has no plug-in-triggered backup, so the habit needs a **recurring calendar reminder** — chosen cadence: **every two weeks** (see how it goes; also run one before big photo imports or major DSM updates). Drive was exFAT out of the box (DSM couldn't cleanly mount it and reported "not ejected safely" — exFAT has no journal); reformatted to ext4 in DSM. The first full backup will be slow on the DS223j's 1 GB RAM — run it overnight. Later: an occasional off-site copy (3-2-1 rule) since one drive at home doesn't cover fire/theft.
- [ ] **Real (non-self-signed) certificate for the NAS's own DSM access** — not crucial, DSM's cert warning is just cosmetic for local access. Leaning toward **`mkcert`** (local CA, install its root cert as trusted on your own devices once, zero external credentials/services involved) over Let's Encrypt DNS-01 through INWX — the latter would need either a scoped INWX API key (check if INWX offers one) or `acme.sh`'s manual mode (no credentials, but manual renewal every ~90 days); full INWX account credentials handed to a script was correctly ruled out as too broad a permission grant for this. - [ ] **Real (non-self-signed) certificate for the NAS's own DSM access** — not crucial, DSM's cert warning is just cosmetic for local access. **Easiest now:** a `nas.home.staffenberger.at` proxy host in NPM (Pi-hole record → `192.168.30.20`) reuses the existing wildcard cert, no extra credentials. Older notes: Leaning toward **`mkcert`** (local CA, install its root cert as trusted on your own devices once, zero external credentials/services involved) over Let's Encrypt DNS-01 through INWX — the latter would need either a scoped INWX API key (check if INWX offers one) or `acme.sh`'s manual mode (no credentials, but manual renewal every ~90 days); full INWX account credentials handed to a script was correctly ruled out as too broad a permission grant for this.
## Ansible — learning, started 2026-09-27 ## Ansible — learning, started 2026-09-27
Setup lives in [`../ansible/`](../ansible/README.md). OctoPrint Pi is the guinea pig; the NUC is the real goal (status.md: Ansible from David's PC, no footprint on the server). Setup lives in [`../ansible/`](../ansible/README.md). OctoPrint Pi is the guinea pig; the OptiPlex homeserver is the real goal (status.md: Ansible from David's PC, no footprint on the server).
- [x] SSH key auth to the Pi: one key per *client device* (`id_ed25519_homelab`), not per server; `~/.ssh/config` entry with `IdentitiesOnly yes`. - [x] SSH key auth to the Pi: one key per *client device* (`id_ed25519_homelab`), not per server; `~/.ssh/config` entry with `IdentitiesOnly yes`.
- [x] `ansible octoprint -m ping` works; `update.yml` run for real on 2026-09-27 (248 packages, kernel → 6.12.109). Gotchas: David's sudo needs a password → run with `-K`; Raspberry Pi OS does **not** create `/var/run/reboot-required` after kernel updates, so the playbook compares the running kernel with the newest installed one of the same flavour. - [x] `ansible octoprint -m ping` works; `update.yml` run for real on 2026-09-27 (248 packages, kernel → 6.12.109). Gotchas: David's sudo needs a password → run with `-K`; Raspberry Pi OS does **not** create `/var/run/reboot-required` after kernel updates, so the playbook compares the running kernel with the newest installed one of the same flavour.
- [ ] **Audit the Pi's hand-made changes** so a rebuild can reproduce them: `apt-mark showmanual`, enabled services, crontabs, `/boot/firmware/config.txt`, installed OctoPrint plugins. **Removed by hand 2026-09-27** (one-off cleanups don't belong in the config playbook — it describes what *should* be there, and an "ensure absent" task would fight future experiments): **Docker** (installed Sep 2026 for a Spoolman test, Spoolman already gone — packages, `docker.list` repo + `docker.asc` key, `/var/lib/docker`, `/var/lib/containerd`, `docker` group) and **nginx** (only the default site, never listening — `haproxy` holds 80/443 and fronts OctoPrint on `127.0.0.1:5000`). Rule going forward: experiment by hand, then either add it to the playbook or remove it; spool-manager-type services belong on the NUC anyway. The Pi 4 (**1 GB RAM**) boots the 32-bit `rpi-v7` kernel instead of the Bookworm default (64-bit `v8` kernel, 32-bit userland) — `config.txt` has an explicit `arm_64bit=0`, which appears to come with the OctoPi image (apt history shows the Pi is essentially a stock OctoPi flash, not years of in-place upgrades). **Decided: leave it.** With 1 GB there's no RAM to gain, and 32-bit userland uses slightly less memory; switching only the kernel isn't worth it before the rebuild. For the rebuild: a **32-bit** OctoPi image is fine on this Pi, using whatever kernel it boots by default. Still wanted from the audit: `config.txt`, enabled services, apt history. - [ ] **Audit the Pi's hand-made changes** so a rebuild can reproduce them: `apt-mark showmanual`, enabled services, crontabs, `/boot/firmware/config.txt`, installed OctoPrint plugins. **Removed by hand 2026-09-27** (one-off cleanups don't belong in the config playbook — it describes what *should* be there, and an "ensure absent" task would fight future experiments): **Docker** (installed Sep 2026 for a Spoolman test, Spoolman already gone — packages, `docker.list` repo + `docker.asc` key, `/var/lib/docker`, `/var/lib/containerd`, `docker` group) and **nginx** (only the default site, never listening — `haproxy` holds 80/443 and fronts OctoPrint on `127.0.0.1:5000`). Rule going forward: experiment by hand, then either add it to the playbook or remove it; spool-manager-type services belong on the homeserver anyway. The Pi 4 (**1 GB RAM**) boots the 32-bit `rpi-v7` kernel instead of the Bookworm default (64-bit `v8` kernel, 32-bit userland) — `config.txt` has an explicit `arm_64bit=0`, which appears to come with the OctoPi image (apt history shows the Pi is essentially a stock OctoPi flash, not years of in-place upgrades). **Decided: leave it.** With 1 GB there's no RAM to gain, and 32-bit userland uses slightly less memory; switching only the kernel isn't worth it before the rebuild. For the rebuild: a **32-bit** OctoPi image is fine on this Pi, using whatever kernel it boots by default. Still wanted from the audit: `config.txt`, enabled services, apt history.
- [x] First *configuration* playbook `playbooks/octoprint.yml` (run 2026-09-27; verified: password SSH refused, timezone set, no Bluetooth device): timezone, ModemManager masked (probes the printer's serial port), Bluetooth off (services + `dtoverlay=disable-bt`), SSH password login off. Wi-Fi stays on — the Pi is on the IoT **Wi-Fi** now, no longer on the cable. nginx removal checked: webcam uses MJPEG (`webcamd`), not the HLS stream (`ffmpeg_hls`, which needed nginx). - [x] First *configuration* playbook `playbooks/octoprint.yml` (run 2026-09-27; verified: password SSH refused, timezone set, no Bluetooth device): timezone, ModemManager masked (probes the printer's serial port), Bluetooth off (services + `dtoverlay=disable-bt`), SSH password login off. Wi-Fi stays on — the Pi is on the IoT **Wi-Fi** now, no longer on the cable. nginx removal checked: webcam uses MJPEG (`webcamd`), not the HLS stream (`ffmpeg_hls`, which needed nginx).
- [ ] **Rebuild procedure** (document once tested): flash current 32-bit OctoPi with Raspberry Pi Imager's OS customisation — hostname, user `david`, IoT Wi-Fi, SSH **public-key only** with `id_ed25519_homelab.pub` (Wi-Fi password stays out of the repo) → `ansible-playbook playbooks/octoprint.yml -K` → restore the OctoPrint backup. Test on a spare SD card. - [ ] **Rebuild procedure** (document once tested): flash current 32-bit OctoPi with Raspberry Pi Imager's OS customisation — hostname, user `david`, IoT Wi-Fi, SSH **public-key only** with `id_ed25519_homelab.pub` (Wi-Fi password stays out of the repo) → `ansible-playbook playbooks/octoprint.yml -K` → restore the OctoPrint backup. Test on a spare SD card.
- [ ] **Automated OctoPrint backups** (deferred — basics first). Rebuild plan: flash card → run playbook → restore OctoPrint backup; OctoPrint's own state (settings, plugins, profiles) comes from its Backup & Restore zip, not Ansible (OctoPrint rewrites its own `config.yaml`, so templating it would fight the UI). Plan: - [ ] **Automated OctoPrint backups** (deferred — basics first). Rebuild plan: flash card → run playbook → restore OctoPrint backup; OctoPrint's own state (settings, plugins, profiles) comes from its Backup & Restore zip, not Ansible (OctoPrint rewrites its own `config.yaml`, so templating it would fight the UI). Plan:
1. Ansible deploys a nightly timer on the Pi running the backup CLI (`octoprint plugins backup:backup` — verify syntax/exclude flags first), excluding uploads/timelapses, keeping the last few. Alternative: the "Backup Scheduler" plugin, but that's UI config outside the repo. 1. Ansible deploys a nightly timer on the Pi running the backup CLI (`octoprint plugins backup:backup` — verify syntax/exclude flags first), excluding uploads/timelapses, keeping the last few. Alternative: the "Backup Scheduler" plugin, but that's UI config outside the repo.
2. Off-device copy must be **pulled** — the `IoT → !IoT` ACL means the Pi can't push to the NAS (keep it that way). Interim: `update.yml` takes a backup and fetches it to David's PC before upgrading. Later: nightly pull from the NUC (Servers → IoT allowed) with a **dedicated key restricted via `rrsync`** to read-only access to the backup folder. 2. Off-device copy must be **pulled** — the `IoT → !IoT` ACL means the Pi can't push to the NAS (keep it that way). Interim: `update.yml` takes a backup and fetches it to David's PC before upgrading. Later: nightly pull from the OptiPlex (Servers → IoT allowed) with a **dedicated key restricted via `rrsync`** to read-only access to the backup folder.
3. Backup zips contain user hashes + API keys — store outside this repo (decide location, e.g. NAS share). 3. Backup zips contain user hashes + API keys — store outside this repo (decide location, e.g. NAS share).
- [ ] Split into roles (`common`, `updates`) + `group_vars`; `ansible-vault` for any secrets (e.g. OctoPrint API key for a "skip while printing" check). - [ ] Split into roles (`common`, `updates`) + `group_vars`; `ansible-vault` for any secrets (e.g. OctoPrint API key for a "skip while printing" check).
- [ ] Add the two personal Linux machines (`community.general.pacman` for Arch-based ones — not fully unattended, Arch updates occasionally need manual steps). - [ ] Add the two personal Linux machines (`community.general.pacman` for Arch-based ones — not fully unattended, Arch updates occasionally need manual steps).
- [ ] Provision the NUC with it (common role + Docker + Compose stacks). - [ ] Provision the OptiPlex with it (common role + Docker + Compose stacks) — see the Homeserver follow-ups.
## Backlog / lower priority ## Backlog / lower priority
- [ ] **Change the Omada Controller auto-backup from daily to weekly** — daily is deliberate for now because so much config is changing; revisit once the switch/AP/VLAN work has settled and changes become infrequent. - [ ] **Change the Omada Controller auto-backup from daily to weekly** — daily is deliberate for now because so much config is changing; revisit once the switch/AP/VLAN work has settled and changes become infrequent.
- [ ] UPS/power protection for the always-on gear (router/switch/homeserver/NAS) — optional cost/complexity tradeoff, not urgent. - [ ] UPS/power protection for the always-on gear (router/switch/homeserver/NAS) — optional cost/complexity tradeoff, not urgent.
- [ ] Basic uptime/service monitoring (e.g. Uptime Kuma) once the NUC exists — nice-to-have, not essential. - [ ] Basic uptime/service monitoring (e.g. Uptime Kuma) on the OptiPlex — nice-to-have, not essential.
## Later ## Later
- [x] **Local DNS + reverse proxy plan**: implemented ahead of schedule, running temporarily on the Synology rather than waiting for the NUC. NPM deployed (SQLite-based), remapped to ports **9080/9443/9081** since DSM's own internal nginx already held 80/443/81 — proxied HTTPS URLs need an explicit `:9443` until this moves to the NUC, where a clean 80/443/81 mapping will work with no conflict (revert `docker/npm/docker-compose.yml`'s `ports:` at that point). Pi-hole Local DNS Records in use (e.g. `pihole.home.staffenberger.at`, `nas.home.staffenberger.at`) rather than the wildcard approach so far — worth switching to the wildcard dnsmasq config later to stop needing a new record per service. - [x] **Local DNS + reverse proxy plan**: first ran temporarily on the Synology; **moved to the OptiPlex on 2026-10-04** with clean 80/443/81 ports (see [`../docker/npm/`](../docker/npm/README.md)). Original notes: NPM deployed (SQLite-based), remapped to ports **9080/9443/9081** since DSM's own internal nginx already held 80/443/81 — proxied HTTPS URLs need an explicit `:9443` until this moves to the NUC, where a clean 80/443/81 mapping will work with no conflict (revert `docker/npm/docker-compose.yml`'s `ports:` at that point). Pi-hole Local DNS Records in use (e.g. `pihole.home.staffenberger.at`, `nas.home.staffenberger.at`) rather than the wildcard approach so far — worth switching to the wildcard dnsmasq config later to stop needing a new record per service.
- [ ] **TLS certs for NPM**: same decision as the NAS's own cert (see above) — leaning `mkcert` over Let's Encrypt DNS-01/INWX for the same credential-scope reasons. Decide once NPM proxy hosts are actually being used for more than testing. - [x] **TLS certs for NPM**: decided and done 2026-10-04 — **Let's Encrypt wildcard** `*.home.staffenberger.at` via NPM's INWX DNS challenge, using a dedicated minimal-rights INWX sub-user (that resolved the credential-scope concern that had pointed toward `mkcert`). Details in [`../docker/npm/`](../docker/npm/README.md).
- [ ] **Off-machine Omada Controller backups to the Synology** — set this up on the NUC after the migration rather than on the desktop. Decided approach: a scheduled **rsync** job on the Controller host copying the Controller's autobackup folder (believed to be `/opt/tplink/EAPController/data/autobackup` inside the container — verify with `docker exec omada-controller ls` once a backup has run) to the NAS. **Not** a NAS folder mounted into the container (avoids the container's startup depending on the NAS, and keeps a failed copy from affecting the Controller). Use a **dedicated Synology user with a small storage quota** and key-based auth, so the credential on the Controller host can only touch that one backup folder — check that the DS223j's volume filesystem supports user quotas, and how DSM allows a non-admin user to do rsync/SFTP (SSH is admin-only by default). Backups contain secrets (Device Account, WireGuard keys, DDNS settings), so restrict the folder to that user and keep the files out of this Git repo. Once the Synology's own Hyper Backup to the 6TB drive exists, these are covered there too. - [ ] **Off-machine Omada Controller backups to the Synology** — superseded by the OptiPlex off-box backup item under Homeserver follow-ups (covers all four stacks); the design notes below still apply. Autobackups are at `~/omada/data/autobackup` on the host (bind mount). Original notes: Decided approach: a scheduled **rsync** job on the Controller host copying the Controller's autobackup folder (believed to be `/opt/tplink/EAPController/data/autobackup` inside the container — verify with `docker exec omada-controller ls` once a backup has run) to the NAS. **Not** a NAS folder mounted into the container (avoids the container's startup depending on the NAS, and keeps a failed copy from affecting the Controller). Use a **dedicated Synology user with a small storage quota** and key-based auth, so the credential on the Controller host can only touch that one backup folder — check that the DS223j's volume filesystem supports user quotas, and how DSM allows a non-admin user to do rsync/SFTP (SSH is admin-only by default). Backups contain secrets (Device Account, WireGuard keys, DDNS settings), so restrict the folder to that user and keep the files out of this Git repo. Once the Synology's own Hyper Backup to the 6TB drive exists, these are covered there too.
- [ ] Migrate Omada Controller from the desktop to the NUC once the NUC is provisioned — see [omada-controller-migration.md](omada-controller-migration.md). - [x] Migrate Omada Controller from the desktop to the homeserver — done 2026-10-04 (OptiPlex), see [omada-controller-migration.md](omada-controller-migration.md).
## Future ideas ## Future ideas