Homelab Update: My Own COPR, Tighter Updates, and a Rewritten Guide

It’s been a few months since I last wrote about the self-hosted stack running on my VPS — Fedora Server, MicroShift, everything deployed through Flux. The cluster has been quietly doing its job, which is the entire point, but a fair amount has changed underneath. This is the round-up: what I optimised, the package repository I ended up building myself, and why I threw away the documentation page and wrote it again from scratch.

The optimisations

None of these were dramatic. They are the kind of thing you do when a system is stable enough that you can afford to look at the parts that merely work rather than the parts that are on fire.

A second disk, added live. The Btrfs pool was at 99% data-allocated. Btrfs takes a new device into a mounted filesystem without downtime — btrfs device add, done, the UUID doesn’t change and /etc/fstab doesn’t need touching. Two things I learned afterwards: creating the same filesystem with both devices in one mkfs.btrfs call would have given me RAID1 metadata instead of DUP, and grub-btrfs quietly stops generating its snapshot menu on multi-device pools.

That second one I fixed at the source. grub-btrfs’s 41_snapshots-btrfs assumed a single root device: on a multi-device pool grub-probe --target=device / prints one device per line, all of them get passed to the follow-up --target=fs_uuid probe, and the UUID comes back empty — so the script aborts and the snapshot submenu silently stops being generated. Taking only the first device is enough, because every member of a Btrfs pool reports the same filesystem UUID. PR #440 was merged on 2026-08-24 and covers both the root and the boot probe, so this one is solved for everybody — no local patch needed any more. Note that the newest tagged release predates the merge, so you need to build from master to get it.

Pods now restart themselves when their config changes. Kubernetes doesn’t restart a pod when a mounted ConfigMap changes, and with envFrom the value never updates at all. My standing note said “remember to kubectl rollout restart”, which is exactly the kind of instruction that gets skipped. Stakater Reloader does it now. One OpenShift wrinkle: the chart defaults to runAsUser: 65534, which restricted-v2 rejects, and omitting the key in your values doesn’t help — Helm merges with chart defaults, so you have to set it to null explicitly to make it disappear.

Application-consistent etcd snapshots. My Btrfs snapshots of /var/lib/microshift are crash-consistent — a picture of the files at one instant, whatever etcd happened to be doing. That’s usually fine, and “usually” isn’t a backup strategy. A daily job now takes a real etcdctl snapshot save and verifies it in a second container, so a corrupt snapshot fails the job instead of sitting in the backup looking healthy.

A few more services: CrowdSec alongside fail2ban, self-hosted Matomo for analytics, DMARC report parsing into a Grafana dashboard, and GeoIP blocking as a native nftables set — roughly 24,000 merged ranges in one kernel-side lookup rather than tens of thousands of firewall rules.

And one workaround deleted. For months I carried a fix for kindnet’s pod CIDR: MicroShift’s own kindnet manifest hard-codes kind’s default 10.244.0.0/16 while MicroShift uses 10.42.0.0/16, and the mismatch breaks masquerading in a way that only surfaces for applications making decisions about client IPs. I’d worked around it with a boot-time patch service, then an override manifest, then a kustomizePaths exception. Then I actually read the script that generates the manifest and found it already computes the correct CLUSTER_CIDR — three lines above where it hard-codes kind’s default for kindnet. Two lines fixed it properly. The override, the exception and the service are all gone now, which is the best possible outcome for a workaround: not documented better, just deleted.

Building my own COPR

MicroShift isn’t in Fedora’s repositories. I was getting it from a community COPR project that builds nightly RPMs from the OKD/SCOS payload — until, sometime in May, that nightly moved on to the 5.x stream and stopped producing builds for the 4.y branch I run.

Nothing broke. My cluster kept running the April build perfectly happily. But dnf had nothing to offer any more, and a cluster that can’t receive fixes is a cluster on a timer.

The right fix was upstream, so I wrote it: teach the nightly workflow to build the newest 4.y release branch alongside 5.x. That’s PR #234. It’s still open — and even if it merges tomorrow it isn’t sufficient on its own, because the workflow pushes into a COPR project owned by the organisation, and creating that project needs an admin. A merged PR plus waiting-on-someone-else is not a supply chain my uptime should depend on.

So I ran the pipeline myself:

You’re welcome to enable the COPR, with the obvious caveat that it exists to serve one homelab.

The interesting differences from the upstream pipeline turned out to be small but non-obvious. The release branch can’t be a fixed matrix, because the project branches a release before the OKD payload for that stream exists — so the workflow resolves at runtime which is the newest release-4.y that actually has a payload, on both architectures. The CNI plugin build is skipped because it’s wired against epel-10 and Fedora ships the package anyway. And verification happens by enabling the fresh repo in clean fedora:43/fedora:44 containers and letting dnf resolve the whole set, rather than upstream’s bootc image test, which can’t work with Fedora RPMs.

One thing worth stating plainly, because it’s easy to assume wrongly: the GitHub Actions runner builds a source RPM and submits it. COPR’s build workers compile the binaries. My server only consumes the finished repository. Nothing is compiled on the runner, and nothing is compiled on the server.

The part that actually mattered

Getting builds back took an afternoon. Then I nearly wrecked the cluster with them.

dnf5-automatic runs on my server every morning, applies updates unattended and reboots when needed. That’s deliberate — it’s a single-node homelab, and I’d rather have security patches applied on time than perfectly scheduled. It had been harmless for months, because the upstream nightly wasn’t producing 4.y builds and dnf had nothing to install.

My own COPR produces a build every night. And the MicroShift RPM restarts crio and microshift in its %post scriptlet. So from the moment my repository went live, the arrangement was: every night, unattended, the cluster restarts onto a build nobody has looked at.

I caught it before it fired, which was luck rather than diligence.

The fix is to make the packages invisible to the updater by disabling the repositories rather than excluding package names. Repository granularity is the right one here: an exclude list has to name every package — microshift, microshift-greenboot, microshift-selinux, microshift-kindnet, plus cri-o and cri-tools from a different repository — and the failure mode of missing one is a cluster restart at 6am. Updates now run through a single script that switches the repositories on with --enablerepo for the duration of its own transaction: snapshot, one transaction, re-apply manifests, health check, pod comparison, Telegram.

Three things bit me while writing that script, and all three failed silently:

  • --allowerasing, or MicroShift is skipped. MicroShift needs cri-tools >= 1.35.0. Fedora packages cri-tools versioned, and the variants exclude each other. Without --allowerasing dnf doesn’t fail — it drops MicroShift from the transaction and installs everything else.
  • dnf install doesn’t upgrade. For an already-installed package, dnf5 reports “already installed” and moves on. Since cri-o lives in a repository I’d just disabled, nothing else was ever going to update it.
  • Disabled repositories serve stale metadata. Nothing refreshes them any more, so repoquery answers from yesterday’s cache. My script picked yesterday’s build as “the newest”, concluded it was already installed, and exited successfully having done nothing at all.

There’s a fourth, still unexplained: MicroShift applies its manifests at every start, but whether that updates existing resources is unreliable. The same server-side apply, logged as successful, took effect for one DaemonSet and not another — leaving a kube-proxy running for 86 days on an image from an RPM I’d replaced in April. My script re-applies the paths itself now.

Rewriting the guide

Which brings me to the documentation. The homelab page was written as a migration walkthrough: a changelog at the top, then nine numbered steps in the order I happened to build things. That structure made sense while the migration was happening. Fifteen months of patches later it had four changelog entries stacked above the actual content, the TLS sync script printed twice in two different outdated versions, and kindnet explained in three places — two of which described a workaround I’d since deleted.

The deeper problem was the ordering. “Step 4, Step 5, Step 6” is the sequence I built things in, which is of no use to anyone reading it — including me at seven in the morning when something is broken.

So I rewrote it from scratch, organised by layer instead:

  • The host — what exists before Kubernetes and survives it: SSH, storage, snapshots, firewall, VPN, intrusion defence
  • The cluster — MicroShift, GitOps, TLS
  • The services — the applications themselves
  • Operations — scripts, timers, and the update discipline above
  • Backup & disaster recovery — with both rollback and full-loss scenarios
  • What I got wrong — every mistake that cost me hours, in one place

The changelog is gone. It was archaeology, and the substance from it is folded into the sections it belongs to. What replaced it, and what I think is the most useful section on the page, is What I Got Wrong — thirteen entries now, and they share a shape worth naming: the system kept working, so nothing alerted, and I found out weeks later. The kube-proxy running three months on a stale image. The certificate sync script with a hard-coded namespace list. The dnf transaction silently dropping the one package it was supposed to install.

That’s the actual lesson from this round of work. Not the COPR, not the packaging — the fact that most of what went wrong here was invisible by construction, and the fix each time was to make the system tell me rather than to be more careful.