Self-Hosted Cloud on a VPS: Fedora Server + MicroShift + GitOps


A complete walkthrough of running a Kubernetes-based self-hosted stack on a single VPS — WireGuard VPN, GitOps with Flux, Btrfs snapshots, wildcard TLS, and a full application suite including Nextcloud, Vaultwarden, Paperless, a mail stack, and monitoring. Everything bootstrapped from scratch. Every step documented.

This is written for a Linux administrator who wants to replicate the setup. Personal identifiers are replaced with generic placeholders; the parts that are genuinely mine and useful to you — the COPR project, the fork, the upstream patches — are linked.

Rewritten from scratch: 2026-08-27. The first version of this page grew as a migration diary, with a changelog on top and nine numbered steps in the order I happened to build things. That order stopped being useful once the migration was over. This version is organised by layer — host, cluster, services, operations — because that is how you reason about it when something breaks at seven in the morning.


Table of Contents


What runs here

graph TD
  Internet["Internet"]
  Admin["Admin / family<br/>via WireGuard"]

  subgraph VPS["Fedora Server 43 · 8 vCPU · 16 GiB · Btrfs pool (2 devices)"]
    WG["wg0 · 10.0.0.1/24<br/>firewalld zone: internal"]

    subgraph K8S["MicroShift 4.22 · OKD/SCOS"]
      RTR["openshift-ingress<br/>HAProxy router · 80/443"]

      subgraph APPS["Applications"]
        nc["nextcloud + collabora"]
        paper["paperless"]
        vault["vaultwarden"]
        hp["homepage"]
        matomo["matomo"]
        mcp["mcp-* (3 servers)"]
      end

      mail["mailstack<br/>postfix · dovecot · rspamd · clamav<br/>hostNetwork: 25/587/993"]
      acme["acme-dns<br/>externalIP :53"]
      pihole["pihole<br/>DNS on 10.0.0.1:53"]

      subgraph INFRA["Platform"]
        cert["cert-manager"]
        flux["flux-system"]
        mon["monitoring · alloy"]
        cs["crowdsec"]
        rel["reloader"]
        etcdbk["etcd-backup"]
      end
    end
  end

  Pi["Raspberry Pi · 10.0.0.12<br/>btrbk backup target"]
  GH["GitHub<br/>homelab repo"]
  GC["Grafana Cloud"]
  LE["Let's Encrypt"]

  Internet -->|"80/443"| RTR
  Internet -->|"25/587/993"| mail
  Internet -->|"53 DNS-01"| acme
  Admin --> WG
  WG --> pihole
  WG --> RTR

  RTR --> APPS
  flux -->|"pull"| GH
  cert -->|"DNS-01"| acme
  cert -->|"ACME"| LE
  mon -->|"push"| GC
  WG --> Pi

GitOps: All Kubernetes manifests live in a GitHub repository (github.com/youruser/homelab). Flux CD watches the repo and applies changes automatically. A GitHub Action creates a Btrfs snapshot before every push.

TLS: One wildcard certificate from Let’s Encrypt covers all domains. cert-manager + acme-dns handle DNS-01 challenges. A systemd timer syncs the certificate to every app namespace weekly.

Backups: btrbk creates hourly snapshots of all Btrfs subvolumes and sends them via SSH to a Raspberry Pi. grub-btrfs registers every snapshot in the GRUB menu for easy rollback.


Why these building blocks

Standard Kubernetes (kubeadm, k3s, k0s) on a single VPS works, but MicroShift brings some specific advantages for an edge/single-node scenario:

  • Minimal footprint: Ships without the heavy components (etcd cluster, controller-manager HA). Uses CRI-O and an embedded etcd.
  • OpenShift semantics: Security Context Constraints (SCC) instead of PodSecurityAdmission. More expressive, but requires explicit SCC bindings for every workload.
  • HAProxy ingress included: openshift-ingress/router-default handles TLS termination out of the box.
  • OVN-Kubernetes by default — but we replace it with kube-kindnet (simpler, single-node appropriate).

The tradeoff is real: every service account that needs more than the default gets an explicit SCC binding, which is boilerplate you don’t write on vanilla Kubernetes. In exchange you get a permission model that says what it means, and an ingress controller you didn’t have to install.


Why Flux and not Argo CD. Flux is a set of controllers with no UI and no server component of its own. On a single node that is an advantage: less to run, less to secure, and the only interface is git — which means the audit log is git log.

Why Btrfs and not LVM + ext4. Snapshots are the entire point. Copy-on-write snapshots of the whole system, taken in under a second, bootable from the GRUB menu, and shippable to another machine as an incremental stream. That combination is what makes an unattended dnf update on a production box a reasonable thing to do.

Why WireGuard for administration. The Kubernetes API, the Pihole admin interface, kubelet — none of that is exposed publicly. The firewall’s internal zone is bound to the VPN interface, so the administrative surface simply does not exist from the internet.


Before you start

ResourceValue
OSFedora Server 43 (fresh install)
vCPUs8
RAM16 GiB
Disk≥ 512 GiB (Btrfs) — see the note on growing the pool in Layer 1
NetworkPublic IPv4, ports 22/80/443/51820 reachable
AccessRoot SSH with key auth
DNSA domain you control at a registrar (for ACME and ingress hostnames)
GitHubAccount with a private homelab repository
Grafana CloudFree tier account (for monitoring)
TelegramBot token + chat ID (for notifications — optional but recommended)

The Raspberry Pi (remote backup target) is optional but strongly recommended for disaster recovery.


Layer 1 — The Host

Everything in this layer exists before Kubernetes does, and survives it. If the cluster is broken, these are the parts that still work — which is why SSH, snapshots and the firewall are set up first.

Base system

SSH Hardening

Move SSH to a non-standard port and disable password auth:

sed -i 's/^#Port 22/Port 222/' /etc/ssh/sshd_config
grep -q "^PermitRootLogin" /etc/ssh/sshd_config \
  && sed -i 's/^PermitRootLogin.*/PermitRootLogin prohibit-password/' /etc/ssh/sshd_config \
  || echo "PermitRootLogin prohibit-password" >> /etc/ssh/sshd_config
grep -q "^PasswordAuthentication" /etc/ssh/sshd_config \
  && sed -i 's/^PasswordAuthentication.*/PasswordAuthentication no/' /etc/ssh/sshd_config \
  || echo "PasswordAuthentication no" >> /etc/ssh/sshd_config

# SELinux: allow sshd on port 222
dnf install -y policycoreutils-python-utils
semanage port -a -t ssh_port_t -p tcp 222

# Copy your authorized_keys, then:
systemctl restart sshd

Check that it took effect — mine didn’t. On an Anaconda-installed Fedora the installer drops /etc/ssh/sshd_config.d/01-permitrootlogin.conf containing PermitRootLogin yes. The main config’s Include /etc/ssh/sshd_config.d/*.conf sits near the top (line 15 here), the sed above writes at line 40, and sshd uses the first value it obtains for a keyword. The drop-in wins; the sed changes the file and nothing else. The drop-in says so itself: “Remove this file to opt-out.”

sshd -T | grep -E '^(port|permitrootlogin|passwordauthentication)'

Why this is worth fixing rather than shrugging at. Right now it changes nothing: with PasswordAuthentication no and KbdInteractiveAuthentication no, root can only get in with a key, and a password attempt is refused outright:

$ ssh -o PreferredAuthentications=password -o PubkeyAuthentication=no root@host
Permission denied (publickey,gssapi-keyex,gssapi-with-mic).

But prohibit-password disables password and keyboard-interactive login for root specifically, independent of the global setting. PermitRootLogin yes does not — it delegates root’s protection entirely to PasswordAuthentication no. The day anyone flips that to yes, temporarily, to let some user in, root becomes password-loggable in the same moment. You lose a layer without being told.

If you want the documented state, delete the drop-in and reload sshd.

Firewall exceptions for port 222 come with the firewall setup below.

etckeeper

Version-control /etc from day one:

dnf install -y etckeeper
etckeeper init
etckeeper commit "initial fedora server setup"

All subsequent /etc changes should be followed by etckeeper commit "<description>". The git log becomes your authoritative change history.

Base Packages

dnf install -y \
  vim-enhanced htop jq \
  podman btrfs-progs \
  fail2ban wireguard-tools \
  snapper btrbk \
  dnf5-plugin-automatic \
  policycoreutils-python-utils

Sysctl Tuning

cat > /etc/sysctl.d/99-microshift.conf <<'EOF'
fs.inotify.max_user_watches = 524288
fs.inotify.max_user_instances = 16384
EOF

cat > /etc/sysctl.d/90-redis.conf <<'EOF'
vm.overcommit_memory = 1
EOF

cat > /etc/sysctl.d/90-wireguard.conf <<'EOF'
net.ipv4.ip_forward = 1
EOF

sysctl --system

vm.overcommit_memory=1 is required by Redis (used by Rspamd and Nextcloud). The inotify limits are needed by MicroShift/Flux at scale.

Telegram Notifications (optional)

Create a Telegram bot via @BotFather, start a chat and note your chat ID (use @userinfobot). Store credentials in /etc/telegramrc:

install -m 600 /dev/stdin /etc/telegramrc <<'EOF'
TOKEN=<YOUR_BOT_TOKEN>
CHATID=<YOUR_CHAT_ID>
EOF

sendtelegram.sh — reads /etc/telegramrc and sends a message via the Telegram Bot API:

#!/bin/bash
# Usage: sendtelegram.sh [-c configfile] [-t token] [-i chatid] [-m message]

while getopts ":c:t:i:p:m:v" opt; do
    case "$opt" in
        c) CONFIGFILE=$OPTARG ;;  t) TOKEN_ARG=$OPTARG ;;
        i) CHATID_ARG=$OPTARG ;;  m) TEXT=$OPTARG ;;
        p) PARSEMODE_ARG=$OPTARG ;; v) VERBOSE=1 ;;
    esac
done

if [ -n "$CONFIGFILE" ]; then . "$CONFIGFILE"
elif [ -f /etc/telegramrc ]; then . /etc/telegramrc; fi

if [ -n "$TOKEN_ARG" ]; then TOKEN=$TOKEN_ARG; fi
if [ -n "$CHATID_ARG" ]; then CHATID=$CHATID_ARG; fi

URL="https://api.telegram.org/bot$TOKEN/sendMessage"
CMDARGS="chat_id=$CHATID&disable_web_page_preview=1&text=$TEXT"
[ -n "${PARSEMODE_ARG:-}" ] && CMDARGS="${CMDARGS}&parse_mode=$PARSEMODE_ARG"
curl -s --max-time 10 -d "$CMDARGS" "$URL" > /dev/null
install -m 755 sendtelegram.sh /usr/local/bin/sendtelegram.sh
sendtelegram.sh -m "Base system ready"

All other scripts call sendtelegram.sh or source /etc/telegramrc directly.

Mask passim

passim is a P2P cache for fwupd — useless on a headless VPS:

systemctl mask passim
etckeeper commit "base system: ssh/222, packages, sysctl, telegram, dnf-automatic"
snapper -c root create -d "basis-system-fertig"

Storage layout

Why Separate Subvolumes?

Granular subvolumes allow btrbk to snapshot and send only what changed. You can exclude volatile data (/var/cache, /var/tmp) from backups while still including critical data (/var/lib/microshift, PVCs, GitOps working tree).

Subvolume Creation

Fedora installs / as a root subvolume. Mount the top-level and add the rest:

mkdir -p /mnt/btrfs-top
mount -o subvolid=5 /dev/vda3 /mnt/btrfs-top

for sv in var home var_log var_cache var_tmp \
          var_lib_microshift var_lib_pvc var_lib_containers var_lib_kubelet \
          data btrbk_snapshots; do
  btrfs subvolume create /mnt/btrfs-top/$sv
done

umount /mnt/btrfs-top

/etc/fstab Extension

Get your Btrfs UUID with blkid /dev/vda3, then add to /etc/fstab:

UUID=<btrfs-uuid> /home               btrfs  subvol=home,compress=zstd:1                  0 0
UUID=<btrfs-uuid> /var                btrfs  subvol=var,compress=zstd:1                   0 0
UUID=<btrfs-uuid> /var/cache          btrfs  subvol=var_cache,compress=zstd:1             0 0
UUID=<btrfs-uuid> /var/log            btrfs  subvol=var_log,compress=zstd:1               0 0
UUID=<btrfs-uuid> /var/tmp            btrfs  subvol=var_tmp,compress=zstd:1               0 0
UUID=<btrfs-uuid> /var/lib/microshift btrfs  subvol=var_lib_microshift,compress=zstd,noatime 0 0
UUID=<btrfs-uuid> /var/lib/pvc        btrfs  subvol=var_lib_pvc,compress=zstd,noatime     0 0
UUID=<btrfs-uuid> /var/lib/containers btrfs  subvol=var_lib_containers,compress=zstd,noatime 0 0
UUID=<btrfs-uuid> /var/lib/kubelet    btrfs  subvol=var_lib_kubelet,compress=zstd,noatime 0 0
UUID=<btrfs-uuid> /data               btrfs  subvol=data,compress=zstd:1                  0 0
UUID=<btrfs-uuid> /mnt/btrfs-top      btrfs  subvolid=5,compress=zstd:1,noauto            0 0

noatime on the four /var/lib/* subvolumes matters more than it looks: etcd, the container store and the PVCs are write-heavy, and on a copy-on-write filesystem every access-time update is another write. The top-level mount is noauto — it is only needed for btrbk and for reaching snapshots directly.

mkdir -p /var/lib/{microshift,pvc,containers,kubelet} /data /mnt/btrfs-top
systemctl daemon-reload && mount -a

SELinux Context for PVCs

The local-path-provisioner init container (busybox) cannot run chcon. Set the SELinux label from the host:

semanage fcontext -a -t container_file_t "/var/lib/pvc(/.*)?"
restorecon -Rv /var/lib/pvc

SELinux context for the script directory

/data is a data subvolume, so anything you drop there inherits a label systemd refuses to execute. Every maintenance script in this guide lives in /data/scripts/, so the rule has to exist before the first one is written:

mkdir -p /data/scripts
semanage fcontext -a -t bin_t "/data/scripts(/.*)?"
restorecon -RFv /data/scripts

Without it, every systemd unit that runs one of these scripts dies after a few milliseconds with status=203/EXEC and an AVC denial in the audit log — no output, no useful error. The -F is not optional either: restorecon without it leaves a label that was once set explicitly (“customized by admin”) in place.

The same trap arrives from the other direction: podman run -v /data/scripts:/x:Z relabels the directory recursively to container_file_t with private MCS categories, silently disabling every timer that starts a script from there. Use :z (lowercase) or copy the file out instead.

# after any suspicion:
stat -c '%C %n' /data/scripts/*.sh | grep -v bin_t

Growing the Pool Later

Btrfs adds devices to a live filesystem — no migration, no reformat, the UUID stays put and /etc/fstab needs no change because it mounts by filesystem UUID:

wipefs -n /dev/vdb        # must print nothing
btrfs device add /dev/vdb /
btrfs filesystem show /   # two devids now

Two things worth knowing before you do this:

  • mkfs.btrfs with two devices is not the same thing. Given two devices in one call, mkfs defaults metadata to RAID1. Adding the second device afterwards keeps Metadata: DUP / Data: single. If you ever rebuild this filesystem from a backup, create it with the first device only and add the second afterwards, or you silently end up with a different layout.
  • grub-btrfs used to break on multi-device pools — fixed upstream. 41_snapshots-btrfs assumed one root device; on a pool with two, grub-probe --target=device / returns one device per line, all of them reach the follow-up --target=fs_uuid probe, and the UUID comes back empty. The script aborts, the snapshot submenu stops being generated, and the only symptom is a generic error in the daemon log. Taking the first device is sufficient — every member of the pool reports the same filesystem UUID. PR #440 (merged 2026-08-24) does exactly that, for both the root and the boot probe. Build from master: the newest tagged release is older than the merge.

⚠️ Do NOT Enable Btrfs Quotas

This is a hard-won lesson: Btrfs quotas and etcd are incompatible.

When quotas are enabled, every snapshot deletion (btrbk and snapper-cleanup run hourly/daily) triggers a qgroup rescan. With hundreds of snapshots across many subvolumes, this rescan can block Btrfs metadata operations — including fsync — for up to 15 minutes.

etcd fails with DeadlineExceeded if fsync stalls for more than 5 seconds → kube-apiserver hangs → MicroShift crashes. The MicroShift restart loop persists until the rescan completes, then crashes again at the next btrbk run.

Symptom: MicroShift crashes at the same time every hour, exactly 15 minutes after btrbk runs. journalctl -u microshift -n 100 shows etcd fsync errors.

Fix: btrfs quota disable /

btrfs-list (a useful subvolume overview tool) works without quotas — the REFER/EXCL columns show empty, but all other data is present.

Why Btrfs has this problem (and ZFS doesn’t): ZFS encodes block ownership in the block itself at write time (birth_txg). Space accounting is O(1) per write, always consistent, and snapshot deletion is processed asynchronously via a per-snapshot “deadlist” — no global rescan, never blocking. Btrfs introduced snapshots without built-in ownership metadata, so qgroups must walk backreferences retroactively to determine who owns what — expensive and blocking.

What about simple quotas (squota)? Since kernel 6.7, Btrfs offers an alternative mode (btrfs quota enable --simple /) that attributes extents permanently to their creating subvolume — no backref walking, no rescan, O(1) per operation. This is safe for etcd. The trade-off: snapshots show ~0 exclusive usage (all extents stay attributed to the original subvolume), so snapshot size measurement is not possible. For this setup, quotas remain disabled — squota doesn’t provide useful space accounting for snapshots either.

full qgroupsquotadisabledZFS
How ownership is determinedbackref walk (expensive)creator subvolume (O(1))—birth_txg (O(1))
When accounting happensretroactively (rescan)inline—inline, atomic
Snapshot deleteblocking rescanno rescan—async deadlist
Numbers always correctno (inconsistent flag)no (snapshots ~0)—yes
etcd-safenoyesyes—
Snapshot size measurableyesnonoyes

Snapshots and backup

Three mechanisms with three different jobs: snapper for fast local rollback, grub-btrfs to make those rollbacks reachable from the boot menu, and btrbk to get the data off the machine entirely.

snapper — Timeline Snapshots

Eight subvolumes get timeline snapshots — all with identical retention (24h/8d/5w, numbered limit 20):

# Root
snapper -c root create-config /
snapper -c root set-config \
  TIMELINE_CREATE=yes TIMELINE_CLEANUP=yes \
  TIMELINE_LIMIT_HOURLY=24 TIMELINE_LIMIT_DAILY=8 \
  TIMELINE_LIMIT_WEEKLY=5 TIMELINE_LIMIT_MONTHLY=0 TIMELINE_LIMIT_YEARLY=0 \
  NUMBER_CLEANUP=yes NUMBER_LIMIT=20

# All other subvolumes — same retention
for CFG_SUBVOL in \
  "home:/home" "data:/data" \
  "var_lib_pvc:/var/lib/pvc" "var_lib_microshift:/var/lib/microshift" \
  "var:/var" "var_log:/var/log" "var_lib_containers:/var/lib/containers"; do
  CFG="${CFG_SUBVOL%%:*}"
  SUBVOL="${CFG_SUBVOL##*:}"
  snapper -c "$CFG" create-config "$SUBVOL"
  snapper -c "$CFG" set-config \
    TIMELINE_CREATE=yes TIMELINE_CLEANUP=yes \
    TIMELINE_LIMIT_HOURLY=24 TIMELINE_LIMIT_DAILY=8 \
    TIMELINE_LIMIT_WEEKLY=5 TIMELINE_LIMIT_MONTHLY=0 TIMELINE_LIMIT_YEARLY=0 \
    NUMBER_CLEANUP=yes NUMBER_LIMIT=20
done

systemctl enable --now snapper-timeline.timer snapper-cleanup.timer

SELinux trap, and this one bites on exactly this setup. snapper create-config creates a nested .snapshots subvolume that inherits the parent’s SELinux type. For /var/lib/pvc that is container_file_t — the label we deliberately set two sections above — and snapperd is not allowed to write there. It fails with mkdir errno:13, and the config looks created but never produces a snapshot. Add a more specific rule for each .snapshots and relabel:

for d in / /home /data /var /var/log /var/lib/pvc /var/lib/microshift /var/lib/containers; do
  s="${d%/}/.snapshots"
  semanage fcontext -a -t snapperd_data_t "${s}(/.*)?" 2>/dev/null
  restorecon -RFv "$s"
done

# verify — all eight must report snapperd_data_t
stat -c '%C %n' /.snapshots /home/.snapshots /data/.snapshots /var/.snapshots \
  /var/log/.snapshots /var/lib/{pvc,microshift,containers}/.snapshots

snap-all — convenience script that creates a numbered snapshot across all eight configs at once (used before risky changes):

cat > /data/scripts/snap-all <<'EOF'
#!/bin/bash
set -euo pipefail
DESC="${1:?Verwendung: snap-all <beschreibung>}"
for cfg in root home data var_lib_pvc var_lib_microshift var var_log var_lib_containers; do
  printf "  %-22s ... " "$cfg"
  snapper -c "$cfg" create --cleanup-algorithm number --description "$DESC"
  echo "ok"
done
EOF
chmod 755 /data/scripts/snap-all
restorecon -F /data/scripts/snap-all
ln -sf /data/scripts/snap-all /usr/local/bin/snap-all

Usage: snap-all "before-risky-change" — creates one snapshot per config, all with --cleanup-algorithm number so they count against NUMBER_LIMIT and are eventually pruned automatically.

grub-btrfs — Rollback from GRUB

grub-btrfs is not in Fedora repos — build from source:

dnf install -y make gettext
git clone https://github.com/Antynea/grub-btrfs /root/git/grub-btrfs
cd /root/git/grub-btrfs && make install

# Fedora-specific paths
sed -i \
  -e 's|#GRUB_BTRFS_GRUB_DIRNAME=.*|GRUB_BTRFS_GRUB_DIRNAME="/boot/grub2"|' \
  -e 's|#GRUB_BTRFS_SCRIPT_CHECK=.*|GRUB_BTRFS_SCRIPT_CHECK=grub2-script-check|' \
  /etc/default/grub-btrfs/config

systemctl enable --now grub-btrfsd.service
grub2-mkconfig -o /boot/grub2/grub.cfg

The daemon watches /.snapshots via inotify and adds new snapshots to the GRUB menu automatically.

Cloning master rather than checking out a tag is deliberate: the multi-device fix (PR #440) is merged but not yet in a tagged release. If your pool has more than one device and you install from a tarball, the snapshot submenu will silently stop being generated — see the note in the storage section.

btrbk — Snapshots + Remote Backup

cat > /etc/btrbk/btrbk.conf <<'EOF'
timestamp_format        long
# Local: keep only the latest snapshot as parent reference for incremental send
# snapper handles local rollback, btrbk local snapshots are minimal
snapshot_preserve_min   latest
snapshot_preserve       2h 1d 0w
# Remote (Pi): full retention history
target_preserve_min     1h
target_preserve         24h 8d 5w

ssh_identity            /root/.ssh/id_ed25519

volume /mnt/btrfs-top
  snapshot_dir  btrbk_snapshots

  subvolume root
    snapshot_create  always
    target ssh://<backup-user>@10.0.0.12/backup/btrfs/server/root
  subvolume var
    snapshot_create  always
    target ssh://<backup-user>@10.0.0.12/backup/btrfs/server/var
  subvolume home
    snapshot_create  always
    target ssh://<backup-user>@10.0.0.12/backup/btrfs/server/home
  subvolume var_lib_pvc
    snapshot_create  always
    target ssh://<backup-user>@10.0.0.12/backup/btrfs/server/var_lib_pvc
  subvolume var_lib_microshift
    snapshot_create  always
    target ssh://<backup-user>@10.0.0.12/backup/btrfs/server/var_lib_microshift
  subvolume var_lib_containers
    snapshot_create  always
    target ssh://<backup-user>@10.0.0.12/backup/btrfs/server/var_lib_containers
  subvolume var_log
    snapshot_create  always
    target ssh://<backup-user>@10.0.0.12/backup/btrfs/server/var_log
  subvolume data
    snapshot_create  always
    target ssh://<backup-user>@10.0.0.12/backup/btrfs/server/data
EOF

systemctl enable --now btrbk.timer

One target per subvolume, not one for the whole volume. btrbk would accept a single volume-level target, but then everything lands in one directory and you cannot tell the streams apart. The per-subvolume layout is what the restore procedure walks: ls /backup/btrfs/server/<subvolume>/ | sort | tail -1 gives you the newest snapshot of exactly one subvolume.

The remote target (Raspberry Pi) is only reachable once WireGuard is configured. Until then, btrbk creates local snapshots only.

/boot Backup

cat > /data/scripts/boot-backup.sh <<'EOF'
#!/bin/bash
set -euo pipefail
DEST=/var/lib/boot-backup
mkdir -p "$DEST/boot" "$DEST/efi"
rsync -aAX --delete /boot/    "$DEST/boot/"
rsync -aAX --delete /boot/efi/ "$DEST/efi/"
sfdisk --dump /dev/vda        > "$DEST/partition-table.sfdisk"
sgdisk --backup="$DEST/partition-table.sgdisk" /dev/vda
EOF
chmod 755 /data/scripts/boot-backup.sh
restorecon -F /data/scripts/boot-backup.sh
ln -sf /data/scripts/boot-backup.sh /usr/local/sbin/boot-backup.sh

systemctl enable --now boot-backup.timer
# /etc/systemd/system/boot-backup.service
[Unit]
Description=Backup /boot and partition table
[Service]
Type=oneshot
ExecStart=/usr/local/sbin/boot-backup.sh
# /etc/systemd/system/boot-backup.timer
[Unit]
Description=Daily /boot and partition table backup
[Timer]
OnCalendar=daily
Persistent=true
[Install]
WantedBy=timers.target

Plus a drop-in so a backup also runs straight after every unattended update — that is the moment /boot actually changes, and the moment you most want a copy of the previous kernel:

# /etc/systemd/system/dnf5-automatic.service.d/post-backup.conf
[Service]
ExecStartPost=/usr/local/sbin/boot-backup.sh

/boot is XFS on its own partition, so Btrfs snapshots do not cover it. This rsync copy lands in /var/lib/boot-backup/, which sits on the var subvolume — and therefore is snapshotted and shipped to the Pi.


Firewall and VPN

firewalld Zones

The setup uses three zones:

ZoneInterface/SourcePurpose
FedoraServerens3 (public NIC)External traffic, explicit whitelist
internalwg0 (WireGuard)VPN clients — trusted, full access
trusted10.42.0.0/16 (Pod CIDR)Pod-to-pod + kubelet traffic
# FedoraServer zone — remove defaults, add only what's needed
firewall-cmd --permanent --zone=FedoraServer --remove-service=cockpit
firewall-cmd --permanent --zone=FedoraServer --remove-service=dhcpv6-client

# Custom SSH service on port 222
firewall-cmd --permanent --new-service=myssh
firewall-cmd --permanent --service=myssh --add-port=222/tcp
firewall-cmd --permanent --zone=FedoraServer --add-service=myssh
firewall-cmd --permanent --zone=FedoraServer --remove-service=ssh

# HTTP(S) for ingress
firewall-cmd --permanent --zone=FedoraServer --add-port=80/tcp
firewall-cmd --permanent --zone=FedoraServer --add-port=443/tcp

# Mail ports
for p in 25 587 993; do
  firewall-cmd --permanent --zone=FedoraServer --add-port=${p}/tcp
done

# WireGuard
firewall-cmd --permanent --zone=FedoraServer --add-port=51820/udp

# acme-dns (DNS-01 challenge server)
firewall-cmd --permanent --zone=FedoraServer --add-port=53/tcp
firewall-cmd --permanent --zone=FedoraServer --add-port=53/udp

# NAT for WireGuard clients
firewall-cmd --permanent --zone=FedoraServer --add-masquerade

# internal zone — wg0 interface
firewall-cmd --permanent --zone=internal --add-interface=wg0
# 6443 = Kubernetes API, 10250 = kubelet, 8000 = Paperless, 11334 = Rspamd UI,
# 888/3012 = Vaultwarden. These are direct-access ports for administration; the
# same services are also reachable through the ingress on 443.
for p in 25 53 80 222 443 587 888 993 3012 6443 8000 10250 11334; do
  firewall-cmd --permanent --zone=internal --add-port=${p}/tcp
done
firewall-cmd --permanent --zone=internal --add-port=53/udp

# trusted zone — Pod CIDR
firewall-cmd --permanent --zone=trusted --add-source=10.42.0.0/16
firewall-cmd --permanent --zone=trusted --add-source=169.254.169.1/32
firewall-cmd --permanent --zone=trusted --add-port=30000-32767/tcp

Cross-Zone Forwarding Policies

firewalld requires explicit policies for cross-zone forwarding — plain rules don’t cover forwarded packets:

# WireGuard clients → Internet (NAT)
firewall-cmd --permanent --new-policy=wg-to-internet
firewall-cmd --permanent --policy=wg-to-internet --add-ingress-zone=internal
firewall-cmd --permanent --policy=wg-to-internet --add-egress-zone=FedoraServer
firewall-cmd --permanent --policy=wg-to-internet --set-target=ACCEPT

# WireGuard clients → Pods (via HAProxy DNAT)
firewall-cmd --permanent --new-policy=wg-to-cluster
firewall-cmd --permanent --policy=wg-to-cluster --add-ingress-zone=internal
firewall-cmd --permanent --policy=wg-to-cluster --add-egress-zone=trusted
firewall-cmd --permanent --policy=wg-to-cluster --set-target=ACCEPT

firewall-cmd --reload

Critical: The --add-interface=wg0 for the internal zone must be --permanent. Without it, after a reload the policy zone assignment breaks and VPN clients lose cluster access.

WireGuard Server

cd /etc/wireguard && umask 077
wg genkey | tee server.key | wg pubkey > server.pub

/etc/wireguard/wg0.conf:

[Interface]
Address    = 10.0.0.1/24
ListenPort = 51820
PrivateKey = <CONTENTS OF server.key>

[Peer]
# Example client
PublicKey  = <client-pubkey>
AllowedIPs = 10.0.0.2/32

# Add one [Peer] block per device

Client configuration (use the WireGuard app):

  • Endpoint: yourserver.example.com:51820
  • AllowedIPs: 0.0.0.0/0 (route all traffic through VPN)
  • DNS: 10.0.0.1 (Pihole — deployed later; fall back to 1.1.1.1 initially)
chmod 600 /etc/wireguard/wg0.conf
systemctl enable --now wg-quick@wg0

No PostUp/PostDown rules. Most WireGuard guides add iptables -A FORWARD -i %i -j ACCEPT here. Don’t — forwarding is already handled by the firewalld policies above, and the internal zone carries forward: yes. Adding raw iptables rules on top means two systems managing the same chain, and the one you’ll remember to check when something breaks is the wrong one.

Intrusion defence

fail2ban

Phase 1: Base config (SSH only, before mailstack)
cat > /etc/fail2ban/jail.local <<'EOF'
[DEFAULT]
bantime  = 86400
findtime = 259200
maxretry = 3
backend  = auto
action   = nftables[type=multiport]

[sshd]
enabled  = true
port     = 222
logpath  = %(sshd_log)s
EOF

systemctl enable --now fail2ban

Why nftables and not firewallcmd-rich-rules. The obvious action on a firewalld system manages each ban as its own rich rule. That works until the ban list gets long. With ~2000 banned addresses, stopping fail2ban means ~2000 individual firewall-cmd --remove-rich-rule calls — and if firewalld was reloaded in between without fail2ban noticing, most of those rules no longer exist, so each call waits for a timeout. Mine took 90 seconds to stop and was then killed with SIGABRT.

fail2ban 1.1 ships action.d/nftables.conf, which keeps one nftables set per jail (addr-set-<jail> in table inet f2b-table). A ban is one nft add element; a stop is one set deletion regardless of size. And because inet f2b-table is a separate table from inet firewalld, a firewall-cmd --reload leaves the bans alone.

firewallcmd-ipset looks like the middle ground but is not: it goes through firewall-cmd --direct, which is deprecated in firewalld 2.x and incompatible with the nftables backend.

Phase 2: After mailstack deployment

The mail logs live inside a PVC mounted at /var/lib/pvc/<pvc-uuid>_mailstack_mail-logs/. Get the UUID:

kubectl get pvc -n mailstack mail-logs -o jsonpath='{.spec.volumeName}'

Add to /etc/fail2ban/jail.local (replace <PVC_UUID> with the value above):

[postfix]
enabled  = true
port     = smtp,submission
backend  = polling
logpath  = /var/lib/pvc/<PVC_UUID>_mailstack_mail-logs/postfix.log

[postfix-sasl]
enabled  = true
port     = smtp,submission
backend  = polling
logpath  = /var/lib/pvc/<PVC_UUID>_mailstack_mail-logs/postfix.log

[dovecot]
enabled  = true
port     = imaps
backend  = polling
logpath  = /var/lib/pvc/<PVC_UUID>_mailstack_mail-logs/dovecot.log

[postfix-sasl-unknown]
enabled        = true
filter         = postfix-sasl-unknown
action         = nftables[type=allports]
ignorecommand  = /usr/local/bin/fail2ban-sasl-mysql-check.sh <ip>
backend        = polling
logpath        = /var/lib/pvc/<PVC_UUID>_mailstack_mail-logs/postfix.log
maxretry       = 1
bantime        = 604800
findtime       = 86400

[sshd-unknown]
enabled        = true
filter         = sshd-unknown
action         = nftables[type=allports]
backend        = systemd
port           = 222
maxretry       = 1
bantime        = 604800
findtime       = 86400

[dovecot-unknown]
enabled        = true
filter         = dovecot-unknown
action         = nftables[type=allports]
backend        = polling
logpath        = /var/lib/pvc/<PVC_UUID>_mailstack_mail-logs/dovecot.log
maxretry       = 1
bantime        = 604800
findtime       = 86400

[postfix-rcpt-unknown]
enabled        = true
filter         = postfix-rcpt-unknown
action         = nftables[type=allports]
backend        = polling
logpath        = /var/lib/pvc/<PVC_UUID>_mailstack_mail-logs/postfix.log
maxretry       = 3
bantime        = 604800
findtime       = 3600

The three *-unknown jails each need their own filter in /etc/fail2ban/filter.d/ — they exist to catch attempts against accounts that do not exist, which the stock sshd/dovecot/postfix filters do not distinguish. Someone guessing usernames is not a user who mistyped a password, so they get maxretry = 1 and a week-long ban across all ports:

# /etc/fail2ban/filter.d/sshd-unknown.conf
[INCLUDES]
before = common.conf
[Definition]
_daemon = sshd(?:-session)?
prefregex = ^<F-MLFID>%(__prefix_line)s</F-MLFID><F-CONTENT>.+</F-CONTENT>$
failregex = ^[Ii]nvalid user <F-USER>\S+</F-USER> from <HOST>(?: port \d+)?\s*$
journalmatch = _SYSTEMD_UNIT=sshd.service + _COMM=sshd + _COMM=sshd-session
# /etc/fail2ban/filter.d/dovecot-unknown.conf
[INCLUDES]
before = common.conf
[Definition]
failregex = auth-worker\(<F-USER>[^,]+</F-USER>,<HOST>\)(?:<[^>]+>)*: request \[\d+\]: Info: sql: unknown user
datepattern = %%b %%d %%H:%%M:%%S
              {^LN-BEG}
# /etc/fail2ban/filter.d/postfix-rcpt-unknown.conf
[INCLUDES]
before = common.conf
[Definition]
failregex = reject: RCPT from \S+\[<HOST>\][^:]*: 550 5\.1\.1 \S+: Recipient address rejected: User unknown in virtual mailbox table
datepattern = %%b %%d %%H:%%M:%%S
              {^LN-BEG}
postfix-sasl-unknown — smart SASL jail

This custom jail bans on the first failed SASL login, but only if the attempted username does not exist in the mailbox database. Legitimate users who mistype their password from a new IP are not banned; automated scanners using random addresses are blocked immediately for 7 days.

Custom filter /etc/fail2ban/filter.d/postfix-sasl-unknown.conf:

[INCLUDES]
before = common.conf

[Definition]
failregex = warning: \S+\[<HOST>\](?::\d+)?: SASL \S+ authentication failed: .*, sasl_username=<F-USER>\S+</F-USER>
ignoreregex =
datepattern = %%b %%d %%H:%%M:%%S
              {^LN-BEG}

MySQL check script /usr/local/bin/fail2ban-sasl-mysql-check.sh (return 0 = do NOT ban, return 1 = ban):

#!/bin/bash
IP="$1"
LOG="/var/lib/pvc/<PVC_UUID>_mailstack_mail-logs/postfix.log"
source /etc/fail2ban/.mysql-sasl-check.conf

USER=$(grep -F "[$IP]" "$LOG" | grep "SASL.*authentication failed" | \
    grep -oP 'sasl_username=\K\S+' | tail -1)

[ -z "$USER" ] && exit 1
echo "$USER" | grep -qP '^[a-zA-Z0-9._%+\-]+@[a-zA-Z0-9.\-]+\.[a-zA-Z]{2,}$' || exit 1

SAFE_USER="${USER//\'/\'\'}"
COUNT=$(mysql -h "$MYSQL_HOST" -u "$MYSQL_USER" -p"$MYSQL_PASS" "$MYSQL_DB" -sN \
    -e "SELECT COUNT(*) FROM mailbox WHERE username='$SAFE_USER' AND active=1" 2>/dev/null)

[ "$COUNT" = "1" ] && exit 0 || exit 1
chmod 755 /usr/local/bin/fail2ban-sasl-mysql-check.sh

MySQL credentials in /etc/fail2ban/.mysql-sasl-check.conf (chmod 600). The MYSQL_HOST is the cluster IP of the maildb service:

# MYSQL_HOST: kubectl get svc -n mailstack maildb -o jsonpath='{.spec.clusterIP}'
install -m 600 /dev/stdin /etc/fail2ban/.mysql-sasl-check.conf <<'EOF'
MYSQL_HOST=<maildb-cluster-ip>
MYSQL_USER=postfix
MYSQL_PASS=<postfix-db-password>
MYSQL_DB=postfix
EOF

systemctl restart fail2ban
Auto-unban WireGuard clients

Family members mistype their mail password. fail2ban does its job and bans them. The address it bans is not the tunnel address 10.0.0.x — with hostNetwork: true on the mail pods, Postfix and Dovecot see the client’s real public IP, and that is what ends up in the jail. So this script walks the current WireGuard peers, takes their endpoint addresses, and unbans those:

cat > /data/scripts/unban-fail2ban-clients.sh <<'EOF'
#!/usr/bin/env bash
for ip in $( wg | grep endpoint | sed -e "s#endpoint: \(.*\):.*#\1#g" | uniq )
do
    fail2ban-client unban "$ip" | grep -v "0" || true
done
exit 0
EOF
chmod 755 /data/scripts/unban-fail2ban-clients.sh
restorecon -F /data/scripts/unban-fail2ban-clients.sh
ln -sf /data/scripts/unban-fail2ban-clients.sh /usr/local/bin/unban-fail2ban-clients.sh
# /etc/systemd/system/unban-fail2ban.timer
[Unit]
Description=Hourly unban of WireGuard peers from fail2ban
[Timer]
OnCalendar=hourly
Persistent=true
[Install]
WantedBy=timers.target
systemctl enable --now unban-fail2ban.timer

The unit is called unban-fail2ban, the script unban-fail2ban-clients.sh. Mildly annoying, and worth knowing when you go looking in the journal.

GeoIP blocking — noise reduction in front of fail2ban

fail2ban is reactive: an attacker has to reach the service and fail a few times before anything happens. A large share of that traffic comes from a handful of countries that have no legitimate reason to talk to a private homelab. Dropping those at the public interface removes the noise before the log parsers ever see it.

The blocklist is roughly 40,000 CIDR ranges, which is far too many for firewalld source bindings. A native nftables set with flags interval and auto-merge collapses them to about 24,000 entries and gives one kernel-side lookup per packet:

table inet geoblock {
    set geoblock4 {
        type ipv4_addr
        flags interval
        auto-merge
        elements = { ... }        # generated weekly from ipdeny.com zone files
    }

    chain input {
        type filter hook input priority mangle; policy accept;
        iifname != "ens3" accept          # only the public interface
        ct state established,related accept
        tcp dport { 25, 587 } accept      # mail must still arrive from anywhere
        ip saddr @geoblock4 counter drop
    }
}

Three details are load-bearing:

  • Its own table. Like inet f2b-table, inet geoblock is untouched by firewall-cmd --reload.
  • The mail exception. This host receives mail, and senders do not choose their country. Ports 25 and 587 are accepted before the drop rule. IMAPS (993) is deliberately not excepted — that’s your own users, and they can use the VPN.
  • iifname != "ens3" accept first. Without it the set would also apply to the WireGuard interface and the pod network.

A weekly timer regenerates the set from the current zone files. The counter on the drop rule is the quickest way to tell whether it is doing anything at all:

nft list chain inet geoblock input | grep counter


Layer 2 — The Cluster

Installing MicroShift

Install

MicroShift is not in Fedora’s repositories. It comes from a COPR project that builds nightly RPMs from the OKD/SCOS payload — either a community one, or your own if the community build stops covering your release stream (why mine is my own).

dnf copr enable -y <owner>/<microshift-nightly-project>

# microshift-io-dependencies brings in the OpenShift mirror repo (cri-o, cri-tools)
dnf install -y --refresh microshift-io-dependencies

# --allowerasing is required: MicroShift needs cri-tools >= 1.35.0, Fedora packages
# cri-tools versioned, and the variants exclude each other. Without it dnf drops
# MicroShift from the transaction instead of failing.
dnf install -y --allowerasing \
  cri-tools1.35 microshift microshift-kindnet microshift-greenboot microshift-selinux

dnf install -y openshift-clients   # kubectl + oc

Do not leave these repositories enabled if you run unattended updates. The MicroShift RPM restarts crio and microshift in its %post scriptlet, so an automatic update restarts your cluster without warning. See Operations.

Configuration: Replace CNI and Storage

MicroShift defaults to OVN-Kubernetes (CNI) and TopoLVM (storage). Both are replaced:

# Base config
install -m 644 /dev/stdin /etc/microshift/config.yaml <<'EOF'
apiServer:
  logLevel: Warning
  auditLog:
    maxFileSize: 200
    maxFiles: 3
    maxFileAge: 7
    profile: None
EOF

# Drop-ins: MicroShift merges everything under config.d/ over config.yaml
install -d /etc/microshift/config.d
install -m 644 /dev/stdin /etc/microshift/config.d/00-disableDefaultCNI.yaml <<'EOF'
network:
  cniPlugin: "none"
EOF

install -m 644 /dev/stdin /etc/microshift/config.d/01-disableTopoLVM.yaml <<'EOF'
storage:
  driver: none
  optionalCsiComponents:
    - none
EOF

auditLog belongs under apiServer, not at the top level. Put it at the root and MicroShift parses the file without complaint and ignores the block — your audit log keeps its defaults and nothing tells you. Verify what actually took effect:

microshift show-config | grep -A5 auditLog

Splitting the CNI and storage switches into config.d/ drop-ins rather than one file is a convenience: each is a single decision you might want to revisit on its own, and a drop-in can be removed without editing around it.

MicroShift ships internal kindnet manifests under /usr/lib/microshift/manifests.d/000-microshift-kindnet/. cniPlugin: none disables OVN-K but not this internal kindnet set — which is what you want, as long as it carries the right pod CIDR.

For a long time it did not: the manifest hard-coded POD_SUBNET=10.244.0.0/16 (kind’s default) while MicroShift’s pod CIDR is 10.42.0.0/16. Builds carrying the fix need no intervention at all. Check before you work around anything:

grep -A1 POD_SUBNET /usr/lib/microshift/manifests.d/000-microshift-kindnet/*.yaml
# want: 10.42.0.0/16

If your build still has the old value, the workaround is a kustomizePaths override in /etc/microshift/config.d/ that excludes the built-in kindnet directory, plus your own kindnet manifest in /etc/microshift/manifests/. Remember that kustomizePaths replaces the default list rather than adding to it, so you have to spell out the paths you still want — including 000-microshift-kube-proxy.

CNI: kube-kindnet

Why POD_SUBNET has to match. kindnet’s KIND-MASQ-AGENT chain uses this value to decide which traffic to masquerade. If it doesn’t match MicroShift’s pod CIDR, requests from specific source ranges get masqueraded when they shouldn’t — which stays invisible until you deploy something that makes decisions based on the client IP, such as an IP allowlist. Then it fails in a way that looks like an application bug.

If you do need to ship your own manifest, it is two documents separated by --- (a Namespace and the DaemonSet). The part that matters is the container’s env block — HOST_IP and POD_IP come from the downward API, POD_SUBNET is the one you have to get right:

env:
  - name: HOST_IP
    valueFrom:
      fieldRef:
        fieldPath: status.hostIP
  - name: POD_IP
    valueFrom:
      fieldRef:
        fieldPath: status.podIP
  - name: POD_SUBNET
    value: "10.42.0.0/16"     # MicroShift's pod CIDR, NOT kind's 10.244.0.0/16

Verify against the running DaemonSet rather than against your manifest — the built-in one wins if you didn’t actually exclude it:

kubectl -n kube-kindnet get ds kube-kindnet-ds \
  -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="POD_SUBNET")].value}'

Storage: local-path-provisioner

Use the standard rancher/local-path-provisioner manifest with two Fedora-specific changes:

  1. Full registry paths: Fedora blocks short image names. rancher/local-path-provisioner:v0.x → docker.io/rancher/local-path-provisioner:v0.x, busybox:latest → docker.io/library/busybox:latest.
  2. Remove chcon from the setup script: busybox doesn’t have chcon. The SELinux label for /var/lib/pvc was set by semanage in the storage layout — no runtime chcon needed.

Add SCC bindings for privileged and hostmount-anyuid in the same manifest.

Start MicroShift

systemctl enable --now microshift

export KUBECONFIG=/var/lib/microshift/resources/kubeadmin/kubeconfig
echo 'export KUBECONFIG=/var/lib/microshift/resources/kubeadmin/kubeconfig' >> /root/.bashrc
echo 'alias k=kubectl' >> /root/.bashrc

kubectl get pods -A -w

Expected system pods once stable — seven, no more:

NamespacePodRole
openshift-ingressrouter-defaultHAProxy ingress, terminates TLS
openshift-dnsdns-default, node-resolvercluster DNS
openshift-service-caservice-casigns internal service certificates
kube-proxykube-proxyservice networking
kube-kindnetkube-kindnet-dsCNI
local-path-storagelocal-path-provisionerPVCs

kube-system stays empty, and that surprises people who know OpenShift. There is no csi-snapshot-controller here: the storage config above sets driver: none and optionalCsiComponents: [none], so the optional CSI pieces are never bootstrapped. Volume snapshots are a Btrfs concern on this box, not a Kubernetes one — the PVC directory sits on its own subvolume that snapper and btrbk already cover, which is both cheaper and restorable from outside a broken cluster.


GitOps with Flux CD

SSH Key for GitHub

ssh-keygen -t ed25519 -C "fedora-server" -f /root/.ssh/id_ed25519 -N ""
cat /root/.ssh/id_ed25519.pub
# → Add as Deploy Key (with write access) to github.com/youruser/homelab

SCC Bindings for Flux

MicroShift’s SCCs reject Flux’s default fsGroup: 1337. Create bindings before bootstrapping, so Flux pods can start:

cat > /etc/microshift/manifests/flux-scc.yaml <<'EOF'
apiVersion: v1
kind: Namespace
metadata:
  name: flux-system
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: flux-source-controller-scc
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: system:openshift:scc:privileged
subjects:
  - kind: ServiceAccount
    name: source-controller
    namespace: flux-system
---
# Repeat for: kustomize-controller, helm-controller, notification-controller
EOF

systemctl restart microshift

Flux Bootstrap

# NOTE: the official install.sh fails on Fedora with bash syntax errors.
# Install from the release tarball instead:
FLUX_VERSION=2.9.4
ARCH=$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/')
curl -sLo /tmp/flux.tar.gz \
  "https://github.com/fluxcd/flux2/releases/download/v${FLUX_VERSION}/flux_${FLUX_VERSION}_linux_${ARCH}.tar.gz"
tar -xzf /tmp/flux.tar.gz -C /tmp flux
install -m 755 /tmp/flux /usr/local/bin/flux

flux bootstrap git \
  --url=ssh://git@github.com/youruser/homelab \
  --branch=main \
  --path=configuration/ \
  --private-key-file=/root/.ssh/id_ed25519 \
  --silent

Clone the working tree to /data (backed by Btrfs, included in backups):

cd /data && git clone git@github.com:youruser/homelab.git .

Telegram Notifications for Flux Events

kubectl create secret generic telegram-secret -n flux-system \
  --from-literal=token="$(awk -F= '/^TOKEN=/{print $2}' /etc/telegramrc)"

Then a Provider and an Alert in flux-system:

apiVersion: notification.toolkit.fluxcd.io/v1beta3
kind: Provider
metadata:
  name: telegram
  namespace: flux-system
spec:
  type: telegram
  channel: "<YOUR_CHAT_ID>"
  address: https://api.telegram.org
  secretRef:
    name: telegram-secret
---
apiVersion: notification.toolkit.fluxcd.io/v1beta3
kind: Alert
metadata:
  name: telegram-alert
  namespace: flux-system
spec:
  providerRef:
    name: telegram
  eventSeverity: info
  eventSources:
    - kind: Kustomization
      name: '*'
    - kind: HelmRelease
      name: '*'
  summary: "Flux GitOps Update"
  exclusionList:
    - "Dependencies do not meet ready condition"
    - "dependency .* is not ready"
    - "retrying in"

The exclusionList is what makes this usable. With eventSeverity: info you get told about every successful reconcile — which is the point — but you also get every transient message from the dependsOn gating, and those fire in bursts whenever anything changes. Three patterns remove the noise without hiding real failures. Without them, the channel becomes something you mute, which defeats the exercise.

Pre-Update Btrfs Snapshot via GitHub Action

Before every push to main, a GitHub Action SSHes to the server and creates a Btrfs snapshot. This gives you a clean rollback point if a Flux reconciliation breaks something.

Restricted SSH key (command-restricted, no PTY):

ssh-keygen -t ed25519 -f /root/.ssh/github-actions -N "" -C "gh-actions-snapshot"

KEY=$(cat /root/.ssh/github-actions.pub)
echo "command=\"/usr/local/bin/pre-flux-snapshot.sh\",no-port-forwarding,no-X11-forwarding,no-agent-forwarding,no-pty $KEY" \
  >> /root/.ssh/authorized_keys
cat > /data/scripts/pre-flux-snapshot.sh <<'EOF'
#!/bin/bash
DESCRIPTION="pre-flux-update-$(date +%Y%m%d-%H%M%S)"
flux suspend ks --all -n flux-system 2>/dev/null
snap-all "$DESCRIPTION"
flux resume ks --all -n flux-system 2>/dev/null
echo "Snapshots created: $DESCRIPTION"
EOF
chmod 755 /data/scripts/pre-flux-snapshot.sh
restorecon -F /data/scripts/pre-flux-snapshot.sh
ln -sf /data/scripts/pre-flux-snapshot.sh /usr/local/bin/pre-flux-snapshot.sh

Store the private key as FEDORA_SSH_KEY secret in the GitHub repo.

Renovate Bot — Automatic Image Updates

Renovate runs as a GitHub Action every 6 hours and opens PRs for new container image tags. renovate.json lives in the repository root, not under configuration/ — it configures the bot, it is not something Flux should apply.

{
  "extends": ["config:recommended"],
  "schedule": ["at any time"],
  "prHourlyLimit": 5,
  "fetchChangeLogs": "pr",
  "ignoreTests": true,

  "kubernetes": { "managerFilePatterns": ["/\\.yaml$/"] },
  "flux":       { "managerFilePatterns": ["/\\.yaml$/"] },

  "packageRules": [
    {
      "description": "MCP servers — merge immediately",
      "matchFileNames": ["configuration/mcp-openshift/**"],
      "minimumReleaseAge": "0 hours",
      "automerge": true,
      "automergeType": "pr"
    },
    {
      "description": "Mail stack — open a PR, I merge by hand",
      "matchFileNames": ["configuration/mailstack/**"],
      "minimumReleaseAge": "0 hours",
      "automerge": false
    }
  ]
}

Scope the rules by path, not by update type. The tempting configuration is “automerge patch and minor everywhere, hold majors” — but the risk of an update has little to do with its version number. A patch release of the mail stack can cost me mail; a major bump of an MCP server costs me nothing, because if it breaks I notice immediately and nobody else is affected. So the rules match matchFileNames per namespace and each one decides for itself.

Two more things worth knowing:

  • ignoreTests: true is necessary here because there is no CI on the GitOps repo. Without it Renovate waits forever for status checks that will never arrive, and nothing ever automerges.
  • Digest pinning on floating tags. redis:8 or nextcloud:34-apache never change name, so Renovate has no version to compare. With pinDigests it pins the current digest and opens a PR when it moves — which is how you get patch updates on a floating tag without imagePullPolicy: Always. One caveat learned the hard way: Docker Hub publishes a multi-arch index progressively, and an automerge caught an intermediate state with no linux/amd64 entry, producing ImagePullBackOff. For that image the rule now tracks the version tag and waits three hours.

The Btrfs snapshot GitHub Action runs before Renovate’s auto-merges, so every automated image update has a rollback point.


Wildcard TLS with cert-manager + acme-dns

Architecture

architecture-beta
  group certns(cloud)[cert_manager namespace]

  service cname(internet)[acme challenge CNAME]
  service acmedns(server)[acme_dns port 53] in certns
  service issuer(server)[ClusterIssuer DNS 01] in certns
  service secret(disk)[wildcard tls Secret] in certns
  service sync(server)[sync wildcard tls]
  service appns(disk)[app namespace secrets]

  cname:R --> L:acmedns
  issuer:T --> B:acmedns
  issuer:B --> T:secret
  secret:B --> T:sync
  sync:B --> T:appns

Why acme-dns instead of direct DNS-01? Most domain registrars don’t offer a usable DNS API for cert-manager. acme-dns is a minimal DNS server that only handles TXT records for ACME challenges — you point a CNAME at it once, and cert-manager does the rest via a stable API.

Deploy acme-dns via Flux

Create configuration/acme-dns/ manifests (Deployment, Services, PVC, ConfigMap) and apply via Flux. Key points:

  • Two Services, not one. acme-dns-dns carries ports 53/tcp+udp with externalIPs: [<YOUR_SERVER_IP>]; acme-dns carries the registration API on port 80 as a plain ClusterIP. Keeping them apart means the public one exposes only DNS.
  • externalIPs rather than hostPort. kube-proxy creates a single DNAT rule on the public IP, and the Service endpoint is managed independently of the pod lifecycle — so a RollingUpdate works, which it does not with hostPort (the new pod cannot bind while the old one holds it).
  • PVC for the SQLite credential database. Lose it and every CNAME you set below is worthless.

Register one account per apex domain — the API takes no arguments and returns a random subdomain each time:

POD=$(kubectl get pod -n acme-dns -l app=acme-dns -o name)
for _ in example.com example.net; do
  kubectl exec -n acme-dns "$POD" -- \
    curl -s -X POST http://localhost/register -H 'Content-Type: application/json' -d '{}'
  echo
done
# Save the JSON — username, password, subdomain and fulldomain, one set per domain

Store as a Kubernetes secret:

install -m 600 <json-file> /etc/acme-dns/acmedns.json
kubectl create secret generic acmedns-credentials -n cert-manager \
  --from-file=acmedns.json=/etc/acme-dns/acmedns.json

CNAME Records (one-time)

At your registrar, for each apex domain:

_acme-challenge.example.com   CNAME   <fulldomain-from-json>

cert-manager: ClusterIssuer and Certificate

Deploy cert-manager via Helm (Flux HelmRelease), then the two resources that do the work. The ClusterIssuer is where acme-dns actually gets wired in — host points at the in-cluster Service, so the challenge never leaves the node:

apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-production
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: <your-email>
    privateKeySecretRef:
      name: letsencrypt-production-key
    solvers:
      - dns01:
          acmeDNS:
            host: http://acme-dns.acme-dns.svc.cluster.local
            accountSecretRef:
              name: acmedns-credentials
              key: acmedns.json

Then one Certificate covering every domain — apex and wildcard each need their own entry, a wildcard does not match the bare domain:

apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: wildcard-tls
  namespace: cert-manager
spec:
  secretName: wildcard-tls
  issuerRef:
    name: letsencrypt-production
    kind: ClusterIssuer
  dnsNames:
    - example.com
    - "*.example.com"
    - example.net
    - "*.example.net"

Set up a letsencrypt-staging issuer against acme-staging-v02.api.letsencrypt.org as well and use it first. Let’s Encrypt’s rate limits for the production endpoint are low enough that a few failed attempts will lock you out for a week.

Sync to App Namespaces

Kubernetes requires TLS secrets to be in the same namespace as the Ingress, and under the name that Ingress references. A weekly timer copies the wildcard secret out of cert-manager into every namespace that needs it.

The targets are not configured in the script — it reads them from the Ingress objects at runtime. The full script is in Operations; set up the timer here:

# script in /data/scripts/, symlink into /usr/local/bin/
chmod 755 /data/scripts/sync-wildcard-tls.sh
restorecon -F /data/scripts/sync-wildcard-tls.sh
ln -sf /data/scripts/sync-wildcard-tls.sh /usr/local/bin/sync-wildcard-tls.sh

# systemd timer: weekly
systemctl enable --now sync-wildcard-tls.timer
/usr/local/bin/sync-wildcard-tls.sh   # run immediately

When adding a new service with a TLS ingress, the only requirement is that the secret name ends in -tls. Run the script once and it picks the new namespace up on its own.


Layer 3 — The Services

All services follow the same pattern:

  1. Manifests in configuration/<name>/ + Kustomization CR configuration/<name>-ks.yaml
  2. Secrets created manually with kubectl create secret (never in git)
  3. flux reconcile kustomization <name> --with-source
  4. Verify with kubectl get pods -n <namespace>

Here are the notable per-service details:

Pihole

Runs as the DNS server for all WireGuard clients. Uses hostPort: 53 on the WireGuard interface so clients can point directly to 10.0.0.1 as their DNS server.

# configuration/pihole/deployment.yaml (excerpt)
strategy:
  type: Recreate    # prevents Pending state when hostPort: 53 is already bound
containers:
- name: pihole
  image: docker.io/pihole/pihole:2026.07.2
  env:
  - name: FTLCONF_dns_upstreams
    value: "8.8.8.8;8.8.4.4"
  - name: FTLCONF_dns_listeningMode
    value: "ALL"
  - name: FTLCONF_webserver_api_password
    valueFrom:
      secretKeyRef:
        name: pihole-password
        key: FTLCONF_webserver_api_password
  ports:
  - containerPort: 53
    hostPort: 53
    hostIP: 10.0.0.1        # bind to WireGuard interface only
    protocol: UDP
    name: dns-udp
  - containerPort: 53
    hostPort: 53
    hostIP: 10.0.0.1
    protocol: TCP
    name: dns-tcp
  - containerPort: 80
    name: http
  securityContext:
    privileged: true        # required for FTL (NET_ADMIN)
  volumeMounts:
  - name: pihole-data
    mountPath: /etc/pihole
  - name: dnsmasq-custom
    mountPath: /etc/dnsmasq.d/99-custom.conf
    subPath: 99-custom.conf
# configuration/pihole/ingress.yaml
metadata:
  annotations:
    route.openshift.io/termination: "edge"
    haproxy.router.openshift.io/ip_whitelist: "10.0.0.0/24 <CORP_IP_RANGE>"
spec:
  tls:
  - hosts: [dnsconfig.example.com]
    secretName: pihole-tls
  rules:
  - host: dnsconfig.example.com
# configuration/pihole/scc.yaml — required for privileged workloads
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
  name: pihole-privileged-scc
roleRef:
  apiGroup: rbac.authorization.k8s.io
  kind: ClusterRole
  name: system:openshift:scc:privileged
subjects:
- kind: ServiceAccount
  name: default
  namespace: pihole

After deploying: add local DNS entries in Pihole for all services (so service.example.com resolves to 10.0.0.1 for VPN clients, bypassing public DNS).

Vaultwarden

Small footprint: 2 Gi PVC, 100 Mi memory limit. Access restricted to VPN + any additional IP ranges.

# configuration/vaultwarden/deployment.yaml (excerpt)
containers:
- name: vaultwarden
  image: docker.io/vaultwarden/server:1.37.2
  env:
  - name: WEBSOCKET_ENABLED
    value: "true"
  - name: ROCKET_ADDRESS
    value: "0.0.0.0"
  resources:
    limits:
      memory: 100Mi
  volumeMounts:
  - name: data
    mountPath: /data
  strategy:
    type: Recreate   # important: only one writer at a time
# configuration/vaultwarden/ingress.yaml (excerpt)
annotations:
  haproxy.router.openshift.io/ip_whitelist: "10.0.0.0/24 <CORP_IP_RANGE>"
  haproxy.router.openshift.io/timeout: "300s"

Admin token: openssl rand -base64 48 | tr -d '\n'

Paperless-NGX

  • Four PVCs: data (PostgreSQL), media, consume, export
  • Critical for restore: always restore paperless-data and paperless-media from the same point in time — they must stay in sync

Nextcloud + MariaDB

# configuration/nextcloud/deployment.yaml (excerpt)
containers:
- name: nextcloud
  image: docker.io/nextcloud:34-apache
  env:
  - name: MYSQL_HOST
    value: mariadb
  - name: MYSQL_DATABASE
    value: owncloud
  - name: MYSQL_USER
    value: owncloud
  - name: MYSQL_PASSWORD
    valueFrom:
      secretKeyRef:
        name: nextcloud-secret
        key: MYSQL_PASSWORD
  - name: NEXTCLOUD_TRUSTED_DOMAINS
    value: cloud.example.com
  - name: REDIS_HOST
    value: redis
  - name: PHP_UPLOAD_LIMIT
    value: 10G
  - name: PHP_MEMORY_LIMIT
    value: 1G
  - name: APACHE_DISABLE_REWRITE_IP
    value: "1"
  - name: TRUSTED_PROXIES
    value: 10.42.0.0/16    # pod CIDR — required for HAProxy reverse-proxy headers
  resources:
    limits:
      memory: 2Gi
# configuration/nextcloud/ingress.yaml
annotations:
  haproxy.router.openshift.io/timeout: "300s"
  haproxy.router.openshift.io/proxy-body-size: "10g"
  haproxy.router.openshift.io/hsts_header: max-age=15552000;includeSubDomains;preload

Two CronJobs in the same file — Nextcloud’s background job runner (every 5 min) and automatic app updates (daily at 03:00):

# configuration/nextcloud/cronjob.yaml
---
apiVersion: batch/v1
kind: CronJob
metadata:
  name: nextcloud-cron
spec:
  schedule: "*/5 * * * *"
  concurrencyPolicy: Forbid
  jobTemplate:
    spec:
      template:
        spec:
          securityContext:
            runAsUser: 33    # www-data
            runAsGroup: 33
          containers:
          - name: nextcloud-cron
            image: docker.io/nextcloud:34-apache
            command: ["php", "-f", "/var/www/html/cron.php"]
            volumeMounts:
            - {name: html, mountPath: /var/www/html}
            - {name: data, mountPath: /var/www/html/data}
---
apiVersion: batch/v1
kind: CronJob
metadata:
  name: nextcloud-app-update
spec:
  schedule: "0 3 * * *"
  concurrencyPolicy: Forbid
  jobTemplate:
    spec:
      template:
        spec:
          securityContext:
            runAsUser: 33
            runAsGroup: 33
          containers:
          - name: nextcloud-app-update
            image: docker.io/nextcloud:34-apache
            command: ["php", "-f", "/var/www/html/occ", "app:update", "--all"]

Mailstack

The most complex service — seven interdependent components:

architecture-beta
  group mailstack(cloud)[mailstack namespace]

  service db(database)[MariaDB] in mailstack
  service pa(server)[PostfixAdmin] in mailstack
  service postfix(server)[Postfix SMTP] in mailstack
  service dovecot(server)[Dovecot IMAP] in mailstack
  service rspamd(server)[Rspamd] in mailstack
  service redis(database)[Redis] in mailstack
  service clamav(server)[ClamAV] in mailstack

  db:T --> B:pa
  db:L --> R:postfix
  db:R --> L:dovecot
  postfix:T -- B:rspamd
  postfix:B --> T:dovecot
  rspamd:L --> R:redis
  rspamd:T --> B:clamav

Postfix: Uses hostNetwork: true so that connections arrive at Postfix with the real client IP — critical for fail2ban to see the actual source address. hostPort alone routes connections through the pod bridge (kindnet), which replaces the source IP with the bridge address (10.42.0.1). Config files (main.cf, master.cf, MySQL maps) are mounted from a ConfigMap. Mail logs are written to a shared PVC (mail-logs) that fail2ban reads from the host.

# configuration/mailstack/postfix-deployment.yaml (excerpt)
strategy:
  type: Recreate
template:
  spec:
    hostNetwork: true
    dnsPolicy: ClusterFirstWithHostNet
    containers:
    - name: postfix
      image: docker.io/boky/postfix:latest
      command: ["/usr/sbin/postfix", "-c", "/etc/postfix", "start-fg"]
      securityContext:
        privileged: true
      volumeMounts:
      - {name: config, mountPath: /etc/postfix/main.cf, subPath: main.cf}
      - {name: tls,    mountPath: /etc/ssl/mail,         readOnly: true}
      - {name: mail-logs, mountPath: /var/log/mail}
    - name: log-tailer   # sidecar: tails postfix.log to stdout for kubectl logs
      image: docker.io/library/alpine:3
      command: ["/bin/sh", "-c", "tail -F /var/log/mail/postfix.log"]
      volumeMounts:
      - {name: mail-logs, mountPath: /var/log/mail}
    - name: logrotate   # sidecar: daily rotation of postfix.log
      image: docker.io/library/alpine:3
      command: ["/bin/sh", "-c"]
      args: ["apk add logrotate -q && while true; do logrotate ... || { echo 'ERROR'; exit 1; }; sleep 86400; done"]
      volumeMounts:
      - {name: mail-logs, mountPath: /var/log/mail}

Dovecot 2.4 breaking changes (if migrating from 2.3):

  • MySQL connection config can no longer use %{env:VAR} in the mysql{} block — mount the credentials as a Kubernetes Secret file instead
  • encrypt = dovecot:SHA512-CRYPT renamed to encrypt = system
  • ssl_protocols removed, ssl_min_protocol is the new knob
# configuration/mailstack/dovecot-deployment.yaml (excerpt)
strategy:
  type: Recreate
template:
  spec:
    hostNetwork: true
    dnsPolicy: ClusterFirstWithHostNet
    containers:
    - name: dovecot
      image: docker.io/dovecot/dovecot:latest-root
      ports:
      - {containerPort: 993}   # metadata only — hostNetwork binds to host stack directly
      securityContext:
        privileged: true
      volumeMounts:
      - {name: config,    mountPath: /etc/dovecot/conf.d/10-auth.conf, subPath: 10-auth.conf}
      - {name: db-secret, mountPath: /etc/dovecot/conf.d/01-db-connection.conf,
         subPath: 01-db-connection.conf, readOnly: true}   # MySQL credentials as secret file
      - {name: vmail,     mountPath: /home/vmail}
      - {name: tls,       mountPath: /etc/dovecot/ssl, readOnly: true}
      - {name: mail-logs, mountPath: /var/log/mail}
    - name: log-tailer   # sidecar: tails dovecot.log to stdout for kubectl logs
      image: docker.io/library/alpine:3
      command: ["/bin/sh", "-c", "tail -F /var/log/mail/dovecot.log"]
      volumeMounts:
      - {name: mail-logs, mountPath: /var/log/mail}

Rspamd:

# configuration/mailstack/rspamd-deployment.yaml (excerpt)
metadata:
  annotations:
    prometheus.io/scrape: "true"    # Alloy picks this up for Grafana Cloud
    prometheus.io/port: "11334"
    prometheus.io/path: "/metrics"
containers:
- name: rspamd
  image: docker.io/rspamd/rspamd:4.1.5
  securityContext:
    privileged: true
  ports:
  - {containerPort: 11332, name: milter, hostPort: 11332}
  - {containerPort: 11334, name: controller}
  resources:
    requests:
      memory: 128Mi
    limits:
      memory: 768Mi    # raised from 512Mi — startup spike causes OOMKill on node reboot
  livenessProbe:
    httpGet: {path: /ping, port: 11334}   # use httpGet, NOT tcpSocket
  readinessProbe:
    httpGet: {path: /ping, port: 11334}
  volumeMounts:
  - {name: config,    mountPath: /etc/rspamd/local.d}
  - {name: dkim-keys, mountPath: /etc/rspamd/dkim, readOnly: true}

Key Rspamd config (local.d/dkim_signing.conf): set try_fallback = false to prevent signing with the wrong key when a domain’s DKIM key is missing. Set secure_ip to the pod bridge IP (10.42.0.1) so Grafana Alloy can scrape /metrics without authentication.

ClamAV: Run freshclam as a shell loop sidecar, not as an init container (freshclam takes minutes — init containers block pod startup):

- name: freshclam
  image: docker.io/clamav/clamav:stable_base
  command: ["/bin/sh", "-c"]
  args: ["while true; do freshclam; sleep 3600; done"]

After deploying, activate the fail2ban mail jails (see Intrusion defence, phase 2).

Monitoring (Grafana Cloud)

  • Grafana Alloy runs as a DaemonSet with hostNetwork: true — scrapes metrics from the node and Kubernetes API
  • Grafana Operator manages GrafanaDashboard CRDs and pushes them to Grafana Cloud
  • kube-state-metrics for Kubernetes object metrics — needs a custom SCC for the hostmount-anyuid service account
  • Important: The Grafana Operator reconciles dashboards every ~10 minutes and overwrites any UI changes. Always fix dashboards in the YAML, not the UI.
  • Disable leader election: Add leaderElect: false to the HelmRelease values. On a single-node cluster there is never a competing operator instance — leader election only adds unnecessary risk of crashes when the kube-apiserver has brief pauses (e.g. during etcd compaction).
  • Loki token scope: needs logs:write (created under Stack → Loki → Send Logs, not under Access Policies)

ClickHouse for Rspamd Analytics (optional)

ClickHouse receives per-message metadata from Rspamd and enables detailed spam/ham analysis in Grafana. Rspamd writes to ClickHouse via the HTTP interface; the Grafana datasource uses a suspended CRD approach with a one-time API call to inject credentials.

Matomo (optional)

Self-hosted web analytics with its own MariaDB. Two things are specific to this setup: the code lives in a PVC populated by an init container from the image (so plugins survive restarts and core:update can run), and the image is pinned to a version tag, not a digest. Digest pinning bit me once — Docker Hub updates a multi-arch index progressively during a push, and an automated update caught an intermediate state that had no linux/amd64 entry, producing ImagePullBackOff.

CrowdSec

Crowd-sourced blocklists alongside fail2ban: LAPI and agent in-cluster, the firewall bouncer as a systemd service on the host writing native nftables sets. Details and the reasoning for running both: Adding CrowdSec Next to fail2ban.

Stakater Reloader

Kubernetes does not restart a pod when a mounted ConfigMap or Secret changes — and with envFrom the value never updates at all. Reloader watches both and triggers a rolling restart for workloads that opt in with reloader.stakater.com/auto: "true".

One OpenShift-specific catch: the chart defaults to runAsUser: 65534, which the restricted-v2 SCC rejects because it assigns UIDs from a per-namespace range. Omitting the key in your Helm values does not help — Helm merges with chart defaults, so a missing key means “use the default”. You have to set it to null explicitly so it disappears from the rendered manifest:

reloader:
  deployment:
    securityContext:
      runAsUser: null        # not omitted — explicitly null
      runAsNonRoot: true
      seccompProfile:
        type: RuntimeDefault

DMARC report evaluation

A daily CronJob fetches DMARC aggregate reports from a dedicated mailbox, parses them and feeds a Grafana dashboard — which is how you find out whether your SPF/DKIM alignment actually holds for the receivers that matter.

etcd snapshots

The Btrfs snapshots of /var/lib/microshift are crash-consistent — a picture of the files at one moment, regardless of what etcd was doing. A daily CronJob additionally takes an application-consistent snapshot with etcdctl snapshot save and verifies it in a second container with etcdutl snapshot status, so a broken snapshot fails the job instead of sitting in the backup unnoticed.

It writes to a fixed filename in a PVC rather than rotating: the PVC lives on the subvolume that btrbk already sends hourly to the Pi with 24h 8d 5w retention, so versioning is handled one layer down instead of in two places.

The CronJob needs hostNetwork (etcd is a host process on localhost:2379, not a pod) and a hostPath mount for the client certificates, which together require the privileged SCC.

MCP servers (optional)

Three MCP servers expose the cluster, the document archive and the mailbox to Claude Code — cluster management, read-only document access, and mail. Each runs as a small deployment behind its own ingress; the interesting part is scoping, not deployment. The Paperless one in particular is deliberately read-only: the API token it uses has no write permissions at all, so “read-only” is enforced by the server it talks to rather than by the prompt.


Operations

All maintenance scripts live in /data/scripts/ (Btrfs-backed, included in btrbk backups) with symlinks in /usr/local/bin/. They run as systemd oneshot services with timers:

ScriptTimerPurpose
sendtelegram.sh— (library)Send Telegram messages — called by other scripts
boot-backup.shdaily + post-dnf5-automaticBackup /boot, /boot/efi, partition table
check-backup-btrfs-subvolumes.shdailyAlert if any Btrfs subvolume lacks btrbk or snapper config
unban-fail2ban-clients.shhourly (unban-fail2ban.timer)Unban the WireGuard peers’ endpoint IPs from all fail2ban jails
pre-flux-snapshot.shSSH trigger (GitHub Action)Btrfs snapshot before Flux reconcile
check-flux-update.shweeklyTelegram alert when a new Flux version is available
sync-wildcard-tls.shMon 00:00Copy wildcard TLS secret to all app namespaces
system-check.shdaily 08:00Telegram status: WG peers, pods, mail (24h), fail2ban, disk
podman-image-cleanup.shMon 03:00Remove dangling + unused Podman images, Telegram report
snap-all— (manual, before risky changes)Snapshot all 8 snapper configs at once (root home data var_lib_pvc var_lib_microshift var var_log var_lib_containers)
check-image-updates.shMon 08:00Weekly report: container images vs. upstream releases, in-cluster DaemonSets vs. the manifests shipped in the RPM, and whether a new MicroShift build is waiting
etcd-backup-check.shdaily 04:00Verify the etcd snapshot is fresh, and that the snapshot image still matches the cluster’s etcd version
geoblock.shweeklyRefresh the GeoIP blocklist in a native nftables set
microshift-upgrade-4x.sh— (manual)The only path by which MicroShift gets updated — see below

Automatic updates for the distribution

Fedora’s own packages are applied unattended. On a single-node box that is the right trade: a security fix applied at six in the morning beats one applied when I next remember.

cat > /etc/dnf/automatic.conf <<'EOF'
[commands]
upgrade_type = default
random_sleep = 0
download_updates = yes
apply_updates = yes
reboot = when-needed
reboot_command = "shutdown -r +5 'Reboot after automatic update'"

[emitters]
emit_via = stdio
EOF

systemctl enable --now dnf5-automatic.timer

reboot = when-needed is safe here only because of the layer below it: boot-backup.sh runs as a post-transaction hook, and the Btrfs snapshots make an unattended kernel update reversible from the GRUB menu. Without that, unattended reboots on a box you can’t physically reach would be reckless.

The cluster is a different matter entirely.

Keeping cluster updates out of the automation

dnf5-automatic applies Fedora’s updates unattended and reboots when needed. That is fine for the distribution — and dangerous for the cluster, because the MicroShift RPM restarts crio and microshift in its %post scriptlet. With a nightly package feed enabled, that arrangement restarts your cluster every night onto a build nobody has looked at.

The fix is to make the packages invisible to the updater by disabling their repositories, rather than excluding package names in the updater’s configuration:

dnf config-manager setopt <microshift-repo>.enabled=0
dnf config-manager setopt openshift-mirror-*-beta.enabled=0

Repository-level is the better granularity here. An exclude list has to name every package — microshift, microshift-greenboot, microshift-selinux, microshift-kindnet, plus cri-o and cri-tools from a different repository — and the failure mode of forgetting one is a cluster restart at 6am. setopt also writes to /etc/dnf/repos.override.d/ instead of editing .repo files, one of which is owned by a package and would be overwritten on its next update.

Updates then run through a single script that switches the repositories on with --enablerepo for the duration of its own transaction: snapshot with snap-all, one dnf transaction (MicroShift + cri-o + cri-tools), re-apply the manifests from kustomizePaths, microshift healthcheck, pod comparison against a pre-upgrade baseline, Telegram. Two things that cost me time:

  • --refresh on every repoquery. Once a repository is disabled nothing refreshes its metadata, so queries answer from a stale cache. Mine picked yesterday’s build as “the newest”, concluded it was already installed and exited successfully having done nothing.
  • dnf install does not upgrade. For an already-installed package, dnf5 reports “already installed” and moves on. Pass the exact version from repoquery --upgrades instead.

The reasoning in full: Homelab Update: My Own COPR, Tighter Updates.

system-check.sh

The daily report, and the script that has grown the most. What follows is an excerpt — the gathering of the four numbers that go into every message. The full version also walks the journal for error patterns, watches for OOM kills across boots, distinguishes a reboot caused by the unattended updater from one that wasn’t, and — when something looks off — attaches a short analysis before sending.

The design rule that makes it useful: it sends a message every day, healthy or not. A monitor that only speaks up when something is wrong is indistinguishable from a monitor that has stopped working.

#!/bin/bash
export KUBECONFIG=/var/lib/microshift/resources/kubeadmin/kubeconfig
source /etc/telegramrc

send() {
  curl -s -X POST "https://api.telegram.org/bot${TOKEN}/sendMessage" \
    -d chat_id="${CHATID}" -d parse_mode="HTML" \
    --data-urlencode text="$1" > /dev/null
}

# WireGuard peer status
WG_PEERS=$(wg show wg0 | grep -c "latest handshake")
WG_RECENT=$(wg show wg0 | awk '/latest handshake/ {
  if ($0 ~ /second/ || $0 ~ /minute/) count++
  else if ($0 ~ /hour/) { match($0, /([0-9]+) hour/, a); if (a[1]+0 < 2) count++ }
} END{print count+0}')

# Kubernetes pod health
NOT_RUNNING=$(kubectl get pods -A --no-headers 2>/dev/null \
  | grep -v -E "Running|Completed" | wc -l)
NOT_RUNNING_LIST=$(kubectl get pods -A --no-headers 2>/dev/null \
  | grep -v -E "Running|Completed" \
  | awk '{print $1"/"$2" ("$4")"}' | head -10)
TOTAL_PODS=$(kubectl get pods -A --no-headers 2>/dev/null | grep -c "Running")

# Mail stats (last 24h from postfix logs)
MAIL_IN=$(kubectl logs -n mailstack deploy/mail-postfix --since=24h 2>/dev/null \
  | grep -c "postfix/lmtp.*status=sent" || echo 0)
MAIL_OUT=$(kubectl logs -n mailstack deploy/mail-postfix --since=24h 2>/dev/null \
  | grep -c "postfix/smtp.*status=sent" || echo 0)
MAIL_REJECT=$(kubectl logs -n mailstack deploy/mail-postfix --since=24h 2>/dev/null \
  | grep -c "NOQUEUE: reject" || echo 0)
QUEUE_RAW=$(kubectl exec -n mailstack deploy/mail-postfix -- postqueue -p 2>/dev/null \
  | tail -1)
QUEUE=$(echo "$QUEUE_RAW" | grep -q "empty" && echo "leer" || echo "$QUEUE_RAW")

# fail2ban
BANNED_COUNT=0
for jail in $(fail2ban-client status 2>/dev/null \
  | grep "Jail list:" | sed 's/.*Jail list:\s*//' | tr ', ' '\n' | grep -v '^$'); do
  n=$(fail2ban-client status "$jail" 2>/dev/null \
    | grep "Currently banned:" | awk '{print $NF}')
  BANNED_COUNT=$((BANNED_COUNT + ${n:-0}))
done

DISK_FREE=$(df -h / | awk 'NR==2 {print $4 " von " $2 " frei (" $5 " belegt)"}')

# Build status icons
[ "$WG_RECENT" -ge 1 ] \
  && WG_STATUS="Aktiv: ${WG_RECENT}/${WG_PEERS} Peers (Handshake unter 2h)" \
  && WG_ICON="OK" \
  || { WG_STATUS="Keine aktiven Peers!"; WG_ICON="WARN"; }

[ "$NOT_RUNNING" -eq 0 ] \
  && K8S_STATUS="${TOTAL_PODS} Pods Running" && K8S_ICON="OK" \
  || { K8S_STATUS="${NOT_RUNNING} Pod(s) nicht Running:\n${NOT_RUNNING_LIST}"; K8S_ICON="WARN"; }

[ "$BANNED_COUNT" -gt 0 ] && F2B_ICON="WARN" || F2B_ICON="OK"

send "System-Check $(date '+%d.%m.%Y %H:%M')

WireGuard [${WG_ICON}]
${WG_STATUS}

Kubernetes [${K8S_ICON}]
${K8S_STATUS}

Mail letzte 24h:
  Eingehend zugestellt: ${MAIL_IN}
  Ausgehend gesendet:   ${MAIL_OUT}
  Rejects:              ${MAIL_REJECT}
  Queue:                ${QUEUE}

Fail2ban [${F2B_ICON}]
  Gebannte Clients:     ${BANNED_COUNT}

Speicher:
  ${DISK_FREE}"

Timer — once daily at 08:00:

# /etc/systemd/system/system-check.timer
[Unit]
Description=System Check täglich 08:00
[Timer]
OnCalendar=*-*-* 08:00:00
Persistent=false
[Install]
WantedBy=timers.target

Output: daily system-check message

System-Check 27.08.2026 08:00

WireGuard [OK]
Aktiv: 10/12 Peers (Handshake unter 2h)

Kubernetes [OK]
45 Pods Running

Mail letzte 24h:
  Eingehend zugestellt: 31
  Ausgehend gesendet:   22
  Rejects:               3
  Queue:                leer

Fail2ban [OK]
  Gebannte Clients:      2

Speicher:
  188G von 610G frei (69% belegt)

Anything abnormal adds a section above this one, so the shape of the message tells you whether to read it. The thresholds are deliberately loose — disk at 85 %, more than 20 distinct error patterns in the journal, any OOM kill in 24 hours — because a report that cries wolf gets filtered into a folder nobody opens.


sync-wildcard-tls.sh

Kubernetes requires an Ingress TLS secret to live in the same namespace as the Ingress, so the one wildcard secret from cert-manager has to be copied into every namespace that references it — under the name that namespace’s Ingress expects.

The first version of this script carried a hard-coded namespace list. That’s a list you forget to update, and the failure mode is a browser certificate warning you discover from a family member. It now derives its targets from the cluster instead: every Ingress whose spec.tls[].secretName ends in -tls, excluding the cert-manager namespace itself.

#!/bin/bash
export KUBECONFIG=/var/lib/microshift/resources/kubeadmin/kubeconfig
SRC_NS=cert-manager
SRC_SECRET=wildcard-tls

SRC_JSON="$(kubectl get secret "$SRC_SECRET" -n "$SRC_NS" -o json)" || exit 1

# (namespace, secretName) pairs straight from the Ingress objects
mapfile -t TARGETS < <(kubectl get ingress -A -o json | python3 -c "
import json, sys
for item in json.load(sys.stdin).get('items', []):
    ns = item['metadata']['namespace']
    if ns == 'cert-manager':
        continue
    for tls in item.get('spec', {}).get('tls', []) or []:
        name = tls.get('secretName')
        if name and name.endswith('-tls'):
            print(ns, name)
" | sort -u)

[ "${#TARGETS[@]}" -gt 0 ] || exit 1     # empty result means something is wrong

for entry in "${TARGETS[@]}"; do
  NS="${entry%% *}"; SECRET="${entry##* }"
  echo "$SRC_JSON" | NS="$NS" SECRET="$SECRET" python3 -c "
import json, sys, os
s = json.load(sys.stdin)
s['metadata']['namespace'] = os.environ['NS']
s['metadata']['name'] = os.environ['SECRET']
for k in ['resourceVersion','uid','creationTimestamp','managedFields',
          'annotations','ownerReferences','labels']:
    s['metadata'].pop(k, None)
print(json.dumps(s))
" | kubectl apply -f -
done

Two guards matter here: bail out if the source secret can’t be read, and bail out if the target list comes back empty. Both conditions mean something is broken, and both would otherwise produce a script that exits 0 having done nothing.

pre-flux-snapshot.sh

#!/bin/bash
export KUBECONFIG=/var/lib/microshift/resources/kubeadmin/kubeconfig
DESCRIPTION="vor-flux-update-$(date +%Y%m%d-%H%M%S)"
flux suspend ks --all -n flux-system 2>/dev/null
snap-all "$DESCRIPTION"
flux resume ks --all -n flux-system 2>/dev/null
echo "Snapshots erstellt: $DESCRIPTION"

check-backup-btrfs-subvolumes.sh

Verifies all Btrfs subvolumes are tracked in both btrbk.conf and snapper. Reports gaps via Telegram (triggered by systemd OnFailure):

#!/bin/bash
set -euo pipefail
BTRBK_CONF=/etc/btrbk/btrbk.conf
BTRFS_TOP=/mnt/btrfs-top

EXCLUDED=(btrbk_snapshots var_cache var_tmp var_lib_kubelet)
EXCLUDED_PATTERNS=("root-snap-pre-*")

MOUNTED_HERE=0
if ! mountpoint -q "$BTRFS_TOP" 2>/dev/null; then
  mount "$BTRFS_TOP"; MOUNTED_HERE=1
fi
trap '[[ $MOUNTED_HERE -eq 1 ]] && umount "$BTRFS_TOP" 2>/dev/null || true' EXIT

mapfile -t CURRENT_SUBVOLS < <(
  btrfs subvolume list "$BTRFS_TOP" \
    | awk '$7 == 5 && $NF !~ /\// { print $NF }' | sort)
mapfile -t CONFIGURED < <(
  grep -E '^\s+subvolume\s+' "$BTRBK_CONF" | awk '{print $2}' | sort)

is_excluded() {
  local sv="$1"
  for ex in "${EXCLUDED[@]}"; do [[ "$sv" == "$ex" ]] && return 0; done
  for pat in "${EXCLUDED_PATTERNS[@]}"; do [[ "$sv" == $pat ]] && return 0; done
  return 1
}

MISSING_BTRBK=()
MISSING_SNAPPER=()
for sv in "${CURRENT_SUBVOLS[@]}"; do
  is_excluded "$sv" && continue
  configured=0
  for conf in "${CONFIGURED[@]}"; do [[ "$sv" == "$conf" ]] && configured=1 && break; done
  [[ $configured -eq 0 ]] && MISSING_BTRBK+=("$sv")
  [[ -f "/etc/snapper/configs/$sv" ]] || MISSING_SNAPPER+=("$sv")
done

OVERALL_EXIT=0
[[ ${#MISSING_BTRBK[@]} -gt 0 ]] && { echo "WARN btrbk: ${MISSING_BTRBK[*]}"; OVERALL_EXIT=1; }
[[ ${#MISSING_SNAPPER[@]} -gt 0 ]] && { echo "WARN snapper: ${MISSING_SNAPPER[*]}"; OVERALL_EXIT=1; }
[[ $OVERALL_EXIT -eq 0 ]] && echo "OK: Alle Subvolumes in btrbk und snapper konfiguriert."
exit $OVERALL_EXIT

Backup & Disaster Recovery

Three independent layers protect the system. Each layer has a distinct job and they compose cleanly:

LayerToolScopeWhereRetentionUse case
Local timeline snapshotssnapperAll 8 subvolumesOn-disk .snapshots/24h hourly · 8d daily · 5w weeklyInstant rollback without unmounting
Local btrbk snapshotsbtrbkAll 8 subvolumesbtrbk_snapshots/Latest only (parent reference)Basis for incremental send to Pi
Remote backupbtrbk → PiAll 8 subvolumesRaspberry Pi via WireGuard24h hourly · 8d daily · 5w weeklyFull recovery after VM loss

snapper handles local rollback; btrbk handles the remote copy. Local btrbk snapshots are kept minimal — just one per subvolume as the parent reference so the next send to the Pi remains incremental instead of a full transfer.

What is Backed Up

SubvolumeMountpointsnapperbtrbk → PiContains
root/✅✅System, packages, /etc
home/home✅✅User home directories
data/data✅✅GitOps repo, scripts
var_lib_pvc/var/lib/pvc✅✅Kubernetes PVC data (app data)
var_lib_microshift/var/lib/microshift✅✅Cluster state, kubeconfig
var/var✅✅System state, boot backup
var_log/var/log✅✅Logs
var_lib_containers/var/lib/containers✅✅Container images

Not backed up intentionally: var_cache, var_tmp, var_lib_kubelet — ephemeral or rebuildable.

Pre-Update Snapshots

Before every GitOps update, a GitHub Action SSHs into the server and runs:

flux suspend ks --all -n flux-system
snap-all "vor-flux-update-$(date +%Y%m%d-%H%M%S)"
flux resume ks --all -n flux-system

snap-all creates one numbered snapshot across all 8 snapper configs simultaneously. If a Flux reconcile breaks something, you can roll back any or all subvolumes to the pre-update state within seconds — no mount required, just snapper -c <cfg> rollback <nr>.

Monitoring

check-backup-btrfs-subvolumes.sh runs daily via systemd timer. It verifies that every non-excluded Btrfs subvolume has both a btrbk entry and a snapper config. On failure, OnFailure=check-btrbk-notify.service sends a Telegram alert.

check-backup-btrfs-subvolumes.sh
# → OK: Alle relevanten Subvolumes sind in btrbk.conf konfiguriert.
# → OK: Alle relevanten Subvolumes haben eine snapper-Config.

Scenario A — Rollback (VM still running)

A1. /etc rollback via etckeeper

etckeeper commits every change to /etc as a git commit. Single-file recovery is a one-liner:

# See recent commits
etckeeper vcs log --oneline -10

# Restore a single file
etckeeper vcs checkout HEAD~1 -- /etc/microshift/config.yaml

A2. Root rollback via snapper

The fastest path for a broken system update is the GRUB menu. grub-btrfs adds every snapper snapshot automatically:

1. reboot
2. In the GRUB menu: "Fedora Linux snapshots"
3. Select the snapshot (e.g. "vor-flux-update-20260515-093000")
4. System boots read-only into the snapshot
5. If it works — make the rollback permanent:
snapper rollback   # sets snapshot as new default subvolume
reboot             # boots into the restored, writable system

Without rebooting into the snapshot first:

snapper -c root list
snapper -c root rollback 116   # rolls back root subvolume
reboot

A3. App data rollback from snapper

snapper snapshots are directly accessible as read-only directories — no mounting required:

# For any subvolume, the snapshots are here:
# /var/lib/pvc/.snapshots/<nr>/snapshot/
# /home/.snapshots/<nr>/snapshot/
# etc.

# Example: restore a PVC directory from snapshot 42
snapper -c var_lib_pvc list
kubectl scale deploy vaultwarden -n vaultwarden --replicas=0
cp -a /var/lib/pvc/.snapshots/42/snapshot/<pvc-dir>/. /var/lib/pvc/<pvc-dir>/
kubectl scale deploy vaultwarden -n vaultwarden --replicas=1

For apps with multiple PVCs (Nextcloud, Paperless), always restore all related PVCs in the same operation to avoid split-brain.

A4. Individual subvolume from btrbk snapshot

If the snapper snapshot is too recent or you need a specific point in the btrbk window:

mount /mnt/btrfs-top
btrbk list snapshots   # find the right snapshot name

systemctl stop microshift

# Save current state
btrfs subvolume snapshot /mnt/btrfs-top/var_lib_pvc \
  /mnt/btrfs-top/btrbk_snapshots/var_lib_pvc.recovery-backup

# Replace with snapshot
umount /var/lib/pvc
btrfs subvolume delete /mnt/btrfs-top/var_lib_pvc
btrfs subvolume snapshot \
  /mnt/btrfs-top/btrbk_snapshots/var_lib_pvc.20260515T0200 \
  /mnt/btrfs-top/var_lib_pvc
mount /var/lib/pvc

systemctl start microshift
umount /mnt/btrfs-top

Scenario B — Full VM Loss (restore from Pi)

B1. Provision new VM, boot Fedora LiveISO

dnf install -y btrfs-progs rsync gdisk

B2. Restore partition table

# Fetch partition table from Pi backup
scp <backup-user>@10.0.0.12:/backup/btrfs/server/var/<snapshot>/lib/boot-backup/partition-table.sfdisk /tmp/

# Restore (same disk size)
sfdisk /dev/vda < /tmp/partition-table.sfdisk

B3. Create filesystems with original UUIDs

mkfs.vfat -F32 /dev/vda1
mkfs.xfs -f /dev/vda2
xfs_admin -U <boot-uuid> /dev/vda2

# Btrfs: FIRST device only — see the note below
mkfs.btrfs -f -L "<fs-label>" -U <btrfs-uuid> /dev/vda3

The UUIDs must match the originals — they are baked into /etc/fstab and the GRUB config. Record them now, while the machine still exists: blkid and btrfs filesystem show / are useless after the fact. They belong in your password manager, not only in the backup you are trying to restore.

If your pool has two devices, create it with one and add the other. Passing both to mkfs.btrfs in a single call gives you Metadata: RAID1; adding the second afterwards keeps Metadata: DUP / Data: single, which is what the original had. Silently ending up with a different allocation profile is the kind of thing you discover months later.

mount -o subvolid=5 /dev/vda3 /mnt/restore
btrfs device add /dev/vdb /mnt/restore
btrfs filesystem df /mnt/restore     # Data: single / Metadata: DUP

B4. Mount top-level, receive subvolumes from Pi

/mnt/restore is already mounted from the previous step.

for sv in root var home var_log var_lib_microshift var_lib_pvc var_lib_containers data; do
  SNAP=$(ssh <backup-user>@10.0.0.12 "ls /backup/btrfs/server/${sv}/ | sort | tail -1")
  ssh <backup-user>@10.0.0.12 "btrfs send /backup/btrfs/server/${sv}/${SNAP}" \
    | btrfs receive /mnt/restore/
  # Convert read-only snapshot to writable subvolume
  btrfs subvolume snapshot /mnt/restore/${SNAP} /mnt/restore/${sv}_new
  btrfs subvolume delete /mnt/restore/${SNAP}
  mv /mnt/restore/${sv}_new /mnt/restore/${sv}
done

# Subvolumes that are deliberately NOT backed up still have to exist,
# because fstab mounts them:
for sv in var_cache var_tmp var_lib_kubelet btrbk_snapshots; do
  btrfs subvolume create /mnt/restore/$sv
done

Eight subvolumes come back from the Pi; four are recreated empty. var_cache and var_tmp are rebuildable, var_lib_kubelet is ephemeral pod state, and btrbk_snapshots is just the local snapshot target. Forgetting them means a boot that fails on mount -a.

B5. Restore /boot and reinstall GRUB

mount /dev/vda2 /mnt/boot
mount /dev/vda1 /mnt/efi
rsync -aAX /mnt/restore/var/lib/boot-backup/boot/ /mnt/boot/
rsync -aAX /mnt/restore/var/lib/boot-backup/efi/  /mnt/efi/

mount -o subvol=root /dev/vda3 /mnt/sysroot
# ... mount remaining subvolumes ...
for d in dev dev/pts proc sys run; do mount --bind /$d /mnt/sysroot/$d; done
chroot /mnt/sysroot grub2-install /dev/vda
chroot /mnt/sysroot grub2-mkconfig -o /boot/grub2/grub.cfg

touch /mnt/sysroot/.autorelabel   # trigger SELinux relabel on first boot
umount -R /mnt/sysroot
reboot

After reboot, SELinux relabels all files (a few minutes), then a second reboot brings the system fully up.

Recovery point

The Pi holds 24 hourly + 8 daily + 5 weekly snapshots per subvolume. You can recover to any point within that five-week window by selecting an older snapshot in step B4.



What I Got Wrong

Every item here cost me hours, and most of them share a shape: the system kept working, so nothing alerted, and I found out weeks later.

Btrfs quotas kill etcd. Don’t enable them. Hundreds of snapshots plus hourly btrbk cleanup plus qgroup rescans will crash MicroShift reliably, roughly fifteen minutes after every btrbk run. The symptom looks like a cluster problem, not a filesystem one. Full explanation in Layer 1.

MicroShift SCCs are opt-in. Every service account running a privileged workload needs an explicit binding. Don’t assume privileged is inherited. Check with kubectl auth can-i use scc/privileged --as=system:serviceaccount:<ns>:<sa>.

nginx in OpenShift must run as non-root. The stock nginx:alpine image listens on port 80 and wants root. Use nginxinc/nginx-unprivileged:alpine on 8080; the HAProxy router terminates TLS and proxies with the original Host header.

kindnet’s POD_SUBNET must match MicroShift’s pod CIDR. kindnet defaults to 10.244.0.0/16; MicroShift uses 10.42.0.0/16. A mismatch breaks masquerading in a way that stays invisible until something makes a decision based on the client IP — then it looks like an application bug. I carried a boot-time patch service, then an override manifest, then a kustomizePaths exception, before reading the generator script and finding that it already computed the correct CLUSTER_CIDR three lines above where it hard-coded kind’s default. Two lines fixed it at the source. Check the shipped manifest before building any workaround.

Use hostNetwork: true for mail ports, not hostPort. hostPort routes new connections through the pod network bridge, so the source IP arriving at Postfix or Dovecot is the bridge address — fail2ban is blind to the attacker. hostNetwork: true preserves source IPs end to end. Set dnsPolicy: ClusterFirstWithHostNet so cluster DNS still resolves.

firewalld inter-zone forwarding requires policies. --add-forward-port and zone rules will not move traffic between zones. Create explicit Policy objects with ingress and egress zones.

Dovecot 2.4 is a breaking upgrade from 2.3. Dozens of options renamed, removed, or given new defaults. Budget real time. Specific trap: MySQL passwords in mysql{} blocks cannot use %{env:VAR} — mount them as secret files.

A %post scriptlet can restart your cluster. The MicroShift RPM restarts crio and microshift when it is installed. Combined with unattended updates and a nightly package feed, that means an unannounced cluster restart every night onto a build nobody reviewed. Disable the repositories and update deliberately — see Operations.

dnf install does not upgrade, and disabled repos serve stale metadata. Two separate traps in the same script. For an installed package, dnf5 says “already installed” and does nothing — pass the exact version from repoquery --upgrades. And once a repository is disabled, nothing refreshes its metadata, so repoquery answers from yesterday’s cache; mine concluded the newest build was already installed and exited successfully having done nothing. Use --refresh everywhere.

Applying a manifest is not the same as it taking effect. MicroShift applies the manifests from kustomizePaths at every start, but whether that updates existing resources turned out to be unreliable — the same server-side apply, logged as successful, took effect for one DaemonSet and not another. Result: a kube-proxy that ran 86 days on an image from an RPM I had already replaced, while the logs said the apply succeeded. My upgrade script now re-applies the paths itself.

A hard-coded list is a list you will forget. My TLS sync script carried the target namespaces inline; the failure mode of forgetting an entry is a certificate warning that a family member discovers before you do. Deriving the targets from the cluster’s own Ingress objects removed the class of error entirely. The same reasoning applies to package excludes — disable the repository instead of listing every package name.

Pods don’t restart when a ConfigMap changes — but that no longer has to be manual. This used to be a standing “remember to kubectl rollout restart” note, which is exactly the kind of instruction that gets skipped. Reloader does it now for workloads that opt in; the ones that don’t opt in are listed explicitly, so the exception is visible rather than assumed.


Reference

configuration/
├── kustomization.yaml          ← Flux entry point
├── flux-system/                ← Flux own manifests
├── acme-dns/                   ← acme-dns service
├── cert-manager/               ← cert-manager + ClusterIssuer + Certificate
├── pihole/                     ← DNS server
├── vaultwarden/                ← Password manager
├── paperless/                  ← Document management
├── nextcloud/                  ← File sync + MariaDB
├── collabora/                  ← Online Office
├── mailstack/                  ← Postfix + Dovecot + Rspamd + ClamAV + Redis + MariaDB + PostfixAdmin
├── monitoring/                 ← Alloy + kube-state-metrics
├── monitoring-grafana/         ← Grafana Operator + dashboards
├── homepage/                   ← Static website
├── matomo/                     ← Self-hosted web analytics + MariaDB
├── crowdsec/                   ← CrowdSec LAPI + agent (Helm)
├── reloader/                   ← Restart pods on ConfigMap/Secret change (Helm)
├── dmarc/                      ← Daily DMARC report fetch + parsing
├── etcd-backup/                ← Daily application-consistent etcd snapshot
├── local-path-provisioner/     ← Storage provisioner
└── mcp-openshift/ mcp-paperless/ mcp-email/   ← MCP servers for Claude Code

Every application directory has a sibling <name>-ks.yaml — the Flux Kustomization that points at

Scripts

All maintenance scripts live in /data/scripts/ (Btrfs-backed, included in the btrbk backups) with symlinks in /usr/local/bin/, and run as systemd oneshot services driven by timers. The full inventory is in Operations.

/data is itself the GitOps repository checkout, so the scripts are versioned alongside the manifests — one git push covers both.

My MicroShift packages

MicroShift is not in Fedora’s repositories, and the community nightly that used to build the 4.y stream moved on to 5.x in May 2026. Since then I build my own:

COPR projectmkocustos/microshift-nightly-4x — chroots fedora-43-x86_64, fedora-44-x86_64
Build definitionmkocustos/microshift, branch copr-nightly-4x — one added workflow file on top of the upstream packaging repo
Upstream projectmicroshift-io/microshift

You are welcome to enable the COPR, with the caveat that it exists to serve one homelab: it builds whatever the newest release-4.y branch with an OKD payload happens to be, on my schedule, with no guarantees at all.

Two patches went upstream out of this work:

  • #234 — build the newest 4.y release branch alongside 5.x, which is what my fork’s workflow does. Open.
  • #236 — set kindnet’s POD_SUBNET from the CLUSTER_CIDR the generator script already defines, instead of hard-coding kind’s default. This is the fix that removed the workaround described above.

Worth being precise about where the build happens, because it is easy to assume wrongly: a GitHub Actions runner produces the source RPM and submits it with copr-cli; COPR’s build workers compile the binary RPMs; my server only consumes the finished repository. Nothing is compiled on the runner, and nothing is compiled on the server.