This repository is now the public half of a split: six documents saying what to decide, and one gap. How the cluster is built moved to jeffry/homelab-impl, which is private because its README is an inventory of chart versions and image tags. History starts here deliberately, and not as tidiness. The previous history contained that inventory, and this repository is public — a deletion commit would have removed it from the tree and left it in the log. A fresh root can only carry what is in it. Two pointers rewritten rather than deleted: private-access.md and .loom/README.md both directed a reader to TAILSCALE.md and README.md at the root, which are now private. They now say a fuller reference exists, that it is private, and how to ask — because a public page naming a private thing as its answer is the failure this project has now hit four times. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
4.1 KiB
Get the data back after losing the machine it was on
Two halves that fail differently, and are worth asking about separately — because having one and not the other looks from the outside exactly like having both.
- The database — base backups and continuous archiving, by the operator's plugin, to object storage.
- The volume — a nightly tarball of the disk, by a cron job, to the same place.
Restoring the database without the volume gives you a service listing things that no longer exist. The database holds the metadata; the artifacts are on the disk.
The volume job runs half an hour before the database backup, on purpose. A database referencing an artifact that exists is recoverable. An artifact missing from an older database is not.
Nobody has ever performed a restore here
This is the most important open item and it is stated rather than implied.
Rehearsing one is safe: the operator restores by bootstrapping a new cluster into a throwaway namespace and never touches the running one. Rehearse the tarball too — untar into a scratch volume and confirm the objects are intact.
The failure that reported success for a month
This cluster had no usable backup for over a month while every status field said healthy.
Two faults at once. The archive path named its bucket twice — the endpoint
already selected it — so everything landed one directory deeper than anyone would
look. The tool wrote and read at the same wrong path, so continuous archiving
reported True throughout. And the only backup resource used the deprecated
in-tree method, which cannot succeed against a plugin-configured cluster; it
failed instantly and was never retried.
Status conditions confirm the archive is healthy at whatever location it is actually using, which is not necessarily the one you configured.
So there are three checks and only the third would have caught it:
# 1. conditions — want ContinuousArchiving=True and LastBackupSucceeded=True
kubectl get cluster <name> -n <ns> \
-o jsonpath='{range .status.conditions[*]}{.type}={.status}{"\n"}{end}'
# 2. recovery window — populates on a delay; empty right after a backup is normal
kubectl get objectstore <store> -n <ns> -o jsonpath='{.status.serverRecoveryWindow}{"\n"}'
# 3. THE ONE THAT MATTERS — are the objects where you think they are?
aws --endpoint-url https://<region>.digitaloceanspaces.com s3 ls s3://<bucket>/<prefix>/ --recursive | tail
List the bucket. Checks 1 and 2 reported healthy the entire time.
Four things that are not where the documentation puts them
Use the plugin method. The in-tree field still exists and most posts still show it; on a plugin-configured cluster it can never succeed.
The retention policy is a sibling of the configuration, not inside it. Nesting it is rejected outright — and without it nothing prunes, and the bill grows without bound.
The scheduled-backup cron has six fields. The leading one is seconds. Five fields will not mean what you expect.
One archive path per database, ever. Two sharing a path corrupt each other. This also bites on rebuilds — recreating with the same name reuses the server name at the same path, and the new system ID mixes timelines with the old archive. Clear the prefix, or use a new one.
The tarball has two rough edges, both accepted
It installs its own tools at runtime, because the base image ships neither
tar nor gzip. It fails loudly if that fails, but it adds a dependency on a
package repository being reachable at 02:00.
It is taken from a live volume. Git objects are write-once so this is safe in practice, but a write landing mid-snapshot can be caught partially. For a homelab that beats nightly downtime; scale to zero first if a strictly consistent copy is ever needed.
Checking this is still true
Verified 2026-09-03, when both schedules were created.
kubectl get scheduledbackup,backup -A
kubectl get cronjob -A
kubectl create job -n <ns> --from=cronjob/<name> verify-$(date +%s) # prove it, don't wait for 02:00