Files
homelab-cluster/.loom/published/recovering-data.md
T
jeffryandClaude Opus 5 cc66e227a3 the published surface, on a fresh root
This repository is now the public half of a split: six documents saying what to
decide, and one gap. How the cluster is built moved to jeffry/homelab-impl, which
is private because its README is an inventory of chart versions and image tags.

History starts here deliberately, and not as tidiness. The previous history
contained that inventory, and this repository is public — a deletion commit would
have removed it from the tree and left it in the log. A fresh root can only carry
what is in it.

Two pointers rewritten rather than deleted: private-access.md and .loom/README.md
both directed a reader to TAILSCALE.md and README.md at the root, which are now
private. They now say a fuller reference exists, that it is private, and how to
ask — because a public page naming a private thing as its answer is the failure
this project has now hit four times.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-07 15:19:14 -04:00

4.1 KiB

Get the data back after losing the machine it was on

Two halves that fail differently, and are worth asking about separately — because having one and not the other looks from the outside exactly like having both.

  • The database — base backups and continuous archiving, by the operator's plugin, to object storage.
  • The volume — a nightly tarball of the disk, by a cron job, to the same place.

Restoring the database without the volume gives you a service listing things that no longer exist. The database holds the metadata; the artifacts are on the disk.

The volume job runs half an hour before the database backup, on purpose. A database referencing an artifact that exists is recoverable. An artifact missing from an older database is not.


Nobody has ever performed a restore here

This is the most important open item and it is stated rather than implied.

Rehearsing one is safe: the operator restores by bootstrapping a new cluster into a throwaway namespace and never touches the running one. Rehearse the tarball too — untar into a scratch volume and confirm the objects are intact.

The failure that reported success for a month

This cluster had no usable backup for over a month while every status field said healthy.

Two faults at once. The archive path named its bucket twice — the endpoint already selected it — so everything landed one directory deeper than anyone would look. The tool wrote and read at the same wrong path, so continuous archiving reported True throughout. And the only backup resource used the deprecated in-tree method, which cannot succeed against a plugin-configured cluster; it failed instantly and was never retried.

Status conditions confirm the archive is healthy at whatever location it is actually using, which is not necessarily the one you configured.

So there are three checks and only the third would have caught it:

# 1. conditions — want ContinuousArchiving=True and LastBackupSucceeded=True
kubectl get cluster <name> -n <ns> \
  -o jsonpath='{range .status.conditions[*]}{.type}={.status}{"\n"}{end}'

# 2. recovery window — populates on a delay; empty right after a backup is normal
kubectl get objectstore <store> -n <ns> -o jsonpath='{.status.serverRecoveryWindow}{"\n"}'

# 3. THE ONE THAT MATTERS — are the objects where you think they are?
aws --endpoint-url https://<region>.digitaloceanspaces.com s3 ls s3://<bucket>/<prefix>/ --recursive | tail

List the bucket. Checks 1 and 2 reported healthy the entire time.

Four things that are not where the documentation puts them

Use the plugin method. The in-tree field still exists and most posts still show it; on a plugin-configured cluster it can never succeed.

The retention policy is a sibling of the configuration, not inside it. Nesting it is rejected outright — and without it nothing prunes, and the bill grows without bound.

The scheduled-backup cron has six fields. The leading one is seconds. Five fields will not mean what you expect.

One archive path per database, ever. Two sharing a path corrupt each other. This also bites on rebuilds — recreating with the same name reuses the server name at the same path, and the new system ID mixes timelines with the old archive. Clear the prefix, or use a new one.

The tarball has two rough edges, both accepted

It installs its own tools at runtime, because the base image ships neither tar nor gzip. It fails loudly if that fails, but it adds a dependency on a package repository being reachable at 02:00.

It is taken from a live volume. Git objects are write-once so this is safe in practice, but a write landing mid-snapshot can be caught partially. For a homelab that beats nightly downtime; scale to zero first if a strictly consistent copy is ever needed.


Checking this is still true

Verified 2026-09-03, when both schedules were created.

kubectl get scheduledbackup,backup -A
kubectl get cronjob -A
kubectl create job -n <ns> --from=cronjob/<name> verify-$(date +%s)   # prove it, don't wait for 02:00