Backup & restore
The platform’s own footprint is stateless — every control-plane component reconciles from the Kubernetes API server, so it rebuilds itself on a fresh install. What you must back up is the data plane you bring plus one small in-cluster secret without which encrypted credential data is unrecoverable. This page draws the line between derivable and must-be-restored.
What must be backed up (and what doesn’t)
Section titled “What must be backed up (and what doesn’t)”| Asset | Back up? | Why |
|---|---|---|
| Control-plane Postgres | Yes | The system-of-record: durable runs, prompt/agent versions, datasets + labels, cost rollups, the audit log. Not derivable. |
| The credential-store data (Postgres grants, or the Kubernetes-Secret grants) | Yes | Per-user OBO tool grants. Losing them = every user must re-consent. |
| The KEK (LocalSealer Secret, or the external KMS/transit key) | Yes — critically | Credential ciphertext is inert without the KEK. A Postgres backup without the KEK is unrestorable credential data. |
| Object store (S3-compatible) | Yes (your bucket’s own policy) | Blobs, datasets, replay fixtures, checkpoints. |
| State Layer (Valkey/Redis) | Optional | Session memory + quota accumulators. Sessions fail open on loss; quota windows reset (acceptable). Durable runs are in Postgres, not here. Snapshot only if you want warm session continuity. |
| CRDs / manifests | No (derivable) | Your AgentDeployments and policies live in git / GitOps. Re-apply them; the controller rebuilds agents. |
| Control-plane pods / gateway config | No (derivable) | Reconciled from the CRDs on a fresh install. |
The capability keypair (bff-capability Secret) |
Yes if BYO; else regenerate-safe | The hook never re-keys an existing key. If you generated it in-cluster, back up the Secret so a rebuild keeps existing OBO grants valid; a fresh key invalidates in-flight grants. |
Backing up
Section titled “Backing up”Postgres (control plane + credential grants):
pg_dump "$CONTROLPLANE_DSN" > controlplane-$(date +%F).sql# and the credential-store DSN if it's a separate PostgresThe KEK (LocalSealer) — the one people forget:
There is no fixed name for this Secret: you chose it when you created the CredentialStore, so
find it rather than guessing. Discover the name, then back it up from the platform credential
namespace (ctxmesh unless you set bff.mcp.credentialNamespace):
kubectl get clustercredentialstores,credentialstores -A -o jsonpath=\'{range .items[*]}{.kind}{"/"}{.metadata.name}{"\t"}{.spec.provider.postgres.encryption.localKEKSecretRef.name}{"\n"}{end}'
KEK=<the name that printed>kubectl get secret "$KEK" -n ctxmesh -o yaml > "kek-$KEK-$(date +%F).yaml"If that command prints nothing, no CredentialStore is configured and credentials are held by the
built-in Kubernetes backend — there is no KEK to back up, and the Secrets themselves are the data.
Store the KEK backup separately from the Postgres backup and with tighter access — together they are
the credential data in the clear. For an external KMS/transit custodian (OpenBao transit, cloud
CMK), the “backup” is that service’s own key durability/backup policy — the ciphertext in Postgres
references a key_id that must still exist. Do not delete a per-tenant transit key you still need
(that is crypto-shredding — irreversible by design).
State Layer (only if you want warm sessions): never cp /data. Trigger a snapshot and confirm it
completed before copying:
valkey-cli BGSAVEvalkey-cli INFO persistence # poll rdb_bgsave_in_progress:0 / rdb_last_save_time before snapshottingFor the in-cluster persistent tier, AOF gives restart durability; a network-attached PVC lets the volume follow a reschedule (a ReadWriteOnce local volume pins the pod to one node).
Restoring
Section titled “Restoring”- Fresh install the platform (
helm install) — control plane rebuilds stateless. - Restore Postgres (
psql < dump.sql) before the control plane needs it, or pointCONTROLPLANE_DSNat the restored instance. - Restore the KEK Secret (
kubectl apply -f kek-*.yaml) — credential ciphertext is unreadable until it’s back. For an external custodian, ensure the referenced keys still exist. - Restore the capability keypair Secret if you’re keeping existing OBO grants valid.
- Re-apply your CRDs (GitOps /
kubectl apply) — the controller reconstructs agents, routes, policies, tenants. - Restore the object store / State Layer if you snapshotted them.
Prove the round-trip
Section titled “Prove the round-trip”Back-and-restore is only real if you’ve exercised it. A minimal drill: write a credential grant and some session memory, take a Postgres dump + the KEK backup, tear down, restore into a fresh install, and confirm the grant resolves and the run history is intact. Treat “restore” as a tested runbook, not a theory.