Own infrastructure
Servers, monitoring, encrypted backups and restore checks
The servers my services run on. Backups are encrypted at the source, wait out receiver downtime in a local spool and are delivered once the connection is back, and restores are tested every day rather than on the day of an outage.
Below is the path of one copy, from the server that took it to the daily restore check. Shown separately is what happens while the receiver is unavailable, and what backs each decision.
- Servicesdatabase dumps and configuration on every server
- Encryptionthe copy is encrypted with age at the source
- Spoolsent every 20 minutes, delivery confirmed by a sha256 match
Failure path
- The receiver is unavailable
- the copy stays in the spool
- delivery resumes once the connection is back
- Receiverstores encrypted copies and alone holds the decryption key
- Restore checkrestores the freshest dump daily and compares row counts
Alongside
- Prometheus, Alertmanager, Grafanaservice metrics and alerts
- Standby environmentrunbook and standby host for moving the main bot
Verifiable decisions
Key
Only the receiver holds the private key. The server that made a backup can no longer read it.
Delivery
A delivery counts only once the sender's and the receiver's hashes match. The alert fires on the age of the last confirmed delivery, not on files waiting in the queue.
Restore
Every day a dump is restored into a disposable PostgreSQL of the same major version as production, and row counts are compared with the previous run. The run on 2026-09-23 passed.
In detail
Backups
Every server takes its own backups and encrypts them with age straight away. Only the receiver holds the private key, so a source cannot read even its own older copies. Losing a server does not expose the history of its data.
Delivery through a spool
A separate receiver may be temporarily unavailable. So a copy first lands in a local spool, and a timer sends it every 20 minutes: while the receiver is away, copies wait in the spool and leave once the connection is back.
A delivery counts only after the hash of what was received matches the hash of what was sent. The receiver may accept a truncated stream too, so an “accepted” reply proves nothing. The sender has its own alert on the age of the last confirmed delivery: a queue in the spool is normal, a stale delivery raises the alert.
Restore checks
A backup that has never been restored is a hope, not a backup. Every day a timer decrypts the freshest copies, restores the database dump into a disposable container of the same major PostgreSQL version as production, and compares row counts with the previous successful run.
Outages
Every serious outage gets a write-up, including my own mistakes. One of those write-ups is below.
The backup job that deleted the system
The server answered ping and accepted SSH, yet could not start a single new process. Two hours went into a disk-failure theory. The cause was my own backup delivery script, deployed four minutes before the outage.
- Backup delivery script deployed
- 4 minutes later
- No new process can start
- Recovery from a rescue image