The backup job that deleted the system
The server answered ping and accepted SSH, yet could not start a single new process. Two hours went into a disk-failure theory. The cause was my own backup delivery script, deployed four minutes before the outage.
Symptom
Processes that were already running kept working: SSH accepted connections, the web server answered,
open network connections held, ping went through. The machine looked alive. But starting any new process failed
with No such file or directory: bash, sudo, sed. SSH got through authentication and then dropped
when it tried to start the shell.
The wrong hypothesis
The picture resembles a lost volume: files cannot be read, but whatever is already in memory keeps running. I spent two hours looking for a storage failure. Side signals misled me too: the hosting panel showed a stale last-start time.
The timing was right in front of me: the server died four minutes after a deploy.
Cause
The script that evicts old copies from the send queue selected files with
ls -1t "$dir"/*.*. With nullglob enabled and an empty directory, the pattern disappears entirely and
ls runs with no arguments. With no arguments it lists the current directory, and for a systemd
service the current directory is /.
That list went to rm -f as “the oldest copies”. The oldest entries in the root turned out to be the
symlinks /sbin, /lib and /lib64. rm -f leaves directories alone, but the links were enough:
without the loader, no binary can start.
Fix
The server was recovered from a rescue image: the root volume was mounted so the kernel could replay the unfinished ext4 journal, and the deleted links to the system directories were restored.
Preventing a repeat
- After a sudden outage, check your own latest change first, and only then the infrastructure.
- A pattern that can vanish must never reach a destructive command.
- The outage exposed a gap in the backups: the main bot only had a database dump, without its configuration. The configuration now goes to the standby environment as well.