Infrastructure, automation and reliable systems

I design, automate and run services and infrastructure so they can be verified, restored and understood after a failure.

Selected work

TTVideo

Media download service for Telegram

Send the bot a link and the file arrives in the chat. The heavy lifting runs in queues on separate workers, so one slow request never holds up the rest, and failures are caught by alert rules that live in the repository alongside their tests.

Try it@videott_downloadbot
Status
maintained
Role
Author and operator
Period
2026
Updated
Stack
TypeScript · Node.js · BullMQ · Redis · PostgreSQL · Docker · Prometheus
Verified by
code · tests · incident
How it works
  1. Telegramdelivers updates to a secret-checked webhook
  2. Botverifies the secret and queues a job
  3. Redistwo queues, video and music, plus a cache of sent files
  4. Workers ×3download and ffmpeg; tracks go through a separate engine, one at a time
  5. Chatthe finished file is sent to the user

Alongside

  • PostgreSQLusers and their settings
  • Prometheus and Alertmanagermetrics from every container, rules with tests
  • Webhook watchdogchecks the webhook address on a timer

Verifiable decisions

  1. Work is split by cost. Music tracks get their own queue, one at a time per worker, so they never take video slots.

    code

  2. Every outbound address is checked at DNS resolution time, on every redirect hop. A string-based address check had already been bypassed with an IPv4-mapped IPv6 address.

    tests

  3. A watchdog checks the webhook address on a timer and raises an alert if it changes. It exists because of a real webhook hijacking incident.

    incident

  4. Alert rules are kept in the repository and pass promtool tests before they ship.

    tests

About the project

Own infrastructure

Servers, monitoring, encrypted backups and restore checks

The servers my services run on. Backups are encrypted at the source, wait out receiver downtime in a local spool and are delivered once the connection is back, and restores are tested every day rather than on the day of an outage.

Status
maintained
Period
2026
Verified by:
code, measured run

About the project

How it works
  1. Servicesdatabase dumps and configuration on every server
  2. Encryptionthe copy is encrypted with age at the source
  3. Spoolsent every 20 minutes, delivery confirmed by a sha256 match
  4. Receiverstores encrypted copies and alone holds the decryption key
  5. Restore checkrestores the freshest dump daily and compares row counts

GameBoost

Automatic Windows tuning for games

Status
maintained
Period
2026
Verified by:
tests, release
  1. Every change is written to a state file, and the rollback after a game exits follows it. The state file was fuzzed for 700 iterations without a single exception.

    tests

  2. Release v1.1.0 shipped only after green CI and a review pass with no new defects. 348 tests, 82.5% coverage.

    release

About the project

All projects5 projectsAlso: LogiTrack, Offline geolocation from a map screenshot

Notes

The backup job that deleted the system

The server answered ping and accepted SSH, yet could not start a single new process. Two hours went into a disk-failure theory. The cause was my own backup delivery script, deployed four minutes before the outage.

The week my monitoring showed zero

The workers were alive, the dashboards showed zero downloads, and the only alert that fired pointed in the wrong direction. Why Prometheus service discovery quietly returned an empty list, and why nothing treated that as an error.

How I work

  1. Evidence over confidence

    A check that has never failed has not checked anything yet. Before I trust a green status, I make sure it can turn red.

  2. Rollback before change

    First it is clear how to undo, then what to change. An installation that breaks halfway must be repaired by the next run.

  3. A write-up after every failure

    Symptom, wrong hypothesis, cause and protection against a repeat are written down next to the code, including my own mistakes.

Contact

Got a project?

Available for Telegram bots, backends, deployment and server maintenance.