kodebeat / themes / platform-sre

Does the backup actually come back?

Operational scar tissue turned into tools: the untested restore, the lost evidence, the diagram that no longer matches the cluster.

Before, during and after the incident

Most of these exist because something went wrong once and the tooling to understand it did not exist at the time. Anything not publicly released is marked.

Before

Prove the assumptions while it is cheap. A restore that has never run is a hypothesis, and a diagram nobody checks is fiction.

restores the backup on a schedule and smoke-tests it backup-verify C4 diagrams from Terraform state and a live cluster arch-map architecture documentation with a drift check automap mirrors production traffic and diffs the responses dark-canary timezone-aware cron with DST warnings cron-translate

During

Collect the evidence before the fix destroys it. A support bundle taken in the first five minutes answers questions the post-mortem cannot reconstruct a week later.

one-command Kubernetes support bundle cluster-info-collector network path forensics mtr-toolbox

After

The follow-up work that decides whether the same incident happens twice.

rightsizing as committable patch YAML k8s-rightsizer-report environment differences, with secrets masked by default envdiff keeps a day's port-forwards alive port-forward-manager kubectl, with the sharp edges filed off kubectl-plus failure injection chaos-toolbox

The papers

Every paper carries the methodology behind any number it emits, and what the tool deliberately does not do. Pick them up from the catalogue.

  • Capturing the evidence before the fix destroys ita one-command Kubernetes support bundle
  • Answering the operational questions engineers get wrong under pressureclient-side browser tools for platform engineers
  • Architecture documentation that can be checked line by linea stdlib-only architecture mapper with a drift check
  • Keeping the architecture diagram true to what is deployeda C4 container diagram generated from Terraform state and a live cluster
  • Proving a rewrite behaves identically, on production traffica traffic-mirroring proxy with a structural response differ
  • Turning Kubernetes usage data into a merged requests changea rightsizing report that emits committable patch YAML
  • Proving the backup restores, nightly, with nobody watchinga scheduled restore-and-smoke-test harness
  • Finding the configuration difference between two environments, safelyan environment-variable differ with secret masking on by default
  • Keeping a working day's kubectl port-forwards alivea supervised kubectl port-forward manager
  • The scheduled job that silently skips when the clocks changea cron translator with timezone-aware next runs and DST warnings

The questions in between, in your browser

Cache headers, CronJob schedules, probe timings, regex backtracking, NetworkPolicy reachability: one input, one clear answer, shareable by URL. No accounts, no backend, no telemetry.

Open the toolbox

Working on this?

Tell me what you are looking at and I will tell you honestly whether any of this helps. No pitch attached, and the papers are free either way.

Send a request

Send a request

One message, no follow-up sequence.

Your address is used for this request and nothing else.