A game day run once proves the system survived that afternoon. Firedrill runs the same failures every week against a namespace or a compose stack, scores what came back, and keeps a dated record an auditor can be handed.
shape: monthly subscription, per estate grows from: chaos-toolbox, backup-verify, dark-canary
run 2026-09-21 · 12 failure modes · blast radius within limits
survived 7 degraded 2 failed 3 previous run: 6 · 2 · 4
record: evidence/checkout-prod/2026-09-21
Example output. The last line is the point: a dated record of what was tested and what came back.
Readings from the last run.
The problem
Most teams have run a game day. Almost none have run the same one twice. A pod was killed in a staging namespace before the launch, the service recovered, and the slide has said the platform is resilient ever since. The dependency graph has changed, a sidecar was added, a timeout was tuned, and the claim has not been re-examined.
The tools that inject failure exist and are free: chaos-toolbox, Chaos Mesh, LitmusChaos, the fault-injection services the cloud providers sell. What they produce is an experiment log for the engineer who ran them. What a director, an auditor or an insurer asks for is different: which failure modes the estate survives today, which it does not, and when that was last checked.
DORA asks financial entities and their suppliers for resilience testing with evidence, and NIS2 asks for continuity and incident-handling measures that hold under stress. Both turn an occasional exercise into a recurring obligation, and a recurring obligation needs a record.
Failure scenarios
Each scenario runs in isolation against the real estate, not a mock and not a simulation, and each one carries a declared blast radius and a stop condition.
Terminates a pod or container at random, then measures recovery time and how requests were routed meanwhile.
Consumes memory until the container sits at 90% utilisation. Tests OOM handling and graceful degradation.
Fills the filesystem to 95% capacity. Reveals missing disk alerts and write-path failures.
Injects 500ms of latency on inter-service traffic. Tests timeout configuration and retry logic.
Blocks DNS lookups for targeted services. Exposes missing caches and cascade risks.
Replaces a service certificate with an expired one. Tests TLS error handling and alerting.
Takes a downstream dependency fully offline. Validates circuit breakers and fallback paths.
Isolates a service from the rest of the cluster. Tests split-brain behaviour and leader election.
Pins CPU to 100% on a target container. Checks autoscaling triggers and request queuing.
Invalidates API keys and service-account tokens mid-request. Tests token refresh and auth fallbacks.
Shifts the system clock forward by five minutes. Reveals time-dependent logic and JWT validation gaps.
Injects invalid values into environment variables and config maps. Tests startup validation and safe defaults.
How it works
firedrill init reads a Kubernetes namespace or a docker-compose file and lists the services, their health checks and their dependencies. Anything it cannot infer, the operator records once.
Each failure mode runs in isolation, on the cadence set: a pod killed, a disk filled, latency added to one hop, DNS withheld, a certificate expired, a credential revoked, a dependency made unavailable, a clock skewed. Health checks and error rates are watched throughout, and the run stops the moment a blast radius exceeds the limit declared for it.
Each mode is marked survived, degraded or failed, with the time to first impact and the time to recovery. The score compares run to run because the failures and the thresholds only change when the plan does.
Every run leaves a dated record: what was injected, what was observed, what broke, and the remediation the engine suggests. The record is the deliverable, and it lives in the same dashboard as Lastresort's restore log, so resilience and recovery evidence sit side by side.
Detailed reports
A failed row is not the deliverable. Each mode that breaks produces a report: what was injected, what it took down and when, what that cost, and the remediation to apply.
DNS resolution for payment-service.internal blocked for 30 seconds, with the rest of the namespace left alone.
A cascade. The auth service failed its health checks at T+4s, having no cached resolution to fall back to. The API gateway returned 502 to every request at T+8s. Both stayed down until the injection window closed at T+30s.
One internal name stopped resolving and checkout stopped serving eight seconds later. No alert fired on the auth health checks, so the outage would have been found by a customer rather than by a pager.
dnsmasq or CoreDNS with the cache plugin.T+0s T+4s T+8s T+30s
T+4s auth down · T+8s full outage · T+30s injection ends
Example output. The row in the scorecard is the same drill: the report is what the row expands into.
Who it is for
A game day was run before the last launch and nothing since. The estate has changed and the claim has not been re-examined.
A financial-services customer, an auditor or an insurer asks how resilience is tested and when. A dated log of drills answers the question a slide from last year cannot.
What it costs
One price covers a namespace or a compose stack, every failure mode, on any cadence. Charging per drill would price the thing worth doing more often. The request form asks for the shape of the estate so the tiers fit real ones.
Request access
Requests decide which failure modes and which platforms are supported first.
Free for the whole beta. The first twenty-five teams keep 50% off for twelve months after launch, locked in at signup.
Reasonable objections