AI writes the code.
AI SREs run production.
steadybit proves it survives failure — before it ships.
That stage is Validate. We build it.
01 / what changed
Your engineers merge code they didn’t fully write, at a rate no review process was designed for. The Faros AI Engineering Report 2026 — 22,000 developers, 4,000 teams, two years of data — measured a 34% increase in task throughput per developer and a 210% increase in code-related tasks per team.
Every one of those changes is an availability risk. And now they arrive at agentic speed.
Your pipeline moves faster than it ever has. It does not move safer.
The industry’s answer lives at the far right: detection, correlation, root cause, MTTR. On-call that pages the right person, incidents that explain themselves, recovery that gets faster every quarter.
All of it useful. All of it downstream of the customer already being affected.
And notice what those same tools now promise: catching production issues before customers notice. Prevention is on everyone’s roadmap. Nobody on that side of the pipeline has the mechanism.
Prevention without controlled fault injection is prediction, not proof.
02 / the stage
A stage in your delivery pipeline, before deploy, where a change is proven to survive real failure conditions — dependency loss, latency, resource pressure, zone failure — inside a blast radius you define, with findings that go back to the team that shipped it.
Not an opinion about resilience. Evidence, per change, at the pace your pipeline sets rather than the pace your calendar allows.
Resilience that was never tested isn’t resilience. It’s an assumption that hasn’t been contradicted yet.
It runs as a step in the pipeline you already have. A pull request check before merge. A validation against your pre-prod, preview or canary environment after deploy. Triggered by your CI, by API, or by an agent through MCP.
The run produces an artifact: what was injected, what held, what didn’t, and which team owns the finding. That artifact is what gates the deploy.
200+ integrations, because this has to work on heterogeneous enterprise stacks — Kubernetes and VMware and Kafka and the JVM services nobody wants to touch — not a clean demo cluster.
Failed
What a Validate run gives back to the team that shipped it.
It isn’t incident response. Incident tooling handles the failure you’re already having. This is about the one you haven’t had yet.
It isn’t a game day. A quarterly exercise can’t cover a codebase that changes hourly. If verification only happens when someone schedules it, it happens at human speed — and human speed is exactly the assumption that broke.
It isn’t a chaos engineering programme. A programme is something you staff. A stage is something your pipeline runs.
It isn’t an AI SRE. AI SREs run your production. Validate proves it won’t fall over. Their telemetry feeds our steady-state checks; their incidents become our experiments; our experiments verify their fixes. Different stage, same pipeline.
It isn’t autonomous. Not today. More on that below.
Two of those have an industry building them. The third is the one nobody built.
03 / where we stand
Injecting failure is easy to demo. Stopping safely is the hard part, and that’s the part we build. Pre-flight checks before anything runs. A blast radius you define. Automatic halt the moment your metrics go unhealthy. 200+ integrations, because Validate has to run against Kubernetes and VMware and Kafka and the JVM services nobody wants to touch. On-prem from day one, full feature parity.
This is the unglamorous half of the product, and it is the half that takes years of other people’s outages to get right.
SteadyBuddy, our assistant, is live — and today it’s assistive and read-only. It knows your environment, proposes experiments based on real runs in it, and turns a plain description into an executable one. It suggests; you decide.
Currently running with hand-picked enterprise teams under Steadybit Labs.
Run analysis — every run explaining what held, what didn’t, and what to change — is live in Steadybit Labs today, and reaching general availability in the coming weeks.
What’s still ahead: risk signals triggering the right validations without manual setup. That one is roadmap, and we’ll keep calling it roadmap until it ships.
Every real incident becomes a reproducible experiment. Suffered once, testable forever.
Passed
The same attack, same blast radius, run against the path that follows the recommendation. The unprotected endpoint lost 59% of requests. Behind a circuit breaker, every request returned 200 — and Validate immediately asked for the next failure boundary.
Reactive tools find it. We make sure it doesn’t come back.
Maybe this raises questions — how it scales across hundreds of services, where it fits in your pipeline, what adoption takes. I’d like to hear them.
Send me a message. And if you’d rather see the thing than read about it, I’ll show you a run against a system that looks like yours.
— Benjamin Wilms, co-founder & CEO, steadybit
Get in touch
Tell me what you’re running.
Leave your details and I’ll come back to you personally — usually within a day. If you’d rather see it than read about it, say so and I’ll show you a Validate run against something that looks like your system.
No newsletter, no sequence. A reply from me.