How to Automate Repeated Incident Fixes in 6 Steps
Sang Lee
October 9, 2026

Your team restarted the same Kafka consumer group four times last month. Each time, a different engineer found the fix in a different old Slack thread.
Direct Answer: To automate repeated incident fixes, record the fix that resolved each incident, find the fixes that repeat, and turn each one into a workflow built from what engineers actually ran. Let read-only steps run automatically, and keep a human approval on anything that changes production.
Overview
- Why repeated fixes are the easiest toil to remove
- Six steps from incident history to a safe workflow
- Manual runbooks compared with suggested workflows
- Common mistakes
- FAQ
Why automate repeated incident fixes first?
A repeated fix is toil by definition. Google's SRE book describes toil as manual, repetitive, automatable work that leaves no lasting value, and caps it at 50% of each SRE's time (Google SRE book).
The problem isn't shrinking on its own. In The SRE Report 2026, half of respondents said they spend 34% of their time on repetitive, low-value work, about the same as the year before (Catchpoint).
Repeated fixes are the cheapest place to start because the hard part is done. Someone already found the fix, ran it, and confirmed it worked. You are not designing automation from scratch; you are writing down a known-good sequence.
How to automate repeated incident fixes in 6 steps
1. Record the fix that resolved every incident
Add a required "resolving action" field to every postmortem or incident close-out: the commands, config change, or rollback that ended it. Without this field, step 2 has nothing to search.
2. Find the fixes that repeat, using the Third-Time Rule
The Third-Time Rule: the first time a fix runs by hand, it's a fix. The second time, it's a coincidence. The third time, it's a workflow waiting to be written.
Group the last 90 days of incidents by resolving action and service. Anything that shows up three or more times goes on the shortlist. Sort the shortlist by total engineer time spent, not by count.
3. Split each fix into read-only and state-changing steps
Most fixes start with checks (is consumer lag growing, did a deploy just land) and end with a change (restart, scale, roll back). Label every step before writing anything.
Read-only steps can run on their own from day one. Reversible changes earn autonomy over time, and hard-to-reverse changes always wait for a human.
4. Build the workflow from what engineers actually ran
Pull the commands from the incident channel and terminal history, not from memory. Runbooks written from memory drop the flag that mattered or the check that ruled out a false positive.
5. Gate every state-changing step behind an approval
The workflow runs its checks, posts results in the incident channel, and stops at the first change for a human to approve. One click in Slack is enough. Loosen the gate on reversible steps once the workflow has a clean track record; keep hard-to-reverse steps gated permanently.
6. Review workflows monthly and retire the stale ones
Infrastructure changes faster than workflows. Each month, check which workflows ran, which failed, and which reference services that no longer exist. Delete before you add.
If you're rolling this out alongside an AI SRE agent, our guide to AI SRE implementation in the first 30 days covers the sequencing.
Manual runbooks vs. suggested workflows
The traditional approach asks an engineer to write a runbook after the incident. The newer approach has an AI agent watch how incidents get resolved and propose the workflow itself.
A manual runbook depends on an engineer finding time to write and maintain it. A suggested workflow is drafted from fixes the team already ran, so the backlog of unwritten runbooks stops growing.
How Tony SRE suggests workflows
Tony SRE, the AI incident commander in Vibe OnCall, watches how your team resolves incidents in Slack or Teams. When he sees the same fix applied again, he drafts the workflow and suggests it to the team, instead of waiting for someone to build an agent or write a runbook. Once the team approves it, Tony runs the diagnostic steps on his own and asks before anything sensitive or irreversible.
Common mistakes when automating incident fixes
- Automating a fix you've seen once. One occurrence doesn't tell you whether the fix, or the symptom, will come back. Wait for the third time.
- Automating the symptom and forgetting the cause. If a service needs a restart every week, the workflow buys time. File the bug too.
- Giving the workflow broad credentials. Scope each workflow to the services and actions it needs. Our post on AI agent security for startups explains why access scope is most of the risk.
- Shipping it without telling on-call. If a workflow can run at 3 AM, the on-call engineer needs to know it exists and how to stop it.
FAQ
What counts as a repeated incident fix?
A repeated incident fix is any resolving action your team has run by hand three or more times for the same symptom. Restarts, cache flushes, scaling events, and feature flag rollbacks are the most common.
Should automated incident fixes run without approval?
Read-only checks should; state-changing steps should not, at least at first. Loosen approvals on reversible steps only after a workflow has a clean track record.
Is a workflow the same as a runbook?
No. A runbook is a document a person follows, while a workflow executes the steps itself. Good workflows still link to a runbook so humans can read what the automation does.
How do I find repeated fixes if we don't write postmortems?
Search your incident channels for the commands themselves. Restart and rollback commands are easy to grep, and an AI agent that reads your Slack history can do the search for you.



