Turn it on
1
Open Self-healing
Click Self-healing in the sidebar, in the Prevent section. The page lists what waits for
you, the organization settings, your services and a dry view of what would happen now.
2
Allow it for the organization
Under Organization, tick Allow self-healing for this organization, choose What it may
do and click Save. Actions per hour, whole organization limits the total (1 to 100,
default 10).
3
Configure a service
Under Services, click Configure next to a service. Tick Allow self-healing for this
service, tick the Channels it may use, pin the connectors under Connectors it may use,
and click Save. Nothing acts on a service until its own switch is on.
4
Read the dry view
What would happen now is a dry run over your current alerts and findings. For each candidate
it says whether the action would run on its own, wait for a person, or be refused, and why.
Click Refresh to run it again. It never changes anything.
Choose what it may do
A service can use its own tier, or Same as the organization. A per-action override on a service can only make an action stricter than the tier allows: Never, Ask first, or Automatic where the tier already allows it.
Channels
Each service chooses which kinds of action may start for it.Pin connectors
An action only ever runs through the connector pinned for the service, never one chosen by name or by chance. Pin one per kind under Connectors it may use: AWS ECS, AWS Lambda, AWS EC2 and Auto Scaling and Kubernetes. Connectors live in Settings, on the Infrastructure tab.What it can do
Connector actions run only through the connector pinned for the service. The list of actions is fixed: nothing outside it is ever proposed, approved or started.
Commands that take free text, such as a shell line over SSH or SSM, a Kubernetes manifest or patch, or a Lambda invocation or configuration change, can only run inside a runbook a person wrote and approved. Neither a playbook nor a model can propose them in any tier. A stored runbook with a Run a command inside a pod step on a Kubernetes connector reached through a remote agent is refused, because kubectl would reach the cluster from the platform and bypass the tunnel.
The Kubernetes restart, the scale-outs, the Lambda release, the EC2 reboot, the volume conversion and a pull request on the code repository are not proposed by any playbook: only the planner, when you turn it on, can choose them, and a runbook that uses a connector action among them is judged by these defaults. The pod, task and instance stop and start actions are never proposed: they only decide how a runbook step that uses them is treated.
An action may only touch a resource that belongs to the service (an inventory item attached to it that has not been retired) or a Kubernetes deployment or Auto Scaling group an organization admin listed under Extra targets on the service. Its region has to match the connector’s region, and the account has to match when both are known.
What starts an action
Self-healing starts an action in one of three ways.- A matching runbook. Each approved, unpaused runbook that matches an alert is scored, and at most three are considered. Each is its own candidate with its own cool-down and loop-breaker count. A runbook already run for the alert by a webhook trigger is not run again.
- A playbook. Playbooks read the alert’s own labels, the service’s resources and its deploy history, never free text. The connector playbooks propose only for a service that has the Connector actions channel on, and the certificate playbook only for one that has the Pull requests channel on.
- The planner, when you turn it on (below).
When the ECS restart runs on its own
In the guardrailed tier the restart runs without asking only for an alarm that is a failure of the service. The alarm must be a CloudWatch alarm that:- is in the
ALARMstate (notINSUFFICIENT_DATA, which is a metric that stopped reporting, and notOK); - is not a target tracking alarm (a name starting with
TargetTracking-, which Application Auto Scaling creates for every service); - watches one of the service’s own health metrics: errors (
Errors,5XXError,HTTPCode_Target_5XX_Count,HTTPCode_ELB_5XX_Count), unhealthy hosts (UnHealthyHostCount) or failed health checks (HealthCheckStatus).
ALARM restarts on its own, but an alarm in INSUFFICIENT_DATA or OK, a target tracking alarm, or an alert from another producer waits for a person in every tier.
Pull requests
The Pull requests channel opens a fix pull request through the same path as every other fix request. See Ask for a fix for the GitHub App. The platform never merges a pull request, in any tier. A person merges, or does not.- The service has to link the repository, and your GitHub App installation has to hold it. A code fix goes to the service’s primary code repository and an infrastructure fix to its primary infrastructure repository. A repository found only because its name resembles the service is never used.
- The Auto-open PRs from resolved investigations switch on Integrations, GitHub has to be on. It is off by default.
- Your plan needs fix pull requests, and your monthly AI tokens and AI budget need room.
- The action stays Running while the pull request is prepared and open. It does not take the service’s one in-flight slot, because a pull request waiting for review changes nothing on the service.
- Once the pull request is merged and the service deploys, the 15-minute check starts. If a critical alert appears after it, the service is paused and a person is paged to revert it on GitHub, because the platform does not revert a merged pull request.
- A pull request closed without merging ends the action as cancelled. One not merged or not deployed after 7 days ends as done but unverified.
The planner
Off by default. Under Organization, tick Let a model choose among the actions when no playbook applies. A signal that no playbook and no matching runbook has a candidate for can then be put to a model once. The model has no tools and no free text to act on. It picks from a short list of actions your service settings already allow, aimed at a resource the service owns, and it answers one JSON object ornone. Anything else produces no action and a held row. The target is never the model’s choice: it is resolved from the alert’s labels exactly as a playbook does. The alert text it reads is sanitized and fenced as data.
Whatever the planner picks waits for a person in every tier below full autonomy, even an action a playbook would run on its own. Under full autonomy a choice runs on its own only when it can be undone (a restart, a scale out, a rollout undo, a Lambda release, a pull request). A choice that cannot be undone, such as an EC2 reboot or a volume conversion, still waits. The planner uses your AI budget, and a run asks about at most three signals.
Limits that always apply
- Budgets. Actions per hour per service (default 3, up to 20 in the service settings) and per organization (default 10).
- Concurrency. One action in flight per service by default (At once, up to 5).
- Cool-down. 15 minutes by default (Cool-down, minutes, 5 to 1440) after an action on the same target. For a runbook the target is that runbook on that service.
- Loop breaker. The same action on the same target three times in 24 hours is refused. The service is paused, shown as paused by
loop_breaker, and an alert is filed. An organization admin resumes it after looking at why it keeps happening. - Deploy freeze. An enabled deploy policy with a freeze in the future stops healing on that service. The deploy gate’s override does not lift it. A roll back is never held by a freeze: it is what the freeze exists to allow.
- Maintenance windows. Under Maintenance windows (UTC) on a service, pick the days and a From and To time. Nothing acts on the service during a window. A window that ends before it starts runs past midnight.
- Plan. The feature is checked when a decision is made, again on approval and again when the action starts.
Approve what waits for you
When the tier says a person decides, the action waits in Waiting for you, and a message with Approve and Refuse buttons goes to Slack. It names the action, the service, the target, the tier and how long the request stays open. It goes to the service’s alert route channel, else your organization’s remediation channel, else the warning channel.- Members and above can approve or refuse, on the page or in Slack. Viewers see the list and no buttons.
- Approving checks every rule again first. A pause, a freeze, a closed alert or a plan change since the request is honoured, the person is told why it was not approved, and the request goes back on hold.
- Refuse records who refused and, optionally, why. The same alert does not get the action proposed again for 24 hours.
- A request nobody answers expires after your organization’s approval timeout and the channel is told. Set it in Settings, Preferences, under Automation approval timeout (minutes): the default is 60 minutes, between 5 and 1440.
- An expired request is not posted again while the same alert keeps firing. It is listed under Not asked again with an Ask again button for members and above. An alert that resolves and fires anew asks by itself.
The actions inbox
Click Open the actions inbox on the page, or open/self-healing/actions. It lists every action self-healing decided, and what became of it. Waiting for you comes first, oldest request first because that one is closest to expiring. Everything else follows, newest first. Each part shows at most 100 rows.
- Filter with State and Service. The address keeps your choice.
- States are Waiting for approval, Held, Approved, Checking the target, Running, Verifying, Done, Failed, Refused, Expired and Cancelled.
- Each row names the alert or finding that started it, the action and where it happens, the sentence that says why it is where it is, its state, the tier it ran under, who decided it (a playbook, a runbook match, the planner, a roll back or a schedule), and links to its automation run and pull request.
- Details opens one action with its parameters, what it read before acting, how it is verified, its outcome and the time of each step.
- Approve, Refuse and Ask again work as above, with the same rules.
- The page follows every change, so a decision made in Slack or by someone else moves a row without a reload.
Did it work
Every connector and runbook action is checked for 15 minutes after it finished, from evidence written after it started.
Evidence that is missing or older than the action never counts as fixed. The action ends unverified. A critical alert on the service that was not active when the write was sent, and that appears afterwards, makes the outcome worse whatever was being watched.
- Worse. If the action has a roll back and the read recorded what to put back, the roll back starts at once, through the same pinned connector, under the pause switches. It puts the value back only while the live value is still the one the action set. If a person or an autoscaler changed it since, nothing is written and a person is paged. The service is paused (only an organization admin resumes it) and a Self-healing needs a person alert is filed. An action with no roll back, such as a restart, a reboot or a volume conversion, pauses the service and pages a person.
- No effect. A person is paged.
- Unverified. Recorded, nothing else.
Pause and resume
Members can pause, with a reason, and only an organization admin resumes.- The organization. Under Organization, type a reason next to Pause everything, with a reason and click Pause. The page shows who paused it and why, with a Resume button for admins.
- One service. Open the service (Configure, or View for members) and use Pause this service, with a reason.
- The platform. The operator of SRE Agent can pause self-healing for everyone. The page then says so, and nothing acts until it is lifted.
Full autonomy
Full autonomy is deliberately hard to switch on by accident. An organization admin clicks Enable full autonomy under Organization, or Enable full autonomy for this service on a service, reads the statement the page shows, and types the organization’s name to confirm. SRE Agent records who accepted it, when, the exact text shown, the name typed and the IP address, and writes an audit entry. The acknowledgements are listed on the page under Full autonomy acknowledgements. Return to guardrailed is one click.- It cannot be set from the API or an MCP tool, and the ordinary settings forms refuse the tier.
- It cannot be enabled, and no other setting can be widened, while SRE Agent support is signed in as one of your users.
- Full autonomy can make an incident worse before verification notices. Keep budgets low and maintenance windows set.
The organization statement and the Deletions panel on the page also describe deletions of
unused resources. Deletions are not available yet: the page says so, and nothing is deleted.
Audit and MCP tools
Every policy change and every state change of an action is written to the audit log, for exampleself_healing.policy_updated, self_healing.service_policy_updated, self_healing.full_autonomy_enabled, self_healing.paused, self_healing.resumed and self_healing.action_<state>. Entries list the fields that changed, never a connector credential.
Ten MCP tools give the same reads and decisions. Each is gated like its page: the plan comes first and answers “Self-healing is not included in your plan.”
Every write is a person’s decision, so it needs a personal access token. An API key is nobody, and its call is refused with a sentence that says so. A token’s role ceiling applies: a token capped at member cannot change a policy even when its owner is an organization admin. A policy field set to null or a blank string is refused by name, and nothing changes. Only a service’s
tier accepts null, meaning the organization’s.
No tool enables full autonomy. Both policy tools refuse tier: "full_autonomy" with “Full autonomy is enabled only on the Self-healing page, by an organization admin who types the organization name.” See MCP tools for how to connect a client.
Related
- Write and run runbooks: the runbooks self-healing can start.
- Triage alerts: the alerts that start an action.
- Ask for a fix: the GitHub App behind the Pull requests channel.
- Find and manage your services: the services a policy is set on.
- Plan matrix: which plans include self-healing.