Skip to main content
Self-healing lets SRE Agent act on an incident without waiting for a person: restart or scale a service, roll a deployment back, run an approved runbook, open a fix pull request. It is off for every organization and every service until an organization admin turns it on, and every limit and switch on this page applies in every tier. Once a minute, for each organization that turned it on and has not paused it, self-healing looks at your active alerts. It finds the approved, unpaused runbooks that match and the connector actions its playbooks propose, checks every rule, and then does one of three things: starts the action, asks a person first, or records why it did not. After an action it checks whether the action helped, rolls back what made things worse, and tells a person when it did not help. Organization admins change the policy, resume a pause and enable full autonomy. Members can read the page, pause, and approve or refuse waiting actions. Viewers can read it. Without the plan, the page shows an upgrade card. An organization that was downgraded can still pause self-healing and turn its tier back down.

Turn it on

1

Open Self-healing

Click Self-healing in the sidebar, in the Prevent section. The page lists what waits for you, the organization settings, your services and a dry view of what would happen now.
2

Allow it for the organization

Under Organization, tick Allow self-healing for this organization, choose What it may do and click Save. Actions per hour, whole organization limits the total (1 to 100, default 10).
3

Configure a service

Under Services, click Configure next to a service. Tick Allow self-healing for this service, tick the Channels it may use, pin the connectors under Connectors it may use, and click Save. Nothing acts on a service until its own switch is on.
4

Read the dry view

What would happen now is a dry run over your current alerts and findings. For each candidate it says whether the action would run on its own, wait for a person, or be refused, and why. Click Refresh to run it again. It never changes anything.
Services come from discovery. See Find and manage your services.

Choose what it may do

A service can use its own tier, or Same as the organization. A per-action override on a service can only make an action stricter than the tier allows: Never, Ask first, or Automatic where the tier already allows it.

Channels

Each service chooses which kinds of action may start for it.

Pin connectors

An action only ever runs through the connector pinned for the service, never one chosen by name or by chance. Pin one per kind under Connectors it may use: AWS ECS, AWS Lambda, AWS EC2 and Auto Scaling and Kubernetes. Connectors live in Settings, on the Infrastructure tab.

What it can do

Connector actions run only through the connector pinned for the service. The list of actions is fixed: nothing outside it is ever proposed, approved or started. Commands that take free text, such as a shell line over SSH or SSM, a Kubernetes manifest or patch, or a Lambda invocation or configuration change, can only run inside a runbook a person wrote and approved. Neither a playbook nor a model can propose them in any tier. A stored runbook with a Run a command inside a pod step on a Kubernetes connector reached through a remote agent is refused, because kubectl would reach the cluster from the platform and bypass the tunnel. The Kubernetes restart, the scale-outs, the Lambda release, the EC2 reboot, the volume conversion and a pull request on the code repository are not proposed by any playbook: only the planner, when you turn it on, can choose them, and a runbook that uses a connector action among them is judged by these defaults. The pod, task and instance stop and start actions are never proposed: they only decide how a runbook step that uses them is treated. An action may only touch a resource that belongs to the service (an inventory item attached to it that has not been retired) or a Kubernetes deployment or Auto Scaling group an organization admin listed under Extra targets on the service. Its region has to match the connector’s region, and the account has to match when both are known.

What starts an action

Self-healing starts an action in one of three ways.
  • A matching runbook. Each approved, unpaused runbook that matches an alert is scored, and at most three are considered. Each is its own candidate with its own cool-down and loop-breaker count. A runbook already run for the alert by a webhook trigger is not run again.
  • A playbook. Playbooks read the alert’s own labels, the service’s resources and its deploy history, never free text. The connector playbooks propose only for a service that has the Connector actions channel on, and the certificate playbook only for one that has the Pull requests channel on.
  • The planner, when you turn it on (below).

When the ECS restart runs on its own

In the guardrailed tier the restart runs without asking only for an alarm that is a failure of the service. The alarm must be a CloudWatch alarm that:
  • is in the ALARM state (not INSUFFICIENT_DATA, which is a metric that stopped reporting, and not OK);
  • is not a target tracking alarm (a name starting with TargetTracking-, which Application Auto Scaling creates for every service);
  • watches one of the service’s own health metrics: errors (Errors, 5XXError, HTTPCode_Target_5XX_Count, HTTPCode_ELB_5XX_Count), unhealthy hosts (UnHealthyHostCount) or failed health checks (HealthCheckStatus).
Any other alarm that names the service, such as CPU, memory, latency or an alert from another producer, still gets the restart proposed, and it waits for a person to approve it. Full autonomy lifts the list of health metrics, so a CPU alarm in ALARM restarts on its own, but an alarm in INSUFFICIENT_DATA or OK, a target tracking alarm, or an alert from another producer waits for a person in every tier.

Pull requests

The Pull requests channel opens a fix pull request through the same path as every other fix request. See Ask for a fix for the GitHub App. The platform never merges a pull request, in any tier. A person merges, or does not.
  • The service has to link the repository, and your GitHub App installation has to hold it. A code fix goes to the service’s primary code repository and an infrastructure fix to its primary infrastructure repository. A repository found only because its name resembles the service is never used.
  • The Auto-open PRs from resolved investigations switch on Integrations, GitHub has to be on. It is off by default.
  • Your plan needs fix pull requests, and your monthly AI tokens and AI budget need room.
  • The action stays Running while the pull request is prepared and open. It does not take the service’s one in-flight slot, because a pull request waiting for review changes nothing on the service.
  • Once the pull request is merged and the service deploys, the 15-minute check starts. If a critical alert appears after it, the service is paused and a person is paged to revert it on GitHub, because the platform does not revert a merged pull request.
  • A pull request closed without merging ends the action as cancelled. One not merged or not deployed after 7 days ends as done but unverified.

The planner

Off by default. Under Organization, tick Let a model choose among the actions when no playbook applies. A signal that no playbook and no matching runbook has a candidate for can then be put to a model once. The model has no tools and no free text to act on. It picks from a short list of actions your service settings already allow, aimed at a resource the service owns, and it answers one JSON object or none. Anything else produces no action and a held row. The target is never the model’s choice: it is resolved from the alert’s labels exactly as a playbook does. The alert text it reads is sanitized and fenced as data. Whatever the planner picks waits for a person in every tier below full autonomy, even an action a playbook would run on its own. Under full autonomy a choice runs on its own only when it can be undone (a restart, a scale out, a rollout undo, a Lambda release, a pull request). A choice that cannot be undone, such as an EC2 reboot or a volume conversion, still waits. The planner uses your AI budget, and a run asks about at most three signals.

Limits that always apply

  • Budgets. Actions per hour per service (default 3, up to 20 in the service settings) and per organization (default 10).
  • Concurrency. One action in flight per service by default (At once, up to 5).
  • Cool-down. 15 minutes by default (Cool-down, minutes, 5 to 1440) after an action on the same target. For a runbook the target is that runbook on that service.
  • Loop breaker. The same action on the same target three times in 24 hours is refused. The service is paused, shown as paused by loop_breaker, and an alert is filed. An organization admin resumes it after looking at why it keeps happening.
  • Deploy freeze. An enabled deploy policy with a freeze in the future stops healing on that service. The deploy gate’s override does not lift it. A roll back is never held by a freeze: it is what the freeze exists to allow.
  • Maintenance windows. Under Maintenance windows (UTC) on a service, pick the days and a From and To time. Nothing acts on the service during a window. A window that ends before it starts runs past midnight.
  • Plan. The feature is checked when a decision is made, again on approval and again when the action starts.

Approve what waits for you

When the tier says a person decides, the action waits in Waiting for you, and a message with Approve and Refuse buttons goes to Slack. It names the action, the service, the target, the tier and how long the request stays open. It goes to the service’s alert route channel, else your organization’s remediation channel, else the warning channel.
  • Members and above can approve or refuse, on the page or in Slack. Viewers see the list and no buttons.
  • Approving checks every rule again first. A pause, a freeze, a closed alert or a plan change since the request is honoured, the person is told why it was not approved, and the request goes back on hold.
  • Refuse records who refused and, optionally, why. The same alert does not get the action proposed again for 24 hours.
  • A request nobody answers expires after your organization’s approval timeout and the channel is told. Set it in Settings, Preferences, under Automation approval timeout (minutes): the default is 60 minutes, between 5 and 1440.
  • An expired request is not posted again while the same alert keeps firing. It is listed under Not asked again with an Ask again button for members and above. An alert that resolves and fires anew asks by itself.
A runbook whose approval mode is A person, once already asks before it runs. Self-healing starts it without asking a second time, and the run waits for that person on Automations and in Slack.

The actions inbox

Click Open the actions inbox on the page, or open /self-healing/actions. It lists every action self-healing decided, and what became of it. Waiting for you comes first, oldest request first because that one is closest to expiring. Everything else follows, newest first. Each part shows at most 100 rows.
  • Filter with State and Service. The address keeps your choice.
  • States are Waiting for approval, Held, Approved, Checking the target, Running, Verifying, Done, Failed, Refused, Expired and Cancelled.
  • Each row names the alert or finding that started it, the action and where it happens, the sentence that says why it is where it is, its state, the tier it ran under, who decided it (a playbook, a runbook match, the planner, a roll back or a schedule), and links to its automation run and pull request.
  • Details opens one action with its parameters, what it read before acting, how it is verified, its outcome and the time of each step.
  • Approve, Refuse and Ask again work as above, with the same rules.
  • The page follows every change, so a decision made in Slack or by someone else moves a row without a reload.
A service page shows the same information for one service. See Find and manage your services.

Did it work

Every connector and runbook action is checked for 15 minutes after it finished, from evidence written after it started. Evidence that is missing or older than the action never counts as fixed. The action ends unverified. A critical alert on the service that was not active when the write was sent, and that appears afterwards, makes the outcome worse whatever was being watched.
  • Worse. If the action has a roll back and the read recorded what to put back, the roll back starts at once, through the same pinned connector, under the pause switches. It puts the value back only while the live value is still the one the action set. If a person or an autoscaler changed it since, nothing is written and a person is paged. The service is paused (only an organization admin resumes it) and a Self-healing needs a person alert is filed. An action with no roll back, such as a restart, a reboot or a volume conversion, pauses the service and pages a person.
  • No effect. A person is paged.
  • Unverified. Recorded, nothing else.
When an action fails, stalls for more than two hours, cannot be started, has no effect or made things worse, self-healing files a Self-healing needs a person alert for the service. It reaches your alert routes, mutes and on-call paging like any other alert. Self-healing never acts on its own alerts.

Pause and resume

Members can pause, with a reason, and only an organization admin resumes.
  • The organization. Under Organization, type a reason next to Pause everything, with a reason and click Pause. The page shows who paused it and why, with a Resume button for admins.
  • One service. Open the service (Configure, or View for members) and use Pause this service, with a reason.
  • The platform. The operator of SRE Agent can pause self-healing for everyone. The page then says so, and nothing acts until it is lifted.
Pausing, turning a scope off and the platform pause cancel the scope’s waiting and running actions. A pause also stops the writes the Control Tower and Config Watcher agents make on their own, such as muting a rule, opening an investigation or applying a configuration fix. Their findings and proposals are still filed.

Full autonomy

Full autonomy is deliberately hard to switch on by accident. An organization admin clicks Enable full autonomy under Organization, or Enable full autonomy for this service on a service, reads the statement the page shows, and types the organization’s name to confirm. SRE Agent records who accepted it, when, the exact text shown, the name typed and the IP address, and writes an audit entry. The acknowledgements are listed on the page under Full autonomy acknowledgements. Return to guardrailed is one click.
  • It cannot be set from the API or an MCP tool, and the ordinary settings forms refuse the tier.
  • It cannot be enabled, and no other setting can be widened, while SRE Agent support is signed in as one of your users.
  • Full autonomy can make an incident worse before verification notices. Keep budgets low and maintenance windows set.
The organization statement and the Deletions panel on the page also describe deletions of unused resources. Deletions are not available yet: the page says so, and nothing is deleted.

Audit and MCP tools

Every policy change and every state change of an action is written to the audit log, for example self_healing.policy_updated, self_healing.service_policy_updated, self_healing.full_autonomy_enabled, self_healing.paused, self_healing.resumed and self_healing.action_<state>. Entries list the fields that changed, never a connector credential. Ten MCP tools give the same reads and decisions. Each is gated like its page: the plan comes first and answers “Self-healing is not included in your plan.” Every write is a person’s decision, so it needs a personal access token. An API key is nobody, and its call is refused with a sentence that says so. A token’s role ceiling applies: a token capped at member cannot change a policy even when its owner is an organization admin. A policy field set to null or a blank string is refused by name, and nothing changes. Only a service’s tier accepts null, meaning the organization’s. No tool enables full autonomy. Both policy tools refuse tier: "full_autonomy" with “Full autonomy is enabled only on the Self-healing page, by an organization admin who types the organization name.” See MCP tools for how to connect a client.