> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sreagent.app/llms.txt
> Use this file to discover all available pages before exploring further.

# Let SRE Agent act on incidents with self-healing

> Turn on self-healing per organization and per service, choose what it may do, approve what waits for you, and pause it at any time.

export const Plan = ({tier}) => <Badge color="blue">{tier} plan</Badge>;

Self-healing lets SRE Agent act on an incident without waiting for a person: restart or scale a service, roll a deployment back, run an approved runbook, open a fix pull request. It is off for every organization and every service until an organization admin turns it on, and every limit and switch on this page applies in every tier.

<Plan tier="Business" />

Once a minute, for each organization that turned it on and has not paused it, self-healing looks at your active alerts. It finds the approved, unpaused runbooks that match and the connector actions its playbooks propose, checks every rule, and then does one of three things: starts the action, asks a person first, or records why it did not. After an action it checks whether the action helped, rolls back what made things worse, and tells a person when it did not help.

Organization admins change the policy, resume a pause and enable full autonomy. Members can read the page, pause, and approve or refuse waiting actions. Viewers can read it. Without the plan, the page shows an upgrade card. An organization that was downgraded can still pause self-healing and turn its tier back down.

## Turn it on

<Steps>
  <Step title="Open Self-healing">
    Click **Self-healing** in the sidebar, in the **Prevent** section. The page lists what waits for
    you, the organization settings, your services and a dry view of what would happen now.
  </Step>

  <Step title="Allow it for the organization">
    Under **Organization**, tick **Allow self-healing for this organization**, choose **What it may
    do** and click **Save**. **Actions per hour, whole organization** limits the total (1 to 100,
    default 10).
  </Step>

  <Step title="Configure a service">
    Under **Services**, click **Configure** next to a service. Tick **Allow self-healing for this
    service**, tick the **Channels** it may use, pin the connectors under **Connectors it may use**,
    and click **Save**. Nothing acts on a service until its own switch is on.
  </Step>

  <Step title="Read the dry view">
    **What would happen now** is a dry run over your current alerts and findings. For each candidate
    it says whether the action would run on its own, wait for a person, or be refused, and why.
    Click **Refresh** to run it again. It never changes anything.
  </Step>
</Steps>

Services come from discovery. See [Find and manage your services](/guides/get-started/services).

## Choose what it may do

| Tier | What it allows |
| - | - |
| **Guardrailed** | The default. Reversible, bounded actions marked automatic run on their own, such as restarting a service or adding capacity within a bound. Everything else waits for a person. The riskiest actions never run. |
| **Propose only** | Every action waits for a person to approve it before anything changes. |
| **Full autonomy** | Every action the connectors allow runs without asking anyone. Budgets, cool-downs, the loop breaker, deploy freezes, maintenance windows and the pause switches still apply. |

A service can use its own tier, or **Same as the organization**. A per-action override on a service can only make an action stricter than the tier allows: **Never**, **Ask first**, or **Automatic** where the tier already allows it.

### Channels

Each service chooses which kinds of action may start for it.

| Channel | What it does |
| - | - |
| **Runbooks** | Runs an approved runbook a person wrote. |
| **Connector actions** | Restarts, scales or rolls back through a pinned connector. |
| **Pull requests** | Opens a fix pull request, a draft unless your GitHub settings say otherwise. Nothing is ever merged. |
| **Platform actions** | Freezes a service's deploys while an SLO has no error budget left or an incident on the status page is still investigating. It needs a setting this page does not offer yet, so leave it off for now. |

### Pin connectors

An action only ever runs through the connector pinned for the service, never one chosen by name or by chance. Pin one per kind under **Connectors it may use**: **AWS ECS**, **AWS Lambda**, **AWS EC2 and Auto Scaling** and **Kubernetes**. Connectors live in **Settings**, on the **Infrastructure** tab.

## What it can do

Connector actions run only through the connector pinned for the service. The list of actions is fixed: nothing outside it is ever proposed, approved or started.

| Action | Where | Default | Put back by |
| - | - | - | - |
| Restart an ECS service (force a new deployment) | ECS | Automatic for a failing alarm on the service's own health, otherwise waits for a person | Nothing to put back |
| Add tasks to an ECS service (at most double, at most 10 more, at most 200) | ECS | Automatic | The task count the read before it recorded, if nobody changed it since |
| Roll an ECS service back to its previous task definition revision | ECS | Waits for a person | The service is rolled forward to the revision it was on |
| Restart a Kubernetes deployment | Kubernetes | Automatic | Nothing to put back |
| Add replicas to a Kubernetes deployment (same bounds) | Kubernetes | Automatic | The replica count the read recorded |
| Roll a Kubernetes deployment back one revision | Kubernetes | Waits for a person | Nothing: it is the roll back |
| Raise a Lambda function's reserved concurrency (by half, at least 1, at most 100 more; a planner choice may go up to double, still at most 100 more) | Lambda | Waits for a person | The reservation the read recorded |
| Return a Lambda function to the shared pool | Lambda | Waits for a person | The reservation the read recorded |
| Reboot an EC2 instance | EC2 | Waits for a person | Cannot be undone |
| Add instances to an Auto Scaling group (same bounds) | Auto Scaling | Automatic | The capacity the read recorded |
| Convert an EBS volume from gp2 to gp3 | EBS | Waits for a person | Cannot be undone |
| Delete one pod of a Kubernetes deployment, stop one ECS task, stop or start an EC2 instance | Kubernetes, ECS, EC2 | Full autonomy only | None |
| Open a fix pull request on the service's code or infrastructure repository | GitHub | Automatic, with the **Pull requests** channel | Nothing is merged for you |
| Run an approved runbook | Runbooks | The strictest of its steps: a runbook whose steps all run on their own (or that only reads) is automatic with the **Runbooks** channel, and one with a step that waits for a person, or a free-text command, waits for a person | Depends on the runbook |

Commands that take free text, such as a shell line over SSH or SSM, a Kubernetes manifest or patch, or a Lambda invocation or configuration change, can only run inside a runbook a person wrote and approved. Neither a playbook nor a model can propose them in any tier. A stored runbook with a **Run a command inside a pod** step on a Kubernetes connector reached through a remote agent is refused, because kubectl would reach the cluster from the platform and bypass the tunnel.

The Kubernetes restart, the scale-outs, the Lambda release, the EC2 reboot, the volume conversion and a pull request on the code repository are not proposed by any playbook: only the planner, when you turn it on, can choose them, and a runbook that uses a connector action among them is judged by these defaults. The pod, task and instance stop and start actions are never proposed: they only decide how a runbook step that uses them is treated.

An action may only touch a resource that belongs to the service (an inventory item attached to it that has not been retired) or a Kubernetes deployment or Auto Scaling group an organization admin listed under **Extra targets** on the service. Its region has to match the connector's region, and the account has to match when both are known.

## What starts an action

Self-healing starts an action in one of three ways.

* **A matching runbook.** Each approved, unpaused runbook that matches an alert is scored, and at most three are considered. Each is its own candidate with its own cool-down and loop-breaker count. A runbook already run for the alert by a webhook trigger is not run again.
* **A playbook.** Playbooks read the alert's own labels, the service's resources and its deploy history, never free text. The connector playbooks propose only for a service that has the **Connector actions** channel on, and the certificate playbook only for one that has the **Pull requests** channel on.
* **The planner,** when you turn it on (below).

| When | Proposed |
| - | - |
| An alarm names an ECS service the service owns, and the deploy record holds a rollout of it that finished in the last hour with the revision it replaced | Roll the ECS service back to that previous revision. It waits for a person by default, and the request says which two revisions. |
| The same alarm with no such rollout | Restart the ECS service. |
| An alert names a Kubernetes deployment the service owns, after a deploy in the last hour | Roll the deployment back one revision. It waits for a person by default. Nothing is proposed without a deploy. |
| A `Throttles` alarm on a Lambda function the service owns | Raise the function's reserved concurrency. The action ends quietly when the function has no reservation. |
| An expiring certificate finding for a certificate the service owns, when the service links an infrastructure repository | Open an infrastructure fix pull request, for a service with the **Pull requests** channel on. |
| A Control Tower deploy regression finding for a service that owns exactly one ECS service or one Kubernetes deployment | The roll back of that one. It waits for a person by default. |

### When the ECS restart runs on its own

In the guardrailed tier the restart runs without asking only for an alarm that is a failure of the service. The alarm must be a CloudWatch alarm that:

* is in the `ALARM` state (not `INSUFFICIENT_DATA`, which is a metric that stopped reporting, and not `OK`);
* is not a target tracking alarm (a name starting with `TargetTracking-`, which Application Auto Scaling creates for every service);
* watches one of the service's own health metrics: errors (`Errors`, `5XXError`, `HTTPCode_Target_5XX_Count`, `HTTPCode_ELB_5XX_Count`), unhealthy hosts (`UnHealthyHostCount`) or failed health checks (`HealthCheckStatus`).

Any other alarm that names the service, such as CPU, memory, latency or an alert from another producer, still gets the restart proposed, and it waits for a person to approve it. Full autonomy lifts the list of health metrics, so a CPU alarm in `ALARM` restarts on its own, but an alarm in `INSUFFICIENT_DATA` or `OK`, a target tracking alarm, or an alert from another producer waits for a person in every tier.

### Pull requests

The **Pull requests** channel opens a fix pull request through the same path as every other fix request. See [Ask for a fix](/guides/fix/fix-requests) for the GitHub App. **The platform never merges a pull request, in any tier.** A person merges, or does not.

* The service has to link the repository, and your GitHub App installation has to hold it. A code fix goes to the service's primary code repository and an infrastructure fix to its primary infrastructure repository. A repository found only because its name resembles the service is never used.
* The **Auto-open PRs from resolved investigations** switch on **Integrations**, **GitHub** has to be on. It is off by default.
* Your plan needs fix pull requests, and your monthly AI tokens and AI budget need room.
* The action stays **Running** while the pull request is prepared and open. It does not take the service's one in-flight slot, because a pull request waiting for review changes nothing on the service.
* Once the pull request is merged and the service deploys, the 15-minute check starts. If a critical alert appears after it, the service is paused and a person is paged to revert it on GitHub, because the platform does not revert a merged pull request.
* A pull request closed without merging ends the action as cancelled. One not merged or not deployed after 7 days ends as done but unverified.

### The planner

Off by default. Under **Organization**, tick **Let a model choose among the actions when no playbook applies**. A signal that no playbook and no matching runbook has a candidate for can then be put to a model once.

The model has no tools and no free text to act on. It picks from a short list of actions your service settings already allow, aimed at a resource the service owns, and it answers one JSON object or `none`. Anything else produces no action and a held row. The target is never the model's choice: it is resolved from the alert's labels exactly as a playbook does. The alert text it reads is sanitized and fenced as data.

Whatever the planner picks waits for a person in every tier below full autonomy, even an action a playbook would run on its own. Under full autonomy a choice runs on its own only when it can be undone (a restart, a scale out, a rollout undo, a Lambda release, a pull request). A choice that cannot be undone, such as an EC2 reboot or a volume conversion, still waits. The planner uses your AI budget, and a run asks about at most three signals.

## Limits that always apply

* **Budgets.** Actions per hour per service (default 3, up to 20 in the service settings) and per organization (default 10).
* **Concurrency.** One action in flight per service by default (**At once**, up to 5).
* **Cool-down.** 15 minutes by default (**Cool-down, minutes**, 5 to 1440) after an action on the same target. For a runbook the target is that runbook on that service.
* **Loop breaker.** The same action on the same target three times in 24 hours is refused. The service is paused, shown as paused by `loop_breaker`, and an alert is filed. An organization admin resumes it after looking at why it keeps happening.
* **Deploy freeze.** An enabled deploy policy with a freeze in the future stops healing on that service. The deploy gate's override does not lift it. A roll back is never held by a freeze: it is what the freeze exists to allow.
* **Maintenance windows.** Under **Maintenance windows (UTC)** on a service, pick the days and a **From** and **To** time. Nothing acts on the service during a window. A window that ends before it starts runs past midnight.
* **Plan.** The feature is checked when a decision is made, again on approval and again when the action starts.

## Approve what waits for you

When the tier says a person decides, the action waits in **Waiting for you**, and a message with **Approve** and **Refuse** buttons goes to Slack. It names the action, the service, the target, the tier and how long the request stays open. It goes to the service's alert route channel, else your organization's remediation channel, else the warning channel.

* Members and above can approve or refuse, on the page or in Slack. Viewers see the list and no buttons.
* Approving checks every rule again first. A pause, a freeze, a closed alert or a plan change since the request is honoured, the person is told why it was not approved, and the request goes back on hold.
* **Refuse** records who refused and, optionally, why. The same alert does not get the action proposed again for 24 hours.
* A request nobody answers expires after your organization's approval timeout and the channel is told. Set it in **Settings**, **Preferences**, under **Automation approval timeout (minutes)**: the default is 60 minutes, between 5 and 1440.
* An expired request is not posted again while the same alert keeps firing. It is listed under **Not asked again** with an **Ask again** button for members and above. An alert that resolves and fires anew asks by itself.

A runbook whose approval mode is **A person, once** already asks before it runs. Self-healing starts it without asking a second time, and the run waits for that person on **Automations** and in Slack.

## The actions inbox

Click **Open the actions inbox** on the page, or open `/self-healing/actions`. It lists every action self-healing decided, and what became of it. **Waiting for you** comes first, oldest request first because that one is closest to expiring. Everything else follows, newest first. Each part shows at most 100 rows.

* **Filter** with **State** and **Service**. The address keeps your choice.
* **States** are **Waiting for approval**, **Held**, **Approved**, **Checking the target**, **Running**, **Verifying**, **Done**, **Failed**, **Refused**, **Expired** and **Cancelled**.
* **Each row** names the alert or finding that started it, the action and where it happens, the sentence that says why it is where it is, its state, the tier it ran under, who decided it (a playbook, a runbook match, the planner, a roll back or a schedule), and links to its automation run and pull request.
* **Details** opens one action with its parameters, what it read before acting, how it is verified, its outcome and the time of each step.
* **Approve**, **Refuse** and **Ask again** work as above, with the same rules.
* The page follows every change, so a decision made in Slack or by someone else moves a row without a reload.

A service page shows the same information for one service. See [Find and manage your services](/guides/get-started/services).

## Did it work

Every connector and runbook action is checked for 15 minutes after it finished, from evidence written after it started.

| Watching | Fixed | No effect | Worse |
| - | - | - | - |
| The alert that started it | It resolved after the action | Still active after the window | Its severity rose |
| The SLO an SLO alert names | Its fast burn rate is under the threshold in a later computation | Still over the threshold | Its burn rate rose clearly |
| The synthetic check a synthetic alert names | A result after the action passed | Still failing after the window | Not used |
| The service, when nothing more specific applies | No high or critical alert is active, and one that was has resolved | A high or critical alert is still active | Not used |

Evidence that is missing or older than the action never counts as fixed. The action ends **unverified**. A critical alert on the service that was not active when the write was sent, and that appears afterwards, makes the outcome worse whatever was being watched.

* **Worse.** If the action has a roll back and the read recorded what to put back, the roll back starts at once, through the same pinned connector, under the pause switches. It puts the value back only while the live value is still the one the action set. If a person or an autoscaler changed it since, nothing is written and a person is paged. The service is paused (only an organization admin resumes it) and a **Self-healing needs a person** alert is filed. An action with no roll back, such as a restart, a reboot or a volume conversion, pauses the service and pages a person.
* **No effect.** A person is paged.
* **Unverified.** Recorded, nothing else.

When an action fails, stalls for more than two hours, cannot be started, has no effect or made things worse, self-healing files a **Self-healing needs a person** alert for the service. It reaches your alert routes, mutes and on-call paging like any other alert. Self-healing never acts on its own alerts.

## Pause and resume

Members can pause, with a reason, and only an organization admin resumes.

* **The organization.** Under **Organization**, type a reason next to **Pause everything, with a reason** and click **Pause**. The page shows who paused it and why, with a **Resume** button for admins.
* **One service.** Open the service (**Configure**, or **View** for members) and use **Pause this service, with a reason**.
* **The platform.** The operator of SRE Agent can pause self-healing for everyone. The page then says so, and nothing acts until it is lifted.

Pausing, turning a scope off and the platform pause cancel the scope's waiting and running actions. A pause also stops the writes the Control Tower and Config Watcher agents make on their own, such as muting a rule, opening an investigation or applying a configuration fix. Their findings and proposals are still filed.

## Full autonomy

Full autonomy is deliberately hard to switch on by accident. An organization admin clicks **Enable full autonomy** under **Organization**, or **Enable full autonomy for this service** on a service, reads the statement the page shows, and types the organization's name to confirm.

SRE Agent records who accepted it, when, the exact text shown, the name typed and the IP address, and writes an audit entry. The acknowledgements are listed on the page under **Full autonomy acknowledgements**. **Return to guardrailed** is one click.

* It cannot be set from the API or an MCP tool, and the ordinary settings forms refuse the tier.
* It cannot be enabled, and no other setting can be widened, while SRE Agent support is signed in as one of your users.
* Full autonomy can make an incident worse before verification notices. Keep budgets low and maintenance windows set.

<Note>
  The organization statement and the **Deletions** panel on the page also describe deletions of
  unused resources. Deletions are not available yet: the page says so, and nothing is deleted.
</Note>

## Audit and MCP tools

Every policy change and every state change of an action is written to the audit log, for example `self_healing.policy_updated`, `self_healing.service_policy_updated`, `self_healing.full_autonomy_enabled`, `self_healing.paused`, `self_healing.resumed` and `self_healing.action_<state>`. Entries list the fields that changed, never a connector credential.

Ten MCP tools give the same reads and decisions. Each is gated like its page: the plan comes first and answers "Self-healing is not included in your plan."

| Tool | Role needed | What it does |
| - | - | - |
| `get_self_healing_settings` | Viewer | The platform pause, the organization policy, each service's policy and the full autonomy acknowledgements. |
| `list_healing_actions` | Viewer | The inbox, newest first, with the reason sentence. Filter by `state` and `service_id`, `limit` 1 to 200. |
| `get_healing_action` | Viewer | One action with its parameters, what it read, how it is verified, its outcome and links. |
| `preview_self_healing` | Viewer | The dry view. Writes nothing. |
| `approve_healing_action` | Member | Approve a waiting request. Every rule is checked again. |
| `refuse_healing_action` | Member | Refuse a waiting request, with an optional reason. |
| `pause_self_healing` | Member | Pause the organization, or one service with `service_id`. A reason is required. |
| `resume_self_healing` | Org admin | Resume the organization or one service. Needs `confirm: true`. |
| `update_self_healing_policy` | Org admin | The organization's `enabled`, `tier`, `planner_enabled` and `max_actions_per_hour`. |
| `update_service_healing_policy` | Org admin | One service's switch, tier, channels, connector pins, action overrides, extra targets, windows and budgets. |

Every write is a person's decision, so it needs a personal access token. An API key is nobody, and its call is refused with a sentence that says so. A token's role ceiling applies: a token capped at member cannot change a policy even when its owner is an organization admin. A policy field set to null or a blank string is refused by name, and nothing changes. Only a service's `tier` accepts null, meaning the organization's.

No tool enables full autonomy. Both policy tools refuse `tier: "full_autonomy"` with "Full autonomy is enabled only on the Self-healing page, by an organization admin who types the organization name." See [MCP tools](/api-reference/mcp) for how to connect a client.

## Related

* [Write and run runbooks](/guides/prevent/runbooks): the runbooks self-healing can start.
* [Triage alerts](/guides/respond/alerts): the alerts that start an action.
* [Ask for a fix](/guides/fix/fix-requests): the GitHub App behind the Pull requests channel.
* [Find and manage your services](/guides/get-started/services): the services a policy is set on.
* [Plan matrix](/guides/reference/plan-matrix): which plans include self-healing.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.