> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sreagent.app/llms.txt
> Use this file to discover all available pages before exploring further.

# Let self-healing act beyond services with Areas

> Switch on self-healing for your own records and for infrastructure findings one area at a time, set each area's tier, budget and pins, undo what it did, and pause it.

export const Plan = ({tier}) => <Badge color="blue">{tier} plan</Badge>;

[Self-healing](/guides/fix/self-healing) starts with actions on a service: restart it, roll it back, run a runbook. **Areas** extend it to the rest of the product. Each of ten areas can act on your own records (mute a noisy alert rule, label a new card, draft a report) or, through a connector you pin, on infrastructure findings (deactivate an unused access key, convert a volume). Every area is off until an organization admin turns it on, and every action it takes appears in the actions inbox and on the area's own page (a failed deploy's ECS or Kubernetes roll back is a service action and is listed under Self-healing, but a pause of the Deployments area holds it).

<Plan tier="Business" />

Areas need the same plan as self-healing, plus the feature of the page they act on, listed in the table below. An area your plan lacks says which feature is missing and shows nothing else, but an area you switched on before a downgrade can still be paused.

Organization admins turn an area on and change it. Members can read it, pause it, approve or refuse what waits, and undo most actions. Viewers see a summary.

## The areas

| Area | What it acts on | Also needs |
| - | - | - |
| **Control Tower** | Noisy alert rules and slow error-budget burns | The Control Tower feature |
| **Suggestions** | Telemetry bindings and suggested alert rules | Nothing beyond self-healing |
| **Board** | New cards in Triage | The ops board |
| **Reports** | Incident and weekly reports | Nothing beyond self-healing |
| **Status page** | Incidents that SRE Agent opened from an alert | The status page |
| **On-call** | Gaps in a schedule and people whose pages keep failing | Nothing beyond self-healing |
| **SLOs** | An SLO's target | Nothing beyond self-healing |
| **Deployments** | Failed deploys and a service's deploy freeze | Nothing beyond self-healing |
| **Security** | IAM access keys, IAM policies and container images with fixable vulnerabilities | The infrastructure suite |
| **FinOps** | Volumes, idle non-production resources and over-provisioned resources | The infrastructure suite |

The **Areas** section also has a first card, **Self-healing**, for the service actions on the [self-healing page](/guides/fix/self-healing). It has no switch of its own: the **Organization** and **Services** sections of the page switch it.

## Turn an area on

<Steps>
  <Step title="Allow self-healing for the organization">
    An area only acts while **Allow self-healing for this organization** is ticked under
    **Organization** on the **Self-healing** page and the organization is not paused. See [Turn it
    on](/guides/fix/self-healing#turn-it-on).
  </Step>

  <Step title="Open the area">
    On the same page, scroll to **Areas**. Each area has a card with its status: **Off**, **On**,
    **Paused** or **Not in your plan**. The link in the card's heading, such as **Open Board**,
    takes you to the page the area acts on.
  </Step>

  <Step title="Switch it on and choose its limits">
    Tick **Allow** followed by the area's name and **to act**, for example **Allow Board to act**.
    Choose **What it may do**, set **Actions per hour**, adjust the per-action choices, pin
    connectors and a repository where the card offers them, and click **Save** followed by the
    area's name, for example **Save Board**.
  </Step>
</Steps>

**What would happen now**, the dry view, also covers the actions of the areas you turned on. It never changes anything.

Each card holds:

* **What it may do.** **Same as the organization**, **Guardrailed** or **Propose only**. An area never has full autonomy of its own: under full autonomy it follows the organization's acknowledgement. See [Full autonomy](/guides/fix/self-healing#full-autonomy).
* **Actions per hour.** From 1 to 20. The default is 3, and 10 for the Board.
* **What it may do on its own.** One choice per action: the default, **Never**, or **Ask first**. A choice can only make an action stricter than the tier allows, so **Automatic** is never offered. A note under each action says what it does by default, that it always waits for a person where that is so, that it cannot be undone, and who can approve it.
* **Connector.** For the actions that act on a resource, the enabled connectors of your organization of that type. An action runs only through the connector pinned here, never one chosen by name or by chance. Once you have connectors in more than one region, each region gets its own pin, and a finding in a region with no pin is held with a reason that names the region. Actions on a resource a service owns use that service's own pins, set in the **Services** section.
* **Infrastructure repository.** For the areas that open pull requests on a repository of their own, one of the repositories your GitHub installation holds. Nothing is ever merged.
* **Settings.** The **Weekly management report day** on Reports, and the **Non-production tags** on FinOps.
* **The model's own cap.** On Control Tower and Suggestions, the card shows the aggressivity of the model that works in that area and links to where you change it. It is separate from the tier.
* **Pause.** See [Pause an area](#pause-an-area).

## What an action does

Every action in an area follows the same three steps. It reads the current state and records what it will put back. It checks the rules a second time and writes. A minute later it reads the result back and ends as in place or not in place.

The rules on [Limits that always apply](/guides/fix/self-healing#limits-that-always-apply) hold for an area too, with these differences:

* For an action on your own records, or on a resource the area pins a connector for, the area's switch and pause take the place of the service's. A pause on the service a record belongs to also stops the action.
* An action on a resource a service owns (a Lambda alias roll back, a deploy freeze, an image bump or right-sizing pull request, scaling an ECS service to zero) also needs that service's self-healing switch and the matching channel, and counts against the service's budget as well as the area's.
* One action on your own records or area-pinned resources is in flight per area at a time; a service's actions follow the service's own limit.
* The cool-down is 15 minutes per target, or the service's own cool-down for an action on its resources. On the Board it is per action and target, so a card can be assigned, labelled and moved in one pass but is not assigned twice in 15 minutes.
* A deploy freeze on a service does not hold the actions that only change your own records, such as a status page update, a report, an SLO target, an on-call override or a card. It still holds connector and pull request actions that name the service, except roll backs and the freeze itself.
* An action in an area never overwrites what a person did in the meantime.

Actions on your own records write directly and use no credential. Two things can still leave the platform: an investigation sends the incident's data to your AI provider like any other investigation, and a card paired with a ticket tracker pushes an assignment or a move to that tracker like any other change to the card.

### Undo

A finished action on your own records that can be taken back (a mute, a region fill, a card change, a draft report, an on-call override, an SLO target or a deploy freeze) has an **Undo** button in the actions inbox and in its details. A person whose role cannot approve that action sees who can instead. There is no Undo for an action that cannot be undone, such as an investigation, a post to Slack or a public status update, and none for an action taken through a connector or a pull request: take those back in AWS, in Kubernetes or by closing the pull request. Undo is itself an action: it goes through the same rules, so a pause refuses it, and an action is undone at most once. An undo never overwrites a change a person made since. It leaves the card, override, target or freeze as it is, says so, and ends done.

When the check after a connector action finds things worse, for example a new critical alert on the service, SRE Agent puts back what it changed on its own where it can (it reactivates a key, starts an instance, restores a task count or rolls an ECS service forward), pauses where it acted and pages a person.

### Who approves

Approving an action, and undoing one, takes the role the page it writes to asks for the same change. An organization admin approves the on-call covers, the status page update, the deploy freeze and the Suggestions region fill. A member approves everything else. A member still sees an admin's requests and can refuse them, and the page says who can approve one.

### What it never acts on

* Signals a model wrote, such as Control Tower agent findings and the alerts the Control Tower filed.
* Text on a card, finding, suggestion, incident or report. Area actions read ids, statuses, schedules, numbers and the detectors' own computed fields, never titles, descriptions or comments. The one exception is a report's list of action items, which becomes the title of a card only after a person approved it.
* Alerts a person acknowledged or muted, alerts that are flapping, and alerts or findings that are no longer open.
* Anything it cannot pin to exactly one target the finding or record itself names, an alert or finding that names more than one service, and, for an action on a service's resource, a resource that service does not own.
* A service marked ignored. A card that names it is not assigned or moved.
* A status page incident a person opened.
* The demonstration organization, which only gets the dry view.

## Control Tower

The Control Tower's detectors file findings. Two kinds have an action, and only for findings a detector wrote.

| Action | When | Default | Undo |
| - | - | - | - |
| **Mute a noisy alert rule** | A noisy rule finding: a rule that fires and clears on its own with nobody acting. The mute lasts 15 minutes to 7 days, from the finding, and a day when it gives none. Nothing is proposed while a mute of the rule is in force. | Automatic in guardrailed | Deletes exactly the mute self-healing created. A mute a person made on the same rule stays. |
| **Open an investigation for a slow burn** | A trend finding: an SLO spending its error budget faster than its target allows. It is not proposed once your monthly AI budget is spent or the investigation limit for the minute is reached. | Waits for a person; automatic under full autonomy only | Nothing to put back, which is why a person decides unless the organization runs full autonomy. |

Unlike a mute a person writes, which matches any alert whose title, source or fingerprint contains its pattern, a mute that self-healing creates matches only an alert whose fingerprint is exactly the rule's, so it never silences another rule's alert.

## Suggestions

| Action | When | Default | Undo |
| - | - | - | - |
| **Fill in a telemetry binding's region** | Config Watch proposed a region for a binding that has none, and its discovery proof is less than an hour old. The region is recomputed from fresh discovery when decided, again before it runs and again inside the write. | Automatic in guardrailed | Puts the region back to empty, only while it is still the region self-healing wrote. A region a person set since stays. |
| **Open a pull request that applies a suggested alert rule** | A suggestion to improve a rule, on the infrastructure repository the area pins. | Waits for a person | Close the pull request. |

## Board

For a card in the **Triage** column created in the last 24 hours, three actions, each written as the system:

| Action | When | Default | Undo |
| - | - | - | - |
| **Assign a new card to the person on call** | The card names a service, has no assignee, and the service's alert route names a schedule with somebody on call. With no route, no schedule or nobody on call, the action is held with that reason. | Automatic in guardrailed | Removes only the person it added. |
| **Label a new card** | Adds `svc:<service slug>`, `source:<where the card came from>` and `priority:<priority>` when the card lacks them. Labels are only added, never removed. | Automatic in guardrailed | Removes exactly the labels it added. |
| **Move a triaged card to To do** | The card has an assignee and a service. It goes to the end of **To do**. | Automatic in guardrailed | Moves it back to Triage, while it is still in To do. |

An assignment notifies the person like any other assignment. An action leaves a card that was archived or moved out of Triage before it ran, and the assignment leaves a card somebody assigned meanwhile.

**Pacing.** A card gets each action at most once every 15 minutes. A new card is assigned first, then labelled, then moved, one step per run, and the three steps use 3 of the Board's 10 actions per hour, so the board finishes about three fully triaged cards an hour. Raise or lower the budget on the Board card. A card older than 24 hours is no longer new and is not acted on, so on a board that receives more new cards than that, the oldest wait unhandled.

## Reports

An action in this area never calls a model itself. The technical and management reports are built from investigations that already exist.

| Action | When | Default | Undo |
| - | - | - | - |
| **Draft a report when an incident resolves** | A status page incident that resolved in the last 6 hours and has a concluded investigation. One technical report per incident, created as a draft. Nothing is drafted while the plan's monthly token allowance is used up. | Automatic in guardrailed | Deletes the draft, and only a report self-healing drafted. |
| **Draft a weekly management report** | On the **Weekly management report day** (Monday unless set) and every day after it until Sunday, for each confirmed service with a concluded investigation in the last 7 days, at most 50. One report per service per week. | Automatic in guardrailed | Deletes the draft. |
| **Post a finished report's link to Slack** | A ready report that self-healing drafted in the last 24 hours, for every type except customer reports. It needs a connected Slack workspace and is sent once per report. | Waits for a person; automatic under full autonomy only | Nothing: a message in Slack cannot be taken back. |
| **Post a customer report's link to Slack** | The same, for a customer report. | Waits for a person in every tier | Nothing. |
| **File a report's action items as cards** | A ready report whose text has `- [ ]` lines. Each line (at most 10, cut to 200 characters) becomes a card in Triage. It needs the ticket board. Once per report. | Waits for a person in every tier | Archives the cards nobody has taken up (still in Triage with nobody assigned). |

A drafted report appears in **Reports** as generating and becomes ready when the report worker has filled it, a minute or more later. A post is sent after the action is recorded, so if Slack refuses it the action ends as not in place and the report stays unposted.

## Status page

| Action | When | Default | Undo |
| - | - | - | - |
| **Tell the status page the cause is identified** | An incident SRE Agent opened from an alert (never one a person wrote) that still reads investigating while its investigation has concluded and names a cause. The incident moves to identified with one update. | Waits for a person; automatic under full autonomy only | Nothing: a public update cannot be unpublished. |

The update is written by the platform from the linked alert's severity and its service label, for example "We have identified the cause of the issue affecting checkout and are working to restore service." Nothing from the investigation is read into it, no model drafts it, and the update is attributed to nobody. An organization admin approves it.

## On-call

Both actions add an override to a schedule. Both wait for a person in every tier, full autonomy included, and the override records the approving person as its creator. An organization admin approves them.

| Action | When | Undo |
| - | - | - |
| **Cover a gap in an on-call schedule** | A stretch in the next 72 hours that nobody covers on an enabled schedule. The person is the one who held the shift just before the gap, else the first person rostered on another enabled schedule. A person whose pages are failing is never chosen. | Deletes exactly the override it created, only while nobody has changed it. |
| **Cover a shift whose person cannot be reached** | A person with two or more failed page deliveries in the last 24 hours who holds a shift in the next 72 hours. The next rostered person who is active and not failing too covers their stretch. | Deletes exactly the override it created, only while nobody has changed it. |

The person and the stretch are checked again when the action is approved and when it runs, so a gap that was covered meanwhile, or a person who left, stops it.

## SLOs

| Action | When | Default | Undo |
| - | - | - | - |
| **Move an SLO's target one step** | An active, warning or breached SLO at least 28 days old, with 28 days of recent measurements and no public incident for its service in that time. **Too loose:** the error budget never fell below 80 percent and compliance is a full step above the target. **Too tight:** the budget ran out in three of the last four weekly windows. | Waits for a person in every tier | Puts the previous target back, only while it is still the one the action wrote. |

The ladder is 90, 95, 99, 99.5, 99.9, 99.95 and 99.99, and the target moves to the next step, never two. The advice is read again when the action is approved and when it runs, and a target a person edited meanwhile is left alone.

## Deployments

The area has to be on for self-healing to read failed deploys and error budgets at all. Its own actions are the Lambda alias roll back and the deploy freeze. A failed ECS or Kubernetes deploy is rolled back by the service's own self-healing: the roll back follows the service's tier, budget and pause, and is listed under **Self-healing** in the actions inbox. Because it answers a failed deploy, which this area reads, pausing the Deployments area, or turning it off, holds it too, as does pausing the service or the organization.

Failed deploys are still read while the area is paused. The roll back is held at the gate with the pause as its reason, and the dry view shows it as paused. A roll back that is waiting or running when you pause is cancelled, and one approved or started later is held. Resume the area and, within a minute, each held roll back, and one the pause cancelled, is decided again while the deploy is still inside the two-hour window: by default it then waits for a person in the actions inbox, as it would have without the pause. A roll back that an alert or a Control Tower deploy regression finding asks for is not this area's: pausing the area does not stop it, and the service's and the organization's switches do.

| Action | When | Default | Undo |
| - | - | - | - |
| **Roll a deploy back** (service self-healing) | A deploy that succeeded in the last two hours, whose post-deploy synthetic checks failed, and that nothing has been deployed over since. An ECS deploy becomes the ECS roll back to the revision it replaced, a Lambda alias deploy becomes the alias roll back below, and a Kubernetes deploy becomes a roll out undo when the deploy record names the namespace and deployment. | Waits for a person | No Undo button. If things get worse after the ECS roll back, SRE Agent rolls the service forward again on its own, unless a person has paused the Deployments area, the service or the organization. The Kubernetes roll back has no automatic undo. |
| **Move a Lambda alias back to its previous version** | Proposed by the roll back above. The alias must still point at the version the deploy set, must not send a share of requests to another version, and an earlier published version must exist. | Waits for a person | Nothing: move the alias again by hand. |
| **Freeze a service's deploys** | An SLO of a confirmed service whose error budget is spent: 24 hours. An incident SRE Agent opened from an alert that is still investigating: 4 hours. It sets the freeze on the service's existing, enabled deploy policy and never creates one. | Waits for a person | Clears the freeze, only while its reason still names this action, and puts back an earlier freeze of a person's. |

Every roll back needs the service's self-healing switch, its **Connector actions** channel and its pinned connector. The freeze needs the service's switch and its **Platform actions** channel, and an organization admin approves it. A deploy freeze never holds a roll back or the freeze itself.

## Security

Pin an **AWS IAM** connector on the area for the key action, and an infrastructure repository for the policy pull request.

| Action | When | Default | Undo |
| - | - | - | - |
| **Deactivate an unused IAM access key** | An open stale access key finding from an IAM scan that completed in the last 48 hours, one proposal per key. The key is read again when the action runs: it must be active and not used in the last 90 days, or the action ends as in use. A key one of your own connectors or data sources signs with is never deactivated. | Waits for a person | No Undo button: reactivate the key in AWS. SRE Agent reactivates it on its own only when things get worse after the action, and only while it is still inactive. |
| **Open a pull request that narrows an over-broad IAM policy** | An open wildcard policy or unused permissions finding, on the area's pinned infrastructure repository. | Waits for a person | Close the pull request. |
| **Open a pull request that bumps an image** | A container image whose newest completed scan lists open critical or high vulnerabilities with a fixed version, on the code repository of the one service whose live inventory runs the image. | Waits for a person | Close the pull request. |

Rotating a key is not automated. A deactivation is verified by the next daily IAM inventory.

## FinOps

Pin an **AWS EC2 and Auto Scaling** connector on the area for the gp3 conversion and the instance stop, and one per region once your connectors span more than one region. Scaling to zero uses the owning service's pinned ECS connector, and the right-sizing pull request opens on the owning service's linked infrastructure repository.

| Action | When | Default | Undo |
| - | - | - | - |
| **Convert a gp2 volume to gp3** | An open gp2 to gp3 finding. The volume must still be gp2 when it is read. | Waits for a person | Nothing: a volume type change cannot be undone from here. |
| **Stop an idle instance** | An open idle instance finding, only for an instance whose tags match one of the area's **Non-production tags** rules. | Waits for a person | No Undo button: start the instance in AWS. SRE Agent starts it on its own only when things get worse after the action, and only while it is still stopped. |
| **Scale an idle ECS service to zero** | An open idle workload finding on an ECS service whose tags match a non-production rule and that runs at least one task. It uses the owning service's pinned ECS connector. | Waits for a person | No Undo button: set the task count in AWS. SRE Agent puts it back on its own only when things get worse after the action, and only while it is still zero. |
| **Open a pull request that right-sizes a resource** | An open oversized workload, over-reserved Fargate service or over-provisioned Lambda finding, for the one service that owns the resource and has the **Pull requests** channel on. | Waits for a person | Close the pull request. |

**Non-production tags** are set by an organization admin: a tag key, with the values that mean non-production, for example `env` with `staging` and `dev`. A key is compared exactly and a value without regard to case. With no rules nothing is non-production, so nothing is stopped or scaled to zero. The tags are read when the action runs, and a production resource is left alone without paging anyone. The finding's monthly saving is copied onto the action so the inbox can total what was done. A finding is verified by the next scan.

A pause or a maintenance window on any service that owns the resource stops the action, not only on the service the finding names.

## Pause an area

A member can pause an area with a reason, and only an organization admin resumes it.

* **On the page.** On the area's card, type a reason next to **Pause this area, with a reason** and click **Pause**. The card shows who paused it and why, with **Resume** for admins.
* **Over MCP.** `pause_autonomy_area` and `resume_autonomy_area`, below.

Pausing cancels every pending approval and every running action of the area. For Deployments that includes the ECS and Kubernetes roll backs that failed deploys start. Pausing the whole organization, or the platform, stops the areas too, and a pause on the service a record belongs to stops the actions on that record. An organization whose plan no longer includes an area can still pause it.

**Paging is never paused.** Paging, escalation and on-call notifications read none of these switches. Pausing self-healing for the platform, the organization, a service or an area never silences a page.

## See what an area did

* **On the area's own page.** **What the platform did here** lists the area's ten latest actions, with **All of** the area's name **in the actions inbox**. The panel is on the Control Tower, Suggestions, Board, Reports, Status page, On-call, SLOs, Deployments, Security and FinOps pages. An organization without the plan sees no panel.
* **In the actions inbox.** Open `/self-healing/actions` and pick the area in **Area**. The address keeps your choice, as `?area=` followed by the area's key.

Each row says what the action is, where it happens, why it is where it is, its state and who decided it. See [The actions inbox](/guides/fix/self-healing#the-actions-inbox).

## MCP tools

Four tools read and change the same settings as the Areas section. Each is gated like its page: the plan comes first and answers "Self-healing is not included in your plan."

| Tool | Role needed | What it does |
| - | - | - |
| `get_autonomy_areas` | Viewer | Each area's switch, status, tier, budget, pause, overrides, pinned connectors, pinned repository, settings and the actions it may take. An area the plan lacks says what is missing and nothing else. |
| `update_autonomy_area` | Org admin | One area's `enabled`, `tier`, `max_actions_per_hour`, `action_overrides`, `connector_ids`, `repos` and `settings`. The maps are merged over what is stored, and a null value removes its key. |
| `pause_autonomy_area` | Member | Pause one area. A reason is required. |
| `resume_autonomy_area` | Org admin | Resume one area. Needs `confirm: true`. |

The area key is one of `overseer` (Control Tower), `suggestions`, `board`, `reports`, `status_page`, `oncall`, `slo`, `deployments`, `finops` or `security`. Self-healing for services is switched with its own tools, and an unknown key answers "Autonomy area not found." `list_healing_actions` also takes an `area` filter.

Every write is a person's decision, so it needs a personal access token. An API key is nobody, and its call is refused with a sentence that says so. A token's role ceiling applies. `update_autonomy_area` refuses `tier: "full_autonomy"` with "Full autonomy is enabled only on the Self-healing page, by an organization admin who types the organization name." It also refuses a connector pin the page does not offer for the area. Writes are audited as `self_healing.area_policy_updated`, `self_healing.area_paused` and `self_healing.area_resumed`. See [MCP tools](/api-reference/mcp) for how to connect a client.

## Related

* [Let SRE Agent act on incidents with self-healing](/guides/fix/self-healing): the service actions, tiers, limits and full autonomy.
* [Find and manage your services](/guides/get-started/services): the service pins an area's resource actions use.
* [Triage alerts](/guides/respond/alerts): the alerts the Control Tower and the Board work from.
* [Plan matrix](/guides/reference/plan-matrix): which plans include self-healing.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.