Skip to main content
Self-healing starts with actions on a service: restart it, roll it back, run a runbook. Areas extend it to the rest of the product. Each of ten areas can act on your own records (mute a noisy alert rule, label a new card, draft a report) or, through a connector you pin, on infrastructure findings (deactivate an unused access key, convert a volume). Every area is off until an organization admin turns it on, and every action it takes appears in the actions inbox and on the area’s own page (a failed deploy’s ECS or Kubernetes roll back is a service action and is listed under Self-healing, but a pause of the Deployments area holds it). Areas need the same plan as self-healing, plus the feature of the page they act on, listed in the table below. An area your plan lacks says which feature is missing and shows nothing else, but an area you switched on before a downgrade can still be paused. Organization admins turn an area on and change it. Members can read it, pause it, approve or refuse what waits, and undo most actions. Viewers see a summary.

The areas

The Areas section also has a first card, Self-healing, for the service actions on the self-healing page. It has no switch of its own: the Organization and Services sections of the page switch it.

Turn an area on

1

Allow self-healing for the organization

An area only acts while Allow self-healing for this organization is ticked under Organization on the Self-healing page and the organization is not paused. See Turn it on.
2

Open the area

On the same page, scroll to Areas. Each area has a card with its status: Off, On, Paused or Not in your plan. The link in the card’s heading, such as Open Board, takes you to the page the area acts on.
3

Switch it on and choose its limits

Tick Allow followed by the area’s name and to act, for example Allow Board to act. Choose What it may do, set Actions per hour, adjust the per-action choices, pin connectors and a repository where the card offers them, and click Save followed by the area’s name, for example Save Board.
What would happen now, the dry view, also covers the actions of the areas you turned on. It never changes anything. Each card holds:
  • What it may do. Same as the organization, Guardrailed or Propose only. An area never has full autonomy of its own: under full autonomy it follows the organization’s acknowledgement. See Full autonomy.
  • Actions per hour. From 1 to 20. The default is 3, and 10 for the Board.
  • What it may do on its own. One choice per action: the default, Never, or Ask first. A choice can only make an action stricter than the tier allows, so Automatic is never offered. A note under each action says what it does by default, that it always waits for a person where that is so, that it cannot be undone, and who can approve it.
  • Connector. For the actions that act on a resource, the enabled connectors of your organization of that type. An action runs only through the connector pinned here, never one chosen by name or by chance. Once you have connectors in more than one region, each region gets its own pin, and a finding in a region with no pin is held with a reason that names the region. Actions on a resource a service owns use that service’s own pins, set in the Services section.
  • Infrastructure repository. For the areas that open pull requests on a repository of their own, one of the repositories your GitHub installation holds. Nothing is ever merged.
  • Settings. The Weekly management report day on Reports, and the Non-production tags on FinOps.
  • The model’s own cap. On Control Tower and Suggestions, the card shows the aggressivity of the model that works in that area and links to where you change it. It is separate from the tier.
  • Pause. See Pause an area.

What an action does

Every action in an area follows the same three steps. It reads the current state and records what it will put back. It checks the rules a second time and writes. A minute later it reads the result back and ends as in place or not in place. The rules on Limits that always apply hold for an area too, with these differences:
  • For an action on your own records, or on a resource the area pins a connector for, the area’s switch and pause take the place of the service’s. A pause on the service a record belongs to also stops the action.
  • An action on a resource a service owns (a Lambda alias roll back, a deploy freeze, an image bump or right-sizing pull request, scaling an ECS service to zero) also needs that service’s self-healing switch and the matching channel, and counts against the service’s budget as well as the area’s.
  • One action on your own records or area-pinned resources is in flight per area at a time; a service’s actions follow the service’s own limit.
  • The cool-down is 15 minutes per target, or the service’s own cool-down for an action on its resources. On the Board it is per action and target, so a card can be assigned, labelled and moved in one pass but is not assigned twice in 15 minutes.
  • A deploy freeze on a service does not hold the actions that only change your own records, such as a status page update, a report, an SLO target, an on-call override or a card. It still holds connector and pull request actions that name the service, except roll backs and the freeze itself.
  • An action in an area never overwrites what a person did in the meantime.
Actions on your own records write directly and use no credential. Two things can still leave the platform: an investigation sends the incident’s data to your AI provider like any other investigation, and a card paired with a ticket tracker pushes an assignment or a move to that tracker like any other change to the card.

Undo

A finished action on your own records that can be taken back (a mute, a region fill, a card change, a draft report, an on-call override, an SLO target or a deploy freeze) has an Undo button in the actions inbox and in its details. A person whose role cannot approve that action sees who can instead. There is no Undo for an action that cannot be undone, such as an investigation, a post to Slack or a public status update, and none for an action taken through a connector or a pull request: take those back in AWS, in Kubernetes or by closing the pull request. Undo is itself an action: it goes through the same rules, so a pause refuses it, and an action is undone at most once. An undo never overwrites a change a person made since. It leaves the card, override, target or freeze as it is, says so, and ends done. When the check after a connector action finds things worse, for example a new critical alert on the service, SRE Agent puts back what it changed on its own where it can (it reactivates a key, starts an instance, restores a task count or rolls an ECS service forward), pauses where it acted and pages a person.

Who approves

Approving an action, and undoing one, takes the role the page it writes to asks for the same change. An organization admin approves the on-call covers, the status page update, the deploy freeze and the Suggestions region fill. A member approves everything else. A member still sees an admin’s requests and can refuse them, and the page says who can approve one.

What it never acts on

  • Signals a model wrote, such as Control Tower agent findings and the alerts the Control Tower filed.
  • Text on a card, finding, suggestion, incident or report. Area actions read ids, statuses, schedules, numbers and the detectors’ own computed fields, never titles, descriptions or comments. The one exception is a report’s list of action items, which becomes the title of a card only after a person approved it.
  • Alerts a person acknowledged or muted, alerts that are flapping, and alerts or findings that are no longer open.
  • Anything it cannot pin to exactly one target the finding or record itself names, an alert or finding that names more than one service, and, for an action on a service’s resource, a resource that service does not own.
  • A service marked ignored. A card that names it is not assigned or moved.
  • A status page incident a person opened.
  • The demonstration organization, which only gets the dry view.

Control Tower

The Control Tower’s detectors file findings. Two kinds have an action, and only for findings a detector wrote. Unlike a mute a person writes, which matches any alert whose title, source or fingerprint contains its pattern, a mute that self-healing creates matches only an alert whose fingerprint is exactly the rule’s, so it never silences another rule’s alert.

Suggestions

Board

For a card in the Triage column created in the last 24 hours, three actions, each written as the system: An assignment notifies the person like any other assignment. An action leaves a card that was archived or moved out of Triage before it ran, and the assignment leaves a card somebody assigned meanwhile. Pacing. A card gets each action at most once every 15 minutes. A new card is assigned first, then labelled, then moved, one step per run, and the three steps use 3 of the Board’s 10 actions per hour, so the board finishes about three fully triaged cards an hour. Raise or lower the budget on the Board card. A card older than 24 hours is no longer new and is not acted on, so on a board that receives more new cards than that, the oldest wait unhandled.

Reports

An action in this area never calls a model itself. The technical and management reports are built from investigations that already exist. A drafted report appears in Reports as generating and becomes ready when the report worker has filled it, a minute or more later. A post is sent after the action is recorded, so if Slack refuses it the action ends as not in place and the report stays unposted.

Status page

The update is written by the platform from the linked alert’s severity and its service label, for example “We have identified the cause of the issue affecting checkout and are working to restore service.” Nothing from the investigation is read into it, no model drafts it, and the update is attributed to nobody. An organization admin approves it.

On-call

Both actions add an override to a schedule. Both wait for a person in every tier, full autonomy included, and the override records the approving person as its creator. An organization admin approves them. The person and the stretch are checked again when the action is approved and when it runs, so a gap that was covered meanwhile, or a person who left, stops it.

SLOs

The ladder is 90, 95, 99, 99.5, 99.9, 99.95 and 99.99, and the target moves to the next step, never two. The advice is read again when the action is approved and when it runs, and a target a person edited meanwhile is left alone.

Deployments

The area has to be on for self-healing to read failed deploys and error budgets at all. Its own actions are the Lambda alias roll back and the deploy freeze. A failed ECS or Kubernetes deploy is rolled back by the service’s own self-healing: the roll back follows the service’s tier, budget and pause, and is listed under Self-healing in the actions inbox. Because it answers a failed deploy, which this area reads, pausing the Deployments area, or turning it off, holds it too, as does pausing the service or the organization. Failed deploys are still read while the area is paused. The roll back is held at the gate with the pause as its reason, and the dry view shows it as paused. A roll back that is waiting or running when you pause is cancelled, and one approved or started later is held. Resume the area and, within a minute, each held roll back, and one the pause cancelled, is decided again while the deploy is still inside the two-hour window: by default it then waits for a person in the actions inbox, as it would have without the pause. A roll back that an alert or a Control Tower deploy regression finding asks for is not this area’s: pausing the area does not stop it, and the service’s and the organization’s switches do. Every roll back needs the service’s self-healing switch, its Connector actions channel and its pinned connector. The freeze needs the service’s switch and its Platform actions channel, and an organization admin approves it. A deploy freeze never holds a roll back or the freeze itself.

Security

Pin an AWS IAM connector on the area for the key action, and an infrastructure repository for the policy pull request. Rotating a key is not automated. A deactivation is verified by the next daily IAM inventory.

FinOps

Pin an AWS EC2 and Auto Scaling connector on the area for the gp3 conversion and the instance stop, and one per region once your connectors span more than one region. Scaling to zero uses the owning service’s pinned ECS connector, and the right-sizing pull request opens on the owning service’s linked infrastructure repository. Non-production tags are set by an organization admin: a tag key, with the values that mean non-production, for example env with staging and dev. A key is compared exactly and a value without regard to case. With no rules nothing is non-production, so nothing is stopped or scaled to zero. The tags are read when the action runs, and a production resource is left alone without paging anyone. The finding’s monthly saving is copied onto the action so the inbox can total what was done. A finding is verified by the next scan. A pause or a maintenance window on any service that owns the resource stops the action, not only on the service the finding names.

Pause an area

A member can pause an area with a reason, and only an organization admin resumes it.
  • On the page. On the area’s card, type a reason next to Pause this area, with a reason and click Pause. The card shows who paused it and why, with Resume for admins.
  • Over MCP. pause_autonomy_area and resume_autonomy_area, below.
Pausing cancels every pending approval and every running action of the area. For Deployments that includes the ECS and Kubernetes roll backs that failed deploys start. Pausing the whole organization, or the platform, stops the areas too, and a pause on the service a record belongs to stops the actions on that record. An organization whose plan no longer includes an area can still pause it. Paging is never paused. Paging, escalation and on-call notifications read none of these switches. Pausing self-healing for the platform, the organization, a service or an area never silences a page.

See what an area did

  • On the area’s own page. What the platform did here lists the area’s ten latest actions, with All of the area’s name in the actions inbox. The panel is on the Control Tower, Suggestions, Board, Reports, Status page, On-call, SLOs, Deployments, Security and FinOps pages. An organization without the plan sees no panel.
  • In the actions inbox. Open /self-healing/actions and pick the area in Area. The address keeps your choice, as ?area= followed by the area’s key.
Each row says what the action is, where it happens, why it is where it is, its state and who decided it. See The actions inbox.

MCP tools

Four tools read and change the same settings as the Areas section. Each is gated like its page: the plan comes first and answers “Self-healing is not included in your plan.” The area key is one of overseer (Control Tower), suggestions, board, reports, status_page, oncall, slo, deployments, finops or security. Self-healing for services is switched with its own tools, and an unknown key answers “Autonomy area not found.” list_healing_actions also takes an area filter. Every write is a person’s decision, so it needs a personal access token. An API key is nobody, and its call is refused with a sentence that says so. A token’s role ceiling applies. update_autonomy_area refuses tier: "full_autonomy" with “Full autonomy is enabled only on the Self-healing page, by an organization admin who types the organization name.” It also refuses a connector pin the page does not offer for the area. Writes are audited as self_healing.area_policy_updated, self_healing.area_paused and self_healing.area_resumed. See MCP tools for how to connect a client.