The areas
The Areas section also has a first card, Self-healing, for the service actions on the self-healing page. It has no switch of its own: the Organization and Services sections of the page switch it.
Turn an area on
1
Allow self-healing for the organization
An area only acts while Allow self-healing for this organization is ticked under
Organization on the Self-healing page and the organization is not paused. See Turn it
on.
2
Open the area
On the same page, scroll to Areas. Each area has a card with its status: Off, On,
Paused or Not in your plan. The link in the card’s heading, such as Open Board,
takes you to the page the area acts on.
3
Switch it on and choose its limits
Tick Allow followed by the area’s name and to act, for example Allow Board to act.
Choose What it may do, set Actions per hour, adjust the per-action choices, pin
connectors and a repository where the card offers them, and click Save followed by the
area’s name, for example Save Board.
- What it may do. Same as the organization, Guardrailed or Propose only. An area never has full autonomy of its own: under full autonomy it follows the organization’s acknowledgement. See Full autonomy.
- Actions per hour. From 1 to 20. The default is 3, and 10 for the Board.
- What it may do on its own. One choice per action: the default, Never, or Ask first. A choice can only make an action stricter than the tier allows, so Automatic is never offered. A note under each action says what it does by default, that it always waits for a person where that is so, that it cannot be undone, and who can approve it.
- Connector. For the actions that act on a resource, the enabled connectors of your organization of that type. An action runs only through the connector pinned here, never one chosen by name or by chance. Once you have connectors in more than one region, each region gets its own pin, and a finding in a region with no pin is held with a reason that names the region. Actions on a resource a service owns use that service’s own pins, set in the Services section.
- Infrastructure repository. For the areas that open pull requests on a repository of their own, one of the repositories your GitHub installation holds. Nothing is ever merged.
- Settings. The Weekly management report day on Reports, and the Non-production tags on FinOps.
- The model’s own cap. On Control Tower and Suggestions, the card shows the aggressivity of the model that works in that area and links to where you change it. It is separate from the tier.
- Pause. See Pause an area.
What an action does
Every action in an area follows the same three steps. It reads the current state and records what it will put back. It checks the rules a second time and writes. A minute later it reads the result back and ends as in place or not in place. The rules on Limits that always apply hold for an area too, with these differences:- For an action on your own records, or on a resource the area pins a connector for, the area’s switch and pause take the place of the service’s. A pause on the service a record belongs to also stops the action.
- An action on a resource a service owns (a Lambda alias roll back, a deploy freeze, an image bump or right-sizing pull request, scaling an ECS service to zero) also needs that service’s self-healing switch and the matching channel, and counts against the service’s budget as well as the area’s.
- One action on your own records or area-pinned resources is in flight per area at a time; a service’s actions follow the service’s own limit.
- The cool-down is 15 minutes per target, or the service’s own cool-down for an action on its resources. On the Board it is per action and target, so a card can be assigned, labelled and moved in one pass but is not assigned twice in 15 minutes.
- A deploy freeze on a service does not hold the actions that only change your own records, such as a status page update, a report, an SLO target, an on-call override or a card. It still holds connector and pull request actions that name the service, except roll backs and the freeze itself.
- An action in an area never overwrites what a person did in the meantime.
Undo
A finished action on your own records that can be taken back (a mute, a region fill, a card change, a draft report, an on-call override, an SLO target or a deploy freeze) has an Undo button in the actions inbox and in its details. A person whose role cannot approve that action sees who can instead. There is no Undo for an action that cannot be undone, such as an investigation, a post to Slack or a public status update, and none for an action taken through a connector or a pull request: take those back in AWS, in Kubernetes or by closing the pull request. Undo is itself an action: it goes through the same rules, so a pause refuses it, and an action is undone at most once. An undo never overwrites a change a person made since. It leaves the card, override, target or freeze as it is, says so, and ends done. When the check after a connector action finds things worse, for example a new critical alert on the service, SRE Agent puts back what it changed on its own where it can (it reactivates a key, starts an instance, restores a task count or rolls an ECS service forward), pauses where it acted and pages a person.Who approves
Approving an action, and undoing one, takes the role the page it writes to asks for the same change. An organization admin approves the on-call covers, the status page update, the deploy freeze and the Suggestions region fill. A member approves everything else. A member still sees an admin’s requests and can refuse them, and the page says who can approve one.What it never acts on
- Signals a model wrote, such as Control Tower agent findings and the alerts the Control Tower filed.
- Text on a card, finding, suggestion, incident or report. Area actions read ids, statuses, schedules, numbers and the detectors’ own computed fields, never titles, descriptions or comments. The one exception is a report’s list of action items, which becomes the title of a card only after a person approved it.
- Alerts a person acknowledged or muted, alerts that are flapping, and alerts or findings that are no longer open.
- Anything it cannot pin to exactly one target the finding or record itself names, an alert or finding that names more than one service, and, for an action on a service’s resource, a resource that service does not own.
- A service marked ignored. A card that names it is not assigned or moved.
- A status page incident a person opened.
- The demonstration organization, which only gets the dry view.
Control Tower
The Control Tower’s detectors file findings. Two kinds have an action, and only for findings a detector wrote.
Unlike a mute a person writes, which matches any alert whose title, source or fingerprint contains its pattern, a mute that self-healing creates matches only an alert whose fingerprint is exactly the rule’s, so it never silences another rule’s alert.
Suggestions
Board
For a card in the Triage column created in the last 24 hours, three actions, each written as the system:
An assignment notifies the person like any other assignment. An action leaves a card that was archived or moved out of Triage before it ran, and the assignment leaves a card somebody assigned meanwhile.
Pacing. A card gets each action at most once every 15 minutes. A new card is assigned first, then labelled, then moved, one step per run, and the three steps use 3 of the Board’s 10 actions per hour, so the board finishes about three fully triaged cards an hour. Raise or lower the budget on the Board card. A card older than 24 hours is no longer new and is not acted on, so on a board that receives more new cards than that, the oldest wait unhandled.
Reports
An action in this area never calls a model itself. The technical and management reports are built from investigations that already exist.
A drafted report appears in Reports as generating and becomes ready when the report worker has filled it, a minute or more later. A post is sent after the action is recorded, so if Slack refuses it the action ends as not in place and the report stays unposted.
Status page
The update is written by the platform from the linked alert’s severity and its service label, for example “We have identified the cause of the issue affecting checkout and are working to restore service.” Nothing from the investigation is read into it, no model drafts it, and the update is attributed to nobody. An organization admin approves it.
On-call
Both actions add an override to a schedule. Both wait for a person in every tier, full autonomy included, and the override records the approving person as its creator. An organization admin approves them.
The person and the stretch are checked again when the action is approved and when it runs, so a gap that was covered meanwhile, or a person who left, stops it.
SLOs
The ladder is 90, 95, 99, 99.5, 99.9, 99.95 and 99.99, and the target moves to the next step, never two. The advice is read again when the action is approved and when it runs, and a target a person edited meanwhile is left alone.
Deployments
The area has to be on for self-healing to read failed deploys and error budgets at all. Its own actions are the Lambda alias roll back and the deploy freeze. A failed ECS or Kubernetes deploy is rolled back by the service’s own self-healing: the roll back follows the service’s tier, budget and pause, and is listed under Self-healing in the actions inbox. Because it answers a failed deploy, which this area reads, pausing the Deployments area, or turning it off, holds it too, as does pausing the service or the organization. Failed deploys are still read while the area is paused. The roll back is held at the gate with the pause as its reason, and the dry view shows it as paused. A roll back that is waiting or running when you pause is cancelled, and one approved or started later is held. Resume the area and, within a minute, each held roll back, and one the pause cancelled, is decided again while the deploy is still inside the two-hour window: by default it then waits for a person in the actions inbox, as it would have without the pause. A roll back that an alert or a Control Tower deploy regression finding asks for is not this area’s: pausing the area does not stop it, and the service’s and the organization’s switches do.
Every roll back needs the service’s self-healing switch, its Connector actions channel and its pinned connector. The freeze needs the service’s switch and its Platform actions channel, and an organization admin approves it. A deploy freeze never holds a roll back or the freeze itself.
Security
Pin an AWS IAM connector on the area for the key action, and an infrastructure repository for the policy pull request.
Rotating a key is not automated. A deactivation is verified by the next daily IAM inventory.
FinOps
Pin an AWS EC2 and Auto Scaling connector on the area for the gp3 conversion and the instance stop, and one per region once your connectors span more than one region. Scaling to zero uses the owning service’s pinned ECS connector, and the right-sizing pull request opens on the owning service’s linked infrastructure repository.
Non-production tags are set by an organization admin: a tag key, with the values that mean non-production, for example
env with staging and dev. A key is compared exactly and a value without regard to case. With no rules nothing is non-production, so nothing is stopped or scaled to zero. The tags are read when the action runs, and a production resource is left alone without paging anyone. The finding’s monthly saving is copied onto the action so the inbox can total what was done. A finding is verified by the next scan.
A pause or a maintenance window on any service that owns the resource stops the action, not only on the service the finding names.
Pause an area
A member can pause an area with a reason, and only an organization admin resumes it.- On the page. On the area’s card, type a reason next to Pause this area, with a reason and click Pause. The card shows who paused it and why, with Resume for admins.
- Over MCP.
pause_autonomy_areaandresume_autonomy_area, below.
See what an area did
- On the area’s own page. What the platform did here lists the area’s ten latest actions, with All of the area’s name in the actions inbox. The panel is on the Control Tower, Suggestions, Board, Reports, Status page, On-call, SLOs, Deployments, Security and FinOps pages. An organization without the plan sees no panel.
- In the actions inbox. Open
/self-healing/actionsand pick the area in Area. The address keeps your choice, as?area=followed by the area’s key.
MCP tools
Four tools read and change the same settings as the Areas section. Each is gated like its page: the plan comes first and answers “Self-healing is not included in your plan.”
The area key is one of
overseer (Control Tower), suggestions, board, reports, status_page, oncall, slo, deployments, finops or security. Self-healing for services is switched with its own tools, and an unknown key answers “Autonomy area not found.” list_healing_actions also takes an area filter.
Every write is a person’s decision, so it needs a personal access token. An API key is nobody, and its call is refused with a sentence that says so. A token’s role ceiling applies. update_autonomy_area refuses tier: "full_autonomy" with “Full autonomy is enabled only on the Self-healing page, by an organization admin who types the organization name.” It also refuses a connector pin the page does not offer for the area. Writes are audited as self_healing.area_policy_updated, self_healing.area_paused and self_healing.area_resumed. See MCP tools for how to connect a client.
Related
- Let SRE Agent act on incidents with self-healing: the service actions, tiers, limits and full autonomy.
- Find and manage your services: the service pins an area’s resource actions use.
- Triage alerts: the alerts the Control Tower and the Board work from.
- Plan matrix: which plans include self-healing.