> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sreagent.app/llms.txt
> Use this file to discover all available pages before exploring further.

# Runbook actions and step reference

> Look up what each connector can do, how templates and fan-out work, how to verify a step's output, and why a step fails.

export const Plan = ({tier}) => <Badge color="blue">{tier} plan</Badge>;

This is the reference behind [Write and run runbooks](/guides/prevent/runbooks). The in-app **Connector reference** page (the **Every action, with its parameters** link in the step editor) lists every action with its parameters and shows which connectors your organization has configured.

<Plan tier="Pro" />

## Use values from the alert and earlier steps

Actions accept templates. An unresolved reference fails the step with an error and never sends a blank value.

| Template | Resolves to |
| - | - |
| `{{alert.title}}`, `{{alert.severity}}`, `{{alert.source}}` | The alert that triggered the run |
| `{{alert.labels.<key>}}`, `{{alert.annotations.<key>}}` | A label or annotation on that alert |
| `{{step_N.output}}`, `{{step_N.<field>}}` | The output of step N, or a field an `input` step collected (steps count from 1) |
| `{{param.<name>}}` | A recipe parameter |

One runbook can serve many near-identical workloads this way. A step that names `{{alert.labels.service}}` restarts whichever service alerted. An alert from a synthetic check linked to a service carries that service as a label, and an SLO breach carries `sli_service`, the service its SLI is tagged with, so choose that value deliberately.

## Fan out over many items

A discovery step (`get_unhealthy_pods`, `get_unhealthy_deployments`, or an AWS `list_*` action) finds items, and a later write step can act on each one. Tick **Fan out: run this action once per item found by the discovery step** and set the identifying field to `*`. If discovery finds nothing, the step does nothing. If discovery finds many items, the step pauses and asks for approval with the count, in the **A person at each risky step** mode. There is also a limit on how many items one fan-out may touch, and above it the step refuses instead of asking, in every mode. **After the fan-out** can re-scan once the writes finish and fail the step unless fewer, or none, remain, with a wait of up to 600 seconds for changes AWS applies in the background.

## Verify what a step says

A step passes when its connector answers without an error. To fail on what the answer says, set **Verify Output** on the step: pick an operator and, for most, a value. The operators are `contains`, `not contains`, `equals`, `not equals`, `regex`, `items empty`, `items present`, `count eq` and `count lt`. The item and count operators read what the step found rather than what it printed, so "the re-scan must find nothing" is `items empty` with no text to type.

Verify with the same signal that triggered the runbook. If a synthetic check started it, finish with a step that runs that check and expects success.

## What each connector can do

Steps run through the connectors you configure under **Settings > Infrastructure**. The four AWS connector types share one action set.

| Target | Read actions | Write actions |
| - | - | - |
| `kubernetes` | `describe_resource`, `get_logs`, `get_resource`, `get_unhealthy_deployments`, `get_unhealthy_pods`, `http_health_check`, `rollout_status`, `top_nodes`, `top_pods` | `apply_manifest`, `create_resource`, `patch_resource`, `delete_resource`, `delete_pod`, `restart_deployment`, `restart_unhealthy_deployments`, `rollback_deployment`, `scale_deployment`, `cordon_node`, `uncordon_node`, `drain_node`, `exec_command` |
| `ssh` | none | `run_command`, `upload_file`, `download_file` |
| `aws_ecs`, `aws_lambda`, `aws_ec2`, `aws_eks` | `aws_read_call`, `describe_service`, `describe_instance`, `describe_volume`, `describe_auto_scaling_group`, `get_function`, `list_services`, `list_tasks`, `list_volumes`, `list_functions`, `list_auto_scaling_groups`, `list_log_groups` | `force_new_deployment`, `restart_task`, `update_service_desired_count`, `invoke_function`, `update_function_config`, `set_function_concurrency`, `remove_function_concurrency`, `modify_volume`, `reboot_instance`, `start_instance`, `stop_instance`, `set_desired_capacity`, `set_log_group_retention`, `ssm_send_command` |
| `prometheus` | `query`, `health_check` | none |
| `cloudwatch` | `get_logs`, `get_metrics` | none |
| `synthetic` | `run_check` | none |

`prometheus`, `cloudwatch` and `synthetic` use your existing data sources and checks, so they need no connector. Actions that run arbitrary commands (`exec_command`, `run_command`, `ssm_send_command`) should be marked high risk.

Two actions look alike and are not. `restart_deployment` restarts one named deployment and refuses an empty name. `restart_unhealthy_deployments` restarts every unhealthy deployment in the namespace, or in the whole cluster if you give no namespace.

## When a step fails

| Message | Cause |
| - | - |
| No enabled connector found for type | The organization has no enabled connector of that type. Data sources do not count. |
| Unresolved reference(s) in action | A `{{...}}` template points at a step or field that does not exist. |
| `AccessDenied` from AWS | The connector's role lacks the action. Compare it with the IAM permissions Validate lists. |
| `agent_unavailable` | The target is reached through a remote agent that is offline. |
| The runbook is not approved | Only approved runbooks run. Dry run it, then approve it. |

## Rules of thumb

* If the platform cannot do a step, use an `input` or `approval_gate` step and hand off to a person.
* Put the destructive step behind exactly one approval, and restart the traffic-facing component last, after the services behind it.
* Prefer a recipe over hand-written steps whenever one fits.

## Related

* [Connect AWS with a read-only role](/guides/get-started/connect-aws): give a runbook connector read access.
* [Monitor endpoints with synthetic checks](/guides/prevent/synthetic-checks): verify a step with a synthetic check.
* [Troubleshooting](/guides/reference/troubleshooting): more refused-step messages and fixes.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.