> ## Documentation Index
> Fetch the complete documentation index at: https://docs.sreagent.app/llms.txt
> Use this file to discover all available pages before exploring further.

# Find and manage your services

> See how SRE Agent discovers your services, confirm, merge or ignore them, bind telemetry and repositories to each one, and read a service's page.

export const Plan = ({tier}) => <Badge color="blue">{tier} plan</Badge>;

A service is something you run: a checkout API, a worker, a Lambda function. SRE Agent finds them from your alert labels, your cloud and cluster inventory, deploy events, SLIs and the repositories your GitHub App can read. **Services** lists them, and a service's own page shows everything SRE Agent holds about it.

<Plan tier="Free" />

The list, the pages and discovery are on every plan. Telemetry signals, binding health and owner teams belong to the infrastructure suite, and linking repositories needs the GitHub integration. Both are on Business and above. A locked section shows an upgrade card. See [Plan matrix](/guides/reference/plan-matrix).

<Steps>
  <Step title="Open Services">
    Click **Services** in the **Prevent** section of the sidebar. Services appear once alerts,
    inventory, deploys or repositories name them.
  </Step>

  <Step title="Confirm the ones that are real">
    A service starts as **discovered**. Open it and click **Confirm**, or tick several rows and
    click **Confirm selected**. Confirmed services are offered first in every service picker.
  </Step>

  <Step title="Fix the duplicates">
    If the same service is listed twice under different names, an organization admin merges one into
    the other from its page. See below.
  </Step>
</Steps>

## How services are found

Discovery runs for every organization once an hour and reads only what you have already connected. Each name it meets is a candidate from one of these places:

* alert labels of the last 30 days;
* ECS services, Lambda functions and Kubernetes workloads in the hourly inventory;
* deploy events, SLIs, synthetic checks and telemetry bindings;
* the repositories your GitHub installation can read.

There is no button that creates a service by hand. A service exists because something named it. A candidate is resolved in this order, and discovery never merges on a similarity score:

1. A name that is already an alias of a service marks that service as seen. An ignored service is left as it is.
2. A name whose canonical key equals the canonical key of an existing service or alias is added as an alias of that service. These are the only merges discovery makes without a person. An ignored service absorbs the new spelling as an alias and stays ignored.
3. Otherwise a **discovered** service is created. Its slug is the canonical key and its name is the first spelling seen.

The canonical key is the name in lower case, with a Kubernetes pod suffix removed, then at most one environment word and at most one role word removed:

* Environment words at the end after `-` or `_`: `prod`, `production`, `staging`, `stage`, `dev`, `development`, `qa`, `test`, `uat`, `preprod`.
* Environment words at the start before `-` or `_`, only when no ending matched: `prod`, `staging`, `stage`, `dev`, `qa`, `uat`, `preprod`.
* Role words at the end after `-` or `_`: `service`, `svc`.

So `Payments-Prod`, `prod-payments` and `payments-svc-prod` are one service, `payments`, while `production-line` and `test-runner` are left alone. Only `-` and `_` separate words, so `payments.prod` is a different name.

Anything weaker than that is a suggestion. The service's **Aliases** section lists similar services, computed when you open it, and an organization admin can accept one.

### Which label names a service

An alert names a service by the first non-blank text among its `service`, `app` and `job` labels, in that order, trimmed and compared in lower case. The `namespace` label never names a service. A `service` label that is present wins even when it resolves to nothing: the alert is then not attached to the service its `app` label names.

## Read the list

The list shows confirmed and discovered services under **Active**. The tabs **Confirmed**, **Discovered** and **Ignored** narrow it, and **Search services** matches name and slug. Another organization's services never appear. Each row shows:

* the status, and with the infrastructure suite the owner team;
* **Telemetry**: the signal counts and the binding health (below);
* **Repos**: the repositories linked to it and their role;
* **Open alerts** and **Last deploy**.

### Binding health

A service is listed because something named it, not because anything works. With the infrastructure suite, each row says what SRE Agent found about the data sources and selectors bound to it. The worst state across a service's signals is the one shown.

| Shown | Meaning |
| - | - |
| **No telemetry bound** | No data source is bound to the service yet. |
| **Source failing** | The last call to a data source it reads failed, and nothing has succeeded since. The sanitized error is shown. |
| **Throttled, retrying** | The last call was rate limited. A rate limit is never shown as a failure: SRE Agent retries on its own. |
| **Source disabled** | A data source it reads is switched off. |
| **Stale** | A data source has not answered for over 24 hours, a selector you wrote was last checked more than a day ago, or discovery has not re-confirmed a discovered selector for 3 days. |
| **No data** | The last check of a selector you wrote found a value it names missing from the source, for example a typo in a label value or a metric name. |
| **Not checked** | A data source has not answered a query yet, or a selector you wrote has not been confirmed, or only partly. Each signal says why. |
| **Sources reachable** | The sources answered within 24 hours, and each selector was confirmed: a recent check found data for the whole selector, or discovery re-confirmed it. |

A selector you wrote is checked like a discovered one, on the same hourly probe, with the read that answers the whole of it where one exists. Prometheus, Loki, Datadog and Tempo selectors can be checked. A New Relic metrics selector is only partly checked, and AWS selectors (CloudWatch, X-Ray) and New Relic logs are not checked at all. "Sources reachable" never promises that a particular query returns rows.

## Work with one service

Click a service's name to open its page. The sections are **Overview**, **Telemetry**, **Resources**, **Alerts**, **SLOs**, **Deploys**, **Routing**, **Runbooks and synthetic checks**, **Healing**, **Aliases** and **Repositories**. Each shows what belongs to this service and is empty when nothing does.

| Action | Who can do it | Where |
| - | - | - |
| Read the list and a page | Everyone in the organization, including viewers | **Services** |
| Rename it or edit its description | Member | **Overview**: change **Name** or **Description** and click **Rename** |
| Confirm a discovered service | Member | **Confirm** on the page, or **Confirm selected** on the list |
| Set the owner team | Member, with the infrastructure suite | **Overview**: choose a team and click **Set owner** |
| Add, edit or remove a telemetry signal | Member, with the infrastructure suite | **Telemetry** |
| Link or unlink a repository | Member, with the GitHub integration | **Repositories** |
| Ignore a service, restore an ignored one, merge, accept a suggestion | Organization admin | **Ignore** or **Restore** on the page, **Merge** and **Merge into it** under **Aliases** |

### Ignore and restore

**Ignore** removes a service from every picker, and discovery will not recreate it. An ignored service sits on the **Ignored** tab. **Restore** brings it back, and only an organization admin can restore, so a member can never undo an admin's decision. Confirming an ignored service is refused.

### Merge duplicates

Open the service you want to fold away. Under **Aliases**, choose the service to keep next to **Merge this service into another** and click **Merge**, or click **Merge into it** on a suggestion under **These might be the same service**. Merging cannot be undone, and it is refused when either service is ignored. Restore it first.

* The absorbed name becomes an alias of the kept service, and its other aliases move.
* Every alert, SLI, synthetic check, deploy, deploy policy, alert route, card, repository mapping, self-healing action and telemetry binding that pointed at it points at the kept service.
* Repository links, signals and resources move unless the kept service already has the same one. For two signals on the same data source and kind, one you wrote beats a discovered one.
* The kept service's empty owner team and description are filled from the absorbed one. A value the kept service already has always wins.

### Bind telemetry

<Plan tier="Business" />

A signal says where one service's metrics, logs or traces are in one data source and how to select them. Discovery fills them in by reading label values from your sources. You can also add your own. A signal you wrote is never replaced by a discovered one.

<Steps>
  <Step title="Pick the data source">
    On the service's **Telemetry** section, choose a **Data source** that can hold a signal.
  </Step>

  <Step title="Write the selector">
    In **Selector**, enter a JSON object in the shape that source type reads (below). The source
    type's provider checks the shape when you save.
  </Step>

  <Step title="Save">
    Click **Save signal**. **Edit** and **Remove** are on each signal, and **Open in Explorer**
    opens the Explorer on the service and the signal's tab; for a Prometheus, Loki, Tempo or Datadog
    signal the tab starts from its selector.
  </Step>
</Steps>

| Data source type | Signal | Selector keys |
| - | - | - |
| Prometheus | metrics | `metric` (required), `labels` |
| Loki | logs | `query`, a stream selector with an equality matcher |
| Tempo | traces | `service_name` |
| Datadog Metrics | metrics | `metric`, `tags` |
| Datadog Logs | logs | `query` |
| New Relic Metrics | metrics | `metric`, `attributes` |
| New Relic Logs | logs | `query` |
| CloudWatch Metrics | metrics | `metric_selectors` |
| CloudWatch Logs | logs | `log_groups` |
| X-Ray | traces | `trace_service_names` |

For example, a Prometheus selector is `{"metric": "http_requests_total", "labels": {"job": "checkout"}}`, a Loki selector is `{"query": "{app=\"checkout\"}"}` and a CloudWatch Logs selector is `{"log_groups": ["/ecs/checkout"]}`.

Open a service in the Explorer with `/explore?service=<slug>` and the Prometheus, Loki, Tempo and Datadog tabs start from the selector recorded for the service on the source the tab reads. Nothing runs until you run it.

### Link repositories

<Plan tier="Business" />

A service has **code** repositories and **infra** repositories, each with an optional path prefix for a monorepo. One repository per role is **primary**. On **Repositories**, choose a **Repository**, a **Role** and a **Path prefix (monorepo)**, tick **Primary** if it is the main one, and click **Link**. The list holds only repositories your GitHub App can read. If it is empty, connect the GitHub App on the **Integrations** page first. **Unlink** removes a link.

Discovery links repositories too. It links a code repository when the repository's short name is an alias of the service, when a deploy event for the service names it, or when a repository mapping names the service. The first one becomes primary, and a primary you set is never replaced. It links an infrastructure repository only on evidence, such as a repository named `<alias>-infra` or `<alias>-terraform`. Links found this way are marked as discovered, never confirmed.

Which repository a service has is the service's link. How a fix may write to it, such as the folders it may change, stays on the repository's fix settings. See [Ask for a fix](/guides/fix/fix-requests).

### Self-healing on a service

The **Healing** section shows the service's self-healing switch, tier, channels and pause, what waits for a person with a link to decide in the actions inbox, and its ten most recent actions. An organization without self-healing sees an upgrade card. See [Let SRE Agent act on incidents with self-healing](/guides/fix/self-healing).

## Where services show up

Every alert, SLI, synthetic check, deploy, alert route, deploy policy, card and repository mapping records which service it belongs to, however the name is spelled. As a result:

* A fix request, card or Slack command that names a service (its slug or any alias) goes to that service's primary code repository when your GitHub App can read it.
* The fix brief for an alert names the service the alert belongs to.
* A new card on a service files under the service's owner team unless you choose one.
* Every form with a service field offers your services first, confirmed before discovered, and still accepts any name. The Alerts page filters by service.

## MCP tools

The same actions are available over MCP. MCP needs the Business plan or above, which also includes the infrastructure suite and the GitHub integration, so every tool below is available wherever MCP is.

| Tool | Role needed |
| - | - |
| `list_services`, `get_service` | Viewer |
| `update_service`, `confirm_service` | Member |
| `ignore_service`, `restore_service`, `merge_services` | Org admin |
| `link_service_repo`, `unlink_service_repo` | Member |
| `set_service_signal`, `delete_service_signal` | Member |

`merge_services`, `unlink_service_repo` and `delete_service_signal` need `confirm: true`. `merge_services` takes `id`, the service folded away, and `into_id`, the one kept. `get_active_alerts` and `request_fix` also accept `service_id` in place of `service`. See [MCP tools](/api-reference/mcp) for how to connect a client.

## Related

* [Connect your data](/guides/get-started/connect-your-data): the data sources a signal points at.
* [Triage alerts](/guides/respond/alerts): the alerts a service collects.
* [Manage people and roles](/guides/administer/organizations-and-roles): who is a member and who is an organization admin.
* [Ask for a fix](/guides/fix/fix-requests): how a fix finds a service's repository.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.