Skip to main content
A service is something you run: a checkout API, a worker, a Lambda function. SRE Agent finds them from your alert labels, your cloud and cluster inventory, deploy events, SLIs and the repositories your GitHub App can read. Services lists them, and a service’s own page shows everything SRE Agent holds about it. The list, the pages and discovery are on every plan. Telemetry signals, binding health and owner teams belong to the infrastructure suite, and linking repositories needs the GitHub integration. Both are on Business and above. A locked section shows an upgrade card. See Plan matrix.
1

Open Services

Click Services in the Prevent section of the sidebar. Services appear once alerts, inventory, deploys or repositories name them.
2

Confirm the ones that are real

A service starts as discovered. Open it and click Confirm, or tick several rows and click Confirm selected. Confirmed services are offered first in every service picker.
3

Fix the duplicates

If the same service is listed twice under different names, an organization admin merges one into the other from its page. See below.

How services are found

Discovery runs for every organization once an hour and reads only what you have already connected. Each name it meets is a candidate from one of these places:
  • alert labels of the last 30 days;
  • ECS services, Lambda functions and Kubernetes workloads in the hourly inventory;
  • deploy events, SLIs, synthetic checks and telemetry bindings;
  • the repositories your GitHub installation can read.
There is no button that creates a service by hand. A service exists because something named it. A candidate is resolved in this order, and discovery never merges on a similarity score:
  1. A name that is already an alias of a service marks that service as seen. An ignored service is left as it is.
  2. A name whose canonical key equals the canonical key of an existing service or alias is added as an alias of that service. These are the only merges discovery makes without a person. An ignored service absorbs the new spelling as an alias and stays ignored.
  3. Otherwise a discovered service is created. Its slug is the canonical key and its name is the first spelling seen.
The canonical key is the name in lower case, with a Kubernetes pod suffix removed, then at most one environment word and at most one role word removed:
  • Environment words at the end after - or _: prod, production, staging, stage, dev, development, qa, test, uat, preprod.
  • Environment words at the start before - or _, only when no ending matched: prod, staging, stage, dev, qa, uat, preprod.
  • Role words at the end after - or _: service, svc.
So Payments-Prod, prod-payments and payments-svc-prod are one service, payments, while production-line and test-runner are left alone. Only - and _ separate words, so payments.prod is a different name. Anything weaker than that is a suggestion. The service’s Aliases section lists similar services, computed when you open it, and an organization admin can accept one.

Which label names a service

An alert names a service by the first non-blank text among its service, app and job labels, in that order, trimmed and compared in lower case. The namespace label never names a service. A service label that is present wins even when it resolves to nothing: the alert is then not attached to the service its app label names.

Read the list

The list shows confirmed and discovered services under Active. The tabs Confirmed, Discovered and Ignored narrow it, and Search services matches name and slug. Another organization’s services never appear. Each row shows:
  • the status, and with the infrastructure suite the owner team;
  • Telemetry: the signal counts and the binding health (below);
  • Repos: the repositories linked to it and their role;
  • Open alerts and Last deploy.

Binding health

A service is listed because something named it, not because anything works. With the infrastructure suite, each row says what SRE Agent found about the data sources and selectors bound to it. The worst state across a service’s signals is the one shown. A selector you wrote is checked like a discovered one, on the same hourly probe, with the read that answers the whole of it where one exists. Prometheus, Loki, Datadog and Tempo selectors can be checked. A New Relic metrics selector is only partly checked, and AWS selectors (CloudWatch, X-Ray) and New Relic logs are not checked at all. “Sources reachable” never promises that a particular query returns rows.

Work with one service

Click a service’s name to open its page. The sections are Overview, Telemetry, Resources, Alerts, SLOs, Deploys, Routing, Runbooks and synthetic checks, Healing, Aliases and Repositories. Each shows what belongs to this service and is empty when nothing does.

Ignore and restore

Ignore removes a service from every picker, and discovery will not recreate it. An ignored service sits on the Ignored tab. Restore brings it back, and only an organization admin can restore, so a member can never undo an admin’s decision. Confirming an ignored service is refused.

Merge duplicates

Open the service you want to fold away. Under Aliases, choose the service to keep next to Merge this service into another and click Merge, or click Merge into it on a suggestion under These might be the same service. Merging cannot be undone, and it is refused when either service is ignored. Restore it first.
  • The absorbed name becomes an alias of the kept service, and its other aliases move.
  • Every alert, SLI, synthetic check, deploy, deploy policy, alert route, card, repository mapping, self-healing action and telemetry binding that pointed at it points at the kept service.
  • Repository links, signals and resources move unless the kept service already has the same one. For two signals on the same data source and kind, one you wrote beats a discovered one.
  • The kept service’s empty owner team and description are filled from the absorbed one. A value the kept service already has always wins.

Bind telemetry

A signal says where one service’s metrics, logs or traces are in one data source and how to select them. Discovery fills them in by reading label values from your sources. You can also add your own. A signal you wrote is never replaced by a discovered one.
1

Pick the data source

On the service’s Telemetry section, choose a Data source that can hold a signal.
2

Write the selector

In Selector, enter a JSON object in the shape that source type reads (below). The source type’s provider checks the shape when you save.
3

Save

Click Save signal. Edit and Remove are on each signal, and Open in Explorer opens the Explorer on the service and the signal’s tab; for a Prometheus, Loki, Tempo or Datadog signal the tab starts from its selector.
For example, a Prometheus selector is {"metric": "http_requests_total", "labels": {"job": "checkout"}}, a Loki selector is {"query": "{app=\"checkout\"}"} and a CloudWatch Logs selector is {"log_groups": ["/ecs/checkout"]}. Open a service in the Explorer with /explore?service=<slug> and the Prometheus, Loki, Tempo and Datadog tabs start from the selector recorded for the service on the source the tab reads. Nothing runs until you run it. A service has code repositories and infra repositories, each with an optional path prefix for a monorepo. One repository per role is primary. On Repositories, choose a Repository, a Role and a Path prefix (monorepo), tick Primary if it is the main one, and click Link. The list holds only repositories your GitHub App can read. If it is empty, connect the GitHub App on the Integrations page first. Unlink removes a link. Discovery links repositories too. It links a code repository when the repository’s short name is an alias of the service, when a deploy event for the service names it, or when a repository mapping names the service. The first one becomes primary, and a primary you set is never replaced. It links an infrastructure repository only on evidence, such as a repository named <alias>-infra or <alias>-terraform. Links found this way are marked as discovered, never confirmed. Which repository a service has is the service’s link. How a fix may write to it, such as the folders it may change, stays on the repository’s fix settings. See Ask for a fix.

Self-healing on a service

The Healing section shows the service’s self-healing switch, tier, channels and pause, what waits for a person with a link to decide in the actions inbox, and its ten most recent actions. An organization without self-healing sees an upgrade card. See Let SRE Agent act on incidents with self-healing.

Where services show up

Every alert, SLI, synthetic check, deploy, alert route, deploy policy, card and repository mapping records which service it belongs to, however the name is spelled. As a result:
  • A fix request, card or Slack command that names a service (its slug or any alias) goes to that service’s primary code repository when your GitHub App can read it.
  • The fix brief for an alert names the service the alert belongs to.
  • A new card on a service files under the service’s owner team unless you choose one.
  • Every form with a service field offers your services first, confirmed before discovered, and still accepts any name. The Alerts page filters by service.

MCP tools

The same actions are available over MCP. MCP needs the Business plan or above, which also includes the infrastructure suite and the GitHub integration, so every tool below is available wherever MCP is. merge_services, unlink_service_repo and delete_service_signal need confirm: true. merge_services takes id, the service folded away, and into_id, the one kept. get_active_alerts and request_fix also accept service_id in place of service. See MCP tools for how to connect a client.