Building your own DIY agent for incident resolution?

Building your own DIY agent for incident resolution?

Building your own DIY agent for incident resolution?

Claude Code for SRE

Claude Code or an AI SRE agent: build or buy?

Hyground is an AI SRE agent that runs inside your own cluster. It starts investigating when an alert fires, uses only the access your admin grants, keeps a record of every tool call and draws on your runbooks. Claude Code is a coding agent that runs on an engineer's machine, with everything that machine can reach.

Claude Code for SRE

Claude Code or an AI SRE agent: build or buy?

Hyground is an AI SRE agent that runs inside your own cluster. It starts investigating when an alert fires, uses only the access your admin grants, keeps a record of every tool call and draws on your runbooks. Claude Code is a coding agent that runs on an engineer's machine, with everything that machine can reach.

Claude Code for SRE

Claude Code or an AI SRE agent: build or buy?

Hyground is an AI SRE agent that runs inside your own cluster. It starts investigating when an alert fires, uses only the access your admin grants, keeps a record of every tool call and draws on your runbooks. Claude Code is a coding agent that runs on an engineer's machine, with everything that machine can reach.

Side by side

Hyground vs Claude Code: an AI SRE agent and a coding agent

Hyground is built to be on call. Claude Code is built to write code with an engineer at the keyboard.

Claude Code details from Anthropic's documentation, read 25 September 2026. Research preview and beta labels are Anthropic's own.
What on-call needsHygroundClaude Code
What it can reachOnly the systems your admin connects, with the access your admin sets, never an engineer's personal accounts. Credentials stay in the adapters and never enter the agent's sandbox.Its shell commands run with the engineer's own access. Even inside its optional sandbox, the default lets commands read credential files such as ~/.aws/credentials and ~/.ssh, and they inherit the tokens in the shell. Permission rules and, in auto mode, a classifier decide what runs.
Starts an investigationOn its own when an alert webhook arrives, on a schedule, or from Slack, Teams, email or the web app.When an engineer prompts it, or from a script or CI job. Routines accept an API call from your monitoring tool, as a research preview.
Where it runsInside your own Kubernetes cluster, installed with a Helm chart: EKS, AKS, GKE, OpenShift, K3s or on-premises.A laptop, a CI runner, or a cloud session on Anthropic's infrastructure. Cloud sessions on your own runners are in beta.
What reaches the model providerPrompts and tool results go only to the model you choose, and Hyground never receives them. With a self-hosted model, no operational data leaves your own infrastructure.Prompts and tool output go to Anthropic, Amazon Bedrock, Google Vertex AI or Microsoft Foundry. Anthropic's own API offers global or US-only inference.
Changes to productionRead-only by default. Nothing in your infrastructure changes unless you enable it.Set by permission rules. Auto mode blocks production deploys by default, and Anthropic says it doesn't guarantee safety.
Knowledge of your systemsRunbooks and docs from Confluence, Git and file uploads, team skills, and recall of past sessions.CLAUDE.md files, skills and whatever the engineer brings to the session.
Record of what happenedEvery message, tool call and result is stored per session and can be exported. The platform, not the agent, writes the audit log.Local transcripts per user, cloud session history, OpenTelemetry export, and a Compliance API on Enterprise.
Who can use itThe whole team, through Slack, Teams, email or the web app, behind your single sign-on.Engineers with a Claude seat or API access and a configured client.
Which modelsAny provider through LiteLLM, Claude included, or a model you host yourself.Claude only. Anthropic doesn't support routing Claude Code to other models.
Pricing modelPriced on infrastructure size, not seats. Model tokens are paid to your provider.Per seat, or per token.

The security risk

Claude Code runs with your laptop’s access to production

On-call means reading logs, alerts and tickets that other people and systems write, on a machine that holds production credentials.

Why Hyground

What an AI SRE agent does that a coding agent doesn’t

The parts that on-call needs, already built and running in your cluster.

Build vs buy

What a Claude-based SRE agent costs to build and run

Tokens show up on the invoice, but most of the cost of a self-built agent is the engineering around it.

CostHygroundBuild on Claude
Model tokensPer token to the model provider you choose, on your own contract.Per token to Anthropic or your cloud provider. On SREGym, Claude Code used 1.81 times the tokens of a purpose-built SRE agent on the same model.
Building itA pilot in your cluster. Our engineers build any other connector that you need with you during onboarding.Our estimate: 5 to 7 senior engineers for about a year before it's ready for production.
Keeping it runningHyground ships the releases and maintains the integrations.A permanent team for model upgrades, MCP and API changes, and new integrations.
Checking its answersWe benchmark open-weight models continuously and validate each model provider.A test set of your past incidents, run again whenever the model or the prompts change.
Security reviewOne platform, read-only by default and ISO/IEC 27001:2022 certified, with a jailbreak guard covering six attack categories, a separate alert-injection judge, and secrets redacted before the model sees them.A new system with read access to production, reviewed from scratch. Its inputs, such as logs, alerts and tickets, are untrusted.
Pricing modelPriced on infrastructure size, not seats. Quote on request.Tokens, hosting and engineering salaries.

The build path

What Anthropic’s SRE agent cookbooks build, and what production adds

Anthropic publishes two SRE agent recipes, one on the Claude Agent SDK and one on Claude Managed Agents. Both are good starting points, and both say what they leave to you.

From Anthropic's Claude Cookbook and Managed Agents documentation, read 25 September 2026.
CookbookThe site reliability agentSRE incident responder
Built onClaude Agent SDKClaude Managed Agents, in beta
Published16 February 202610 April 2026
What it doesInvestigates an incident, finds the cause, applies a fix and writes it up, on its own.Takes a page, reads the logs and opens a pull request that cites the runbook it followed.
Test systemA local Docker stack with one injected fault: a database connection pool cut from 20 connections to 1. Alerts and deploys are simulated.PagerDuty, GitHub and Datadog are mocked with local files, and the infrastructure repo is a single manifest.
GuardrailsShell commands limited to docker, validation hooks, and a separate step so a person can review the diagnosis before any fix.A skill that forbids patching live resources and routes every fix through a pull request.
Left to youMCP servers for Kubernetes, Slack, Confluence and your observability stack, plus the knowledge of your infrastructure that the agent doesn't have.Swapping the mocks for GitHub MCP, a Slack approval button and live logs. Services reachable only inside your network need custom tools, and MCP tunnels to them are a research preview.

What we'd add before calling it ready for on-call

The Managed Agents cookbook says that after those three swaps "it's ready for on-call". We'd add four things first: a service identity with scoped access instead of API keys in the environment, checks on inbound alerts for prompt injection, retrieval over your runbooks and past incidents, and tests against your own past incidents that run again whenever the model changes.

The decision

When to run Hyground, and when to build on Claude

A note on Claude Code

What Claude Code does better

For work at the keyboard, Claude Code is the better tool. It writes the fix as a pull request or a config change for review, and it connects to Datadog, Sentry, PagerDuty and Grafana through their official MCP servers. On SREGym, a July 2026 benchmark from UIUC and the University of Toronto, it had the highest end-to-end success rate of the agents tested, ahead of Codex and a purpose-built SRE agent on the same model. Anthropic’s routines, still a research preview, let a monitoring tool start it.

FAQ

Claude Code and AI SRE agents: common questions

What's the difference between an AI SRE agent and Claude Code?

An AI SRE agent like Hyground is built to be on call: it starts from an alert, works with the access your admin sets, keeps a record of every step and draws on your runbooks. Claude Code is a coding agent that investigates when an engineer prompts it, with that engineer's access, and it's at its best writing the fix.

Can I use Claude Code as an SRE agent?

With an engineer beside it, yes. To leave it on call, you'd build the alert triggers, a service identity, an audit trail and knowledge of your systems yourself. Anthropic's routines, which let a monitoring tool start a session, are still a research preview. It connects to Datadog, Sentry, PagerDuty and Grafana over MCP.

Is Claude Code a security risk in production?

It can be. Claude Code runs with what the engineer's machine can reach, and even its optional sandbox lets commands read credential files such as ~/.aws/credentials and ~/.ssh by default. Anthropic provides deny rules, managed settings and a classifier that blocks production deploys by default, but it says auto mode doesn't guarantee safety and advises against piping untrusted content to Claude. Logs, alerts and tickets are exactly that kind of content.

What does Anthropic's site reliability agent cookbook build, and is it enough for production?

It builds an Agent SDK agent that investigates a local Docker stack, finds one injected fault and fixes it. Anthropic writes that production would need MCP servers for Kubernetes, Slack, Confluence and your observability stack, plus knowledge of your infrastructure that the agent doesn't have. It works as a template for your own build.

Claude Agent SDK or Claude Code: which should an SRE agent be built on?

The Agent SDK is the Claude Code agent loop as a Python or TypeScript library, so it's the one to build a service on. Claude Code is the interactive tool. Either way, you host long-lived agent processes and build the triggers, access control and records yourself.

Can Hyground use Claude models?

Yes. Hyground reaches its model through LiteLLM, so it can run on Claude through Anthropic, Amazon Bedrock or Google Vertex AI, or on OpenAI, Gemini, Azure OpenAI or a model you host yourself. Full feature depth needs a state-of-the-art model.

Does Hyground work with Claude Code?

Yes. Hyground has an optional MCP server that Claude Code, Cursor, GitHub Copilot and Kiro can connect to. Engineers can start and read Hyground investigations from their coding agent, and each session stays attributed to their single sign-on login.

Can Claude Code run air-gapped or on-premises?

The CLI runs on your own machines, but it always needs a Claude endpoint: Anthropic's API, Amazon Bedrock, Google Vertex AI or Microsoft Foundry. Anthropic doesn't support other models and documents no offline mode. Hyground runs inside your cluster and can run fully air-gapped with a self-hosted model.

How much does it cost to build an AI SRE agent on Claude?

Tokens are the smaller part of the cost. Our estimate is 5 to 7 senior engineers for about a year to reach production, then a permanent team to keep up with model, MCP and API changes. Hyground is priced on infrastructure size, not seats, and you pay model tokens to your provider directly.

Does Anthropic use Claude for SRE in production?

Anthropic's engineers use Claude Code during incidents and report real wins, such as answering control-flow questions during incidents three times as fast. An Anthropic reliability engineer told QCon London in March 2026 that Claude doesn't run their incident response, because it confuses correlation with causation.

Further reading

More on building versus buying an SRE agent

See Hyground in action