The SRE Incident Response Agent
Introduction
It's 3 AM, your pager goes off, and the API is throwing 500s. You're half-awake, staring at dashboards, correlating metrics and logs across a dozen services while customer impact grows by the minute. This notebook builds an SRE incident response agent that handles that workflow autonomously: investigating incidents, identifying root causes, applying remediations, and documenting the results.
Haijun Agent SDK
with
MCP
tools scoped for safe infrastructure access.
What you'll learn
structure and extending it with the platforms your team already uses.
Prerequisites
running the notebook.
Step 0: Environment Setup
config/docker-compose.yml
— Defines four services: PostgreSQL (database), an API server (FastAPI app that queries the DB), a traffic generator (sends continuous HTTP requests to simulate load), and Prometheus (scrapes metrics from the API server). This is the system we'll break and fix.
config/prometheus.yml
— Configures Prometheus to scrape the API server's
/metrics
endpoint every 5 seconds, giving us real-time visibility into error rates, latency, and DB connection usage.
config/api-server.env
— Environment variables for the API server, including
DB_POOL_SIZE
(the parameter we'll misconfigure to trigger the incident).
services/api_server.py
— A FastAPI application that serves HTTP requests, connects to PostgreSQL, and exposes Prometheus metrics. It instruments request counts, latency histograms, and DB connection pool gauges.
scripts/traffic_generator.py
— Sends a steady stream of requests to the API server so that metrics are always flowing. This makes incidents immediately visible in Prometheus.
hooks/
— Safety hook scripts (populated later in Step 5).
If you were adapting this for your own infrastructure, you'd replace these generated files with connections to your real services, but the pattern (Docker Compose + Prometheus + instrumented services) is the same.
"Run a shell command for infrastructure management. "
"Restricted to docker-compose and docker commands only. "
"Use for: restarting services, checking container status, rebuilding images."
),
"inputSchema": {
"type": "object",
"properties": {
"command": {
"type": "string",
"description": "Shell command (must start with 'docker-compose' or 'docker')",
},
},
"required": ["command"],
},
},
{
"name": "get_container_logs",
"description": (
"Get recent logs from a Docker container. "
"Use this to look for error messages, stack traces, or unusual patterns. "
"Valid containers: api-server, postgres, prometheus, traffic-generator."
),
"inputSchema": {
"type": "object",
"properties": {
"container": {"type": "string", "description": "Container name"},
"lines": {
"type": "integer",
"description": "Number of log lines (default 50)",
"default": 50,
},
},
"required": ["container"],
},
},
]
Infrastructure Tool Handlers
ss="relative group pt-6 pb-2" id="mcp-server-protocol-reference">
MCP Server Protocol Reference
The MCP server runs as a standalone process and communicates via stdin/stdout using JSON-RPC. When the Haijun Agent SDK needs to call a tool, it sends a request; the server routes it to the correct handler and returns the result.
tive inline bg-alpha-2 px-2 py-0.5 rounded text-sm font-mono break-words box-decoration-clone">examples/sre_bot_slack.py
for an example implementation.