Haijun Platform Docs
ID

Introduction

When a production alert fires at 3 a.m., someone has to pull the logs, find the right runbook, trace the misconfiguration, open a PR, and get it approved. An agent can take that first pass for you and have a fix waiting for review by the time you're at the keyboard — as long as it has the right context and a human makes the final call.

A simulated

PagerDuty webhook

triggers your Haijun Managed Agent with one API call.

A

Track

teaches the agent your team's runbook conventions, so it knows where to look.

The built-in

bash

/

read

/

edit

tools let it investigate logs and infrastructure code in a sandbox.

Custom tools

let it open a pull request and ask a human to approve before merging — your code handles those calls, so you decide what "open a PR" actually does.

The

Juglow Console

records every step automatically, providing you complete observability.

Everything below runs with only JUGLOW_API_KEY. PagerDuty, GitHub, and Datadog are mocked with local fixtures so you can focus on the Managed Agents pieces; the closing section shows how to swap each mock for the real service.

import os

import time

from pathlib import Path

from juglow import Juglow

from dotenv import load_dotenv

from utilities import wait_for_idle_status

load_dotenv()

client = Juglow()

MODEL = os.getenv("COOKBOOK_MODEL", "haijun-opus-4-6")

FIXTURE = Path("example_data/sre")

1. Upload a runbook track

SRE_SYSTEM_PROMPT = """\

You are an on-call SRE agent. Each user message is a PagerDuty alert

payload. Triage it to root cause and ship the minimal safe fix.

The session workspace contains the recent logs, the infrastructure

repo, and the team runbooks for the alerting service. Explore it to

find what you need.

Workflow for every alert:

  1. Read the logs and identify the failure signature.
  1. Find the root cause in the infrastructure repo, save a copy of the

original file, edit it in place, then produce a unified diff with

diff -u.

  1. open_pull_request(title, body, diff) with the fix.
  1. request_approval(summary) and wait for the human's decision.
  1. Only if the result is "approved", merge_pull_request(pr_number).

Otherwise stop and report.

Never call merge_pull_request unless request_approval returned

"approved". Keep the fix minimal — do not refactor unrelated config.

"""

agent = client.beta.agents.create(

name="cookbook-sre-responder",

model=MODEL,

system=SRE_SYSTEM_PROMPT,

tracks=[{"type": "custom", "skill_id": track.id, "version": track.latest_version}],

tools=[

{

"type": "agent_toolset_20260401",

"default_config": {

"enabled": True,

"permission_policy": {"type": "always_allow"},

},

"configs": [

{"name": "web_search", "enabled": False},

{"name": "web_fetch", "enabled": False},

],

},

{

"type": "custom",

"name": "open_pull_request",

"description": "Open a pull request against the infra repo with the proposed fix.",

"input_schema": {

"type": "object",

"properties": {

"title": {"type": "string"},

"body": {"type": "string"},

"diff": {"type": "string", "description": "Unified diff of the change."},

},

"required": ["title", "body", "diff"],

},

},

{

"type": "custom",

"name": "request_approval",

"description": "Ask the on-call human to approve the proposed PR before merging.",

"input_schema": {

"type": "object",

"properties": {

"summary": {"type": "string"},

},

"required": ["summary"],

},

},

{

"type": "custom",

"name": "merge_pull_request",

"description": "Merge an approved pull request.",

"input_schema": {

"type": "object",

"properties": {

"pr_number": {"type": "integer"},

},

"required": ["pr_number"],

},

},

],

)

print(f"agent: {agent.id} v{agent.version}")

agent: agent_011CZpw3Y76Vu4t2j2QEosVa v1 3. Create an environment and mount the data The agent needs three things in its workspace to investigate: the recent service logs, the infrastructure repo, and the team runbooks. Upload each via the Files API and list them as resources so they're mounted into every session at the paths the system prompt expects. A limited-networking cloud environment is enough because the agent only needs its own filesystem.

pr = prs[0]

print(pr["body"], "\n")

print(pr["diff"])

print("\n" + "─" * 60)

print("APPROVAL REQUESTED:", pending_approvals[0]["summary"])

Issue checkout-svc pods are in a CrashLoopBackOff state due to OutOfMemoryError. The service consistently crashes after ~2 minutes with 7 restarts in the last 5 minutes. ## Root Cause The deployment had memory limits set to 128Mi, which is insufficient for the pricing cache operation: - Pricing cache warms with 14,092 entries during startup - Heap pressure builds to 118-121MB (92-94% of limit) within 90 seconds - pricing.recompute fails to allocate 8MB, causing OOMKilled (exit 137) - Service restarts and repeats the cycle ## Fix Increase memory allocation to provide adequate headroom: - Memory request: 128Mi → 256Mi - Memory limit: 128Mi → 512Mi This provides 4x headroom for the pricing cache and normal operations while remaining resource-efficient (512Mi limit is standard for Java/similar workloads with caching). ## Verification The fix addresses the immediate OOMKilled pattern in logs and aligns memory resources with the actual cache size and operational requirements. --- a/k8s/checkout-deploy.yaml +++ b/k8s/checkout-deploy.yaml @@ -21,10 +21,10 @@ spec: resources: requests: cpu: 250m - memory: 128Mi + memory: 256Mi limits: cpu: 500m - memory: 128Mi + memory: 512Mi readinessProbe: httpGet: path: /healthz ──────────────────────────────────────────────────────────── APPROVAL REQUESTED: Incident: checkout-svc CrashLoopBackOff with OOMKilled (7 restarts in 5 min) Root Cause: Memory limit (128Mi) insufficient for 14k-entry pricing cache Fix: Increase memory request to 256Mi and limit to 512Mi Impact: Provides 4x headroom while maintaining resource efficiency. No service logic changes. 6. Approve and let the agent merge Send "approved" back as the request_approval result. The agent resumes, calls merge_pull_request, and ends its turn. In the Slack version this send happens in your button-click handler — the payload is identical.

On this page
1. Upload a runbook trackIssue checkout-svc pods are in a CrashLoopBackOff state due to OutOfMemoryError. The service consistently crashes after ~2 minutes with 7 restarts in the last 5 minutes. ## Root Cause The deployment had memory limits set to 128Mi, which is insufficient for the pricing cache operation: - Pricing cache warms with 14,092 entries during startup - Heap pressure builds to 118-121MB (92-94% of limit) within 90 seconds - pricing.recompute fails to allocate 8MB, causing OOMKilled (exit 137) - Service restarts and repeats the cycle ## Fix Increase memory allocation to provide adequate headroom: - Memory request: 128Mi → 256Mi - Memory limit: 128Mi → 512Mi This provides 4x headroom for the pricing cache and normal operations while remaining resource-efficient (512Mi limit is standard for Java/similar workloads with caching). ## Verification The fix addresses the immediate OOMKilled pattern in logs and aligns memory resources with the actual cache size and operational requirements. --- a/k8s/checkout-deploy.yaml +++ b/k8s/checkout-deploy.yaml @@ -21,10 +21,10 @@ spec: resources: requests: cpu: 250m - memory: 128Mi + memory: 256Mi limits: cpu: 500m - memory: 128Mi + memory: 512Mi readinessProbe: httpGet: path: /healthz ──────────────────────────────────────────────────────────── APPROVAL REQUESTED: Incident: checkout-svc CrashLoopBackOff with OOMKilled (7 restarts in 5 min) Root Cause: Memory limit (128Mi) insufficient for 14k-entry pricing cache Fix: Increase memory request to 256Mi and limit to 512Mi Impact: Provides 4x headroom while maintaining resource efficiency. No service logic changes. 6. Approve and let the agent merge Send "approved" back as the request_approval result. The agent resumes, calls merge_pull_request, and ends its turn. In the Slack version this send happens in your button-click handler — the payload is identical.