Haijun Platform Docs
ID

AI adoption is starting to take a familiar shape: you build on AI to stay at the intelligence frontier, but scaling your product makes token costs unsustainable. The zeitgeist refers to this phenomenon as a shift from "tokenmaxxing" to "budgetmaxxing." Yet, the defensible solution isn't to downgrade your models or shrink the scope of your AI use cases. Instead, we see successful teams focus on optimizing the architecture surrounding their models to reach the Pareto frontier of intelligence and cost.

This cookbook lays out the checklist Juglow's Applied AI team runs builders through to bring their API spend into line without sacrificing product quality. Here's a bird's eye view:

Input token management

– let the model discover context via tools

Agent-loop efficiency

– keep multi-turn context from compounding

Output token management

– hand the model tighter generation constraints

Batch API

– defer non-interactive workloads to an async queue

Model selection and effort

– find the cheapest tier that still clears your quality bar

Notice how model selection sits near the bottom. It's the easiest lever to pull but directly constrains the intelligence of your product. The goal of this cookbook is to maintain your intelligence ceiling while recouping cost where the dollars are not driving any material gains in product performance.

6">

%%capture

%pip install --upgrade juglow matplotlib pillow python-dotenv

import base64

import copy

import csv

import io

import json

import textwrap

import time

from datetime import datetime

from functools import partial

from pathlib import Path

import juglow

import matplotlib.pyplot as plt

from dotenv import load_dotenv

from IPython.display import HTML, display

from PIL import Image

load_dotenv()

client = juglow.Juglow()

With that in place, define the base models and pricing schemas. We'll use Opus as the default workhorse for this notebook, but your use case might warrant a different intelligence ceiling – the baseline eval in the next section will help us identify the right starting point. Here we just load in some helper functions to make our eval results easily legible.

model $/M in $/M out haijun-fable-5 10.00 50.00 haijun-opus-5 5.00 25.00 <- our default haijun-sonnet-5 2.00 10.00 haijun-haiku-4-5 1.00 5.00 # Shared eval output format def results_table(rows, , numbered=False, ref="baseline"): """rows: [(label, trials)], each trial {"correct", "n", "per_task", "turns"}. The first row is the cost baseline""" base = sum(t["per_task"] for t in rows[0][1]) / len(rows[0][1]) head = [ "config", "trials (pass/n)", "mean pass", "requests", "$ / task", "$ / 10k tasks", f"vs {ref}", ] if numbered: head.insert(0, "#") body = [] for i, (label, trials) in enumerate(rows, 1): n = trials[0]["n"] passes = [t["correct"] for t in trials] mean = sum(passes) / (len(passes) n) cost = sum(t["per_task"] for t in trials) / len(trials) turns = sum(t["turns"] for t in trials) / len(trials) delta = "—" if i == 1 else f"{cost / base - 1:+.0%}" vals = ([f"[{i:2d}]"] if numbered else []) + [ label, " · ".join(f"{p}/{n}" for p in passes), f"{mean:.0%}", f"{turns:.0f}", f"${cost:.4f}", f"${cost * 10_000:,.0f}", delta, ] tds = "".join( f'{v}' for j, v in enumerate(vals) ) body.append(f"{tds}") ths = "".join(f"{h}" for h in head) return HTML(f"

{ths}{''.join(body)}
") _TRIAL_COLS = "{cfg:<34}{trial:>6}{passed:>9}{turns:>10}{ptask:>11}{p10k:>15} {misses}" def trial_header(): print( _TRIAL_COLS.format( cfg="config", trial="trial", passed="pass/n", turns="requests", ptask="$ / task", p10k="$ / 10k tasks", misses="misses", ), flush=True, ) def trial_row(label, trial, res, note=""): """res: one trial dict — {"correct", "n", "per_task", "turns", "misses"}""" misses = ", ".join(res.get("misses") or []) or "—" print( _TRIAL_COLS.format( cfg=label, trial=trial, passed=f"{res['correct']}/{res['n']}", turns=res["turns"], ptask=f"${res['per_task']:.4f}", p10k=f"${res['per_task'] * 10_000:,.0f}", misses=misses, ) + (f" ({note})" if note else ""), flush=True, ) def context_header(width=44): print( f"{'':<32s} {'context window':<{width}s} {'cache_w':>10s} {'cache_r':>10s} tokens in context", flush=True, ) def context_bar(label, seen, prev=None, cache_w=None, cache_r=None, , full=180_000, width=44): """Per-turn context size: solid bar for tokens the model saw. prev shows the turn-over-turn delta; cache_w/cache_r log how much of the turn was a cache write vs a cache read.""" solid = min(width, round(seen / full width)) bar = "█" * solid delta = f" ({seen - prev:+,})" if prev is not None else "" cache = f"{cache_w:>10,} {cache_r:>10,} " if cache_w is not None else "" print(f"{label:<32s} {bar:<{width}s} {cache}{seen:>8,}{delta}", flush=True) To anchor this cost optimization exercise in reality, we'll run through a use case that embodies what our Applied AI team might face in the field: a claims adjuster agent for Acme Insurance, a fictional auto-insurance company. The cell below loads synthetic data from assets/, which serves as environment context for our agent to work with. It mocks a real production backend without all of the underlying infra.

ty is set out in Section 6 of the underwriting manual below. "

"Standard investigation on every claim: retrieve the claim record, the policy, "

"the policyholder's claim history, and the SIU fraud-indicator score before "

"deciding; pull the estimate, photos, or additional documents only when the "

"decision turns on them. Adjudicate strictly per the manual: decide claims within "

"your authority yourself, and escalate only where the manual requires it: "

"route_to_supervisor for the supervisor review or higher-authority approval the "

"manual specifies, refer_to_siu for the manual's mandatory SIU referrals. Conclude "

"every adjudication by calling exactly one of approve_claim, deny_claim, "

"route_to_supervisor, or refer_to_siu.\n\n"

"=== UNDERWRITING MANUAL ===\n"

f"{POLICY_MANUAL}\n"

"=== END MANUAL ==="

)

def _tool(name, description, params):

return {

"name": name,

"description": description,

"input_schema": {

"type": "object",

"properties": {k: {"type": "string", "description": v} for k, v in params.items()},

"required": list(params),

},

}

TOOLS = [

_tool("get_claim", "Retrieve the full claim record.", {"claim_id": "Claim ID, e.g. CLM-001"}),

_tool("get_policy", "Retrieve the policy bound to a claim.", {"claim_id": "Claim ID"}),

_tool(

"get_customer_history",

"List the policyholder's prior claims.",

{"policyholder_id": "Policyholder ID"},

),

_tool(

"lookup_repair_estimate",

"Get the itemized repair estimate for a claim.",

{"claim_id": "Claim ID"},

),

_tool(

"check_fraud_signals",

"Run SIU fraud-indicator scoring on a claim.",

{"claim_id": "Claim ID"},

),

_tool(

"get_damage_photos", "Retrieve damage-photo metadata for a claim.", {"claim_id": "Claim ID"}

),

_tool(

"calculate_payout",

"Compute payout after deductible and depreciation.",

{"claim_id": "Claim ID"},

),

_tool(

"request_docs",

"Request additional documentation from the insured.",

{"claim_id": "Claim ID", "docs": "Comma-separated doc types"},

),

_tool(

"approve_claim",

"Record an APPROVE decision.",

{"claim_id": "Claim ID", "amount": "Payout in USD"},

),

_tool(

"deny_claim",

"Record a DENY decision with an exclusion code.",

{"claim_id": "Claim ID", "code": "Exclusion code"},

),

_tool(

"route_to_supervisor",

"Send the claim up for the supervisor review or higher-authority approval the manual "

"requires (fraud-indicator score in the 3-5 band; payment above your authority).",

{"claim_id": "Claim ID", "reason": "Why"},

),

_tool(

"refer_to_siu",

"File the mandatory referral to the fraud unit (SIU, form SIU-REF-1) when the manual requires it: "

"fraud-indicator score of 6 or higher, or any Section 5.2 mandatory trigger.",

{"claim_id": "Claim ID", "reason": "Why"},

),

]

TERMINAL = {

"approve_claim": "APPROVE",

"deny_claim": "DENY",

"route_to_supervisor": "SUPERVISOR",

"refer_to_siu": "FRAUD",

}

n_tokens = client.messages.count_tokens(

model=MODEL,

system=SYSTEM_PROMPT,

tools=TOOLS,

messages=[{"role": "user", "content": "hi"}],

).input_tokens

print(f"--- system prompt ({len(SYSTEM_PROMPT):,} chars) ---\n")

print(SYSTEM_PROMPT[:280] + "...\n")

print(f"--- tools ({len(json.dumps(TOOLS)):,} chars) ---\n")

print(textwrap.fill(", ".join(t["name"] for t in TOOLS), 84))

print(f"\nStatic prefix (system + tools): {n_tokens:,} tokens")

--- system prompt (30,395 chars) --- You are a Senior Adjuster in Acme Insurance's auto claims unit; your payment and denial authority is set out in Section 6 of the underwriting manual below. Standard investigation on every claim: retrieve the claim record, the policy, the policyholder's claim history, and the SIU ... --- tools (3,342 chars) --- get_claim, get_policy, get_customer_history, lookup_repair_estimate, check_fraud_signals, get_damage_photos, calculate_payout, request_docs, approve_claim, deny_claim, route_to_supervisor, refer_to_siu Static prefix (system + tools): 13,027 tokens Before you start optimizing First ensure that the agent actually works, regardless of whether or not the unit economics are scalable. If you can't make an agent that solves your problem, you definitely won't be able to make a cost-optimal one. Run the agent on Opus or Fable to start; you should maintain a high intelligence ceiling and pull on other cost levers before sacrificing the agent's overall reasoning capabilities.

d [overflow-wrap:anywhere]" style="padding-top:12px;padding-inline:12px;padding-bottom:12px;tab-size:4">

tool_search = {"type": "tool_search_tool_regex_20251119", "name": "tool_search_tool_regex"}

always_on = {"get_claim", "approve_claim", "deny_claim", "route_to_supervisor", "refer_to_siu"}

tools_full = TOOLS + [read_manual]

deferred_tools = [tool_search] + [

({**t, "defer_loading": True} if t["name"] not in always_on else t) for t in tools_full

]

count_tokens rejects server tools like tool_search, so read billed input off a 1-token request

lean = client.messages.create(

model=MODEL,

max_tokens=1,

system=SYSTEM_LEAN,

tools=deferred_tools,

messages=[{"role": "user", "content": f"Claim {claim['id']}: {claim['summary']}"}],

).usage

print(

f"{'defer tools via tool_search':30s} {lean.input_tokens:>6,} tokens "

f"(+{clean.input_tokens - lean.input_tokens:,} tokens saved)"

)

defer tools via tool_search 1,696 tokens (+429 tokens saved) Keep in mind that a smaller prefix doesn't always translate to a cheaper cost per task. By deferring context so it's progressively disclosed, the model may incur more tokens trying to find that context than if it had just read it upfront. Prefix cost savings depend on the shape of your agent trajectory and should be validated against your own eval. You can see the impact of these changes on our claims agent in the §Putting it all together section.

ent:-12ch"> fontsize=8,

color="C1",

va="bottom",

linespacing=1.5,

)

ax.scatter([], [], label="dominated", **GREY)

if var:

ax.scatter([], [], label="added lever (dominated)", **DIAMOND)

ax.set(

xscale="log",

xlabel="$ per 10k tasks (log)",

ylabel=f"mean pass (of {N})",

title=title,

xlim=(xlo, xhi),

ylim=(ylo, yhi),

yticks=[t for t in range(0, N + 1, 2) if t >= ylo],

)

ticks = [

t

for t in (50, 100, 200, 500, 1_000, 2_000, 5_000, 10_000, 20_000, 50_000)

if min(xs) * 0.5 <= t <= max(xs) * 1.5

]

ax.set_xticks(ticks, [f"${t / 1000:g}k" if t >= 1000 else f"${t}" for t in ticks])

ax.minorticks_off()

ax.grid(alpha=0.15, which="both")

ax.legend(loc="lower right", frameon=False, fontsize=9, bbox_to_anchor=(1, 0.12))

plt.tight_layout()

plt.show()

pareto_plot(SWEEP)

![Output image](/cookbook/images/notebooks/cost-optimization-cost-optimization/cost-optimization-cost-optimization_cell67_out0_77861878.png