Haijun Platform Docs
ID

Introduction

Most agent workloads have two very different jobs inside them: a small amount of planning and judgment, and a large amount of mechanical reading and doing. Web research is the extreme case, and it's the example this notebook uses: verifying twenty facts against their authoritative sources means pulling hundreds of thousands of tokens of web pages through a model, and at frontier rates that reading bill dominates.

amount, with 84-98% of the team's input tokens billed at the worker rate.

By the end of this cookbook, you'll be able to:

Meter each thread with the typed per-thread cumulative

usage

The same economics apply to any workload where a cheap model can do the token-heavy leg: document review, log analysis, codebase sweeps.

is notebook uses the simplest possible team — one worker type — because the point here is the cost structure, not the team design.

Setup

last radius you want for

that input — and the coordinator reading the reports has no

tools at all.

tools=[

{

"type": "agent_toolset_20260401",

"default_config": {"enabled": False},

"configs": [

{"name": "web_search", "enabled": True},

{"name": "web_fetch", "enabled": True},

],

}

],

system=(

"You are a search worker researching one focused sub-question for "

"a coordinator. Use web_search and web_fetch to find the answer. "

"Be thorough: try multiple query phrasings, follow promising "

"links, and cross-check facts across sources. Report back with "

"the specific answer you found and the evidence (URLs, quotes) "

"that supports it. If you could not find a definitive answer, say "

"exactly what you did find and what remains uncertain. Always "

"finish by calling submit_result."

),

betas=BETAS,

)

coordinator = client.beta.agents.create(

name="search-coordinator",

model=COORDINATOR_MODEL,

multiagent={

"type": "coordinator",

"agents": [{"type": "agent", "id": worker.id}],

},

system=(

"You are coordinating a team of search workers to answer a hard "

"web-research question. Your workers have web_search and "

"web_fetch; you do not. Break the question into focused "

"sub-questions and delegate each to a worker via create_agent. "

"Run several workers in parallel on independent sub-questions, "

"and ALWAYS call wait_for_agents after spawning before drawing "

"any conclusion. When a worker reports, decide whether its "

"findings answer the sub-question or whether to send a follow-up "

"with send_to_agent. If a worker returns an infrastructure error "

"(rate limit, timeout) instead of findings, re-assign the same "

"sub-question to a fresh worker. Once you have enough evidence, "

"synthesize the workers' findings into a single final answer to "

"the original question."

),

betas=BETAS,

)

print(f"worker {worker.id}")

print(f"coordinator {coordinator.id}")

2. Run a research question

;input": 2.0, "output": 10.0},

}

def total_input(u):

cache = u.cache_creation # None on threads with no cache activity

return (

u.input_tokens

+ u.cache_read_input_tokens

+ (cache.ephemeral_5m_input_tokens if cache else 0)

+ (cache.ephemeral_1h_input_tokens if cache else 0)

)

def counterfactual_cost(u, model):

"""Re-price a thread's tokens at another model's rates (the what-if)."""

p = PRICES[model]

cache = u.cache_creation

return (

u.input_tokens * p["input"]

+ (cache.ephemeral_5m_input_tokens if cache else 0) * p["input"] * 1.25

+ (cache.ephemeral_1h_input_tokens if cache else 0) * p["input"] * 2.0

+ u.cache_read_input_tokens * p["input"] * 0.1

+ u.output_tokens * p["output"]

) / 1e6

def dollars(list_cost):

return float(list_cost.amount) / 100 # amount is an integer string in cents

def report(session_id):

session_usage = client.beta.sessions.retrieve(session_id, betas=BETAS).usage

threads = list(client.beta.sessions.threads.list(session_id, betas=BETAS))

primary = next(t for t in threads if t.parent_thread_id is None)

workers = [t for t in threads if t.parent_thread_id is not None]

workers_in = sum(total_input(t.usage) for t in workers)

print(

f" primary thread ({primary.agent.model.id}): "

f"{total_input(primary.usage):>9,} in / {primary.usage.output_tokens:>6,} out"

f" -> ${dollars(primary.usage.list_cost):.2f}"

)

if workers:

print(

f" {len(workers)} worker(s): {workers_in:>9,} in / "

f"{sum(t.usage.output_tokens for t in workers):>6,} out"

f" -> ${sum(dollars(t.usage.list_cost) for t in workers):.2f}"

)

print(

f" workers' share of input: {workers_in / (workers_in + total_input(primary.usage)):.0%}"

)

total = dollars(session_usage.list_cost)

print(f" total cost (session usage.list_cost): ${total:.2f}")

return total, threads

print("split team (fable coordinator + sonnet workers):")

split_cost, split_threads = report(session.id)

print("\nsolo frontier agent:")

solo_cost, _ = report(solo_session.id)

The counterfactual that isolates the rate split: this run's team

workload with every token billed at the frontier rate.

frontier_team_cost = sum(counterfactual_cost(t.usage, COORDINATOR_MODEL) for t in split_threads)

print(f"\nsolo / split cost ratio on this pair of runs: {solo_cost / split_cost:.1f}x")

print(f"the split team's workload at all-frontier rates: ${frontier_team_cost:.2f}")

The two runs did nearly the same reading — that's the point of matching the verification standard. What differs is the rate the reading billed at and the shape of the work: the team's twenty lookups ran as parallel worker threads at the cheap rate, while the solo agent ground through them serially in one frontier-priced context. On the authors' runs 84-98% of the team's input tokens billed at the worker rate. Token volumes vary run to run, so treat any single printed ratio as one sample; the structure is the stable part.

On this page
Setup2. Run a research question