Introduction
Most agent workloads have two very different jobs inside them: a small amount of planning and judgment, and a large amount of mechanical reading and doing. Web research is the extreme case, and it's the example this notebook uses: verifying twenty facts against their authoritative sources means pulling hundreds of thousands of tokens of web pages through a model, and at frontier rates that reading bill dominates.
amount, with 84-98% of the team's input tokens billed at the worker rate.
By the end of this cookbook, you'll be able to:
Meter each thread with the typed per-thread cumulative
usage
The same economics apply to any workload where a cheap model can do the token-heavy leg: document review, log analysis, codebase sweeps.
is notebook uses the simplest possible team — one worker type — because the point here is the cost structure, not the team design.
Setup
last radius you want for
that input — and the coordinator reading the reports has no
tools at all.
tools=[
{
"type": "agent_toolset_20260401",
"default_config": {"enabled": False},
"configs": [
{"name": "web_search", "enabled": True},
{"name": "web_fetch", "enabled": True},
],
}
],
system=(
"You are a search worker researching one focused sub-question for "
"a coordinator. Use web_search and web_fetch to find the answer. "
"Be thorough: try multiple query phrasings, follow promising "
"links, and cross-check facts across sources. Report back with "
"the specific answer you found and the evidence (URLs, quotes) "
"that supports it. If you could not find a definitive answer, say "
"exactly what you did find and what remains uncertain. Always "
"finish by calling submit_result."
),
betas=BETAS,
)
coordinator = client.beta.agents.create(
name="search-coordinator",
model=COORDINATOR_MODEL,
multiagent={
"type": "coordinator",
"agents": [{"type": "agent", "id": worker.id}],
},
system=(
"You are coordinating a team of search workers to answer a hard "
"web-research question. Your workers have web_search and "
"web_fetch; you do not. Break the question into focused "
"sub-questions and delegate each to a worker via create_agent. "
"Run several workers in parallel on independent sub-questions, "
"and ALWAYS call wait_for_agents after spawning before drawing "
"any conclusion. When a worker reports, decide whether its "
"findings answer the sub-question or whether to send a follow-up "
"with send_to_agent. If a worker returns an infrastructure error "
"(rate limit, timeout) instead of findings, re-assign the same "
"sub-question to a fresh worker. Once you have enough evidence, "
"synthesize the workers' findings into a single final answer to "
"the original question."
),
betas=BETAS,
)
print(f"worker {worker.id}")
print(f"coordinator {coordinator.id}")
2. Run a research question
;input": 2.0, "output": 10.0},
}
def total_input(u):
cache = u.cache_creation # None on threads with no cache activity
return (
u.input_tokens
+ u.cache_read_input_tokens
+ (cache.ephemeral_5m_input_tokens if cache else 0)
+ (cache.ephemeral_1h_input_tokens if cache else 0)
)
def counterfactual_cost(u, model):
"""Re-price a thread's tokens at another model's rates (the what-if)."""
p = PRICES[model]
cache = u.cache_creation
return (
u.input_tokens * p["input"]
+ (cache.ephemeral_5m_input_tokens if cache else 0) * p["input"] * 1.25
+ (cache.ephemeral_1h_input_tokens if cache else 0) * p["input"] * 2.0
+ u.cache_read_input_tokens * p["input"] * 0.1
+ u.output_tokens * p["output"]
) / 1e6
def dollars(list_cost):
return float(list_cost.amount) / 100 # amount is an integer string in cents
def report(session_id):
session_usage = client.beta.sessions.retrieve(session_id, betas=BETAS).usage
threads = list(client.beta.sessions.threads.list(session_id, betas=BETAS))
primary = next(t for t in threads if t.parent_thread_id is None)
workers = [t for t in threads if t.parent_thread_id is not None]
workers_in = sum(total_input(t.usage) for t in workers)
print(
f" primary thread ({primary.agent.model.id}): "
f"{total_input(primary.usage):>9,} in / {primary.usage.output_tokens:>6,} out"
f" -> ${dollars(primary.usage.list_cost):.2f}"
)
if workers:
print(
f" {len(workers)} worker(s): {workers_in:>9,} in / "
f"{sum(t.usage.output_tokens for t in workers):>6,} out"
f" -> ${sum(dollars(t.usage.list_cost) for t in workers):.2f}"
)
print(
f" workers' share of input: {workers_in / (workers_in + total_input(primary.usage)):.0%}"
)
total = dollars(session_usage.list_cost)
print(f" total cost (session usage.list_cost): ${total:.2f}")
return total, threads
print("split team (fable coordinator + sonnet workers):")
split_cost, split_threads = report(session.id)
print("\nsolo frontier agent:")
solo_cost, _ = report(solo_session.id)
The counterfactual that isolates the rate split: this run's team
workload with every token billed at the frontier rate.
frontier_team_cost = sum(counterfactual_cost(t.usage, COORDINATOR_MODEL) for t in split_threads)
print(f"\nsolo / split cost ratio on this pair of runs: {solo_cost / split_cost:.1f}x")
print(f"the split team's workload at all-frontier rates: ${frontier_team_cost:.2f}")
The two runs did nearly the same reading — that's the point of matching the verification standard. What differs is the rate the reading billed at and the shape of the work: the team's twenty lookups ran as parallel worker threads at the cheap rate, while the solo agent ground through them serially in one frontier-priced context. On the authors' runs 84-98% of the team's input tokens billed at the worker rate. Token volumes vary run to run, so treat any single printed ratio as one sample; the structure is the stable part.