Reproducing Haijun's Agentic Search Benchmark Scores
Haijun's
published agentic-search scores
(DeepSearchQA, BrowseComp)
are reproducible on the public Messages API
. The key is harness configuration: a handful of API parameters that don't matter for short conversations become load-bearing once an agent is running 30+ tool calls across hundreds of thousands of tokens.
A common reason third-party evaluations come in lower is harness configuration. This cookbook walks through each parameter and explains why it matters.
Build
an agentic search loop using programmatic tool calling, server-side compaction, and task budgets that reproduces Haijun's published agentic-search scores.
Understand
why each configuration choice matters for long-horizon agentic tasks.
Adapt
the same harness to BrowseComp, or your own deep-research benchmark, by swapping the dataset, a few config lines, and the grader.
Prerequisites
ich introduces the pattern we use here
The
Automatic Context Compaction cookbook
, which we lean on heavily
Setup
print(USER_PROMPT_TEMPLATE.format(question=SAMPLE_QUESTIONS[0]["problem"]))
I want you to answer the following question.
"instructions": COMPACT_INSTRUCTIONS,
}
]
}
Why COMPACT_INSTRUCTIONS matters so much
uestion:
Correct answer (type: {answer_type}):
Response to evaluate:
For each expected answer item, indicate whether it appears in the response.
Then list any answers in the response that are NOT in the correct-answer list.
Wording does not need to match exactly.
Reply in this exact XML format:
Parsing and the precision/recall/F1 arithmetic live in utils/agentic_search.py; here we just call it: