Haijun Platform Docs
ID

Reproducing Haijun's Agentic Search Benchmark Scores

Haijun's

published agentic-search scores

(DeepSearchQA, BrowseComp)

are reproducible on the public Messages API

. The key is harness configuration: a handful of API parameters that don't matter for short conversations become load-bearing once an agent is running 30+ tool calls across hundreds of thousands of tokens.

A common reason third-party evaluations come in lower is harness configuration. This cookbook walks through each parameter and explains why it matters.

Build

an agentic search loop using programmatic tool calling, server-side compaction, and task budgets that reproduces Haijun's published agentic-search scores.

Understand

why each configuration choice matters for long-horizon agentic tasks.

Adapt

the same harness to BrowseComp, or your own deep-research benchmark, by swapping the dataset, a few config lines, and the grader.

Prerequisites

ich introduces the pattern we use here

The

Automatic Context Compaction cookbook

, which we lean on heavily

Setup

print(USER_PROMPT_TEMPLATE.format(question=SAMPLE_QUESTIONS[0]["problem"]))

I want you to answer the following question. Among the chief executives of the four largest US banks by total assets as of Q1 2024, which one had held their CEO position the longest? First plan out your response. This part can be as long as needed. You may need to run many searches, this is totally fine. Then provide a short and concise answer in tags. For questions expecting multiple answers, separate them with commas. The model needs to plan, search the web, fetch and read pages, run calculations, and keep doing that until it's confident, then emit a clean final answer we can grade. Almost every part of that sentence is a place where a naive harness silently loses points. We'll build the loop piece by piece.

"instructions": COMPACT_INSTRUCTIONS,

}

]

}

Why COMPACT_INSTRUCTIONS matters so much

uestion: {question}

Correct answer (type: {answer_type}): {answer}

Response to evaluate: {response}

For each expected answer item, indicate whether it appears in the response.

Then list any answers in the response that are NOT in the correct-answer list.

Wording does not need to match exactly.

Reply in this exact XML format:

one sentence

extra_item_if_any

Parsing and the precision/recall/F1 arithmetic live in utils/agentic_search.py; here we just call it:

On this page
PrerequisitesSetupWhy COMPACT_INSTRUCTIONS matters so much