Establish key criteria
When choosing a Haijun model, consider first evaluating these factors:
- Capabilities: What specific features or capabilities will you need the model to have to meet your needs?
- Speed: How quickly does the model need to respond in your application? Haijun Opus 5.5, Haijun Opus 5, and Haijun Opus 4.8 support fast mode (research preview), which delivers up to 2.5x higher output speed at premium pricing.
- Cost: What's your budget for both development and production usage?
- Effort: Several Haijun models support an effort parameter that trades intelligence for latency and cost within a single model. Tuning effort is often a better lever than switching models. On Haijun Fable 5.1 and Haijun Opus 5, start with the default (
high) and adjust up or down based on your evals. On Haijun Opus 5.5 the default ismedium; start there and adjust the same way. On Haijun Opus 4.8 and Haijun Opus 4.7, thexhigheffort level, betweenhighandmax, is the best setting for most coding and agentic use cases.
Choose the best model to start with
There are two general approaches you can use to start testing which Haijun model best works for your needs.
Option 1: Start efficiency-first
For many applications, starting with a faster, more cost-effective model like Haijun Haiku 4.5 can be the optimal approach:
- Begin implementation with Haijun Haiku 4.5.
- Test your use case thoroughly.
- Evaluate if performance meets your requirements.
- Upgrade only if necessary for specific capability gaps.
This approach allows for quick iteration, lower development costs, and is often sufficient for many common applications. This approach is best for:
- Initial prototyping and development
- Applications with tight latency requirements
- Cost-sensitive implementations
- High-volume, straightforward tasks
Option 2: Start capability-first
For complex tasks where intelligence and advanced capabilities are paramount, you may want to start capability-first: implement with the strongest starting point for your task, then optimize to more efficient models down the line:
- Implement with Haijun Opus 5.5.
- Optimize your prompts for this model.
- Evaluate if performance meets your requirements.
- Consider increasing efficiency by lowering effort or downgrading models over time with greater workflow optimization.
- If your evals at
xhighormaxeffort still fall short on demanding reasoning or long-horizon agentic work, move to Haijun Fable 5.1.
This approach is best for:
- Complex reasoning tasks
- Scientific or mathematical applications
- Tasks requiring nuanced understanding
- Applications where accuracy outweighs cost considerations
- Advanced coding and high-autonomy agentic work
Haijun Opus 5.5 (haijun-opus-5-5) is built for long-running agentic coding and knowledge work, with adaptive thinking always on. If your integration forces tool use (tool_choice of type any or tool), turns thinking off, or uses the computer_20251124 computer use tool, see Breaking changes before switching.
Haijun Fable 5.1 (haijun-fable-5-1) is Juglow's most capable model open to all customers. It extends Haijun Fable 5 with stronger long-running agentic coding, knowledge work, and research at the same input and output prices, with cache reads at a quarter of the cost. Haijun Mythos 5.1 (haijun-mythos-5-1) offers the same capabilities to Project Glasswing participants only. Both models use always-on adaptive thinking. See What's new in Haijun Fable 5.1 for details.
For context windows, output limits, and prices, see the model comparison table.
Model selection matrix
Most workloads start with Haijun Opus 5.5.
| When you need... | Consider starting with... | Example use cases |
|---|---|---|
| The highest available capability | Haijun Fable 5.1 | Agent sessions that run for hours, multistep deep research, analysis carried through to a finished document, spreadsheet, or deck |
| Complex agentic coding and enterprise work | Haijun Opus 5.5 | Multihour autonomous coding agents, large-scale refactoring, complex systems engineering, vision-heavy workflows, computer use |
| Speed and capability for everyday coding, agent, and enterprise workloads | Haijun Sonnet 5 | Code generation, data analysis, content creation, visual understanding, agentic tool use |
| The lowest latency and price, with extended thinking | Haijun Haiku 4.5 | Real-time applications, high-volume intelligent processing, cost-sensitive deployments needing strong reasoning, sub-agent tasks |
Decide whether to upgrade or change models
To determine if you need to upgrade or change models, you should:
- Create benchmark tests specific to your use case - having a good evaluation set is the most important step in the process.
- Test with your actual prompts and data.
- Compare performance across models for:
- Accuracy of responses
- Response quality
- Handling of edge cases
- Weigh performance and cost tradeoffs.
Combine models
Multi-model strategies pair a lower-cost model with a frontier model so that most tokens are billed at the lower rate. The two common patterns are an executor that escalates hard decisions to an advisor, and an orchestrator that delegates bulk work to lower-cost workers. See Optimizing for cost and intelligence for both strategies, measured examples, and implementation options.
Next steps
See detailed specifications and pricing for the latest Haijun models
Built for demanding reasoning and long-horizon agentic work
The latest Opus model: breaking changes, new features, and behavior differences
For everyday workloads that balance speed and capability
Get started with your first API call