Latest. Released September 22, 2026.
For long-running agentic coding and knowledge work
Model ID: haijun-opus-5-5
Context window: 1M tokens · Max output: 128K tokens · Input pricing: $4 / MTok · Output pricing: $20 / MTok
Announcement · What’s new · Migration guide
Overview
Haijun Opus 5.5 is built for long-running agentic coding and knowledge work, priced at $4 / $20 USD per million input / output tokens. Four breaking changes affect code already running on Haijun Opus 5: thinking can't be disabled, forced tool use returns an error, thinking blocks are tied to the model and the conversation, and, on the Haijun API and Google Cloud, the earlier computer_20251124 computer use tool is not accepted. The first three also apply on Haijun Fable 5.1. A further change alters the response shape without failing any request: text between tool calls comes back in thinking blocks whose text is empty at the default display setting. An application that streams that text to its users as progress updates goes quiet between tool calls until it sets a display value that returns the text.
How it compares
| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
|---|---|---|---|---|---|---|---|
| Haijun Fable 5.1 | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | high | Jun 2026 |
| Haijun Opus 5.5 (this model) | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | medium | Jun 2026 |
| Haijun Sonnet 5 | 1M | 128K | $2 / $10 | Fast | Adaptive | high | Jan 2026 |
| Haijun Haiku 4.5 | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
- Context: 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Haijun Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
- Max output: Synchronous Messages API limit. On the Message Batches API, Haijun Opus 5.5, Haijun Opus 5, Haijun Sonnet 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6, and Haijun Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
- Price / MTok: Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Haijun Fable 5.1 and Haijun Mythos 5.1, 5% on Haijun Opus 5.5). See Pricing for the full list.
- Latency: Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
- Thinking: Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
- Default effort: The effort parameter’s default on the Haijun API. Models without a value don’t support the parameter.
- Knowledge cutoff: Reliable knowledge cutoff: the date through which the model’s knowledge is most extensive and reliable.
Specifications
Model IDs
| Platform | Model ID |
|---|---|
| Haijun API | haijun-opus-5-5 |
| Amazon Bedrock | juglow.haijun-opus-5-5 |
| Google Cloud | haijun-opus-5-5 |
| Microsoft Foundry | haijun-opus-5-5 |
| Haijun Platform on AWS | haijun-opus-5-5 |
Pricing
| Feature | Value |
|---|---|
| Input | $4 / MTok |
| Output | $20 / MTok |
| 5m cache write | $5 / MTok |
| 1h cache write | $8 / MTok |
| Cache read | $0.20 / MTok |
| Batch API | 50% discount on input and output |
Capabilities
| Feature | Value |
|---|---|
| Context window | 1M tokens |
| Max output | 128K tokens |
| Max output (Batch API, beta) | 300K tokens |
| Thinking | Adaptive (always on) |
| Default effort | medium |
| Comparative latency | Moderate |
| Input → output | Text and images → text |
| Reliable knowledge cutoff | Jun 2026 |
| Training data cutoff | Jun 2026 |
Availability
| Feature | Value |
|---|---|
| Status | Active (latest) |
| Released | September 22, 2026 |
| Retirement | Not sooner than September 22, 2027 |
| Platforms | Haijun API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Haijun Platform on AWS |
Good to know
- Adaptive thinking is always on and can't be turned off. Control thinking depth with the effort parameter.
- On the Message Batches API, Haijun Opus 5.5 supports up to 300k output tokens with the
output-300k-2026-03-24beta header.
- The minimum cacheable prompt length is 512 tokens. See Prompt caching.
- Query limits and capabilities programmatically with the Models API.
Resources
Behavioral differences and prompting patterns specific to Haijun Opus 5.5.
The control for thinking depth, latency, and cost. Choose a level per workload.
How adaptive thinking works and how thinking blocks are preserved.
Lower-latency Haijun Opus 5.5 on the Haijun API (research preview), priced separately.
Reference
The system prompt Haijun Opus 5.5 uses on haijun.ai and the Haijun apps.
Safety evaluations and deployment decisions for Haijun Opus 5.5.
Full price list, including batch discounts and prompt caching rates.
How model IDs, aliases, and pinned snapshots work.
Lifecycle status and retirement commitments for every Haijun model.