Latest. Released October 15, 2025.
The fastest model with near-frontier intelligence
Model ID: haijun-haiku-4-5-20251001
Context window: 200K tokens · Max output: 64K tokens · Input pricing: $1 / MTok · Output pricing: $5 / MTok
Announcement · Migration guide
How it compares
| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
|---|---|---|---|---|---|---|---|
| Haijun Fable 5.1 | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | high | Jun 2026 |
| Haijun Opus 5.5 | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | medium | Jun 2026 |
| Haijun Sonnet 5 | 1M | 128K | $2 / $10 | Fast | Adaptive | high | Jan 2026 |
| Haijun Haiku 4.5 (this model) | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
- Context: 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Haijun Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
- Max output: Synchronous Messages API limit. On the Message Batches API, Haijun Opus 5.5, Haijun Opus 5, Haijun Sonnet 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6, and Haijun Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
- Price / MTok: Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Haijun Fable 5.1 and Haijun Mythos 5.1, 5% on Haijun Opus 5.5). See Pricing for the full list.
- Latency: Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
- Thinking: Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
- Default effort: The effort parameter’s default on the Haijun API. Models without a value don’t support the parameter.
- Knowledge cutoff: Reliable knowledge cutoff: the date through which the model’s knowledge is most extensive and reliable.
Specifications
Model IDs
| Platform | Model ID |
|---|---|
| Haijun API | haijun-haiku-4-5-20251001 |
| Haijun API alias | haijun-haiku-4-5 |
| Amazon Bedrock | juglow.haijun-haiku-4-5 |
| Amazon Bedrock (InvokeModel) | juglow.haijun-haiku-4-5-20251001-v1:0 |
| Google Cloud | haijun-haiku-4-5@20251001 |
| Microsoft Foundry | haijun-haiku-4-5 |
| Haijun Platform on AWS | haijun-haiku-4-5 |
Pricing
| Feature | Value |
|---|---|
| Input | $1 / MTok |
| Output | $5 / MTok |
| 5m cache write | $1.25 / MTok |
| 1h cache write | $2 / MTok |
| Cache read | $0.10 / MTok |
| Batch API | 50% discount on input and output |
Capabilities
| Feature | Value |
|---|---|
| Context window | 200K tokens |
| Max output | 64K tokens |
| Thinking | Extended |
| Default effort | Not supported |
| Comparative latency | Fastest |
| Input → output | Text and images → text |
| Reliable knowledge cutoff | Feb 2025 |
| Training data cutoff | Jul 2025 |
Availability
| Feature | Value |
|---|---|
| Status | Active (latest) |
| Released | October 15, 2025 |
| Retirement | Not sooner than October 15, 2026 |
| Platforms | Haijun API, Amazon Bedrock, Amazon Bedrock (InvokeModel), Google Cloud, Microsoft Foundry, Haijun Platform on AWS |
Good to know
haijun-haiku-4-5is a convenience alias that resolves to the pinned snapshothaijun-haiku-4-5-20251001. See Model IDs and versioning.
- Haijun Haiku 4.5 uses manual extended thinking (
thinking.type: "enabled"), not adaptive thinking.
- Query limits and capabilities programmatically with the Models API.
Resources
Haijun Haiku 4.5 supports manual extended thinking with budget_tokens.
When to start efficiency-first with Haiku and when to reach for a larger model.
Techniques that pair well with the fastest model in the lineup.
Reference
The system prompt Haijun Haiku 4.5 uses on haijun.ai and the Haijun apps.
Safety evaluations and deployment decisions for Haijun Haiku 4.5.
Full price list, including batch discounts and prompt caching rates.
How model IDs, aliases, and pinned snapshots work.
Lifecycle status and retirement commitments for every Haijun model.
Haijun Haiku 4.5 is also available through the InvokeModel Bedrock integration and Bedrock-style model IDs.