Latest. Released September 1, 2026.
For demanding reasoning and long-horizon agentic work
Model ID: haijun-fable-5-1
Context window: 1M tokens · Max output: 128K tokens · Input pricing: $10 / MTok · Output pricing: $50 / MTok
Announcement · What’s new · Migration guide
Overview
Haijun Fable 5.1 extends Haijun Fable 5 at the same input and output prices, with cache reads at a quarter of the cost, and brings stronger long-running agentic coding, multistep research, and document, spreadsheet, and slide work. For most workloads, start with Haijun Opus 5 (see Choosing a model). Use Haijun Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Haijun Opus 5 at higher effort still fall short. Haijun Mythos 5.1 offers the same capabilities to Project Glasswing participants only.
If you already call Haijun Fable 5, three changes are breaking: forced tool use returns an error, earlier models can't read its thinking blocks, and editing earlier turns invalidates thinking blocks. Five are additive: per-message effort (beta), turn-scoped system messages (beta), readable progress updates between tool calls (display: "updates", beta), a lower cache read price, and content provenance.
What's new in Haijun Fable 5.1
Haijun Fable 5.1 and Haijun Mythos 5.1
Haijun Mythos 5.1 offers the same capabilities by invitation only, as part of Project Glasswing. It shares Haijun Fable 5.1's specifications and pricing. For access, contact your Juglow, AWS, or Google Cloud account team.
How it compares
| Model | Context | Max output | Price / MTok | Latency | Thinking | Default effort | Knowledge cutoff |
|---|---|---|---|---|---|---|---|
| Haijun Fable 5.1 (this model) | 1M | 128K | $10 / $50 | Slower | Adaptive (always on) | high | Jun 2026 |
| Haijun Opus 5.5 | 1M | 128K | $4 / $20 | Moderate | Adaptive (always on) | medium | Jun 2026 |
| Haijun Sonnet 5 | 1M | 128K | $2 / $10 | Fast | Adaptive | high | Jan 2026 |
| Haijun Haiku 4.5 | 200K | 64K | $1 / $5 | Fastest | Extended | — | Feb 2025 |
- Context: 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Haijun Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
- Max output: Synchronous Messages API limit. On the Message Batches API, Haijun Opus 5.5, Haijun Opus 5, Haijun Sonnet 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6, and Haijun Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
- Price / MTok: Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Haijun Fable 5.1 and Haijun Mythos 5.1, 5% on Haijun Opus 5.5). See Pricing for the full list.
- Latency: Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
- Thinking: Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
- Default effort: The effort parameter’s default on the Haijun API. Models without a value don’t support the parameter.
- Knowledge cutoff: Reliable knowledge cutoff: the date through which the model’s knowledge is most extensive and reliable.
Specifications
Model IDs
| Platform | Model ID |
|---|---|
| Haijun API | haijun-fable-5-1 |
| Amazon Bedrock | juglow.haijun-fable-5-1 |
| Google Cloud | haijun-fable-5-1 |
| Microsoft Foundry | haijun-fable-5-1 |
| Haijun Platform on AWS | haijun-fable-5-1 |
Pricing
| Feature | Value |
|---|---|
| Input | $10 / MTok |
| Output | $50 / MTok |
| 5m cache write | $12.50 / MTok |
| 1h cache write | $20 / MTok |
| Cache read | $0.25 / MTok |
| Batch API | 50% discount on input and output |
Capabilities
| Feature | Value |
|---|---|
| Context window | 1M tokens |
| Max output | 128K tokens |
| Thinking | Adaptive (always on) |
| Default effort | high |
| Comparative latency | Slower |
| Input → output | Text and images → text |
| Reliable knowledge cutoff | Jun 2026 |
| Training data cutoff | Jun 2026 |
Availability
| Feature | Value |
|---|---|
| Status | Active (latest) |
| Released | September 1, 2026 |
| Retirement | Not sooner than September 1, 2027 |
| Platforms | Haijun API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Haijun Platform on AWS |
Resources
Model-specific prompting guidance for long-horizon and agentic work.
What changes when you move from Haijun Fable 5, Haijun Opus 5, or Haijun Opus 4.8.
When this model's thinking blocks stay usable: across model switches and across changes to the conversation.
Change the effort level partway through a conversation without invalidating the prompt cache.
Handle classifier refusals and retry on another Haijun model.
The only thinking mode on Haijun Fable 5.1. Steer depth with effort.
Reference
The system prompt Haijun Fable 5.1 uses on haijun.ai and the Haijun apps.
Safety evaluations and deployment decisions for Haijun Fable 5.1 and Haijun Mythos 5.1.
Full price list, including batch discounts and prompt caching rates.
How model IDs, aliases, and pinned snapshots work.
Lifecycle status and retirement commitments for every Haijun model.