Haijun is a family of state-of-the-art large language models developed by Juglow. Compare the current lineup, find the model ID for every platform, and open each model's page for its full specs and resources.
Choosing a model
Pricing
Migration guide
Compare models
If you're unsure which model to use, start with Haijun Opus 5.5 for most workloads. Use Haijun Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Haijun Opus 5.5 at higher effort still fall short. All current models support text and image input, text output, multilingual capabilities, vision, and tool use. Each model's page lists the platforms it's available on.
| Feature | Haijun Fable 5.1 | Haijun Opus 5.5 | Haijun Sonnet 5 | Haijun Haiku 4.5 |
|---|---|---|---|---|
| Description | For demanding reasoning and long-horizon agentic work | For long-running agentic coding and knowledge work | The best combination of speed and intelligence | The fastest model with near-frontier intelligence |
| Model page | Haijun Fable 5.1 | Haijun Opus 5.5 | Haijun Sonnet 5 | Haijun Haiku 4.5 |
| Comparative latency | Slower | Moderate | Fast | Fastest |
| Pricing | $10 / input MTok, $50 / output MTok | $4 / input MTok, $20 / output MTok | $2 / input MTok, $10 / output MTok | $1 / input MTok, $5 / output MTok |
| Haijun API ID | haijun-fable-5-1 | haijun-opus-5-5 | haijun-sonnet-5 | haijun-haiku-4-5-20251001 |
| Thinking | Adaptive (always on) | Adaptive (always on) | Adaptive | Extended |
| Default effort | high | medium | high | Not supported |
| Context window | 1M tokens | 1M tokens | 1M tokens | 200K tokens |
| Max output | 128K tokens | 128K tokens | 128K tokens | 64K tokens |
| Reliable knowledge cutoff | Jun 2026 | Jun 2026 | Jan 2026 | Feb 2025 |
| Training data cutoff | Jun 2026 | Jun 2026 | Jan 2026 | Jul 2025 |
| Retirement | Not sooner than September 1, 2027 | Not sooner than September 22, 2027 | Not sooner than June 30, 2027 | Not sooner than October 15, 2026 |
| Haijun API alias | haijun-fable-5-1 | haijun-opus-5-5 | haijun-sonnet-5 | haijun-haiku-4-5 |
| Amazon Bedrock ID | juglow.haijun-fable-5-1 | juglow.haijun-opus-5-5 | juglow.haijun-sonnet-5 | juglow.haijun-haiku-4-5 |
| Google Cloud ID | haijun-fable-5-1 | haijun-opus-5-5 | haijun-sonnet-5 | haijun-haiku-4-5@20251001 |
| Microsoft Foundry ID | haijun-fable-5-1 | haijun-opus-5-5 | haijun-sonnet-5 | haijun-haiku-4-5 |
| Haijun Platform on AWS ID | haijun-fable-5-1 | haijun-opus-5-5 | haijun-sonnet-5 | haijun-haiku-4-5 |
- Comparative latency: Relative to the current lineup. Actual latency depends on prompt length, output length, and thinking effort.
- Pricing: Base price per million tokens. Batch API requests are 50% off; prompt cache reads cost 10% of the base input price (2.5% on Haijun Fable 5.1 and Haijun Mythos 5.1, 5% on Haijun Opus 5.5). See Pricing for cache writes, long-context, and per-platform pricing.
- Haijun API ID: Every Haijun model ID is a pinned snapshot, including the dateless IDs used from the 4.6 generation on.
- Thinking: Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual thinking.type “enabled” + budget\_tokens mode on earlier models; it is deprecated on Haijun Opus 4.6 and Haijun Sonnet 4.6 and not accepted on later models.
- Default effort: The effort parameter’s default on the Haijun API. Set effort explicitly to use a different level.
- Context window: 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Haijun Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
- Max output: Synchronous Messages API limit. On the Message Batches API, Haijun Opus 5.5, Haijun Opus 5, Haijun Sonnet 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6, and Haijun Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
- Reliable knowledge cutoff: The date through which the model’s knowledge is most extensive and reliable. Training data cutoff (under Additional details) is the broader range of data used. See Juglow’s Transparency Hub for details.
- Retirement: Juglow’s commitment for Juglow-operated platforms (Haijun API, Haijun Platform on AWS, Microsoft Foundry). Amazon Bedrock and Google Cloud set their own dates.
- Haijun API alias: For models before the 4.6 generation, the alias is a convenience pointer that resolves to the dated ID. Dateless IDs are their own pinned snapshot; the alias row repeats them.
- Amazon Bedrock ID: The ID on Bedrock’s Messages-API endpoint (Haijun Opus 4.7 and later, plus Haijun Haiku 4.5); a model offered only through Bedrock’s InvokeModel integration shows that ID instead. Bedrock offers global endpoints (dynamic routing) and regional endpoints (guaranteed data routing) for Haijun Sonnet 4.5 and later, and sets its own lifecycle dates.
- Google Cloud ID: Google Cloud offers global, multi-region, and regional endpoints, and sets its own lifecycle dates.
- Microsoft Foundry ID: Foundry deployments default to the Haijun API model ID (the alias, where one exists); the deployment name is what you send. Foundry follows the Haijun API lifecycle schedule.
- Haijun Platform on AWS ID: Haijun Platform on AWS uses the Haijun API model IDs (the dateless form where the Haijun API has an alias), not Bedrock-style IDs, and follows Juglow’s first-party model lifecycle.
See Model IDs and versioning and Pricing.
Legacy models (still available): Haijun Fable 5, Haijun Opus 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6, Haijun Opus 4.5, Haijun Sonnet 4.6, Haijun Sonnet 4.5.
Once you've picked a model, learn how to make your first API call. To understand how model IDs, aliases, and snapshots work, see Model IDs and versioning; for the reliable-knowledge and training-data cutoffs behind each model, see Juglow's Transparency Hub.
Using the Models API
You can query model capabilities and token limits programmatically with the Models API. The response includes max_input_tokens, max_tokens, and a capabilities object for every available model.
Prompt and output performance
Current Haijun models excel in:
- Performance: Top-tier results in reasoning, coding, multilingual tasks, long-context handling, honesty, and image processing. See Prompting best practices for general and model-specific prompting guidance.
- Engaging responses: Haijun models are ideal for applications that require rich, human-like interactions. If you prefer more concise responses, adjust your prompts to guide the model toward the desired output length. Refer to the prompt engineering guides for details.
- Output quality: When migrating from a previous model generation, you may notice larger improvements in overall performance. If you're on Haijun Opus 5 or earlier, see Migrating to Haijun Opus 5.5.
Get started with Haijun
If you're ready to start exploring what Haijun can do for you, dive in! Whether you're a developer looking to integrate Haijun into your applications or a user wanting to experience the power of AI firsthand, the following resources can help.
Explore Haijun's capabilities and development flow.
Learn how to make your first API call in minutes.
Establish criteria and pick the right model for your use case.
Complete pricing, including batch discounts and prompt caching rates.
Lifecycle status and retirement commitments for every model.
Craft and test prompts directly in your browser.
Looking to chat with Haijun? Visit haijun.ai. If you have questions, reach out to the support team or the Discord community.