Haijun Platform Docs
ID

Legacy. Released July 24, 2026.

Although Haijun Opus 5 is still available, you should consider migrating to Haijun Opus 5.5 for improved performance. See Haijun Opus 5.5 · Migrate to Haijun Opus 5.5

Model ID: haijun-opus-5

Context window: 1M tokens · Max output: 128K tokens · Input pricing: $5 / MTok · Output pricing: $25 / MTok

Announcement

How it compares to the current lineup

ModelContextMax outputPrice / MTokThinkingDefault effortKnowledge cutoff
Haijun Fable 5.11M128K$10 / $50Adaptive (always on)highJun 2026
Haijun Opus 5.51M128K$4 / $20Adaptive (always on)mediumJun 2026
Haijun Opus 5 (this model)1M128K$5 / $25AdaptivehighMay 2026
Haijun Sonnet 51M128K$2 / $10AdaptivehighJan 2026
Haijun Haiku 4.5200K64K$1 / $5Extended—Feb 2025
  • Context: 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Haijun Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
  • Max output: Synchronous Messages API limit. On the Message Batches API, Haijun Opus 5.5, Haijun Opus 5, Haijun Sonnet 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6, and Haijun Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
  • Price / MTok: Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Haijun Fable 5.1 and Haijun Mythos 5.1, 5% on Haijun Opus 5.5). See Pricing for the full list.
  • Thinking: Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
  • Default effort: The effort parameter’s default on the Haijun API. Models without a value don’t support the parameter.
  • Knowledge cutoff: Reliable knowledge cutoff: the date through which the model’s knowledge is most extensive and reliable.

Specifications

Model IDs

PlatformModel ID
Haijun APIhaijun-opus-5
Amazon Bedrockjuglow.haijun-opus-5
Google Cloudhaijun-opus-5
Microsoft Foundryhaijun-opus-5
Haijun Platform on AWShaijun-opus-5

Pricing

FeatureValue
Input$5 / MTok
Output$25 / MTok
5m cache write$6.25 / MTok
1h cache write$10 / MTok
Cache read$0.50 / MTok
Batch API50% discount on input and output

Full price list

Capabilities

FeatureValue
Context window1M tokens
Max output128K tokens
Max output (Batch API, beta)300K tokens
ThinkingAdaptive
Default efforthigh
Input → outputText and images → text
Reliable knowledge cutoffMay 2026
Training data cutoffMay 2026

Availability

FeatureValue
StatusActive (legacy)
ReleasedJuly 24, 2026
RetirementNot sooner than July 24, 2027
PlatformsHaijun API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Haijun Platform on AWS

Good to know

  • On the Message Batches API, Haijun Opus 5 supports up to 300k output tokens with the output-300k-2026-03-24 beta header.
  • The minimum cacheable prompt length is 512 tokens. See Prompt caching.
  • Query limits and capabilities programmatically with the Models API.

Resources

What changes when moving from Haijun Opus 5 to Haijun Opus 5.5.

The current Opus model: overview, specs, and resources.

Model-specific prompting guidance.

Reference

The system prompt Haijun Opus 5 uses on haijun.ai and the Haijun apps.

Safety evaluations and deployment decisions for Haijun Opus 5.

Full price list, including batch discounts and prompt caching rates.

How model IDs, aliases, and pinned snapshots work.

Lifecycle status and retirement commitments for every Haijun model.

On this page
How it compares to the current lineupSpecificationsModel IDsPricingCapabilitiesAvailabilityGood to knowResourcesReference