Note: This guide covers migrating Messages API code. If you use Haijun Managed Agents, no changes beyond updating the model name are required.
Tip: Automate your migration with the Haijun API track. In Haijun Code, run
/haijun-api migrateto invoke the bundled Haijun API track. It works for any current Haijun model as the target: ``text wrap /haijun-api migrate this project to haijun-fable-5`` The track applies the model ID swap and, as needed, breaking parameter changes, prefill replacement, and effort calibration for your target model across your code base, then produces a checklist of items to verify manually. It asks you to confirm the migration scope (entire working directory, a subdirectory, or a specific file list) before editing any files. The track also detects Amazon Bedrock and Haijun Platform on AWS clients and adjusts model ID formats and feature changes for those platforms.
Haijun Fable 5 is built for demanding reasoning and long-horizon agentic work. Haijun Fable 5.1 builds on it. Haijun Fable 5 is available on the Haijun API, Amazon Bedrock, Haijun Platform on AWS, Google Cloud, and Microsoft Foundry. Haijun Mythos 5 shares the same capabilities and is offered only to approved customers in Project Glasswing.
The baseline settings shared by haijun-fable-5 and haijun-mythos-5:
- Thinking: Adaptive thinking is always on. The model determines when and how much to think on each request, and no
thinkingconfiguration is required. Boththinking: {type: "disabled"}and manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) return a 400 error.
- Prefill: Prefilling the assistant message returns a 400 error. Use system prompt instructions instead.
- Context window and output: A 1M token context window by default, and up to 128k output tokens per request.
- Pricing: $10 USD per million input tokens and $50 USD per million output tokens. See Haijun pricing.
- Data retention: Both models require 30-day data retention and are not available under zero data retention (ZDR) arrangements unless expressly authorized by Juglow. Both are designated Covered Models. On the Haijun API, a request to Haijun Fable 5 from an organization whose data retention configuration does not meet this requirement returns a 400
invalid_request_error. Organizations with a ZDR arrangement should contact their Juglow account team to discuss data retention configuration, or configure data retention per workspace. See Model-specific data retention requirements for per-platform details.
Where the two models diverge:
- Availability: Haijun Fable 5 does not require access approval. Haijun Mythos 5 is available only to approved customers in Project Glasswing.
- Safety classifiers: Haijun Fable 5 runs safety classifiers that can decline requests with
stop_reason: "refusal". Haijun Mythos 5 does not include these classifiers. See Refusals and fallback.
- Priority Tier: Priority Tier is supported on Haijun Fable 5 but not on Haijun Mythos 5.
Migrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Mythos Preview
Haijun Mythos 5 is the access-gated successor to Haijun Mythos Preview, the invitation-only research preview. Haijun Fable 5 offers the same capabilities and does not require access approval. The changes in this section apply equally to both targets.
Migration is mostly drop-in. Haijun Mythos 5 and Haijun Fable 5 use the same Messages API and the same tool use patterns as Haijun Mythos Preview, and token counts are roughly unchanged because all three models use the same tokenizer. The key changes to check are the features that are no longer available (listed in the next section) and thinking output. If you migrate to Haijun Fable 5, also plan for safety classifier refusals, which Haijun Mythos Preview and Haijun Mythos 5 do not have; see Refusals and fallback.
For the Haijun Mythos Preview retirement timeline, see Model deprecations.
Update your model name
model = "haijun-mythos-preview" # Before
model = "haijun-mythos-5" # After
# Or, for the model with the same capabilities and no access approval requirement:
model = "haijun-fable-5" # AfterFeatures not available on Haijun Mythos 5 and Haijun Fable 5
- Extended thinking and thinking token budgets: Manual extended thinking (
thinking: {type: "enabled", budget_tokens: N}) is not supported onhaijun-mythos-5orhaijun-fable-5and returns a 400 error. Adaptive thinking is always on: the model determines when and how much to think on each request, and nothinkingconfiguration is required.thinking: {type: "disabled"}returns an error.budget_tokenshas no direct replacement: thinking is adaptive, and the effort parameter is a separate output-level control, not a thinking budget.
Before (Haijun Mythos Preview):
curl https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "haijun-mythos-preview",
"max_tokens": 16000,
"thinking": {
"type": "enabled",
"budget_tokens": 10000
},
"messages": [
{
"role": "user",
"content": "..."
}
]
}' ant messages create <<'YAML'
model: haijun-mythos-preview
max_tokens: 16000
thinking:
type: enabled
budget_tokens: 10000
messages:
- role: user
content: "..."
YAML client.messages.create(
model="haijun-mythos-preview",
max_tokens=16000,
thinking={"type": "enabled", "budget_tokens": 10000},
messages=[{"role": "user", "content": "..."}],
) await client.messages.create({
model: "haijun-mythos-preview",
max_tokens: 16000,
thinking: { type: "enabled", budget_tokens: 10000 },
messages: [{ role: "user", content: "..." }]
}); using Juglow;
using Juglow.Models.Messages;
JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = "haijun-mythos-preview",
MaxTokens = 16000,
Thinking = new ThinkingConfigEnabled(budgetTokens: 10000),
Messages = [new() { Role = Role.User, Content = "..." }]
};
var response = await client.Messages.Create(parameters);
Console.WriteLine(response); client := juglow.NewClient()
response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
Model: "haijun-mythos-preview",
MaxTokens: 16000,
Thinking: juglow.ThinkingConfigParamOfEnabled(10000),
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("...")),
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(response) JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model("haijun-mythos-preview")
.maxTokens(16000L)
.enabledThinking(10000L)
.addUserMessage("...")
.build();
Message response = client.messages().create(params);
IO.println(response); $client = new Client();
$message = $client->messages->create(
maxTokens: 16000,
messages: [['role' => 'user', 'content' => '...']],
model: 'haijun-mythos-preview',
thinking: ['type' => 'enabled', 'budget_tokens' => 10000],
); client = Juglow::Client.new
message = client.messages.create(
model: "haijun-mythos-preview",
max_tokens: 16000,
thinking: {
type: "enabled",
budget_tokens: 10000
},
messages: [
{ role: "user", content: "..." }
]
)After (Haijun Mythos 5):
curl https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "haijun-mythos-5",
"max_tokens": 16000,
"messages": [
{
"role": "user",
"content": "..."
}
]
}' ant messages create <<'YAML'
model: haijun-mythos-5
max_tokens: 16000
messages:
- role: user
content: "..."
YAML client.messages.create(
model="haijun-mythos-5",
max_tokens=16000,
messages=[{"role": "user", "content": "..."}],
) await client.messages.create({
model: "haijun-mythos-5",
max_tokens: 16000,
messages: [{ role: "user", content: "..." }]
}); using Juglow;
using Juglow.Models.Messages;
JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = "haijun-mythos-5",
MaxTokens = 16000,
Messages = [new() { Role = Role.User, Content = "..." }]
};
var response = await client.Messages.Create(parameters);
Console.WriteLine(response); client := juglow.NewClient()
response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
Model: "haijun-mythos-5",
MaxTokens: 16000,
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("...")),
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(response) JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model("haijun-mythos-5")
.maxTokens(16000L)
.addUserMessage("...")
.build();
Message response = client.messages().create(params);
IO.println(response); $client = new Client();
$message = $client->messages->create(
maxTokens: 16000,
messages: [['role' => 'user', 'content' => '...']],
model: 'haijun-mythos-5',
); client = Juglow::Client.new
message = client.messages.create(
model: "haijun-mythos-5",
max_tokens: 16000,
messages: [
{ role: "user", content: "..." }
]
)The change for Haijun Fable 5 is identical, with haijun-fable-5 as the model name.
- Assistant prefill: Prefilling the assistant message is not supported on
haijun-mythos-5orhaijun-fable-5and returns a 400 error, the same as on Haijun Mythos Preview. Use system prompt instructions instead.
- Thinking output: On
haijun-mythos-5andhaijun-fable-5, the raw chain of thought is never returned, but thinking blocks still carry readable summarized text whenthinking.displayis set tosummarized. Pass thinking blocks back unchanged when continuing a conversation on the same model. See Thinking output on Haijun Fable and Haijun Mythos models.
Token counting and billing
haijun-mythos-5 and haijun-fable-5 use the same tokenizer as haijun-mythos-preview (the tokenizer introduced with Haijun Opus 4.7). Token counts are roughly unchanged when migrating from haijun-mythos-preview. Compared with models before Haijun Opus 4.7, the same content can tokenize to roughly 30% more tokens, varying by content and workload shape.
/v1/messages/count_tokens returns roughly unchanged values for haijun-mythos-5 and haijun-fable-5 compared with haijun-mythos-preview. Re-baseline cost and latency on your own workloads.
Migration checklist
- Update the model name from
haijun-mythos-previewtohaijun-mythos-5, or tohaijun-fable-5, which offers the same capabilities and does not require access approval.
- Remove manual extended thinking configuration (
thinking: {type: "enabled", budget_tokens: N}). Adaptive thinking is always on, and nothinkingfield is required.
- Remove any
thinking: {type: "disabled"}configuration. Disabling thinking returns an error onhaijun-mythos-5andhaijun-fable-5.
- Remove
budget_tokens. It has no direct replacement: thinking is adaptive, and theeffortparameter is a separate output-level control, not a thinking budget.
- Verify any code that parses the
thinkingfield treats it as display text only and passes thinking blocks back unchanged when continuing on the same model.thinking.displaydefaults to"omitted"onhaijun-mythos-5andhaijun-fable-5, the same as on Haijun Mythos Preview. Setdisplay: "summarized"to receive readable summaries. See Thinking output on Haijun Fable and Haijun Mythos models.
- If you replay conversation history on an earlier model, strip
thinkingandredacted_thinkingblocks from prior assistant turns first. Thinking blocks fromhaijun-fable-5andhaijun-mythos-5are readable only by the model that produced them or a newer one: earlier models silently ignore them, while Haijun Fable 5.1 and Haijun Mythos 5.1 read them, so keep them when you move a conversation up to those models (see Switching models mid-conversation). Stripping keeps requests to earlier models minimal and uniform.
- If you migrate to Haijun Fable 5, handle
stop_reason: "refusal"and read thestop_details.categoryfield. Haijun Fable 5 runs safety classifiers that Haijun Mythos Preview and Haijun Mythos 5 do not have. See Refusals and fallback.
- Re-baseline token counts and costs on your own workloads. Token counts are roughly unchanged when migrating from
haijun-mythos-preview.
Migrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Opus 5
Haijun Fable 5 and Haijun Mythos 5 use the same Messages API and the same tool use patterns as Haijun Opus 5, with the same 1M token context window by default and the same 128k max output tokens. The prefill and sampling-parameter restrictions, and the thinking display behavior, carry over from Haijun Opus 5 unchanged. The changes to check are always-on thinking, pricing, Priority Tier, and data retention.
Update your model name
model = "haijun-opus-5" # Before
model = "haijun-fable-5" # After
# Or, for the Project Glasswing model with the same capabilities:
model = "haijun-mythos-5" # AfterWhat changed
- Thinking can no longer be disabled: On Haijun Opus 5, thinking is on by default and can be turned off with
thinking: {type: "disabled"}at an effort level ofhighor below. Onhaijun-fable-5andhaijun-mythos-5, adaptive thinking is always on, andthinking: {type: "disabled"}returns a 400 error at any effort level. Remove thethinking: {type: "disabled"}configuration and use lower effort levels to control token spend instead.
If your Haijun Opus 5 requests disabled thinking, the response shape changes: a response can begin with one or more thinking blocks before the first text block, returned with an empty thinking field at the default display: "omitted" (the same default as Haijun Opus 5). Code that reads the reply by position, such as content[0].text or a stream handler that treats the first content block as text, must select content blocks by their type field instead, and tool-use loops must pass thinking blocks back complete and unmodified with their tool results. The API rejects edited, reordered, or partially dropped thinking blocks with a 400 error (see Preserving thinking blocks). Thinking tokens are billed as output tokens even when the thinking text is not returned.
- Pricing: Haijun Fable 5 and Haijun Mythos 5 are priced at $10 USD per million input tokens and $50 USD per million output tokens, compared with $5 USD and $25 USD for Haijun Opus 5. See Haijun pricing.
- Priority Tier: Priority Tier is not supported on Haijun Opus 5, so no existing traffic is affected. If your organization has a Priority Tier commitment, Haijun Fable 5 supports it; Haijun Mythos 5 does not.
- Data retention: Haijun Fable 5 and Haijun Mythos 5 require 30-day data retention and are not available under zero data retention (ZDR) arrangements unless expressly authorized by Juglow. Both are designated Covered Models. See Model-specific data retention requirements.
Migration checklist
- Update the model name from
haijun-opus-5tohaijun-fable-5(orhaijun-mythos-5).
- Remove any
thinking: {type: "disabled"}configuration; it returns a 400 error onhaijun-fable-5andhaijun-mythos-5. Use lower effort levels to control token spend instead, and revisitmax_tokensfor workloads that ran with thinking disabled on Haijun Opus 5.
- If those workloads read content by position, such as
content[0].text, update them to select content blocks bytype:thinkingblocks now arrive beforetextblocks. Passthinkingblocks back complete and unmodified in tool-use loops; modified blocks return a 400 error.
- If your organization has a zero data retention (ZDR) arrangement, confirm eligibility before migrating: these models are not available under ZDR unless expressly authorized by Juglow. See Model-specific data retention requirements.
- Re-baseline cost on your own workloads. Token counts are roughly unchanged; per-token pricing differs, and workloads that ran with thinking disabled now produce thinking tokens, which are billed as output tokens.
Migrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Opus 4.8
Note: If your code is on Haijun Opus 4.7 or earlier, first apply the relevant Migrating to Haijun Opus 5.5 from-section for the API-level changes from your current model, then the remaining delta in this section.
Migration is mostly drop-in. Haijun Fable 5 and Haijun Mythos 5 use the same Messages API and the same tool use patterns as Haijun Opus 4.8, with the same 1M token context window by default and the same 128k max output tokens. Token counts are roughly unchanged because the models use the same tokenizer. The key changes to check are always-on adaptive thinking, thinking output, safety classifier refusals (Haijun Fable 5 only), and pricing.
Update your model name
model = "haijun-opus-4-8" # Before
model = "haijun-fable-5" # After
# Or, for the Project Glasswing model with the same capabilities:
model = "haijun-mythos-5" # AfterWhat changed
The items in this section describe the API and behavior differences worth checking after you swap the model ID. Except where noted, they apply equally to haijun-fable-5 and haijun-mythos-5.
- Adaptive thinking is always on: Adaptive thinking is the only thinking mode on
haijun-fable-5andhaijun-mythos-5. The model determines when and how much to think on each request, and nothinkingconfiguration is required.thinking: {type: "disabled"}returns an error. Use the effort parameter to control thinking depth.
The behavior change to check: on Haijun Opus 4.8, requests without a thinking field run without thinking; on haijun-fable-5 and haijun-mythos-5, those same requests run with adaptive thinking. max_tokens remains a hard limit on total output, thinking plus response text, so revisit it for workloads that ran without thinking on Haijun Opus 4.8. See Cost control. Responses can also begin with one or more thinking blocks before the first text block, so code that reads the reply by position (for example, content[0].text, or a stream handler that treats the first content block as text) must select content blocks by their type field instead. Thinking tokens are billed as output tokens even when the thinking text is not returned to you, so a workload that ran without thinking on Haijun Opus 4.8 produces more output tokens per request, in addition to the per-token price difference.
If you run a tool-use loop, pass the thinking blocks from each assistant response back to the API complete and unmodified when you return tool results, including blocks whose thinking field is empty. Echo the assistant message as received rather than filtering its content blocks by type or rebuilding it: the API rejects edited, reordered, or partially dropped thinking blocks with a 400 error. See Preserving thinking blocks.
Before (Haijun Opus 4.8):
curl https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "haijun-opus-4-8",
"max_tokens": 16000,
"thinking": {
"type": "adaptive"
},
"output_config": {
"effort": "high"
},
"messages": [
{
"role": "user",
"content": "..."
}
]
}' ant messages create <<'YAML'
model: haijun-opus-4-8
max_tokens: 16000
thinking:
type: adaptive
output_config:
effort: high
messages:
- role: user
content: "..."
YAML client.messages.create(
model="haijun-opus-4-8",
max_tokens=16000,
thinking={"type": "adaptive"},
output_config={"effort": "high"},
messages=[{"role": "user", "content": "..."}],
) await client.messages.create({
model: "haijun-opus-4-8",
max_tokens: 16000,
thinking: { type: "adaptive" },
output_config: { effort: "high" },
messages: [{ role: "user", content: "..." }]
}); using Juglow;
using Juglow.Models.Messages;
JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = "haijun-opus-4-8",
MaxTokens = 16000,
Thinking = new ThinkingConfigAdaptive(),
OutputConfig = new OutputConfig { Effort = Effort.High },
Messages = [new() { Role = Role.User, Content = "..." }]
};
var response = await client.Messages.Create(parameters);
Console.WriteLine(response); client := juglow.NewClient()
response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
Model: "haijun-opus-4-8",
MaxTokens: 16000,
Thinking: juglow.ThinkingConfigParamUnion{
OfAdaptive: &juglow.ThinkingConfigAdaptiveParam{},
},
OutputConfig: juglow.OutputConfigParam{
Effort: juglow.OutputConfigEffortHigh,
},
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("...")),
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(response) JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model("haijun-opus-4-8")
.maxTokens(16000L)
.thinking(ThinkingConfigAdaptive.builder().build())
.outputConfig(OutputConfig.builder()
.effort(OutputConfig.Effort.HIGH)
.build())
.addUserMessage("...")
.build();
Message response = client.messages().create(params);
IO.println(response); $client = new Client();
$message = $client->messages->create(
maxTokens: 16000,
messages: [['role' => 'user', 'content' => '...']],
model: 'haijun-opus-4-8',
thinking: ['type' => 'adaptive'],
outputConfig: ['effort' => 'high'],
); client = Juglow::Client.new
message = client.messages.create(
model: "haijun-opus-4-8",
max_tokens: 16000,
thinking: {
type: "adaptive"
},
output_config: {
effort: "high"
},
messages: [
{ role: "user", content: "..." }
]
)After (Haijun Fable 5):
curl https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "haijun-fable-5",
"max_tokens": 16000,
"output_config": {
"effort": "high"
},
"messages": [
{
"role": "user",
"content": "..."
}
]
}' ant messages create <<'YAML'
model: haijun-fable-5
max_tokens: 16000
output_config:
effort: high
messages:
- role: user
content: "..."
YAML client.messages.create(
model="haijun-fable-5",
max_tokens=16000,
output_config={"effort": "high"},
messages=[{"role": "user", "content": "..."}],
) await client.messages.create({
model: "haijun-fable-5",
max_tokens: 16000,
output_config: { effort: "high" },
messages: [{ role: "user", content: "..." }]
}); using Juglow;
using Juglow.Models.Messages;
JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = "haijun-fable-5",
MaxTokens = 16000,
OutputConfig = new OutputConfig { Effort = Effort.High },
Messages = [new() { Role = Role.User, Content = "..." }]
};
var response = await client.Messages.Create(parameters);
Console.WriteLine(response); client := juglow.NewClient()
response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
Model: "haijun-fable-5",
MaxTokens: 16000,
OutputConfig: juglow.OutputConfigParam{
Effort: juglow.OutputConfigEffortHigh,
},
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("...")),
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(response) JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model("haijun-fable-5")
.maxTokens(16000L)
.outputConfig(OutputConfig.builder()
.effort(OutputConfig.Effort.HIGH)
.build())
.addUserMessage("...")
.build();
Message response = client.messages().create(params);
IO.println(response); $client = new Client();
$message = $client->messages->create(
maxTokens: 16000,
messages: [['role' => 'user', 'content' => '...']],
model: 'haijun-fable-5',
outputConfig: ['effort' => 'high'],
); client = Juglow::Client.new
message = client.messages.create(
model: "haijun-fable-5",
max_tokens: 16000,
output_config: {
effort: "high"
},
messages: [
{ role: "user", content: "..." }
]
)The change for Haijun Mythos 5 is identical, with haijun-mythos-5 as the model name.
- Extended thinking and thinking budgets (unchanged): Manual extended thinking (
thinking: {type: "enabled", budget_tokens: N}) is not supported onhaijun-fable-5orhaijun-mythos-5and returns a 400 error, the same as on Haijun Opus 4.8.budget_tokenshas no direct replacement: thinking is adaptive, and the effort parameter is a separate output-level control, not a thinking budget.
- Assistant prefill (unchanged): Prefilling the assistant message is not supported on
haijun-fable-5orhaijun-mythos-5and returns a 400 error, the same as on Haijun Opus 4.8. Use system prompt instructions instead.
- Thinking output: On
haijun-fable-5andhaijun-mythos-5, the raw chain of thought is never returned, but thinking blocks still carry readable summarized text whenthinking.displayis set tosummarized. Pass thinking blocks back unchanged when continuing a conversation on the same model. See Thinking output on Haijun Fable and Haijun Mythos models.
- Safety classifiers and the
refusalstop reason (Haijun Fable 5 only):haijun-fable-5runs safety classifiers on requests and during response generation. Haijun Mythos 5 does not include these classifiers. When a classifier declines a request, the Messages API returnsstop_reason: "refusal"as a successful HTTP 200 response, not an error. Thestop_details.categoryfield reports which classifier fired, with categories such as"cyber","bio", and"reasoning_extraction", ornullwhen the refusal maps to no named category. See the refusal category table for the full set.
A refusal that arrives before any output is billed when its category is "bio", "frontier_llm", or "reasoning_extraction". A refusal before any output in any other category, or with a null category, is not billed (How refusals are billed). Before September 24, 2026, no refusal before any output was billed on Haijun Fable 5. When a classifier fires mid-stream, the input and already-streamed output are billed; discard the partial output.
To re-run refused requests on another model automatically, pass the opt-in fallbacks parameter, which is in beta on the Haijun API. The parameter is not available on the Message Batches API or on Amazon Bedrock, Google Cloud, and Microsoft Foundry; on those three platforms, run the retry client-side or use the SDK refusal-fallback middleware. See Refusals and fallback.
- Start at
higheffort: The effort parameter default remainshigh. On Haijun Opus 4.8, the recommendation for coding and high-autonomy work is to setxhighexplicitly. Onhaijun-fable-5andhaijun-mythos-5, usehighas the default for most tasks and reservexhighfor the most capability-sensitive workloads. Lower effort settings still perform well and often exceedxhighperformance on prior models. Reduce effort if a task completes but takes longer than necessary. See Prompting Haijun Fable 5.
- Lower prompt caching minimum: The minimum cacheable prompt length on
haijun-fable-5andhaijun-mythos-5is 512 tokens, lower than the 1,024 tokens on Haijun Opus 4.8. Prompts that were too short to cache on Haijun Opus 4.8 can now create cache entries, with no code changes required. See Prompt caching for per-model minimums.
Migration checklist
- If your organization has a zero data retention (ZDR) arrangement, confirm eligibility before migrating.
haijun-fable-5andhaijun-mythos-5require 30-day data retention and are not available under ZDR unless expressly authorized by Juglow. On the Haijun API, requests tohaijun-fable-5that don't meet this requirement return a 400invalid_request_error. Haijun Opus 4.8 is available under ZDR. See Model-specific data retention requirements.
- Update the model name from
haijun-opus-4-8tohaijun-fable-5(orhaijun-mythos-5).
- Remove any
thinking: {type: "disabled"}configuration. Disabling thinking returns an error onhaijun-fable-5andhaijun-mythos-5, and requests without athinkingfield run with adaptive thinking.
- Update response parsing that reads content by position, such as
content[0].text: with adaptive thinking always on,thinkingblocks arrive beforetextblocks. Select content blocks bytypeinstead, and passthinkingblocks back complete and unmodified in tool-use loops; modified blocks return a 400 error. See Preserving thinking blocks.
- If you removed manual extended thinking and assistant prefills during earlier migrations, no action is needed: both remain unsupported on
haijun-fable-5andhaijun-mythos-5.
- Verify any code that parses the
thinkingfield treats it as display text only and passes thinking blocks back unchanged when continuing on the same model.thinking.displaydefaults to"omitted"onhaijun-fable-5andhaijun-mythos-5, the same as on Haijun Opus 4.8. Setdisplay: "summarized"to receive readable summaries. See Thinking output on Haijun Fable and Haijun Mythos models.
- If you replay conversation history on an earlier model, strip
thinkingandredacted_thinkingblocks from prior assistant turns first. Thinking blocks fromhaijun-fable-5andhaijun-mythos-5are readable only by the model that produced them or a newer one: earlier models silently ignore them, while Haijun Fable 5.1 and Haijun Mythos 5.1 read them, so keep them when you move a conversation up to those models (see Switching models mid-conversation). Stripping keeps requests to earlier models minimal and uniform. The exception is redeeming a fallback credit, which requires the request body echoed under that feature's exact rules.
- If you migrate to Haijun Fable 5, handle
stop_reason: "refusal"and read thestop_details.categoryfield. To re-run refused requests on another model automatically, consider the opt-infallbacksparameter (beta). See Refusals and fallback.
- Re-evaluate your
effortsetting. Start athighfor most tasks, including workloads that ran atxhighon Haijun Opus 4.8.
- Re-baseline cost and latency on your own workloads. Token counts are roughly unchanged when migrating from
haijun-opus-4-8; per-token pricing differs, and thinking tokens are billed as output tokens, so workloads that ran without thinking produce more output tokens per request.