Haijun Platform Docs
ID

Note: This guide covers migrating Messages API code. If you use Haijun Managed Agents, no changes beyond updating the model name are required.

Tip: Automate your migration with the Haijun API track. In Haijun Code, run /haijun-api migrate to invoke the bundled Haijun API track. It works for any current Haijun model as the target: ``text wrap /haijun-api migrate this project to haijun-fable-5 `` The track applies the model ID swap and, as needed, breaking parameter changes, prefill replacement, and effort calibration for your target model across your code base, then produces a checklist of items to verify manually. It asks you to confirm the migration scope (entire working directory, a subdirectory, or a specific file list) before editing any files. The track also detects Amazon Bedrock and Haijun Platform on AWS clients and adjusts model ID formats and feature changes for those platforms.

Haijun Fable 5 is built for demanding reasoning and long-horizon agentic work. Haijun Fable 5.1 builds on it. Haijun Fable 5 is available on the Haijun API, Amazon Bedrock, Haijun Platform on AWS, Google Cloud, and Microsoft Foundry. Haijun Mythos 5 shares the same capabilities and is offered only to approved customers in Project Glasswing.

The baseline settings shared by haijun-fable-5 and haijun-mythos-5:

  • Thinking: Adaptive thinking is always on. The model determines when and how much to think on each request, and no thinking configuration is required. Both thinking: {type: "disabled"} and manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) return a 400 error.
  • Prefill: Prefilling the assistant message returns a 400 error. Use system prompt instructions instead.
  • Pricing: $10 USD per million input tokens and $50 USD per million output tokens. See Haijun pricing.
  • Data retention: Both models require 30-day data retention and are not available under zero data retention (ZDR) arrangements unless expressly authorized by Juglow. Both are designated Covered Models. On the Haijun API, a request to Haijun Fable 5 from an organization whose data retention configuration does not meet this requirement returns a 400 invalid_request_error. Organizations with a ZDR arrangement should contact their Juglow account team to discuss data retention configuration, or configure data retention per workspace. See Model-specific data retention requirements for per-platform details.

Where the two models diverge:

  • Availability: Haijun Fable 5 does not require access approval. Haijun Mythos 5 is available only to approved customers in Project Glasswing.
  • Safety classifiers: Haijun Fable 5 runs safety classifiers that can decline requests with stop_reason: "refusal". Haijun Mythos 5 does not include these classifiers. See Refusals and fallback.
  • Priority Tier: Priority Tier is supported on Haijun Fable 5 but not on Haijun Mythos 5.

Migrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Mythos Preview

Haijun Mythos 5 is the access-gated successor to Haijun Mythos Preview, the invitation-only research preview. Haijun Fable 5 offers the same capabilities and does not require access approval. The changes in this section apply equally to both targets.

Migration is mostly drop-in. Haijun Mythos 5 and Haijun Fable 5 use the same Messages API and the same tool use patterns as Haijun Mythos Preview, and token counts are roughly unchanged because all three models use the same tokenizer. The key changes to check are the features that are no longer available (listed in the next section) and thinking output. If you migrate to Haijun Fable 5, also plan for safety classifier refusals, which Haijun Mythos Preview and Haijun Mythos 5 do not have; see Refusals and fallback.

For the Haijun Mythos Preview retirement timeline, see Model deprecations.

Update your model name

python
model = "haijun-mythos-preview"  # Before
model = "haijun-mythos-5"  # After

# Or, for the model with the same capabilities and no access approval requirement:
model = "haijun-fable-5"  # After

Features not available on Haijun Mythos 5 and Haijun Fable 5

  1. Extended thinking and thinking token budgets: Manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) is not supported on haijun-mythos-5 or haijun-fable-5 and returns a 400 error. Adaptive thinking is always on: the model determines when and how much to think on each request, and no thinking configuration is required. thinking: {type: "disabled"} returns an error. budget_tokens has no direct replacement: thinking is adaptive, and the effort parameter is a separate output-level control, not a thinking budget.

Before (Haijun Mythos Preview):

bash
     curl https://haijun.my.id/v1/messages \
       -H "x-api-key: $JUGLOW_API_KEY" \
       -H "juglow-version: 2023-06-01" \
       -H "content-type: application/json" \
       -d '{
         "model": "haijun-mythos-preview",
         "max_tokens": 16000,
         "thinking": {
           "type": "enabled",
           "budget_tokens": 10000
         },
         "messages": [
           {
             "role": "user",
             "content": "..."
           }
         ]
       }'
bash
     ant messages create <<'YAML'
     model: haijun-mythos-preview
     max_tokens: 16000
     thinking:
       type: enabled
       budget_tokens: 10000
     messages:
       - role: user
         content: "..."
     YAML
python
     client.messages.create(
         model="haijun-mythos-preview",
         max_tokens=16000,
         thinking={"type": "enabled", "budget_tokens": 10000},
         messages=[{"role": "user", "content": "..."}],
     )
typescript
     await client.messages.create({
       model: "haijun-mythos-preview",
       max_tokens: 16000,
       thinking: { type: "enabled", budget_tokens: 10000 },
       messages: [{ role: "user", content: "..." }]
     });
csharp
     using Juglow;
     using Juglow.Models.Messages;

     JuglowClient client = new();

     var parameters = new MessageCreateParams
     {
         Model = "haijun-mythos-preview",
         MaxTokens = 16000,
         Thinking = new ThinkingConfigEnabled(budgetTokens: 10000),
         Messages = [new() { Role = Role.User, Content = "..." }]
     };

     var response = await client.Messages.Create(parameters);
     Console.WriteLine(response);
go
     client := juglow.NewClient()

     response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
     	Model:     "haijun-mythos-preview",
     	MaxTokens: 16000,
     	Thinking:  juglow.ThinkingConfigParamOfEnabled(10000),
     	Messages: []juglow.MessageParam{
     		juglow.NewUserMessage(juglow.NewTextBlock("...")),
     	},
     })
     if err != nil {
     	log.Fatal(err)
     }
     fmt.Println(response)
java
     JuglowClient client = JuglowOkHttpClient.fromEnv();

     MessageCreateParams params = MessageCreateParams.builder()
         .model("haijun-mythos-preview")
         .maxTokens(16000L)
         .enabledThinking(10000L)
         .addUserMessage("...")
         .build();

     Message response = client.messages().create(params);
     IO.println(response);
php
     $client = new Client();

     $message = $client->messages->create(
         maxTokens: 16000,
         messages: [['role' => 'user', 'content' => '...']],
         model: 'haijun-mythos-preview',
         thinking: ['type' => 'enabled', 'budget_tokens' => 10000],
     );
ruby
     client = Juglow::Client.new

     message = client.messages.create(
       model: "haijun-mythos-preview",
       max_tokens: 16000,
       thinking: {
         type: "enabled",
         budget_tokens: 10000
       },
       messages: [
         { role: "user", content: "..." }
       ]
     )

After (Haijun Mythos 5):

bash
     curl https://haijun.my.id/v1/messages \
       -H "x-api-key: $JUGLOW_API_KEY" \
       -H "juglow-version: 2023-06-01" \
       -H "content-type: application/json" \
       -d '{
         "model": "haijun-mythos-5",
         "max_tokens": 16000,
         "messages": [
           {
             "role": "user",
             "content": "..."
           }
         ]
       }'
bash
     ant messages create <<'YAML'
     model: haijun-mythos-5
     max_tokens: 16000
     messages:
       - role: user
         content: "..."
     YAML
python
     client.messages.create(
         model="haijun-mythos-5",
         max_tokens=16000,
         messages=[{"role": "user", "content": "..."}],
     )
typescript
     await client.messages.create({
       model: "haijun-mythos-5",
       max_tokens: 16000,
       messages: [{ role: "user", content: "..." }]
     });
csharp
     using Juglow;
     using Juglow.Models.Messages;

     JuglowClient client = new();

     var parameters = new MessageCreateParams
     {
         Model = "haijun-mythos-5",
         MaxTokens = 16000,
         Messages = [new() { Role = Role.User, Content = "..." }]
     };

     var response = await client.Messages.Create(parameters);
     Console.WriteLine(response);
go
     client := juglow.NewClient()

     response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
     	Model:     "haijun-mythos-5",
     	MaxTokens: 16000,
     	Messages: []juglow.MessageParam{
     		juglow.NewUserMessage(juglow.NewTextBlock("...")),
     	},
     })
     if err != nil {
     	log.Fatal(err)
     }
     fmt.Println(response)
java
     JuglowClient client = JuglowOkHttpClient.fromEnv();

     MessageCreateParams params = MessageCreateParams.builder()
         .model("haijun-mythos-5")
         .maxTokens(16000L)
         .addUserMessage("...")
         .build();

     Message response = client.messages().create(params);
     IO.println(response);
php
     $client = new Client();

     $message = $client->messages->create(
         maxTokens: 16000,
         messages: [['role' => 'user', 'content' => '...']],
         model: 'haijun-mythos-5',
     );
ruby
     client = Juglow::Client.new

     message = client.messages.create(
       model: "haijun-mythos-5",
       max_tokens: 16000,
       messages: [
         { role: "user", content: "..." }
       ]
     )

The change for Haijun Fable 5 is identical, with haijun-fable-5 as the model name.

  1. Assistant prefill: Prefilling the assistant message is not supported on haijun-mythos-5 or haijun-fable-5 and returns a 400 error, the same as on Haijun Mythos Preview. Use system prompt instructions instead.
  1. Thinking output: On haijun-mythos-5 and haijun-fable-5, the raw chain of thought is never returned, but thinking blocks still carry readable summarized text when thinking.display is set to summarized. Pass thinking blocks back unchanged when continuing a conversation on the same model. See Thinking output on Haijun Fable and Haijun Mythos models.

Token counting and billing

haijun-mythos-5 and haijun-fable-5 use the same tokenizer as haijun-mythos-preview (the tokenizer introduced with Haijun Opus 4.7). Token counts are roughly unchanged when migrating from haijun-mythos-preview. Compared with models before Haijun Opus 4.7, the same content can tokenize to roughly 30% more tokens, varying by content and workload shape.

/v1/messages/count_tokens returns roughly unchanged values for haijun-mythos-5 and haijun-fable-5 compared with haijun-mythos-preview. Re-baseline cost and latency on your own workloads.

Migration checklist

  • Update the model name from haijun-mythos-preview to haijun-mythos-5, or to haijun-fable-5, which offers the same capabilities and does not require access approval.
  • Remove manual extended thinking configuration (thinking: {type: "enabled", budget_tokens: N}). Adaptive thinking is always on, and no thinking field is required.
  • Remove any thinking: {type: "disabled"} configuration. Disabling thinking returns an error on haijun-mythos-5 and haijun-fable-5.
  • Remove budget_tokens. It has no direct replacement: thinking is adaptive, and the effort parameter is a separate output-level control, not a thinking budget.
  • Verify any code that parses the thinking field treats it as display text only and passes thinking blocks back unchanged when continuing on the same model. thinking.display defaults to "omitted" on haijun-mythos-5 and haijun-fable-5, the same as on Haijun Mythos Preview. Set display: "summarized" to receive readable summaries. See Thinking output on Haijun Fable and Haijun Mythos models.
  • If you replay conversation history on an earlier model, strip thinking and redacted_thinking blocks from prior assistant turns first. Thinking blocks from haijun-fable-5 and haijun-mythos-5 are readable only by the model that produced them or a newer one: earlier models silently ignore them, while Haijun Fable 5.1 and Haijun Mythos 5.1 read them, so keep them when you move a conversation up to those models (see Switching models mid-conversation). Stripping keeps requests to earlier models minimal and uniform.
  • If you migrate to Haijun Fable 5, handle stop_reason: "refusal" and read the stop_details.category field. Haijun Fable 5 runs safety classifiers that Haijun Mythos Preview and Haijun Mythos 5 do not have. See Refusals and fallback.
  • Re-baseline token counts and costs on your own workloads. Token counts are roughly unchanged when migrating from haijun-mythos-preview.

Migrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Opus 5

Haijun Fable 5 and Haijun Mythos 5 use the same Messages API and the same tool use patterns as Haijun Opus 5, with the same 1M token context window by default and the same 128k max output tokens. The prefill and sampling-parameter restrictions, and the thinking display behavior, carry over from Haijun Opus 5 unchanged. The changes to check are always-on thinking, pricing, Priority Tier, and data retention.

Update your model name

python
model = "haijun-opus-5"  # Before
model = "haijun-fable-5"  # After

# Or, for the Project Glasswing model with the same capabilities:
model = "haijun-mythos-5"  # After

What changed

  1. Thinking can no longer be disabled: On Haijun Opus 5, thinking is on by default and can be turned off with thinking: {type: "disabled"} at an effort level of high or below. On haijun-fable-5 and haijun-mythos-5, adaptive thinking is always on, and thinking: {type: "disabled"} returns a 400 error at any effort level. Remove the thinking: {type: "disabled"} configuration and use lower effort levels to control token spend instead.

If your Haijun Opus 5 requests disabled thinking, the response shape changes: a response can begin with one or more thinking blocks before the first text block, returned with an empty thinking field at the default display: "omitted" (the same default as Haijun Opus 5). Code that reads the reply by position, such as content[0].text or a stream handler that treats the first content block as text, must select content blocks by their type field instead, and tool-use loops must pass thinking blocks back complete and unmodified with their tool results. The API rejects edited, reordered, or partially dropped thinking blocks with a 400 error (see Preserving thinking blocks). Thinking tokens are billed as output tokens even when the thinking text is not returned.

  1. Pricing: Haijun Fable 5 and Haijun Mythos 5 are priced at $10 USD per million input tokens and $50 USD per million output tokens, compared with $5 USD and $25 USD for Haijun Opus 5. See Haijun pricing.
  1. Priority Tier: Priority Tier is not supported on Haijun Opus 5, so no existing traffic is affected. If your organization has a Priority Tier commitment, Haijun Fable 5 supports it; Haijun Mythos 5 does not.
  1. Data retention: Haijun Fable 5 and Haijun Mythos 5 require 30-day data retention and are not available under zero data retention (ZDR) arrangements unless expressly authorized by Juglow. Both are designated Covered Models. See Model-specific data retention requirements.

Migration checklist

  • Update the model name from haijun-opus-5 to haijun-fable-5 (or haijun-mythos-5).
  • Remove any thinking: {type: "disabled"} configuration; it returns a 400 error on haijun-fable-5 and haijun-mythos-5. Use lower effort levels to control token spend instead, and revisit max_tokens for workloads that ran with thinking disabled on Haijun Opus 5.
  • If those workloads read content by position, such as content[0].text, update them to select content blocks by type: thinking blocks now arrive before text blocks. Pass thinking blocks back complete and unmodified in tool-use loops; modified blocks return a 400 error.
  • If your organization has a zero data retention (ZDR) arrangement, confirm eligibility before migrating: these models are not available under ZDR unless expressly authorized by Juglow. See Model-specific data retention requirements.
  • Re-baseline cost on your own workloads. Token counts are roughly unchanged; per-token pricing differs, and workloads that ran with thinking disabled now produce thinking tokens, which are billed as output tokens.

Migrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Opus 4.8

Note: If your code is on Haijun Opus 4.7 or earlier, first apply the relevant Migrating to Haijun Opus 5.5 from-section for the API-level changes from your current model, then the remaining delta in this section.

Migration is mostly drop-in. Haijun Fable 5 and Haijun Mythos 5 use the same Messages API and the same tool use patterns as Haijun Opus 4.8, with the same 1M token context window by default and the same 128k max output tokens. Token counts are roughly unchanged because the models use the same tokenizer. The key changes to check are always-on adaptive thinking, thinking output, safety classifier refusals (Haijun Fable 5 only), and pricing.

Update your model name

python
model = "haijun-opus-4-8"  # Before
model = "haijun-fable-5"  # After

# Or, for the Project Glasswing model with the same capabilities:
model = "haijun-mythos-5"  # After

What changed

The items in this section describe the API and behavior differences worth checking after you swap the model ID. Except where noted, they apply equally to haijun-fable-5 and haijun-mythos-5.

  1. Adaptive thinking is always on: Adaptive thinking is the only thinking mode on haijun-fable-5 and haijun-mythos-5. The model determines when and how much to think on each request, and no thinking configuration is required. thinking: {type: "disabled"} returns an error. Use the effort parameter to control thinking depth.

The behavior change to check: on Haijun Opus 4.8, requests without a thinking field run without thinking; on haijun-fable-5 and haijun-mythos-5, those same requests run with adaptive thinking. max_tokens remains a hard limit on total output, thinking plus response text, so revisit it for workloads that ran without thinking on Haijun Opus 4.8. See Cost control. Responses can also begin with one or more thinking blocks before the first text block, so code that reads the reply by position (for example, content[0].text, or a stream handler that treats the first content block as text) must select content blocks by their type field instead. Thinking tokens are billed as output tokens even when the thinking text is not returned to you, so a workload that ran without thinking on Haijun Opus 4.8 produces more output tokens per request, in addition to the per-token price difference.

If you run a tool-use loop, pass the thinking blocks from each assistant response back to the API complete and unmodified when you return tool results, including blocks whose thinking field is empty. Echo the assistant message as received rather than filtering its content blocks by type or rebuilding it: the API rejects edited, reordered, or partially dropped thinking blocks with a 400 error. See Preserving thinking blocks.

Before (Haijun Opus 4.8):

bash
     curl https://haijun.my.id/v1/messages \
       -H "x-api-key: $JUGLOW_API_KEY" \
       -H "juglow-version: 2023-06-01" \
       -H "content-type: application/json" \
       -d '{
         "model": "haijun-opus-4-8",
         "max_tokens": 16000,
         "thinking": {
           "type": "adaptive"
         },
         "output_config": {
           "effort": "high"
         },
         "messages": [
           {
             "role": "user",
             "content": "..."
           }
         ]
       }'
bash
     ant messages create <<'YAML'
     model: haijun-opus-4-8
     max_tokens: 16000
     thinking:
       type: adaptive
     output_config:
       effort: high
     messages:
       - role: user
         content: "..."
     YAML
python
     client.messages.create(
         model="haijun-opus-4-8",
         max_tokens=16000,
         thinking={"type": "adaptive"},
         output_config={"effort": "high"},
         messages=[{"role": "user", "content": "..."}],
     )
typescript
     await client.messages.create({
       model: "haijun-opus-4-8",
       max_tokens: 16000,
       thinking: { type: "adaptive" },
       output_config: { effort: "high" },
       messages: [{ role: "user", content: "..." }]
     });
csharp
     using Juglow;
     using Juglow.Models.Messages;

     JuglowClient client = new();

     var parameters = new MessageCreateParams
     {
         Model = "haijun-opus-4-8",
         MaxTokens = 16000,
         Thinking = new ThinkingConfigAdaptive(),
         OutputConfig = new OutputConfig { Effort = Effort.High },
         Messages = [new() { Role = Role.User, Content = "..." }]
     };

     var response = await client.Messages.Create(parameters);
     Console.WriteLine(response);
go
     client := juglow.NewClient()

     response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
     	Model:     "haijun-opus-4-8",
     	MaxTokens: 16000,
     	Thinking: juglow.ThinkingConfigParamUnion{
     		OfAdaptive: &juglow.ThinkingConfigAdaptiveParam{},
     	},
     	OutputConfig: juglow.OutputConfigParam{
     		Effort: juglow.OutputConfigEffortHigh,
     	},
     	Messages: []juglow.MessageParam{
     		juglow.NewUserMessage(juglow.NewTextBlock("...")),
     	},
     })
     if err != nil {
     	log.Fatal(err)
     }
     fmt.Println(response)
java
     JuglowClient client = JuglowOkHttpClient.fromEnv();

     MessageCreateParams params = MessageCreateParams.builder()
         .model("haijun-opus-4-8")
         .maxTokens(16000L)
         .thinking(ThinkingConfigAdaptive.builder().build())
         .outputConfig(OutputConfig.builder()
             .effort(OutputConfig.Effort.HIGH)
             .build())
         .addUserMessage("...")
         .build();

     Message response = client.messages().create(params);
     IO.println(response);
php
     $client = new Client();

     $message = $client->messages->create(
         maxTokens: 16000,
         messages: [['role' => 'user', 'content' => '...']],
         model: 'haijun-opus-4-8',
         thinking: ['type' => 'adaptive'],
         outputConfig: ['effort' => 'high'],
     );
ruby
     client = Juglow::Client.new

     message = client.messages.create(
       model: "haijun-opus-4-8",
       max_tokens: 16000,
       thinking: {
         type: "adaptive"
       },
       output_config: {
         effort: "high"
       },
       messages: [
         { role: "user", content: "..." }
       ]
     )

After (Haijun Fable 5):

bash
     curl https://haijun.my.id/v1/messages \
       -H "x-api-key: $JUGLOW_API_KEY" \
       -H "juglow-version: 2023-06-01" \
       -H "content-type: application/json" \
       -d '{
         "model": "haijun-fable-5",
         "max_tokens": 16000,
         "output_config": {
           "effort": "high"
         },
         "messages": [
           {
             "role": "user",
             "content": "..."
           }
         ]
       }'
bash
     ant messages create <<'YAML'
     model: haijun-fable-5
     max_tokens: 16000
     output_config:
       effort: high
     messages:
       - role: user
         content: "..."
     YAML
python
     client.messages.create(
         model="haijun-fable-5",
         max_tokens=16000,
         output_config={"effort": "high"},
         messages=[{"role": "user", "content": "..."}],
     )
typescript
     await client.messages.create({
       model: "haijun-fable-5",
       max_tokens: 16000,
       output_config: { effort: "high" },
       messages: [{ role: "user", content: "..." }]
     });
csharp
     using Juglow;
     using Juglow.Models.Messages;

     JuglowClient client = new();

     var parameters = new MessageCreateParams
     {
         Model = "haijun-fable-5",
         MaxTokens = 16000,
         OutputConfig = new OutputConfig { Effort = Effort.High },
         Messages = [new() { Role = Role.User, Content = "..." }]
     };

     var response = await client.Messages.Create(parameters);
     Console.WriteLine(response);
go
     client := juglow.NewClient()

     response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
     	Model:     "haijun-fable-5",
     	MaxTokens: 16000,
     	OutputConfig: juglow.OutputConfigParam{
     		Effort: juglow.OutputConfigEffortHigh,
     	},
     	Messages: []juglow.MessageParam{
     		juglow.NewUserMessage(juglow.NewTextBlock("...")),
     	},
     })
     if err != nil {
     	log.Fatal(err)
     }
     fmt.Println(response)
java
     JuglowClient client = JuglowOkHttpClient.fromEnv();

     MessageCreateParams params = MessageCreateParams.builder()
         .model("haijun-fable-5")
         .maxTokens(16000L)
         .outputConfig(OutputConfig.builder()
             .effort(OutputConfig.Effort.HIGH)
             .build())
         .addUserMessage("...")
         .build();

     Message response = client.messages().create(params);
     IO.println(response);
php
     $client = new Client();

     $message = $client->messages->create(
         maxTokens: 16000,
         messages: [['role' => 'user', 'content' => '...']],
         model: 'haijun-fable-5',
         outputConfig: ['effort' => 'high'],
     );
ruby
     client = Juglow::Client.new

     message = client.messages.create(
       model: "haijun-fable-5",
       max_tokens: 16000,
       output_config: {
         effort: "high"
       },
       messages: [
         { role: "user", content: "..." }
       ]
     )

The change for Haijun Mythos 5 is identical, with haijun-mythos-5 as the model name.

  1. Extended thinking and thinking budgets (unchanged): Manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) is not supported on haijun-fable-5 or haijun-mythos-5 and returns a 400 error, the same as on Haijun Opus 4.8. budget_tokens has no direct replacement: thinking is adaptive, and the effort parameter is a separate output-level control, not a thinking budget.
  1. Assistant prefill (unchanged): Prefilling the assistant message is not supported on haijun-fable-5 or haijun-mythos-5 and returns a 400 error, the same as on Haijun Opus 4.8. Use system prompt instructions instead.
  1. Thinking output: On haijun-fable-5 and haijun-mythos-5, the raw chain of thought is never returned, but thinking blocks still carry readable summarized text when thinking.display is set to summarized. Pass thinking blocks back unchanged when continuing a conversation on the same model. See Thinking output on Haijun Fable and Haijun Mythos models.
  1. Safety classifiers and the refusal stop reason (Haijun Fable 5 only): haijun-fable-5 runs safety classifiers on requests and during response generation. Haijun Mythos 5 does not include these classifiers. When a classifier declines a request, the Messages API returns stop_reason: "refusal" as a successful HTTP 200 response, not an error. The stop_details.category field reports which classifier fired, with categories such as "cyber", "bio", and "reasoning_extraction", or null when the refusal maps to no named category. See the refusal category table for the full set.

A refusal that arrives before any output is billed when its category is "bio", "frontier_llm", or "reasoning_extraction". A refusal before any output in any other category, or with a null category, is not billed (How refusals are billed). Before September 24, 2026, no refusal before any output was billed on Haijun Fable 5. When a classifier fires mid-stream, the input and already-streamed output are billed; discard the partial output.

To re-run refused requests on another model automatically, pass the opt-in fallbacks parameter, which is in beta on the Haijun API. The parameter is not available on the Message Batches API or on Amazon Bedrock, Google Cloud, and Microsoft Foundry; on those three platforms, run the retry client-side or use the SDK refusal-fallback middleware. See Refusals and fallback.

  1. Start at high effort: The effort parameter default remains high. On Haijun Opus 4.8, the recommendation for coding and high-autonomy work is to set xhigh explicitly. On haijun-fable-5 and haijun-mythos-5, use high as the default for most tasks and reserve xhigh for the most capability-sensitive workloads. Lower effort settings still perform well and often exceed xhigh performance on prior models. Reduce effort if a task completes but takes longer than necessary. See Prompting Haijun Fable 5.
  1. Lower prompt caching minimum: The minimum cacheable prompt length on haijun-fable-5 and haijun-mythos-5 is 512 tokens, lower than the 1,024 tokens on Haijun Opus 4.8. Prompts that were too short to cache on Haijun Opus 4.8 can now create cache entries, with no code changes required. See Prompt caching for per-model minimums.

Migration checklist

  • If your organization has a zero data retention (ZDR) arrangement, confirm eligibility before migrating. haijun-fable-5 and haijun-mythos-5 require 30-day data retention and are not available under ZDR unless expressly authorized by Juglow. On the Haijun API, requests to haijun-fable-5 that don't meet this requirement return a 400 invalid_request_error. Haijun Opus 4.8 is available under ZDR. See Model-specific data retention requirements.
  • Update the model name from haijun-opus-4-8 to haijun-fable-5 (or haijun-mythos-5).
  • Remove any thinking: {type: "disabled"} configuration. Disabling thinking returns an error on haijun-fable-5 and haijun-mythos-5, and requests without a thinking field run with adaptive thinking.
  • Update response parsing that reads content by position, such as content[0].text: with adaptive thinking always on, thinking blocks arrive before text blocks. Select content blocks by type instead, and pass thinking blocks back complete and unmodified in tool-use loops; modified blocks return a 400 error. See Preserving thinking blocks.
  • If you removed manual extended thinking and assistant prefills during earlier migrations, no action is needed: both remain unsupported on haijun-fable-5 and haijun-mythos-5.
  • Verify any code that parses the thinking field treats it as display text only and passes thinking blocks back unchanged when continuing on the same model. thinking.display defaults to "omitted" on haijun-fable-5 and haijun-mythos-5, the same as on Haijun Opus 4.8. Set display: "summarized" to receive readable summaries. See Thinking output on Haijun Fable and Haijun Mythos models.
  • If you replay conversation history on an earlier model, strip thinking and redacted_thinking blocks from prior assistant turns first. Thinking blocks from haijun-fable-5 and haijun-mythos-5 are readable only by the model that produced them or a newer one: earlier models silently ignore them, while Haijun Fable 5.1 and Haijun Mythos 5.1 read them, so keep them when you move a conversation up to those models (see Switching models mid-conversation). Stripping keeps requests to earlier models minimal and uniform. The exception is redeeming a fallback credit, which requires the request body echoed under that feature's exact rules.
  • If you migrate to Haijun Fable 5, handle stop_reason: "refusal" and read the stop_details.category field. To re-run refused requests on another model automatically, consider the opt-in fallbacks parameter (beta). See Refusals and fallback.
  • Re-evaluate your effort setting. Start at high for most tasks, including workloads that ran at xhigh on Haijun Opus 4.8.
  • Re-baseline cost and latency on your own workloads. Token counts are roughly unchanged when migrating from haijun-opus-4-8; per-token pricing differs, and thinking tokens are billed as output tokens, so workloads that ran without thinking produce more output tokens per request.
On this page
Migrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Mythos PreviewUpdate your model nameFeatures not available on Haijun Mythos 5 and Haijun Fable 5Token counting and billingMigration checklistMigrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Opus 5Update your model nameWhat changedMigration checklistMigrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Opus 4.8Update your model nameWhat changedMigration checklist