Note: This guide covers migrating Messages API code. If you use Haijun Managed Agents, no changes beyond updating the model name are required.
Tip: Automate your migration with the Haijun API track. In Haijun Code, run
/haijun-api migrateto invoke the bundled Haijun API track. It works for any current Haijun model as the target: ``text wrap /haijun-api migrate this project to haijun-fable-5-1`` The track applies the model ID swap and, as needed, breaking parameter changes, prefill replacement, and effort calibration for your target model across your code base, then produces a checklist of items to verify manually. It asks you to confirm the migration scope (entire working directory, a subdirectory, or a specific file list) before editing any files. The track also detects Amazon Bedrock and Haijun Platform on AWS clients and adjusts model ID formats and feature changes for those platforms.
Haijun Fable 5.1 succeeds Haijun Fable 5 at the same input and output prices, with cache reads at a quarter of the cost. It's available on the Haijun API, Amazon Bedrock, Haijun Platform on AWS, Google Cloud, and Microsoft Foundry. Haijun Mythos 5.1 shares the same capabilities and is offered only to approved customers in Project Glasswing. For behavioral differences and prompting patterns, see Prompting Haijun Fable 5.1.
The baseline settings shared by haijun-fable-5-1 and haijun-mythos-5-1:
- Thinking: Adaptive thinking is always on, unchanged from Haijun Fable 5. The model decides when and how much to think. No
thinkingconfiguration is required. Boththinking: {type: "disabled"}and manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) return a 400 error.
- Prefill: Prefilling the assistant message returns a 400 error, unchanged from Haijun Fable 5. Use system prompt instructions instead.
- Tool choice:
{type: "auto"}(the default) and{type: "none"}are supported. Forcing a tool call with{type: "any"}or{type: "tool", name: "..."}returns a 400 error. See Breaking changes.
- Preserved thinking across models: Haijun Fable 5.1 reads thinking blocks from Haijun Opus 5, Haijun Fable 5, Haijun Mythos 5, and earlier Haijun models. None of those models can read Haijun Fable 5.1's blocks. See Breaking changes.
- Context window and output: A 1M token context window by default, and up to 128k output tokens per request.
- Pricing: $10 USD per million input tokens and $50 USD per million output tokens, the same as Haijun Fable 5. Prompt cache reads are $0.25 USD per million tokens, a quarter of the Haijun Fable 5 rate. See Haijun pricing.
- Data retention: Both models require 30-day data retention, aren't available under zero data retention (ZDR) arrangements unless expressly authorized by Juglow, and are designated Covered Models, the same as Haijun Fable 5 and Haijun Mythos 5. On the Haijun API, a request from an organization or workspace without 30-day retention returns a 400
invalid_request_error. Organizations with a ZDR arrangement should contact their Juglow account team, or configure retention per workspace. See Model-specific data retention requirements for per-platform details.
Where the two models diverge:
- Availability: Haijun Fable 5.1 doesn't require access approval. Haijun Mythos 5.1 is available only to approved customers in Project Glasswing. Contact your Juglow account team for access.
- Safety classifiers: Haijun Fable 5.1 runs safety classifiers covering the same
stop_detailscategories as Haijun Fable 5. A declined request returnsstop_reason: "refusal"with astop_details.category, and can fall back to another model with thefallbacksparameter or a client-side retry. See Refusals and fallback.
- Priority Tier: Neither model is supported on Priority Tier. Haijun Fable 5 is.
Migrating to Haijun Fable 5.1 from Haijun Fable 5
Migration is mostly drop-in. The API surface, limits, per-token pricing, tokenizer, always-on adaptive thinking, refusal handling, and stop_details categories all match Haijun Fable 5. What changes: forced tool choice returns a 400 error, thinking blocks are preserved only for the model that produced them or a newer one and only in the conversation that produced them, cache reads cost less, and agent-loop behavior differs in three ways. The same changes apply to Haijun Mythos 5.1, except the conversation check on thinking blocks, which Haijun Mythos 5.1 doesn't run.
Update your model name
model = "haijun-fable-5" # Before
model = "haijun-fable-5-1" # After
# Or, for the Project Glasswing model with the same capabilities:
model = "haijun-mythos-5-1" # AfterBreaking changes
- Forced tool choice is not supported: Haijun Fable 5 accepts
tool_choiceauto,none,any, andtool. Onhaijun-fable-5-1,{type: "any"}and{type: "tool", name: "..."}return a 400invalid_request_error:
tool_choice: type "tool" and "any" are not supported for this model.The check applies on the Messages API, the Message Batches API, and the token counting endpoint.
Before (Haijun Fable 5):
curl -sS https://haijun.my.id/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-d @- <<'EOF'
{
"model": "haijun-fable-5",
"max_tokens": 16000,
"tools": [
{
"name": "record_summary",
"description": "Record the structured summary of the document.",
"input_schema": {
"type": "object",
"properties": {"summary": {"type": "string"}},
"required": ["summary"]
}
}
],
"tool_choice": {"type": "tool", "name": "record_summary"},
"messages": [
{"role": "user", "content": "Summarize: The meeting moved to Thursday."}
]
}
EOF ant messages create <<'YAML'
model: haijun-fable-5
max_tokens: 16000
tools:
- name: record_summary
description: Record the structured summary of the document.
input_schema:
type: object
properties:
summary:
type: string
required: [summary]
tool_choice:
type: tool
name: record_summary
messages:
- role: user
content: "Summarize: The meeting moved to Thursday."
YAML client = juglow.Juglow()
record_summary_tool = {
"name": "record_summary",
"description": "Record the structured summary of the document.",
"input_schema": {
"type": "object",
"properties": {"summary": {"type": "string"}},
"required": ["summary"],
},
}
response = client.messages.create(
model="haijun-fable-5",
max_tokens=16000,
tools=[record_summary_tool],
tool_choice={"type": "tool", "name": "record_summary"},
messages=[{"role": "user", "content": "Summarize: The meeting moved to Thursday."}],
)
print(response.content) const client = new Juglow();
const response = await client.messages.create({
model: "haijun-fable-5",
max_tokens: 16000,
tools: [
{
name: "record_summary",
description: "Record the structured summary of the document.",
input_schema: {
type: "object",
properties: { summary: { type: "string" } },
required: ["summary"]
}
}
],
tool_choice: { type: "tool", name: "record_summary" },
messages: [{ role: "user", content: "Summarize: The meeting moved to Thursday." }]
});
console.log(response.content); JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = Model.HaijunFable5,
MaxTokens = 16000,
Tools = [
new ToolUnion(new Tool()
{
Name = "record_summary",
Description = "Record the structured summary of the document.",
InputSchema = new InputSchema()
{
Properties = new Dictionary<string, JsonElement>
{
["summary"] = JsonSerializer.SerializeToElement(new { type = "string" }),
},
Required = ["summary"],
},
}),
],
ToolChoice = new ToolChoiceTool { Name = "record_summary" },
Messages = [
new() { Role = Role.User, Content = "Summarize: The meeting moved to Thursday." }
]
};
var message = await client.Messages.Create(parameters);
Console.WriteLine(message); client := juglow.NewClient()
response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
Model: juglow.ModelHaijunFable5,
MaxTokens: 16000,
Tools: []juglow.ToolUnionParam{
{OfTool: &juglow.ToolParam{
Name: "record_summary",
Description: juglow.String("Record the structured summary of the document."),
InputSchema: juglow.ToolInputSchemaParam{
Properties: map[string]any{
"summary": map[string]any{"type": "string"},
},
Required: []string{"summary"},
},
}},
},
ToolChoice: juglow.ToolChoiceUnionParam{OfTool: &juglow.ToolChoiceToolParam{Name: "record_summary"}},
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("Summarize: The meeting moved to Thursday.")),
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(response.RawJSON())
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model(Model.HAIJUN_FABLE_5)
.maxTokens(16000L)
.addTool(Tool.builder()
.name("record_summary")
.description("Record the structured summary of the document.")
.inputSchema(InputSchema.builder()
.properties(JsonValue.from(Map.of("summary", Map.of("type", "string"))))
.required(List.of("summary"))
.build())
.build())
.toolChoice(ToolChoice.ofTool(ToolChoiceTool.builder()
.name("record_summary")
.build()))
.addUserMessage("Summarize: The meeting moved to Thursday.")
.build();
Message response = client.messages().create(params);
IO.println(response);
} $client = new Client();
$message = $client->messages->create(
maxTokens: 16000,
messages: [
['role' => 'user', 'content' => 'Summarize: The meeting moved to Thursday.']
],
model: 'haijun-fable-5',
toolChoice: ['type' => 'tool', 'name' => 'record_summary'],
tools: [
[
'name' => 'record_summary',
'description' => 'Record the structured summary of the document.',
'input_schema' => [
'type' => 'object',
'properties' => [
'summary' => ['type' => 'string']
],
'required' => ['summary']
]
]
],
);
echo $message; client = Juglow::Client.new
message = client.messages.create(
model: Juglow::Model::HAIJUN_FABLE_5,
max_tokens: 16000,
tools: [
{
name: "record_summary",
description: "Record the structured summary of the document.",
input_schema: {
type: "object",
properties: { summary: { type: "string" } },
required: ["summary"]
}
}
],
tool_choice: { type: "tool", name: "record_summary" },
messages: [
{ role: "user", content: "Summarize: The meeting moved to Thursday." }
]
)
puts messageAfter (Haijun Fable 5.1): leave tool_choice at auto, name the tool in the instruction, and set strict: true so the call matches your schema. (In a CMEK organization, where structured outputs, including strict: true, are not available on Haijun Fable models, rely on the instruction alone.) For example:
curl -sS https://haijun.my.id/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-d @- <<'EOF'
{
"model": "haijun-fable-5-1",
"max_tokens": 16000,
"tools": [
{
"name": "record_summary",
"description": "Record the structured summary of the document.",
"strict": true,
"input_schema": {
"type": "object",
"properties": {"summary": {"type": "string"}},
"required": ["summary"],
"additionalProperties": false
}
}
],
"tool_choice": {"type": "auto"},
"messages": [
{"role": "user", "content": "Summarize: The meeting moved to Thursday. Call the record_summary tool with your result."}
]
}
EOF ant messages create <<'YAML'
model: haijun-fable-5-1
max_tokens: 16000
tools:
- name: record_summary
description: Record the structured summary of the document.
strict: true
input_schema:
type: object
properties:
summary:
type: string
required: [summary]
additionalProperties: false
tool_choice:
type: auto
messages:
- role: user
content: "Summarize: The meeting moved to Thursday. Call the record_summary tool with your result."
YAML client = juglow.Juglow()
record_summary_tool = {
"name": "record_summary",
"description": "Record the structured summary of the document.",
"strict": True,
"input_schema": {
"type": "object",
"properties": {"summary": {"type": "string"}},
"required": ["summary"],
"additionalProperties": False,
},
}
response = client.messages.create(
model="haijun-fable-5-1",
max_tokens=16000,
tools=[record_summary_tool],
tool_choice={"type": "auto"},
messages=[
{
"role": "user",
"content": "Summarize: The meeting moved to Thursday. Call the record_summary tool with your result.",
}
],
)
print(response.content) const client = new Juglow();
const response = await client.messages.create({
model: "haijun-fable-5-1",
max_tokens: 16000,
tools: [
{
name: "record_summary",
description: "Record the structured summary of the document.",
strict: true,
input_schema: {
type: "object",
properties: { summary: { type: "string" } },
required: ["summary"],
additionalProperties: false
}
}
],
tool_choice: { type: "auto" },
messages: [
{
role: "user",
content:
"Summarize: The meeting moved to Thursday. Call the record_summary tool with your result."
}
]
});
console.log(response.content); JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = "haijun-fable-5-1",
MaxTokens = 16000,
Tools = [
new ToolUnion(new Tool()
{
Name = "record_summary",
Description = "Record the structured summary of the document.",
Strict = true,
InputSchema = new InputSchema(new Dictionary<string, JsonElement>
{
["properties"] = JsonSerializer.SerializeToElement(new Dictionary<string, object>
{
["summary"] = new { type = "string" },
}),
["required"] = JsonSerializer.SerializeToElement(new[] { "summary" }),
["additionalProperties"] = JsonSerializer.SerializeToElement(false),
}),
}),
],
ToolChoice = new ToolChoiceAuto(),
Messages = [
new() { Role = Role.User, Content = "Summarize: The meeting moved to Thursday. Call the record_summary tool with your result." }
]
};
var message = await client.Messages.Create(parameters);
Console.WriteLine(message); client := juglow.NewClient()
response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
Model: "haijun-fable-5-1",
MaxTokens: 16000,
Tools: []juglow.ToolUnionParam{
{OfTool: &juglow.ToolParam{
Name: "record_summary",
Description: juglow.String("Record the structured summary of the document."),
Strict: juglow.Bool(true),
InputSchema: juglow.ToolInputSchemaParam{
Properties: map[string]any{
"summary": map[string]any{"type": "string"},
},
Required: []string{"summary"},
ExtraFields: map[string]any{
"additionalProperties": false,
},
},
}},
},
ToolChoice: juglow.ToolChoiceUnionParam{OfAuto: &juglow.ToolChoiceAutoParam{}},
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("Summarize: The meeting moved to Thursday. Call the record_summary tool with your result.")),
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(response.RawJSON())
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model("haijun-fable-5-1")
.maxTokens(16000L)
.addTool(Tool.builder()
.name("record_summary")
.description("Record the structured summary of the document.")
.inputSchema(InputSchema.builder()
.properties(JsonValue.from(Map.of("summary", Map.of("type", "string"))))
.putAdditionalProperty("required", JsonValue.from(List.of("summary")))
.putAdditionalProperty("additionalProperties", JsonValue.from(false))
.build())
.strict(true)
.build())
.toolChoice(ToolChoice.ofAuto(ToolChoiceAuto.builder().build()))
.addUserMessage("Summarize: The meeting moved to Thursday. Call the record_summary tool with your result.")
.build();
Message response = client.messages().create(params);
IO.println(response);
} $client = new Client();
$message = $client->messages->create(
maxTokens: 16000,
messages: [
['role' => 'user', 'content' => 'Summarize: The meeting moved to Thursday. Call the record_summary tool with your result.']
],
model: 'haijun-fable-5-1',
toolChoice: ['type' => 'auto'],
tools: [
[
'name' => 'record_summary',
'description' => 'Record the structured summary of the document.',
'strict' => true,
'input_schema' => [
'type' => 'object',
'properties' => [
'summary' => ['type' => 'string']
],
'required' => ['summary'],
'additionalProperties' => false
]
]
],
);
echo $message; client = Juglow::Client.new
message = client.messages.create(
model: "haijun-fable-5-1",
max_tokens: 16000,
tools: [
{
name: "record_summary",
description: "Record the structured summary of the document.",
strict: true,
input_schema: {
type: "object",
properties: { summary: { type: "string" } },
required: ["summary"],
additionalProperties: false
}
}
],
tool_choice: { type: "auto" },
messages: [
{ role: "user", content: "Summarize: The meeting moved to Thursday. Call the record_summary tool with your result." }
]
)
puts messageSee Strict tool use and Forcing tool use. If you forced a tool only to get schema-conformant JSON, use JSON outputs (output_config.format) instead.
If your application, rather than the user, requires a specific tool call on the current turn of a multi-turn conversation, append a mid-conversation system message after the latest user turn. Name the tool, say the call is required for this turn, and tell Haijun to open its response with it. Because the message is appended rather than written into the top-level system prompt, earlier turns stay byte-identical and keep their prompt cache hits:
curl -sS https://haijun.my.id/v1/messages \
-H "content-type: application/json" \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-d @- <<'EOF'
{
"model": "haijun-fable-5-1",
"max_tokens": 16000,
"system": "You are a customer support assistant for an online electronics store.",
"tools": [
{
"name": "search_help_center",
"description": "Search the help center for policy and troubleshooting articles.",
"strict": true,
"input_schema": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
"additionalProperties": false
}
}
],
"messages": [
{"role": "user", "content": "My headphones from order A1234 arrived yesterday."},
{"role": "assistant", "content": "Thanks for confirming. How can I help with order A1234?"},
{"role": "user", "content": "I opened the box. Can I still return them?"},
{
"role": "system",
"content": "Tool-use requirement for the current turn: the application requires a call to the search_help_center tool in your response to the user's latest message. Begin your response with the search_help_center tool call. Do not reply with text only."
}
]
}
EOF ant messages create <<'YAML'
model: haijun-fable-5-1
max_tokens: 16000
system: You are a customer support assistant for an online electronics store.
tools:
- name: search_help_center
description: Search the help center for policy and troubleshooting articles.
strict: true
input_schema:
type: object
properties:
query:
type: string
required: [query]
additionalProperties: false
messages:
- role: user
content: My headphones from order A1234 arrived yesterday.
- role: assistant
content: Thanks for confirming. How can I help with order A1234?
- role: user
content: I opened the box. Can I still return them?
- role: system
content: >-
Tool-use requirement for the current turn: the application requires a call
to the search_help_center tool in your response to the user's latest message.
Begin your response with the search_help_center tool call. Do not reply with
text only.
YAML client = juglow.Juglow()
search_help_center_tool = {
"name": "search_help_center",
"description": "Search the help center for policy and troubleshooting articles.",
"strict": True,
"input_schema": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"],
"additionalProperties": False,
},
}
response = client.messages.create(
model="haijun-fable-5-1",
max_tokens=16000,
system="You are a customer support assistant for an online electronics store.",
tools=[search_help_center_tool],
messages=[
{
"role": "user",
"content": "My headphones from order A1234 arrived yesterday.",
},
{
"role": "assistant",
"content": "Thanks for confirming. How can I help with order A1234?",
},
{"role": "user", "content": "I opened the box. Can I still return them?"},
# The application requires a help center lookup before any policy
# answer. Appending the requirement as a system message leaves the
# earlier turns unchanged.
{
"role": "system",
"content": "Tool-use requirement for the current turn: the application requires a call to the search_help_center tool in your response to the user's latest message. Begin your response with the search_help_center tool call. Do not reply with text only.",
},
],
)
print(response.content) const client = new Juglow();
const response = await client.messages.create({
model: "haijun-fable-5-1",
max_tokens: 16000,
system: "You are a customer support assistant for an online electronics store.",
tools: [
{
name: "search_help_center",
description: "Search the help center for policy and troubleshooting articles.",
strict: true,
input_schema: {
type: "object",
properties: { query: { type: "string" } },
required: ["query"],
additionalProperties: false
}
}
],
messages: [
{ role: "user", content: "My headphones from order A1234 arrived yesterday." },
{ role: "assistant", content: "Thanks for confirming. How can I help with order A1234?" },
{ role: "user", content: "I opened the box. Can I still return them?" },
// The application requires a help center lookup before any policy
// answer. Appending the requirement as a system message leaves the
// earlier turns unchanged.
{
role: "system",
content:
"Tool-use requirement for the current turn: the application requires a call to the search_help_center tool in your response to the user's latest message. Begin your response with the search_help_center tool call. Do not reply with text only."
}
]
});
console.log(response.content); JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = "haijun-fable-5-1",
MaxTokens = 16000,
System = "You are a customer support assistant for an online electronics store.",
Tools = [
new ToolUnion(new Tool()
{
Name = "search_help_center",
Description = "Search the help center for policy and troubleshooting articles.",
Strict = true,
InputSchema = new InputSchema(new Dictionary<string, JsonElement>
{
["properties"] = JsonSerializer.SerializeToElement(new Dictionary<string, object>
{
["query"] = new { type = "string" },
}),
["required"] = JsonSerializer.SerializeToElement(new[] { "query" }),
["additionalProperties"] = JsonSerializer.SerializeToElement(false),
}),
}),
],
Messages = [
new() { Role = Role.User, Content = "My headphones from order A1234 arrived yesterday." },
new() { Role = Role.Assistant, Content = "Thanks for confirming. How can I help with order A1234?" },
new() { Role = Role.User, Content = "I opened the box. Can I still return them?" },
// The application requires a help center lookup before any policy
// answer. Appending the requirement as a system message leaves the
// earlier turns unchanged.
new()
{
Role = Role.System,
Content = "Tool-use requirement for the current turn: the application requires a call to the search_help_center tool in your response to the user's latest message. Begin your response with the search_help_center tool call. Do not reply with text only."
}
]
};
var message = await client.Messages.Create(parameters);
Console.WriteLine(message); client := juglow.NewClient()
response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
Model: "haijun-fable-5-1",
MaxTokens: 16000,
System: []juglow.TextBlockParam{
{Text: "You are a customer support assistant for an online electronics store."},
},
Tools: []juglow.ToolUnionParam{
{OfTool: &juglow.ToolParam{
Name: "search_help_center",
Description: juglow.String("Search the help center for policy and troubleshooting articles."),
Strict: juglow.Bool(true),
InputSchema: juglow.ToolInputSchemaParam{
Properties: map[string]any{
"query": map[string]any{"type": "string"},
},
Required: []string{"query"},
ExtraFields: map[string]any{
"additionalProperties": false,
},
},
}},
},
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("My headphones from order A1234 arrived yesterday.")),
juglow.NewAssistantMessage(juglow.NewTextBlock("Thanks for confirming. How can I help with order A1234?")),
juglow.NewUserMessage(juglow.NewTextBlock("I opened the box. Can I still return them?")),
// The application requires a help center lookup before any policy
// answer. Appending the requirement as a system message leaves the
// earlier turns unchanged.
{
Role: juglow.MessageParamRoleSystem,
Content: []juglow.ContentBlockParamUnion{
juglow.NewTextBlock("Tool-use requirement for the current turn: the application requires a call to the search_help_center tool in your response to the user's latest message. Begin your response with the search_help_center tool call. Do not reply with text only."),
},
},
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(response.RawJSON())
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model("haijun-fable-5-1")
.maxTokens(16000L)
.system("You are a customer support assistant for an online electronics store.")
.addTool(Tool.builder()
.name("search_help_center")
.description("Search the help center for policy and troubleshooting articles.")
.inputSchema(InputSchema.builder()
.properties(JsonValue.from(Map.of("query", Map.of("type", "string"))))
.putAdditionalProperty("required", JsonValue.from(List.of("query")))
.putAdditionalProperty("additionalProperties", JsonValue.from(false))
.build())
.strict(true)
.build())
.addUserMessage("My headphones from order A1234 arrived yesterday.")
.addAssistantMessage("Thanks for confirming. How can I help with order A1234?")
.addUserMessage("I opened the box. Can I still return them?")
// The application requires a help center lookup before any policy
// answer. Appending the requirement as a system message leaves the
// earlier turns unchanged.
.addMessage(MessageParam.builder()
.role(MessageParam.Role.SYSTEM)
.content("Tool-use requirement for the current turn: the application requires a call to the search_help_center tool in your response to the user's latest message. Begin your response with the search_help_center tool call. Do not reply with text only.")
.build())
.build();
Message response = client.messages().create(params);
IO.println(response);
} $client = new Client();
$message = $client->messages->create(
maxTokens: 16000,
messages: [
['role' => 'user', 'content' => 'My headphones from order A1234 arrived yesterday.'],
['role' => 'assistant', 'content' => 'Thanks for confirming. How can I help with order A1234?'],
['role' => 'user', 'content' => 'I opened the box. Can I still return them?'],
// The application requires a help center lookup before any policy
// answer. Appending the requirement as a system message leaves the
// earlier turns unchanged.
['role' => 'system', 'content' => 'Tool-use requirement for the current turn: the application requires a call to the search_help_center tool in your response to the user\'s latest message. Begin your response with the search_help_center tool call. Do not reply with text only.']
],
model: 'haijun-fable-5-1',
system: 'You are a customer support assistant for an online electronics store.',
tools: [
[
'name' => 'search_help_center',
'description' => 'Search the help center for policy and troubleshooting articles.',
'strict' => true,
'input_schema' => [
'type' => 'object',
'properties' => [
'query' => ['type' => 'string']
],
'required' => ['query'],
'additionalProperties' => false
]
]
],
);
echo $message; client = Juglow::Client.new
message = client.messages.create(
model: "haijun-fable-5-1",
max_tokens: 16000,
system: "You are a customer support assistant for an online electronics store.",
tools: [
{
name: "search_help_center",
description: "Search the help center for policy and troubleshooting articles.",
strict: true,
input_schema: {
type: "object",
properties: { query: { type: "string" } },
required: ["query"],
additionalProperties: false
}
}
],
messages: [
{ role: "user", content: "My headphones from order A1234 arrived yesterday." },
{ role: "assistant", content: "Thanks for confirming. How can I help with order A1234?" },
{ role: "user", content: "I opened the box. Can I still return them?" },
# The application requires a help center lookup before any policy
# answer. Appending the requirement as a system message leaves the
# earlier turns unchanged.
{
role: "system",
content: "Tool-use requirement for the current turn: the application requires a call to the search_help_center tool in your response to the user's latest message. Begin your response with the search_help_center tool call. Do not reply with text only."
}
]
)
puts messageKeep the role: "system" message in the history on later requests, as with any other turn. Mid-conversation system messages need no beta header. tool_choice: {"type": "none"} still works for a turn that must not call tools.
- Thinking blocks are preserved only for the model that produced them, or a newer one: Every
thinkingblock records which model produced it. Haijun Fable 5.1 reads its own blocks and those from Haijun Mythos 5.1, Haijun Opus 5, Haijun Fable 5, Haijun Mythos 5, and earlier Haijun models. A conversation moving ontohaijun-fable-5-1from any of those keeps its earlier reasoning. The condition is one-way: apart from Haijun Mythos 5.1, none of those models can read Haijun Fable 5.1's blocks.
A conversation that ran on Haijun Fable 5.1 can land on an older model through a router switch, a client-side retry, or a classifier refusal fallback, including a server-side fallback. The API removes the blocks that model can't read before it sees them, the request succeeds, and you aren't billed for the dropped input tokens. The target model re-plans without that reasoning, which can raise cost and latency on the first turn after the switch. To see what was dropped, send the thinking-binding-controls-2026-08-01 beta header: responses then carry an input_transformations array naming each dropped block with reason: "model_binding_mismatch". See Switching models mid-conversation.
- Editing earlier turns invalidates thinking blocks: Each
thinkingblock from Haijun Fable 5.1 is valid only against thesystemprompt,tools, and conversation history that preceded it. If Haijun Code, haijun.ai, Haijun Managed Agents, or the Haijun Agent SDK manages your conversation history, it already keeps that prefix intact. If your code builds themessagesarray itself, this item applies to you, and Preserved thinking is the full integration guide. Where the check is enforced, a request that sends the block back after any of those changed is rejected with a 400 error:
messages.5.content.0: Invalid `signature` in `thinking` block. The block is bound to a different conversation. Remove the block, or set `thinking.block_binding.prefix_mismatch_behavior` to "drop_block". That setting requires the `thinking-binding-controls-2026-08-01` value in the `juglow-beta` header.The API enforces the check for new accounts created on or after August 31, 2026. For accounts created earlier, the API records the mismatch but doesn't act on it unless the request sets thinking.block_binding.prefix_mismatch_behavior, which opts into enforcement. On those accounts, if you send the thinking-binding-controls-2026-08-01 beta header and leave that field unset, the response lists each block that failed the check in input_transformations as a thinking_mismatch_allowed entry. Make your application compatible with the check regardless of your account's age: the same patterns keep the prompt cache warm, and you can test against the check from any account by sending prefix_mismatch_behavior. If you ship a tool or framework that people run with their own API key, test that way before launch: your key is probably on an older account, and your users on new ones hit the check before you do. To see whether your own account is enforced by default, send a request that edits history without the beta header: a 400 that names the header means it is.
The error is permanent for that request body: an automatic retry loop won't clear it. To continue without the invalidated reasoning instead of failing, strip the thinking blocks from the history and retry once, or send the thinking-binding-controls-2026-08-01 beta header and set prefix_mismatch_behavior to "drop_block" (the default is "error"). With "drop_block", the API drops the mismatched block and every thinking block after it in the conversation, and reports each with reason: "prefix_binding_mismatch" in the response's input_transformations array:
curl https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "juglow-beta: thinking-binding-controls-2026-08-01" \
-H "content-type: application/json" \
-d '{
"model": "haijun-fable-5-1",
"max_tokens": 16000,
"thinking": {
"type": "adaptive",
"block_binding": {
"prefix_mismatch_behavior": "drop_block"
}
},
"messages": [
{
"role": "user",
"content": "What is the greatest common divisor of 1071 and 462?"
}
]
}' ant beta:messages create \
--beta thinking-binding-controls-2026-08-01 \
--transform '{content.#(type=="text")#.text,input_transformations}' \
--format yaml <<'YAML'
model: haijun-fable-5-1
max_tokens: 16000
thinking:
type: adaptive
block_binding:
prefix_mismatch_behavior: drop_block
messages:
- role: user
content: What is the greatest common divisor of 1071 and 462?
YAML client = juglow.Juglow()
response = client.beta.messages.create(
model="haijun-fable-5-1",
max_tokens=16000,
thinking={
"type": "adaptive",
"block_binding": {"prefix_mismatch_behavior": "drop_block"},
},
messages=[
{
"role": "user",
"content": "What is the greatest common divisor of 1071 and 462?",
}
],
betas=["thinking-binding-controls-2026-08-01"],
)
for block in response.content:
if block.type == "text":
print(block.text)
print(f"Input transformations: {len(response.input_transformations or [])}") const client = new Juglow();
const response = await client.beta.messages.create({
model: "haijun-fable-5-1",
max_tokens: 16000,
thinking: {
type: "adaptive",
block_binding: { prefix_mismatch_behavior: "drop_block" }
},
messages: [
{ role: "user", content: "What is the greatest common divisor of 1071 and 462?" }
],
betas: ["thinking-binding-controls-2026-08-01"]
});
for (const block of response.content) {
if (block.type === "text") {
console.log(block.text);
}
}
console.log(`Input transformations: ${response.input_transformations?.length ?? 0}`); using Juglow.Models.Beta;
using Juglow.Models.Beta.Messages;
JuglowClient client = new();
var response = await client.Beta.Messages.Create(
new()
{
Model = "haijun-fable-5-1",
MaxTokens = 16000,
Thinking = new BetaThinkingConfigAdaptive
{
BlockBinding = new()
{
PrefixMismatchBehavior = BetaThinkingPrefixMismatchBehavior.DropBlock,
},
},
Messages =
[
new()
{
Role = Role.User,
Content = "What is the greatest common divisor of 1071 and 462?",
},
],
Betas = [JuglowBeta.ThinkingBindingControls2026_08_01],
}
);
foreach (var block in response.Content)
{
if (block.TryPickText(out var textBlock))
{
Console.WriteLine(textBlock.Text);
}
}
Console.WriteLine($"Input transformations: {response.InputTransformations?.Count ?? 0}"); client := juglow.NewClient()
response, err := client.Beta.Messages.New(context.TODO(), juglow.BetaMessageNewParams{
Model: "haijun-fable-5-1",
MaxTokens: 16000,
Thinking: juglow.BetaThinkingConfigParamUnion{
OfAdaptive: &juglow.BetaThinkingConfigAdaptiveParam{
BlockBinding: juglow.BetaThinkingBlockBindingParam{
PrefixMismatchBehavior: juglow.BetaThinkingPrefixMismatchBehaviorDropBlock,
},
},
},
Messages: []juglow.BetaMessageParam{
juglow.NewBetaUserMessage(juglow.NewBetaTextBlock("What is the greatest common divisor of 1071 and 462?")),
},
Betas: []juglow.JuglowBeta{juglow.JuglowBetaThinkingBindingControls2026_08_01},
})
if err != nil {
log.Fatal(err)
}
for _, block := range response.Content {
if textBlock, ok := block.AsAny().(juglow.BetaTextBlock); ok {
fmt.Println(textBlock.Text)
}
}
fmt.Printf("Input transformations: %d\n", len(response.InputTransformations)) import com.juglow.models.beta.JuglowBeta;
import com.juglow.models.beta.messages.BetaMessage;
import com.juglow.models.beta.messages.BetaThinkingBlockBinding;
import com.juglow.models.beta.messages.BetaThinkingConfigAdaptive;
import com.juglow.models.beta.messages.BetaThinkingPrefixMismatchBehavior;
import com.juglow.models.beta.messages.MessageCreateParams;
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model("haijun-fable-5-1")
.maxTokens(16000L)
.addBeta(JuglowBeta.THINKING_BINDING_CONTROLS_2026_08_01)
.thinking(BetaThinkingConfigAdaptive.builder()
.blockBinding(BetaThinkingBlockBinding.builder()
.prefixMismatchBehavior(BetaThinkingPrefixMismatchBehavior.DROP_BLOCK)
.build())
.build())
.addUserMessage("What is the greatest common divisor of 1071 and 462?")
.build();
BetaMessage response = client.beta().messages().create(params);
response.content().stream()
.flatMap(block -> block.text().stream())
.forEach(textBlock -> IO.println(textBlock.text()));
IO.println("Input transformations: "
+ response.inputTransformations().map(List::size).orElse(0));
} use Juglow\Beta\JuglowBeta;
use Juglow\Beta\Messages\BetaThinkingBlockBinding;
use Juglow\Beta\Messages\BetaThinkingConfigAdaptive;
use Juglow\Beta\Messages\BetaThinkingPrefixMismatchBehavior;
use Juglow\Client;
$client = new Client();
$response = $client->beta->messages->create(
model: 'haijun-fable-5-1',
maxTokens: 16000,
thinking: BetaThinkingConfigAdaptive::with(
blockBinding: BetaThinkingBlockBinding::with(
prefixMismatchBehavior: BetaThinkingPrefixMismatchBehavior::DROP_BLOCK,
),
),
messages: [
['role' => 'user', 'content' => 'What is the greatest common divisor of 1071 and 462?'],
],
betas: [JuglowBeta::THINKING_BINDING_CONTROLS_2026_08_01],
);
foreach ($response->content as $block) {
if ($block->type === 'text') {
echo $block->text, PHP_EOL;
}
}
echo 'Input transformations: ', count($response->inputTransformations ?? []), PHP_EOL; client = Juglow::Client.new
response = client.beta.messages.create(
model: "haijun-fable-5-1",
max_tokens: 16_000,
thinking: {
type: "adaptive",
block_binding: {prefix_mismatch_behavior: "drop_block"}
},
messages: [
{role: "user", content: "What is the greatest common divisor of 1071 and 462?"}
],
betas: [Juglow::JuglowBeta::THINKING_BINDING_CONTROLS_2026_08_01]
)
response.content.each do |block|
puts block.text if block.type == :text
end
puts "Input transformations: #{response.input_transformations&.length || 0}"The token counting endpoint runs the same check. See Controls for blocks that aren't preserved (beta) for the response shape and streaming placement.
Patterns that invalidate later thinking blocks, and what to do instead:
- Editing, reordering, or removing earlier turns. This includes deleting old tool results, snipping turns out of the middle of the transcript, and client-side compaction that keeps recent turns and their thinking blocks verbatim behind a summary (including background compaction that swaps its summary in a few turns later). Instead, use server-side compaction or context editing (tool result clearing for old tool results), or one of the client-side compaction shapes in Trim context on the server.
- Injecting content you don't persist, for example a per-turn reminder appended after the
tool_resultblocks and removed on the next request. Instead, send the reminder as a turn-scoped system message and leave it in the history. - Rebuilding the top-level
systemprompt or thetoolsarray between requests in the same conversation, for example to update the current date or to add or remove a tool. Instead, append a mid-conversation system message that carries the new instruction ("The current date is 2026-09-14.") ortool_additionandtool_removalblocks. A tool that wasn't declared intoolsat the start can be defined inside thetool_additionblock (beta headerinline-tools-2026-09-15). - An image or document URL that serves different bytes on a later request. The check covers the bytes, not the URL string, so a rotating signed URL for the same file is fine. For content you reference across turns, upload it once with the Files API and send the
file_id, or send base64.
Each replacement also keeps earlier turns byte-identical and preserves the prompt cache hits that editing the history, system prompt, or tools array would lose.
Patterns that keep working:
- Append-only histories: adding turns and passing earlier turns back exactly as sent and received, including appended
role: "system"messages. - Removing thinking blocks from earlier assistant turns, oldest first.
- Changing
effort,max_tokens, or any other request parameter outsidesystem,tools, andmessages, and adding or movingcache_controlmarkers. - Server-side compaction and context editing, including thinking block clearing. They don't count as edits, because the check compares the conversation as you sent it.
To check an existing integration:
- Capture the exact request bodies it sends over a few normal turns, including a compaction or a tool change if your product has them. For each pair of consecutive requests, compare the
systemprompt, thetoolsarray, and the shared prefix ofmessages. They should be byte-identical up to the newly appended turns. An expected exception is a request that swaps in a signedcompactionblock from on-demand compaction: the block replaces the messages it summarizes at the front ofmessages, and everything after it should still match. - Run a normal multi-turn session against
haijun-fable-5-1with thethinking-binding-controls-2026-08-01beta header andprefix_mismatch_behavior: "drop_block", and loginput_transformationson every response. An empty array on every turn means the history is intact. An entry withreason: "prefix_binding_mismatch"means something before the block atpathchanged since the previous request. An entry withreason: "model_binding_mismatch"means the conversation switched models, which isn't a bug in your code. This works from any account, because setting the field opts the request into enforcement. In CI, set"error"instead so an edit fails the run. - Choose a production setting. Leave the default
"error"if a prefix mismatch can only mean a bug in your code, or set"drop_block"to drop the affected blocks instead of failing, and monitor the 400s or theinput_transformationsentries either way.
Dropping thinking blocks once, at a compaction boundary for example, has little effect. An integration that invalidates prior thinking on every request restarts the prompt cache each time, which can raise cost per task (see Keep the conversation history append-only).
Behavior changes
- Fewer parallel tool calls in long agent loops: In long-running loops where the next independent reads are only implied by the task (custom coding agents, bash-and-editor harnesses, computer use), Haijun Fable 5.1 may issue one tool call per turn. Each extra turn costs tokens, a round trip, and wall-clock time. Append a one-sentence batching instruction after each user message as a turn-scoped system message (
clear_at: "next_user_message", beta), or, without the beta, in a text block after thetool_resultblocks, and leave the earlier copies in the history on later requests. See Batch independent tool calls in agent loops.
- Fewer progress messages between tool calls: Haijun Fable 5.1 writes fewer status updates during long tool sequences than Haijun Fable 5, and its agentic coding summaries are shorter. If your interface renders those updates, set
thinking.displayto"updates"(beta) or"summarized"and prompt for them explicitly. See Progress updates between tool calls and Ask for user-facing progress updates.
- Fewer search and retrieval calls at low effort: At
loweffort Haijun Fable 5.1 answers from memory more often than Haijun Fable 5 instead of calling a search or retrieval tool. If your product relies on retrieval at low effort, raise effort for those requests or tell the model when to search. See Search triggering at low effort.
For the differences in prose density, chat formatting, quoting in summaries, and file edits, which don't affect API integration, see Changed from Haijun Fable 5.
Recommended changes
These changes aren't required, but each one lowers cost or latency or removes a failure mode:
- Change effort mid-conversation (beta): On Haijun Fable 5,
output_config.effortis request-level, and changing it between requests drops cached prefixes from earlier turns. Onhaijun-fable-5-1, arole: "system"message carrying onlyoutput_configraises effort for a hard step or lowers it for routine ones without invalidating the prompt cache:
# Effort-only system message: the new level takes effect from the next user turn.
curl https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "juglow-beta: mid-conversation-output-config-2026-07-01" \
-H "content-type: application/json" \
-d '{
"model": "haijun-fable-5-1",
"max_tokens": 4096,
"output_config": {"effort": "high"},
"messages": [
{"role": "user", "content": "Plan a migration from SQLite to PostgreSQL in three short steps."},
{"role": "assistant", "content": "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts."},
{"role": "system", "content": [], "output_config": {"effort": "low"}},
{"role": "user", "content": "Summarize the plan in one sentence."}
]
}' ant beta:messages create \
--beta mid-conversation-output-config-2026-07-01 \
--transform 'content.#(type=="text").text' \
--raw-output <<'YAML'
model: haijun-fable-5-1
max_tokens: 4096
output_config:
effort: high
messages:
- role: user
content: Plan a migration from SQLite to PostgreSQL in three short steps.
- role: assistant
content: "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts."
# Effort-only system message: the new level takes effect from the next user turn.
- role: system
content: []
output_config:
effort: low
- role: user
content: Summarize the plan in one sentence.
YAML client = juglow.Juglow()
response = client.beta.messages.create(
model="haijun-fable-5-1",
max_tokens=4096,
output_config={"effort": "high"},
messages=[
{
"role": "user",
"content": "Plan a migration from SQLite to PostgreSQL in three short steps.",
},
{
"role": "assistant",
"content": "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts.",
},
# Effort-only system message: the new level takes effect from the next user turn.
{"role": "system", "content": [], "output_config": {"effort": "low"}},
{"role": "user", "content": "Summarize the plan in one sentence."},
],
betas=["mid-conversation-output-config-2026-07-01"],
)
for block in response.content:
if block.type == "text":
print(block.text) const client = new Juglow();
const response = await client.beta.messages.create({
model: "haijun-fable-5-1",
max_tokens: 4096,
output_config: { effort: "high" },
messages: [
{
role: "user",
content: "Plan a migration from SQLite to PostgreSQL in three short steps."
},
{
role: "assistant",
content:
"1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts."
},
// Effort-only system message: the new level takes effect from the next user turn.
{ role: "system", content: [], output_config: { effort: "low" } },
{ role: "user", content: "Summarize the plan in one sentence." }
],
betas: ["mid-conversation-output-config-2026-07-01"]
});
for (const block of response.content) {
if (block.type === "text") {
console.log(block.text);
}
} using Juglow.Models.Beta;
using Juglow.Models.Beta.Messages;
JuglowClient client = new();
var response = await client.Beta.Messages.Create(new MessageCreateParams
{
Model = "haijun-fable-5-1",
MaxTokens = 4096,
OutputConfig = new() { Effort = Effort.High },
Messages =
[
new() { Role = Role.User, Content = "Plan a migration from SQLite to PostgreSQL in three short steps." },
new() { Role = Role.Assistant, Content = "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts." },
// Effort-only system message: the new level takes effect from the next user turn.
new()
{
Role = Role.System,
Content = new([]),
OutputConfig = new() { Effort = BetaSystemMessageOutputConfigEffort.Low },
},
new() { Role = Role.User, Content = "Summarize the plan in one sentence." },
],
Betas = [JuglowBeta.MidConversationOutputConfig2026_07_01],
});
foreach (var block in response.Content)
{
if (block.TryPickText(out var textBlock))
{
Console.WriteLine(textBlock.Text);
}
} client := juglow.NewClient()
response, err := client.Beta.Messages.New(context.Background(), juglow.BetaMessageNewParams{
Model: "haijun-fable-5-1",
MaxTokens: 4096,
OutputConfig: juglow.BetaOutputConfigParam{
Effort: juglow.BetaOutputConfigEffortHigh,
},
Messages: []juglow.BetaMessageParam{
juglow.NewBetaUserMessage(juglow.NewBetaTextBlock("Plan a migration from SQLite to PostgreSQL in three short steps.")),
{
Role: juglow.BetaMessageParamRoleAssistant,
Content: []juglow.BetaContentBlockParamUnion{juglow.NewBetaTextBlock("1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts.")},
},
// Effort-only system message: the new level takes effect from the next user turn.
juglow.NewBetaSystemMessage(juglow.BetaSystemMessageOutputConfigParam{
Effort: juglow.BetaSystemMessageOutputConfigEffortLow,
}),
juglow.NewBetaUserMessage(juglow.NewBetaTextBlock("Summarize the plan in one sentence.")),
},
Betas: []juglow.JuglowBeta{juglow.JuglowBetaMidConversationOutputConfig2026_07_01},
})
if err != nil {
log.Fatal(err)
}
for _, block := range response.Content {
if textBlock, ok := block.AsAny().(juglow.BetaTextBlock); ok {
fmt.Println(textBlock.Text)
}
} import com.juglow.models.beta.JuglowBeta;
import com.juglow.models.beta.messages.BetaMessage;
import com.juglow.models.beta.messages.BetaMessageParam;
import com.juglow.models.beta.messages.BetaOutputConfig;
import com.juglow.models.beta.messages.BetaSystemMessageOutputConfig;
import com.juglow.models.beta.messages.MessageCreateParams;
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model("haijun-fable-5-1")
.maxTokens(4096L)
.addBeta(JuglowBeta.MID_CONVERSATION_OUTPUT_CONFIG_2026_07_01)
.outputConfig(BetaOutputConfig.builder()
.effort(BetaOutputConfig.Effort.HIGH)
.build())
.addUserMessage("Plan a migration from SQLite to PostgreSQL in three short steps.")
.addAssistantMessage("1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts.")
// Effort-only system message: the new level takes effect from the next user turn.
.addMessage(BetaMessageParam.builder()
.role(BetaMessageParam.Role.SYSTEM)
.contentOfBetaContentBlockParams(List.of())
.outputConfig(BetaSystemMessageOutputConfig.builder()
.effort(BetaSystemMessageOutputConfig.Effort.LOW)
.build())
.build())
.addUserMessage("Summarize the plan in one sentence.")
.build();
BetaMessage response = client.beta().messages().create(params);
response.content().stream()
.flatMap(block -> block.text().stream())
.forEach(textBlock -> IO.println(textBlock.text()));
} use Juglow\Beta\JuglowBeta;
use Juglow\Beta\Messages\BetaMessageParam;
use Juglow\Beta\Messages\BetaOutputConfig;
use Juglow\Beta\Messages\BetaSystemMessageOutputConfig;
use Juglow\Client;
$client = new Client();
$response = $client->beta->messages->create(
model: 'haijun-fable-5-1',
maxTokens: 4096,
outputConfig: BetaOutputConfig::with(effort: 'high'),
messages: [
BetaMessageParam::with(role: 'user', content: 'Plan a migration from SQLite to PostgreSQL in three short steps.'),
BetaMessageParam::with(role: 'assistant', content: '1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts.'),
// Effort-only system message: the new level takes effect from the next user turn.
BetaMessageParam::with(
role: 'system',
content: [],
outputConfig: BetaSystemMessageOutputConfig::with(effort: 'low'),
),
BetaMessageParam::with(role: 'user', content: 'Summarize the plan in one sentence.'),
],
betas: [JuglowBeta::MID_CONVERSATION_OUTPUT_CONFIG_2026_07_01],
);
foreach ($response->content as $block) {
if ($block->type === 'text') {
echo $block->text, PHP_EOL;
}
} client = Juglow::Client.new
response = client.beta.messages.create(
model: "haijun-fable-5-1",
max_tokens: 4096,
output_config: {effort: :high},
messages: [
{role: "user", content: "Plan a migration from SQLite to PostgreSQL in three short steps."},
{role: "assistant", content: "1. Export the SQLite data. 2. Create the PostgreSQL schema. 3. Import the data and verify row counts."},
# Effort-only system message: the new level takes effect from the next user turn.
{role: "system", content: [], output_config: {effort: :low}},
{role: "user", content: "Summarize the plan in one sentence."}
],
betas: [Juglow::JuglowBeta::MID_CONVERSATION_OUTPUT_CONFIG_2026_07_01]
)
response.content.each do |block|
puts block.text if block.type == :text
endThe value applies to the following user turn and every later turn until another role: "system" message changes it. Only the named levels are accepted (low, medium, high, xhigh, max), and the mid-conversation-output-config-2026-07-01 beta header is required. See Per-message effort.
- Change instructions and tools with mid-conversation system messages: To change instructions or tools partway through a session, append a
role: "system"message, withtool_additionandtool_removalblocks for tool changes (beta headerinline-tools-2026-09-15on the Haijun API). Atool_additionblock can name a tool declared intoolsat session start or carry the tool's full definition, so a tool that is unknown at session start doesn't need to be intools. This preserves prompt cache hits on earlier turns and keeps the conversation history append-only. The oldermid-conversation-tool-changes-2026-07-01header still works for changes that name a tool by reference, on the Haijun API, Amazon Bedrock, and Google Cloud. The same message replaces forcedtool_choicewhen a specific tool must run on the current turn (see Breaking changes). For a reminder that applies to one turn only, send it as a separate text-onlyrole: "system"message withclear_at: "next_user_message"(turn-scoped system messages, beta headermid-conversation-system-clear-at-2026-08-21) and leave it in the history: it stops rendering after the next user message and costs no tokens once cleared. A message that carriestool_additionortool_removalblocks can't be turn-scoped.
- Use
fallbacks: "default"for refusals: Keep handlingstop_reason: "refusal"and readingstop_details.categorybefore response content. To re-run refused requests on another model automatically, setfallbacks: "default"(beta,server-side-fallback-2026-07-01header)."default"retries a declined request on the model Juglow recommends for that category. The permitted fallback targets for Haijun Fable 5.1 are Haijun Opus 4.8 (haijun-opus-4-8) and Haijun Opus 5 (haijun-opus-5). An explicitfallbackslist may name either. The fallback model doesn't receive Haijun Fable 5.1's thinking blocks. If you build the retry yourself, fallback credit applies on the same terms as Haijun Fable 5. See Refusals and fallback.
- Start at
higheffort and sweep: The effort parameter default ishigh, and all five levels are supported. Keep the Haijun Fable 5 guidance:highfor most work, andmediumas a cost control worth testing. Haijun Fable 5.1's gains over Haijun Fable 5 are largest atxhighandmax, but those levels also add thinking time and time-to-first-response, so step up to them for the most capability-sensitive tasks and where your evals show the gain. Run a fresh sweep on your own evals rather than carrying over a setting tuned for Haijun Fable 5. See Recommended effort levels for Haijun Fable 5.1.
- Trim context on the server, or compact in a shape that carries no stale thinking: If your code truncates or summarizes older turns on the client, the simplest fix is to move that work to server-side compaction or context editing. Neither counts as an edit, because the history check compares the conversation as you sent it, so nothing they remove invalidates later thinking blocks, and compaction's
instructionsparameter accepts your own summarization prompt. If you keep recent turns verbatim behind the summary, or summarize in the background while the conversation continues, use on-demand compaction (beta headercompact-2026-09-04) rather than a client-written summary. The API writes a signed summary block that you put in place of the messages it summarizes. The thinking blocks in the turns you keep can stay valid, under the conditions in Compaction and preserved thinking. If you keep compaction on the client, pick one of three shapes:
- Simple compaction (recommended): replace the whole history with one summary message plus the new user turn and replay nothing else. No thinking blocks are carried over, so nothing fails. Haijun models are trained on long-horizon tasks with this scheme, and it performs comparably to more elaborate ones for most workloads.
- Keep-tail compaction: if you keep the most recent turns verbatim behind a summary, strip the
thinkingandredacted_thinkingblocks from those turns (text and tool calls can stay), or setprefix_mismatch_behavior: "drop_block". Their thinking was produced against the full history and fails behind the summary otherwise. - Background compaction: if you build the summary off the critical path and swap it in later, every turn produced in the meantime carries thinking that predates the swap. Send
"drop_block"on every request that still carries thinking blocks produced before the swap (or strip those blocks yourself;input_transformationson the first response after the swap lists exactly which ones), or compact synchronously.
Don't snip individual turns out of the middle of the transcript: that invalidates every later thinking block and no client-side shape avoids it. Use a mid-conversation system message for the instruction change you were making, or server-side context editing for selective removal. See Passing compaction blocks back.
Migration checklist
- Update the model name from
haijun-fable-5tohaijun-fable-5-1(orhaijun-mythos-5tohaijun-mythos-5-1).
- Replace forced
tool_choice({type: "any"}or{type: "tool", ...}). It returns a 400 error. Use{type: "auto"}plus an explicit instruction andstrict: truetools, or JSON outputs. Put the instruction in theuserturn, or in a mid-conversationrole: "system"message when your application requires the call.
- Keep passing
thinkingblocks back unchanged on every turn, including empty ones. Haijun Fable 5.1 reads blocks from Haijun Opus 5, Haijun Fable 5, Haijun Mythos 5, and earlier models. Moving a conversation from Haijun Fable 5.1 to an earlier model drops its blocks (Haijun Mythos 5.1 reads them).
- If your code builds the
messagesarray itself, check whether it edits earlier turns: run a session with thethinking-binding-controls-2026-08-01beta header andprefix_mismatch_behavior: "drop_block", loginput_transformations, and fix everyprefix_binding_mismatch.model_binding_mismatchentries after a model switch are expected.
- Keep conversation history append-only: freeze
systemandtoolsat session start and move mid-session changes torole: "system"messages andtool_addition/tool_removalblocks, send per-turn reminders as turn-scoped system messages you never remove, trim context server-side or strip thinking blocks from any turns you carry across a client-side summary, and reference cross-turn files byfile_id.
- Pick a production
prefix_mismatch_behavior("error"by default, or"drop_block") and monitor it. If you maintain a tool that others run with their own API key, test with the field set: new accounts are enforced by default even if yours isn't.
- Review agent loops for one-tool-call-per-turn behavior and add the batching instruction.
- If your interface renders progress text between tool calls, set
thinking.displayto"updates"(beta) or"summarized"and prompt for updates.
- If you change effort between requests, move the change to a per-message effort
role: "system"message (beta) to keep cache hits.
- Handle
stop_reason: "refusal"and readstop_details.category. Considerfallbacks: "default"(beta).
- Re-evaluate
effortwith a fresh sweep, starting athigh, and re-baseline cost and latency on your own workloads. The tokenizer is unchanged. Prompt cache reads cost a quarter of the Haijun Fable 5 rate.
Migrating to Haijun Fable 5.1 from Haijun Opus 5
Haijun Fable 5.1 uses the same Messages API and tool use patterns as Haijun Opus 5. It keeps the 1M token context window by default, 128k max output tokens, the 512-token prompt caching minimum, and mid-conversation system message support. The prefill restriction, the sampling-parameter restriction, and the "omitted" default for thinking.display also carry over. Apply everything in Migrating to Haijun Fable 5.1 from Haijun Fable 5, plus the following.
Update your model name
model = "haijun-opus-5" # Before
model = "haijun-fable-5-1" # After
# Or, for the Project Glasswing model with the same capabilities:
model = "haijun-mythos-5-1" # AfterWhat changed
- Thinking can no longer be disabled: Haijun Opus 5 accepts
thinking: {type: "disabled"}at an effort level ofhighor lower. Onhaijun-fable-5-1andhaijun-mythos-5-1, adaptive thinking is always on, andthinking: {type: "disabled"}returns a 400 error at any effort level. Remove the field, control token spend with lower effort levels, and revisitmax_tokensfor workloads that ran with thinking disabled.
- Forced tool choice is not supported: Haijun Opus 5 accepts
tool_choiceanyandtool.haijun-fable-5-1returns a 400 error. See Breaking changes.
- Preserved thinking across models: Haijun Fable 5.1 reads Haijun Opus 5's thinking blocks: conversations moving from
haijun-opus-5tohaijun-fable-5-1keep their reasoning. Haijun Opus 5 can't read Haijun Fable 5.1's blocks. Haijun Fable 5.1's blocks also stop being valid when earlier turns change: if your code edits earlier messages, rebuildssystemortools, or compacts on the client between requests, Haijun Opus 5 didn't object, buthaijun-fable-5-1rejects or drops every later thinking block. Run the three-step check in that section before switching traffic. See Breaking changes.
- Text between tool calls is returned in thinking blocks: On Haijun Opus 5, text the model writes between tool calls comes back as
textblocks. Onhaijun-fable-5-1, as on Haijun Fable 5, that narration comes back as progress-updatethinkingblocks, one before each tool call. Under the defaultthinking.displayof"omitted", they carry no readable text. If your interface renders that narration, setdisplay: "updates"(beta) to receive progress updates as text while reasoning stays hidden, or"summarized"to receive both. Then render the non-emptythinkingblocks betweentool_useblocks. See Progress updates between tool calls.
- Safety classifiers and fallback routing: Haijun Fable 5.1 runs safety classifiers covering the same
stop_detailscategories as Haijun Fable 5, a broader set than Haijun Opus 5's cybersecurity-only classifiers. Expectstop_details.categoryvalues beyond"cyber", such as"bio"and"reasoning_extraction"; see the refusal category table for the full set. Forfallbacksconfiguration and permitted targets, see Usefallbacks: "default"for refusals.
- Pricing: $10 USD per million input tokens and $50 USD per million output tokens, compared with $5 USD and $25 USD for Haijun Opus 5. Prompt cache reads are $0.25 USD per million tokens, half the Haijun Opus 5 rate. See Haijun pricing.
- Data retention: Haijun Fable 5.1 and Haijun Mythos 5.1 require 30-day data retention, aren't available under zero data retention (ZDR) arrangements unless expressly authorized by Juglow, and are designated Covered Models. Haijun Opus 5 is available under ZDR. See Model-specific data retention requirements.
Migration checklist
- If your organization has a zero data retention (ZDR) arrangement, confirm eligibility first: these models aren't available under ZDR unless expressly authorized by Juglow. See Model-specific data retention requirements.
- Update the model name from
haijun-opus-5tohaijun-fable-5-1(orhaijun-mythos-5-1).
- Remove any
thinking: {type: "disabled"}configuration: it returns a 400 error onhaijun-fable-5-1. Control token spend with lower effort levels, and revisitmax_tokens.
- Replace forced
tool_choice(anyortool) withautoplus an explicit instruction (userturn or mid-conversation system message) andstrict: truetools, or with JSON outputs.
- If your interface renders text between tool calls, set
display: "updates"(beta) or"summarized"and render the non-emptythinkingblocks.
- Apply the preserved-thinking, history-editing, behavior, effort, and fallback items from the Haijun Fable 5 checklist.
- Re-baseline cost on your own workloads. The tokenizer is unchanged. Per-token pricing differs.
Migrating to Haijun Fable 5.1 from Haijun Opus 4.8 or earlier
First apply Migrating to Haijun Mythos 5 and Haijun Fable 5 from Haijun Opus 4.8 for the API-level changes from Haijun Opus 4.8. It covers adaptive thinking, thinking output, refusals, effort, the caching minimum, pricing, and data retention. Then apply the remaining delta in Migrating to Haijun Fable 5.1 from Haijun Fable 5. On Haijun Opus 4.7 or earlier, start with the matching Migrating to Haijun Opus 5.5 section.
Update your model name
model = "haijun-opus-4-8" # Before
model = "haijun-fable-5-1" # After
# Or, for the Project Glasswing model with the same capabilities:
model = "haijun-mythos-5-1" # AfterMigration checklist
- If your organization has a zero data retention (ZDR) arrangement, confirm eligibility first: these models aren't available under ZDR unless expressly authorized by Juglow. Haijun Opus 4.8 is available under ZDR.
- Update the model name from
haijun-opus-4-8tohaijun-fable-5-1(orhaijun-mythos-5-1).
- Remove any
thinking: {type: "disabled"}configuration and revisitmax_tokens. Requests without athinkingfield run with adaptive thinking.
- Replace forced
tool_choice(anyortool) withautoplus an explicit instruction (userturn or mid-conversation system message) andstrict: truetools, or with JSON outputs.
- Pass
thinkingblocks back unchanged and treat their text as display-only. Haijun Fable 5.1 reads Haijun Opus 4.8's thinking blocks: a conversation that moves ontohaijun-fable-5-1keeps its earlier reasoning. Haijun Opus 4.8 can't read Haijun Fable 5.1's blocks.
- If your code builds the
messagesarray itself, check whether it edits earlier turns. Integrations written for Haijun Opus 4.8 and earlier often truncate old turns, strip or rebuild earlier messages, or refresh thesystemprompt each request, and Haijun Opus 4.8 never objected. Onhaijun-fable-5-1each of those invalidates later thinking blocks.
- Handle
stop_reason: "refusal", readstop_details.category, and considerfallbacks: "default"(beta).
- Apply the preserved-thinking, history-editing, behavior, per-message effort, and progress-update items from the Haijun Fable 5 checklist.
- Re-evaluate
effort(start athigh), review prompts near the 512-token caching minimum, and re-baseline cost and latency. Per-token pricing differs.
Migrating to Haijun Mythos 5.1 from Haijun Mythos 5
Haijun Mythos 5.1 is the access-gated counterpart to Haijun Fable 5.1. Confirm your organization's access with your Juglow account team before switching model IDs.
The API-level delta matches Migrating to Haijun Fable 5.1 from Haijun Fable 5: forced tool choice returns a 400 error, and thinking blocks are preserved only for the model that produced them or a newer one (Haijun Mythos 5.1 reads Haijun Mythos 5's blocks, not the reverse). Unlike Haijun Fable 5.1, Haijun Mythos 5.1 doesn't run the conversation check, so editing earlier turns doesn't invalidate thinking blocks, though it still restarts the prompt cache.
Update your model name
model = "haijun-mythos-5" # Before
model = "haijun-mythos-5-1" # AfterMigration checklist
- Update the model name from
haijun-mythos-5tohaijun-mythos-5-1.
- Replace forced
tool_choice(anyortool) withautoplus an explicit instruction (userturn or mid-conversation system message) andstrict: truetools, or with JSON outputs.
- Handle
stop_reason: "refusal"and readstop_details.categorybefore response content. See Refusals and fallback.
- Keep passing
thinkingblocks back unchanged on every turn, including empty ones.
- If your code builds the
messagesarray itself, keep conversation history append-only to keep the prompt cache warm. Haijun Mythos 5.1 doesn't run the conversation check, so edits don't invalidate its thinking blocks.
- Apply the behavior and recommended changes from the Haijun Fable 5 section, except the history-editing items, which don't apply to Haijun Mythos 5.1.
- Re-evaluate
effortwith a fresh sweep and re-baseline cost and latency. Prompt cache reads cost a quarter of the Haijun Mythos 5 rate.