Haijun Platform Docs
ID

Note: This guide covers migrating Messages API code. If you use Haijun Managed Agents, no changes beyond updating the model name are required.

Tip: Automate your migration with the Haijun API track. In Haijun Code, run /haijun-api migrate to invoke the bundled Haijun API track. It works for any current Haijun model as the target: ``text wrap /haijun-api migrate this project to haijun-opus-5-5 `` The track applies the model ID swap and, as needed, breaking parameter changes, prefill replacement, and effort calibration for your target model across your code base, then produces a checklist of items to verify manually. It asks you to confirm the migration scope (entire working directory, a subdirectory, or a specific file list) before editing any files. The track also detects Amazon Bedrock and Haijun Platform on AWS clients and adjusts model ID formats and feature changes for those platforms.

This page lists the code changes for moving to Haijun Opus 5.5 from Haijun Opus 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6 and earlier Opus models, or Haijun Sonnet 5. Every reader needs What every request to Haijun Opus 5.5 must satisfy and Handle thinking in every response. Then go to the section for your current model: its first sentence names the other sections that apply to you. The migration checklist lists every change by starting model.

Haijun Opus 5.5 costs less than Haijun Opus 5 ($4 / $20 USD per million input / output tokens, compared with $5 / $25; see Haijun pricing). For feature support, see What's new in Haijun Opus 5.5. For behavioral differences and model-specific prompting patterns, see Prompting Haijun Opus 5.5.

What every request to Haijun Opus 5.5 must satisfy

Whichever model you are coming from, a request to haijun-opus-5-5 must meet the following. Where an item says a setting is rejected, the API returns a 400 error.

  • Model ID: Use haijun-opus-5-5, a fixed model ID with no date suffix. On Amazon Bedrock, Haijun Platform on AWS, Google Cloud, and Microsoft Foundry, use that platform's model ID; see Availability.
  • Thinking: Send no thinking field, or send thinking: {"type": "adaptive"}, which is equivalent: adaptive thinking is always on. thinking: {"type": "disabled"} and manual thinking budgets (thinking: {"type": "enabled", "budget_tokens": N}) are rejected. See the before and after for thinking.
  • Tool choice: Use tool_choice {"type": "auto"} (the default) or {"type": "none"}. Forcing a tool call with {"type": "any"} or {"type": "tool", "name": "..."} is rejected. See the before and after for tool choice.
  • Sampling parameters: Omit temperature, top_p, and top_k, or leave them at their defaults: any other value is rejected. Use prompting to guide the model's behavior.
  • Prefill: Don't end messages with a prefilled assistant turn: it is rejected. Use structured outputs or system prompt instructions instead.
  • Computer use: On the Haijun API and Google Cloud, declare computer use as the computer_toolset_20260801 toolset; the earlier computer_20251124 tool is rejected there. See the computer use breaking change.
  • Context window: No context-window beta header is needed. The 1M token context window is the default, and a header sent for older models has no effect.

The following request satisfies every item in the list: effort is set, and there is no thinking field. The SDK tabs that print text select it by block type, because thinking blocks come first.

bash
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 4096,
      "messages": [{
        "role": "user",
        "content": "Analyze the trade-offs between microservices and monolithic architectures"
      }],
      "output_config": {
        "effort": "medium"
      }
    }'
bash
  ant messages create \
    --model haijun-opus-5-5 \
    --max-tokens 4096 \
    --output-config '{effort: medium}' \
    --message '{role: user, content: "Analyze the trade-offs between microservices and monolithic architectures"}'
python
  client = juglow.Juglow()

  response = client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=4096,
      messages=[
          {
              "role": "user",
              "content": "Analyze the trade-offs between microservices and monolithic architectures",
          }
      ],
      output_config={"effort": "medium"},
  )

  for block in response.content:
      if block.type == "text":
          print(block.text)
typescript
  const client = new Juglow();

  const response = await client.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 4096,
    messages: [
      {
        role: "user",
        content: "Analyze the trade-offs between microservices and monolithic architectures"
      }
    ],
    output_config: {
      effort: "medium"
    }
  });

  const textBlock = response.content.find(
    (block): block is Juglow.TextBlock => block.type === "text"
  );
  console.log(textBlock?.text);
csharp
  JuglowClient client = new();

  var parameters = new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 4096,
      Messages = [
          new() {
              Role = Role.User,
              Content = "Analyze the trade-offs between microservices and monolithic architectures"
          }
      ],
      OutputConfig = new OutputConfig
      {
          Effort = Effort.Medium
      }
  };

  var message = await client.Messages.Create(parameters);
  Console.WriteLine(message);
go
  client := juglow.NewClient()

  response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 4096,
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(juglow.NewTextBlock("Analyze the trade-offs between microservices and monolithic architectures")),
  	},
  	OutputConfig: juglow.OutputConfigParam{
  		Effort: juglow.OutputConfigEffortMedium,
  	},
  })
  if err != nil {
  	log.Fatal(err)
  }
  for _, block := range response.Content {
  	if textBlock, ok := block.AsAny().(juglow.TextBlock); ok {
  		fmt.Println(textBlock.Text)
  	}
  }
java
  import com.juglow.models.messages.OutputConfig;

  void main() {
      JuglowClient client = JuglowOkHttpClient.fromEnv();

      MessageCreateParams params = MessageCreateParams.builder()
          .model(Model.HAIJUN_OPUS_5_5)
          .maxTokens(4096L)
          .addUserMessage("Analyze the trade-offs between microservices and monolithic architectures")
          .outputConfig(OutputConfig.builder()
              .effort(OutputConfig.Effort.MEDIUM)
              .build())
          .build();

      Message response = client.messages().create(params);
      response.content().stream()
          .flatMap(block -> block.text().stream())
          .forEach(textBlock -> IO.println(textBlock.text()));
  }
php
  $client = new Client();

  $message = $client->messages->create(
      maxTokens: 4096,
      messages: [
          ['role' => 'user', 'content' => 'Analyze the trade-offs between microservices and monolithic architectures']
      ],
      model: 'haijun-opus-5-5',
      outputConfig: ['effort' => 'medium'],
  );

  foreach ($message->content as $block) {
      if ($block->type === 'text') {
          echo $block->text, PHP_EOL;
      }
  }
ruby
  client = Juglow::Client.new

  message = client.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 4096,
    messages: [
      { role: "user", content: "Analyze the trade-offs between microservices and monolithic architectures" }
    ],
    output_config: {
      effort: "medium"
    }
  )

  message.content.each do |block|
    puts block.text if block.type == :text
  end

Handle thinking in every response

Thinking runs on every Haijun Opus 5.5 request, so every response can begin with thinking blocks, and max_tokens covers thinking plus text. If your code already runs with thinking on, items 1 to 3 are likely in place: check items 4 and 5. If it ran without thinking, on any earlier model, each item is a change.

  1. max_tokens covers thinking plus text: On Haijun Opus 4.8 and earlier Opus models, requests without a thinking field run without thinking. Haijun Opus 5 and Haijun Sonnet 5 accept thinking: {"type": "disabled"}. On Haijun Opus 5.5, every request runs with adaptive thinking. max_tokens remains a hard limit on total output, thinking plus response text, so revisit it for workloads that ran without thinking. Thinking tokens are billed as output tokens even when the thinking text is not returned to you, so such a workload can produce more output tokens per request. See Cost control. To spend fewer tokens on thinking, lower the effort level. If you run at xhigh or max effort, set a large max_tokens so the model has room to think and act; start at 64k tokens and tune from there. If your prompts were tuned for running without thinking, see Prompts written for thinking disabled.
  1. Responses begin with thinking blocks: A response can begin with one or more thinking blocks before the first text block. Code that reads the reply by position, such as content[0].text or a stream handler that treats the first content_block_start event as text, breaks on these responses. Select content blocks by their type field instead: read text from the blocks whose type is "text", and branch on the block type when handling stream events.
  1. Return thinking blocks unmodified in tool-use loops: If you run a tool-use loop, pass the thinking blocks from each assistant response back to the API complete and unmodified when you return tool results, including blocks whose thinking field is empty. Echo the assistant message as received rather than filtering its content blocks by type or rebuilding it: the API rejects edited, reordered, or partially dropped thinking blocks with a 400 error. See Preserving thinking blocks.
  1. Thinking text is omitted by default: thinking.display defaults to "omitted", so thinking blocks arrive with an empty thinking field alongside their signature. Treat the thinking field as display text only. To receive readable summaries instead, set thinking.display to "summarized":
python
     thinking = {
         "type": "adaptive",
         "display": "summarized",
     }
typescript
     const thinking = {
       type: "adaptive",
       display: "summarized"
     };
csharp
     var thinking = new ThinkingConfigAdaptive { Display = Display.Summarized };
go
     thinking := juglow.ThinkingConfigParamUnion{
     	OfAdaptive: &juglow.ThinkingConfigAdaptiveParam{
     		Display: juglow.ThinkingConfigAdaptiveDisplaySummarized,
     	},
     }
java
     ThinkingConfigAdaptive thinking = ThinkingConfigAdaptive.builder()
         .display(ThinkingConfigAdaptive.Display.SUMMARIZED)
         .build();
php
     $thinking = ['type' => 'adaptive', 'display' => 'summarized'];
ruby
     thinking = {
       type: "adaptive",
       display: "summarized"
     }

If your product streams reasoning to users, the default appears as a long pause before output begins; set display: "summarized" to restore visible progress during thinking. See Controlling thinking display.

  1. Text between tool calls arrives in thinking blocks: The short notes the model writes between tool calls come back as thinking blocks, which are empty at the default display. See Text between tool calls is returned in thinking blocks.

Migration checklist by starting model

Work down the groups and stop after the one that names your current model: every item up to that point applies to you. If you are on Haijun Opus 5, the first group is the whole list. If you are on Haijun Sonnet 5, apply the first group and the last.

Every starting model

  • Update the model ID to haijun-opus-5-5.
  • Remove thinking: {"type": "disabled"} and thinking: {"type": "enabled", ...}; choose an effort level instead.
  • Set effort explicitly: the default is medium, where Haijun Opus 5's is high.
  • Replace tool_choice types any and tool with auto plus strict tool use or structured outputs.
  • If you use computer use on the Haijun API or Google Cloud, declare computer_toolset_20260801 (no beta header) instead of computer_20251124 and update your agent loop for the toolset. On Amazon Bedrock, keep computer_20251124; check the computer use tool's Compatibility section for other platforms.
  • If a router or fallback can move a conversation from Haijun Opus 5.5 to another model, expect that model to run without Haijun Opus 5.5's thinking blocks (Haijun Fable 5.1 and Haijun Mythos 5.1 on the Haijun API are the exception and keep them). Haijun Opus 5.5 itself reads thinking from Haijun Opus 5 and earlier Opus, Sonnet, and Haiku models, but not from Haijun Fable or Haijun Mythos models.
  • Read content blocks by type, and pass thinking blocks back unmodified in tool-use loops.
  • If your interface renders text between tool calls, set display: "updates" (beta) or "summarized" and render the non-empty thinking blocks.
  • If your code edits earlier turns, the system prompt, or tools mid-conversation, follow Preserved thinking.
  • Handle stop_reason: "refusal" and configure fallback.
  • Re-baseline cost and latency at your chosen effort level.
  • If your code disabled thinking, revisit max_tokens, which covers thinking plus response text; at xhigh or max effort, start at 64k. See Handle thinking in every response.

Haijun Opus 4.8 or earlier

  • Review workloads that ran without a thinking field: on Haijun Opus 5.5 they run with thinking, and thinking can't be disabled. Revisit max_tokens, which remains a hard limit on total output (thinking plus response text), and lower effort where you want less thinking. Thinking tokens are billed as output tokens, so these workloads can produce more output tokens per request.
  • Verify any code that parses the thinking field treats it as display text only. Set display: "summarized" to receive readable summaries.
  • Review prompts near the caching minimum: prompts of 512 tokens or more can create cache entries.
  • If your organization has a Priority Tier commitment, plan capacity separately: Priority Tier is not supported on Haijun Opus 5.5.
  • If you run at xhigh or max effort, raise max_tokens to at least 64k as a starting point.
  • For agentic workloads, consider task budgets (beta) and mid-conversation tool changes (beta).

Haijun Opus 4.7 or earlier

  • Run a fresh effort sweep on your own evals rather than carrying over a setting tuned for an earlier model.
  • Remove any context-window beta header.
  • If you rebuild conversation history to update instructions, consider switching to a mid-conversation system message to preserve prompt cache hits.
  • Verify your stop-reason handling reads stop_details on refusals.
  • If you want fast mode, which Haijun Opus 4.7 rejects, set speed: "fast" with the fast-mode-2026-02-01 beta header on the Haijun API.

Haijun Opus 4.6 or earlier

  • Remove temperature, top_p, and top_k from request payloads.
  • Replace thinking: {"type": "enabled", "budget_tokens": N} with thinking: {"type": "adaptive"} plus the effort parameter, or remove the thinking field entirely; adaptive thinking is always on.
  • If your UI displays thinking content, explicitly opt in to thinking summarization.
  • Re-benchmark end-to-end cost and latency under the updated tokenization.
  • Re-tune max_tokens to account for the updated tokenization, including compaction triggers.
  • Re-test any client-side token-count estimations.
  • If your application sends images, re-budget for high-resolution image support (up to approximately 3x more image tokens per full-resolution image). Downsample before sending if you do not need the additional fidelity.
  • If you consume pointing or bounding-box coordinates from the model, remove any scale-factor conversion; coordinates are 1:1 with actual image pixels on Haijun Opus 4.7 and later models.
  • If your product does legitimate security work, apply to the Cyber Verification Program for access to lower restrictions on cyber content.

Haijun Opus 4.5 or earlier

  • Remove any assistant-message prefills; Haijun Opus 4.6 already rejects them.
  • Verify tool call JSON parsing uses a standard JSON parser.
  • Move from client.beta.messages.create to client.messages.create: adaptive thinking and effort need no beta namespace.
  • Remove the effort-2025-11-24 beta header (the effort parameter does not require it).
  • Remove the fine-grained-tool-streaming-2025-05-14 beta header.
  • Remove the interleaved-thinking-2025-05-14 beta header (adaptive thinking enables interleaved thinking automatically).
  • Migrate output_format to output_config.format (if applicable).

Haijun 4.1 or earlier

  • Update tool versions (text_editor_20250728, code_execution_20260521).
  • Handle the refusal stop reason.
  • Handle the model_context_window_exceeded stop reason.
  • Verify tool string parameter handling for trailing newlines.
  • Remove legacy beta headers (token-efficient-tools-2025-02-19, output-128k-2025-02-19).

Haijun Sonnet 5 only

  • If you rebuild conversation history to update instructions, consider switching to a mid-conversation system message to preserve prompt cache hits.
  • Review prompts near the caching minimum: prompts of 512 tokens or more can create cache entries.

Migrating to Haijun Opus 5.5 from Haijun Opus 5

First work through What every request to Haijun Opus 5.5 must satisfy and Handle thinking in every response. Every starting model needs the changes in this section. They are the request settings that Haijun Opus 5.5 rejects and the response changes that come with it. The checklist for this section is the first group of the migration checklist.

Update your model name

python
model = "haijun-opus-5"  # Before
model = "haijun-opus-5-5"  # After

haijun-opus-5-5 is a fixed model ID with no date suffix, the same scheme as haijun-opus-5. On Amazon Bedrock, Haijun Platform on AWS, Google Cloud, and Microsoft Foundry, use that platform's model ID; see Availability.

Breaking changes

Each change is explained in What's new in Haijun Opus 5.5; this section gives the code change for each.

Thinking can't be disabled

thinking: {"type": "disabled"} and thinking: {"type": "enabled", "budget_tokens": N} both return a 400 error ("thinking.type.disabled" is not supported for this model. or "thinking.type.enabled" is not supported for this model.). Remove the thinking field and pick an effort level; where you disabled thinking to save tokens, use a lower one. Responses then begin with thinking blocks, so select content blocks by type and pass thinking blocks back unmodified with tool results. See Thinking can't be disabled.

Before. Haijun Opus 5 accepts this request, and Haijun Opus 5.5 rejects it with a 400 error:

bash
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5",
      "max_tokens": 16000,
      "thinking": {"type": "disabled"},
      "messages": [{"role": "user", "content": "..."}]
    }'
bash
  ant messages create \
    --model haijun-opus-5 \
    --max-tokens 16000 \
    --thinking '{type: disabled}' \
    --message '{role: user, content: "..."}'
python
  client.messages.create(
      model="haijun-opus-5",
      max_tokens=16000,
      thinking={"type": "disabled"},
      messages=[{"role": "user", "content": "..."}],
  )
typescript
  await client.messages.create({
    model: "haijun-opus-5",
    max_tokens: 16000,
    thinking: { type: "disabled" },
    messages: [{ role: "user", content: "..." }]
  });
csharp
  await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5,
      MaxTokens = 16000,
      Thinking = new ThinkingConfigDisabled(),
      Messages = [new() { Role = Role.User, Content = "..." }],
  });
go
  client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5,
  	MaxTokens: 16000,
  	Thinking: juglow.ThinkingConfigParamUnion{
  		OfDisabled: &juglow.ThinkingConfigDisabledParam{},
  	},
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(juglow.NewTextBlock("...")),
  	},
  })
java
  MessageCreateParams params = MessageCreateParams.builder()
      .model(Model.HAIJUN_OPUS_5)
      .maxTokens(16000L)
      .thinking(ThinkingConfigDisabled.builder().build())
      .addUserMessage("...")
      .build();

  client.messages().create(params);
php
  $client->messages->create(
      model: Model::HAIJUN_OPUS_5,
      maxTokens: 16000,
      thinking: ThinkingConfigDisabled::with(),
      messages: [['role' => 'user', 'content' => '...']],
  );
ruby
  client.messages.create(
    model: Juglow::Model::HAIJUN_OPUS_5,
    max_tokens: 16000,
    thinking: Juglow::ThinkingConfigDisabled.new,
    messages: [{ role: "user", content: "..." }]
  )

After:

bash
  # thinking is always on; effort is the control
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 16000,
      "output_config": {"effort": "low"},
      "messages": [{"role": "user", "content": "..."}]
    }'
bash
  # thinking is always on; effort is the control
  ant messages create \
    --model haijun-opus-5-5 \
    --max-tokens 16000 \
    --output-config '{effort: low}' \
    --message '{role: user, content: "..."}'
python
  client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=16000,
      output_config={"effort": "low"},  # thinking is always on; effort is the control
      messages=[{"role": "user", "content": "..."}],
  )
typescript
  await client.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 16000,
    output_config: { effort: "low" }, // thinking is always on; effort is the control
    messages: [{ role: "user", content: "..." }]
  });
csharp
  await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 16000,
      OutputConfig = new() { Effort = Effort.Low }, // thinking is always on; effort is the control
      Messages = [new() { Role = Role.User, Content = "..." }],
  });
go
  client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 16000,
  	OutputConfig: juglow.OutputConfigParam{
  		Effort: juglow.OutputConfigEffortLow, // thinking is always on; effort is the control
  	},
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(juglow.NewTextBlock("...")),
  	},
  })
java
  MessageCreateParams params = MessageCreateParams.builder()
      .model(Model.HAIJUN_OPUS_5_5)
      .maxTokens(16000L)
      // thinking is always on; effort is the control
      .outputConfig(OutputConfig.builder()
          .effort(OutputConfig.Effort.LOW)
          .build())
      .addUserMessage("...")
      .build();

  client.messages().create(params);
php
  $client->messages->create(
      model: Model::HAIJUN_OPUS_5_5,
      maxTokens: 16000,
      // thinking is always on; effort is the control
      outputConfig: OutputConfig::with(effort: Effort::LOW),
      messages: [['role' => 'user', 'content' => '...']],
  );
ruby
  client.messages.create(
    model: Juglow::Model::HAIJUN_OPUS_5_5,
    max_tokens: 16000,
    # thinking is always on; effort is the control
    output_config: { effort: Juglow::OutputConfig::Effort::LOW },
    messages: [{ role: "user", content: "..." }]
  )

Forced tool use is not supported

tool_choice types any and tool return a 400 error (tool_choice: type "tool" and "any" are not supported for this model.), including on the token counting endpoint. Use auto with strict tool use or structured outputs, and say in the prompt when the tool applies. Strict tool use accepts a subset of JSON Schema, so check each tool's input_schema before you add strict: true. Every object in the schema must set additionalProperties: false; see JSON Schema limitations. See Forced tool use is not supported.

Before. Haijun Opus 5 accepts this request, and Haijun Opus 5.5 rejects it with a 400 error:

bash
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5",
      "max_tokens": 1024,
      "tools": [{
        "name": "get_weather",
        "description": "Get the current weather in a given location",
        "input_schema": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "The city and state, e.g. San Francisco, CA"
            }
          },
          "required": ["location"],
          "additionalProperties": false
        }
      }],
      "tool_choice": {"type": "tool", "name": "get_weather"},
      "messages": [{"role": "user", "content": "What'\''s the weather in Paris?"}]
    }'
bash
  ant messages create <<'YAML'
  model: haijun-opus-5
  max_tokens: 1024
  tools:
    - name: get_weather
      description: Get the current weather in a given location
      input_schema:
        type: object
        properties:
          location:
            type: string
            description: The city and state, e.g. San Francisco, CA
        required: [location]
        additionalProperties: false
  tool_choice:
    type: tool
    name: get_weather
  messages:
    - role: user
      content: What's the weather in Paris?
  YAML
python
  client.messages.create(
      model="haijun-opus-5",
      max_tokens=1024,
      tools=tools,
      tool_choice={"type": "tool", "name": "get_weather"},
      messages=[{"role": "user", "content": "What's the weather in Paris?"}],
  )
typescript
  await client.messages.create({
    model: "haijun-opus-5",
    max_tokens: 1024,
    tools,
    tool_choice: { type: "tool", name: "get_weather" },
    messages: [{ role: "user", content: "What's the weather in Paris?" }]
  });
csharp
  await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5,
      MaxTokens = 1024,
      Tools = [.. tools],
      ToolChoice = new ToolChoiceTool { Name = "get_weather" },
      Messages = [new() { Role = Role.User, Content = "What's the weather in Paris?" }],
  });
go
  client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:      juglow.ModelHaijunOpus5,
  	MaxTokens:  1024,
  	Tools:      tools,
  	ToolChoice: juglow.ToolChoiceParamOfTool("get_weather"),
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(juglow.NewTextBlock("What's the weather in Paris?")),
  	},
  })
java
  MessageCreateParams params = MessageCreateParams.builder()
      .model(Model.HAIJUN_OPUS_5)
      .maxTokens(1024L)
      .tools(tools)
      .toolChoice(ToolChoiceTool.of("get_weather"))
      .addUserMessage("What's the weather in Paris?")
      .build();

  client.messages().create(params);
php
  $client->messages->create(
      model: Model::HAIJUN_OPUS_5,
      maxTokens: 1024,
      tools: $tools,
      toolChoice: ToolChoiceTool::with(name: 'get_weather'),
      messages: [['role' => 'user', 'content' => "What's the weather in Paris?"]],
  );
ruby
  client.messages.create(
    model: Juglow::Model::HAIJUN_OPUS_5,
    max_tokens: 1024,
    tools: tools,
    tool_choice: Juglow::ToolChoiceTool.new(name: "get_weather"),
    messages: [{ role: "user", content: "What's the weather in Paris?" }]
  )

After:

bash
  # strict tool use: every call matches the tool's input_schema
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 1024,
      "tools": [{
        "name": "get_weather",
        "description": "Get the current weather in a given location",
        "input_schema": {
          "type": "object",
          "properties": {
            "location": {
              "type": "string",
              "description": "The city and state, e.g. San Francisco, CA"
            }
          },
          "required": ["location"],
          "additionalProperties": false
        },
        "strict": true
      }],
      "tool_choice": {"type": "auto"},
      "messages": [{
        "role": "user",
        "content": "What'\''s the weather in Paris? Use the get_weather tool."
      }]
    }'
bash
  ant messages create <<'YAML'
  model: haijun-opus-5-5
  max_tokens: 1024
  tools:
    - name: get_weather
      description: Get the current weather in a given location
      input_schema:
        type: object
        properties:
          location:
            type: string
            description: The city and state, e.g. San Francisco, CA
        required: [location]
        additionalProperties: false
      # strict tool use: every call matches the tool's input_schema
      strict: true
  tool_choice:
    type: auto
  messages:
    - role: user
      content: What's the weather in Paris? Use the get_weather tool.
  YAML
python
  client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=1024,
      # strict tool use: every call matches the tool's input_schema
      tools=[{**tool, "strict": True} for tool in tools],
      tool_choice={"type": "auto"},
      messages=[
          {
              "role": "user",
              "content": "What's the weather in Paris? Use the get_weather tool.",
          }
      ],
  )
typescript
  await client.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    // strict tool use: every call matches the tool's input_schema
    tools: tools.map((tool) => ({ ...tool, strict: true })),
    tool_choice: { type: "auto" },
    messages: [
      {
        role: "user",
        content: "What's the weather in Paris? Use the get_weather tool."
      }
    ]
  });
csharp
  await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 1024,
      // strict tool use: every call matches the tool's input_schema
      Tools = [.. tools.Select(tool => tool with { Strict = true })],
      ToolChoice = new ToolChoiceAuto(),
      Messages =
      [
          new()
          {
              Role = Role.User,
              Content = "What's the weather in Paris? Use the get_weather tool.",
          },
      ],
  });
go
  // strict tool use: every call matches the tool's input_schema
  var strictTools []juglow.ToolUnionParam
  for _, tool := range tools {
  	strictTool := *tool.OfTool
  	strictTool.Strict = juglow.Bool(true)
  	strictTools = append(strictTools, juglow.ToolUnionParam{OfTool: &strictTool})
  }
  client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:      juglow.ModelHaijunOpus5_5,
  	MaxTokens:  1024,
  	Tools:      strictTools,
  	ToolChoice: juglow.ToolChoiceUnionParam{OfAuto: &juglow.ToolChoiceAutoParam{}},
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(
  			juglow.NewTextBlock("What's the weather in Paris? Use the get_weather tool."),
  		),
  	},
  })
java
  MessageCreateParams params = MessageCreateParams.builder()
      .model(Model.HAIJUN_OPUS_5_5)
      .maxTokens(1024L)
      // strict tool use: every call matches the tool's input_schema
      .tools(tools.stream()
          .map(tool -> tool.tool()
              .map(customTool -> customTool.toBuilder().strict(true).build())
              .map(ToolUnion::ofTool)
              .orElse(tool))
          .toList())
      .toolChoice(ToolChoiceAuto.builder().build())
      .addUserMessage("What's the weather in Paris? Use the get_weather tool.")
      .build();

  client.messages().create(params);
php
  $client->messages->create(
      model: Model::HAIJUN_OPUS_5_5,
      maxTokens: 1024,
      // strict tool use: every call matches the tool's input_schema
      tools: array_map(fn (Tool $tool) => $tool->withStrict(true), $tools),
      toolChoice: ToolChoiceAuto::with(),
      messages: [
          [
              'role' => 'user',
              'content' => "What's the weather in Paris? Use the get_weather tool.",
          ],
      ],
  );
ruby
  client.messages.create(
    model: Juglow::Model::HAIJUN_OPUS_5_5,
    max_tokens: 1024,
    # strict tool use: every call matches the tool's input_schema
    tools: tools.map { |tool| tool.merge(strict: true) },
    tool_choice: Juglow::ToolChoiceAuto.new,
    messages: [
      { role: "user", content: "What's the weather in Paris? Use the get_weather tool." }
    ]
  )

Thinking blocks are tied to the model and the conversation

On the Haijun API, Haijun Fable 5.1 and Haijun Mythos 5.1 read Haijun Opus 5.5 thinking blocks; no other model does. A router or fallback that moves a conversation from Haijun Opus 5.5 to any other model runs those turns without them. In the other direction, Haijun Opus 5.5 reads thinking blocks from Haijun Opus 5 and earlier Opus, Sonnet, and Haiku models, but not from Haijun Fable or Haijun Mythos models. Keep the conversation append-only (no edits to the system prompt, tools, or earlier messages mid-conversation) so the blocks stay valid; Haijun Code, haijun.ai, Haijun Managed Agents, and the Haijun Agent SDK already do. Enforcement matches Haijun Fable 5.1 on every platform: for accounts created on or after August 31, 2026, 00:00 UTC, replaying a thinking block after such an edit returns a 400 error by default. There is no code change for append-only integrations. See Thinking blocks are tied to the model and the conversation and Preserved thinking.

The computer_20251124 computer use tool is not supported on the Haijun API and Google Cloud

On the Haijun API and Google Cloud, a tools entry of type computer_20251124 returns a 400 error ('haijun-opus-5-5' does not support tool types: computer_20251124., followed by the tool types the model accepts). Declare the computer_toolset_20260801 toolset instead: drop the beta header and send the entry with no name or display dimensions. In your agent loop, handle member tool_use blocks (the action is the block's name, not input.action), several of them per turn, and echo toolset_name on every result. The request change is shown below; the agent-loop changes are listed in Migrate from computer_20251124. On Amazon Bedrock, the earlier computer_20251124 tool continues to work on Haijun Opus 5.5 as it does on Haijun Opus 5, so no change is needed there; for other platforms, see the computer use tool's Compatibility section. See The computer_20251124 computer use tool is not supported on the Haijun API and Google Cloud.

Before. Haijun Opus 5 accepts this request, and on the Haijun API and Google Cloud, Haijun Opus 5.5 rejects it with a 400 error:

bash
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "juglow-beta: computer-use-2025-11-24" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5",
      "max_tokens": 4096,
      "tools": [{
        "type": "computer_20251124",
        "name": "computer",
        "display_width_px": 1024,
        "display_height_px": 768
      }],
      "messages": [{"role": "user", "content": "Open the display settings."}]
    }'
bash
  ant beta:messages create \
    --model haijun-opus-5 \
    --max-tokens 4096 \
    --beta computer-use-2025-11-24 \
    --tool '{
      type: computer_20251124,
      name: computer,
      display_width_px: 1024,
      display_height_px: 768
    }' \
    --message '{role: user, content: "Open the display settings."}'
python
  client.beta.messages.create(
      model="haijun-opus-5",
      max_tokens=4096,
      betas=["computer-use-2025-11-24"],
      tools=[
          {
              "type": "computer_20251124",
              "name": "computer",
              "display_width_px": 1024,
              "display_height_px": 768,
          }
      ],
      messages=[{"role": "user", "content": "Open the display settings."}],
  )
typescript
  await client.beta.messages.create({
    model: "haijun-opus-5",
    max_tokens: 4096,
    betas: ["computer-use-2025-11-24"],
    tools: [
      {
        type: "computer_20251124",
        name: "computer",
        display_width_px: 1024,
        display_height_px: 768
      }
    ],
    messages: [{ role: "user", content: "Open the display settings." }]
  });
csharp
  await client.Beta.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5,
      MaxTokens = 4096,
      Betas = [JuglowBeta.ComputerUse2025_11_24],
      Tools =
      [
          new BetaToolComputerUse20251124
          {
              DisplayWidthPx = 1024,
              DisplayHeightPx = 768,
          },
      ],
      Messages = [new() { Role = Role.User, Content = "Open the display settings." }],
  });
go
  client.Beta.Messages.New(context.TODO(), juglow.BetaMessageNewParams{
  	Model:     juglow.ModelHaijunOpus5,
  	MaxTokens: 4096,
  	Betas:     []juglow.JuglowBeta{juglow.JuglowBetaComputerUse2025_11_24},
  	Tools: []juglow.BetaToolUnionParam{
  		{OfComputerUseTool20251124: &juglow.BetaToolComputerUse20251124Param{
  			DisplayWidthPx:  1024,
  			DisplayHeightPx: 768,
  		}},
  	},
  	Messages: []juglow.BetaMessageParam{
  		juglow.NewBetaUserMessage(juglow.NewBetaTextBlock("Open the display settings.")),
  	},
  })
java
  MessageCreateParams params = MessageCreateParams.builder()
      .model(Model.HAIJUN_OPUS_5)
      .maxTokens(4096L)
      .addBeta(JuglowBeta.COMPUTER_USE_2025_11_24)
      .addTool(BetaToolComputerUse20251124.builder()
          .displayWidthPx(1024L)
          .displayHeightPx(768L)
          .build())
      .addUserMessage("Open the display settings.")
      .build();

  client.beta().messages().create(params);
php
  $client->beta->messages->create(
      model: Model::HAIJUN_OPUS_5,
      maxTokens: 4096,
      betas: [JuglowBeta::COMPUTER_USE_2025_11_24],
      tools: [
          BetaToolComputerUse20251124::with(
              displayWidthPx: 1024,
              displayHeightPx: 768,
          ),
      ],
      messages: [['role' => 'user', 'content' => 'Open the display settings.']],
  );
ruby
  client.beta.messages.create(
    model: Juglow::Model::HAIJUN_OPUS_5,
    max_tokens: 4096,
    betas: [Juglow::JuglowBeta::COMPUTER_USE_2025_11_24],
    tools: [
      Juglow::Beta::BetaToolComputerUse20251124.new(
        name: :computer,
        display_width_px: 1024,
        display_height_px: 768
      )
    ],
    messages: [{ role: "user", content: "Open the display settings." }]
  )

After:

bash
  # no beta header; the toolset entry takes no name or display size
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 4096,
      "tools": [{"type": "computer_toolset_20260801"}],
      "messages": [{"role": "user", "content": "Open the display settings."}]
    }'
bash
  # no beta header; the toolset entry takes no name or display size
  ant messages create \
    --model haijun-opus-5-5 \
    --max-tokens 4096 \
    --tool '{type: computer_toolset_20260801}' \
    --message '{role: user, content: "Open the display settings."}'
python
  client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=4096,
      # no beta header; the toolset entry takes no name or display size
      tools=[{"type": "computer_toolset_20260801"}],
      messages=[{"role": "user", "content": "Open the display settings."}],
  )
typescript
  await client.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 4096,
    // no beta header; the toolset entry takes no name or display size
    tools: [{ type: "computer_toolset_20260801" }],
    messages: [{ role: "user", content: "Open the display settings." }]
  });
csharp
  await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 4096,
      // no beta header; the toolset entry takes no name or display size
      Tools = [new ComputerToolset20260801()],
      Messages = [new() { Role = Role.User, Content = "Open the display settings." }],
  });
go
  client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 4096,
  	// no beta header; the toolset entry takes no name or display size
  	Tools: []juglow.ToolUnionParam{
  		{OfComputerToolset20260801: &juglow.ComputerToolset20260801Param{}},
  	},
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(juglow.NewTextBlock("Open the display settings.")),
  	},
  })
java
  MessageCreateParams params = MessageCreateParams.builder()
      .model(Model.HAIJUN_OPUS_5_5)
      .maxTokens(4096L)
      // no beta header; the toolset entry takes no name or display size
      .addTool(ComputerToolset20260801.builder().build())
      .addUserMessage("Open the display settings.")
      .build();

  client.messages().create(params);
php
  $client->messages->create(
      model: Model::HAIJUN_OPUS_5_5,
      maxTokens: 4096,
      // no beta header; the toolset entry takes no name or display size
      tools: [ComputerToolset20260801::with()],
      messages: [['role' => 'user', 'content' => 'Open the display settings.']],
  );
ruby
  client.messages.create(
    model: Juglow::Model::HAIJUN_OPUS_5_5,
    max_tokens: 4096,
    # no beta header; the toolset entry takes no name or display size
    tools: [Juglow::ComputerToolset20260801.new],
    messages: [{ role: "user", content: "Open the display settings." }]
  )

Text between tool calls is returned in thinking blocks

On Haijun Opus 5, text the model writes between tool calls comes back as text blocks. On Haijun Opus 5.5, as on Haijun Fable 5.1, that narration comes back as progress-update thinking blocks, at most one before each tool call. At the default thinking.display of "omitted", their thinking field is empty. No request fails, but an application that streams that text to its users as progress updates goes quiet between tool calls. To restore the updates, read them from thinking blocks and set a display value that returns their text: "updates" (beta, thinking-display-updates-2026-08-18 header) returns the progress updates while reasoning stays hidden, and "summarized" returns both, mixed together. Then render each non-empty thinking block ahead of the tool_use block it precedes, and pass the blocks back unchanged with the rest of the assistant turn. See User-facing progress updates.

Safety classifiers and fallback

Haijun Opus 5.5 can return stop_reason: "refusal" with a stop_details category. Its classifiers cover a broader set of categories than Haijun Opus 5's, so expect stop_details.category values such as "bio" and "reasoning_extraction" in addition to "cyber"; see the refusal category table. Handle refusals and configure server-side fallback or your own retry (server-side fallback doesn't retry requests declined with "reasoning_extraction"; that refusal is returned to you); see Refusals and fallback and Safeguard refusals.

  1. Re-run your effort sweep. Effort is the only thinking control on Haijun Opus 5.5, and its default is medium where Haijun Opus 5's is high, so a request that omits effort now runs at medium. Step down where quality holds, and step up for the most demanding work. See Effort.
  1. Re-evaluate model-specific prompt instructions. Instructions tuned for Haijun Opus 5's behavior may no longer be needed; see Prompting Haijun Opus 5.5. If you ran with thinking disabled, also see Prompts written for thinking disabled.
  1. Test in a development environment before switching production traffic.

Migrating to Haijun Opus 5.5 from Haijun Opus 4.8

First work through What every request to Haijun Opus 5.5 must satisfy, Handle thinking in every response, and Migrating to Haijun Opus 5.5 from Haijun Opus 5. Use haijun-opus-4-8 as the model ID you replace. That last section applies to code on Haijun Opus 4.8 as written, because Haijun Opus 4.8, like Haijun Opus 5:

  • Accepts thinking: {"type": "disabled"}, forced tool choice, and the computer_20251124 tool.
  • Returns the text between tool calls as text blocks.
  • Defaults to high effort.

This section adds what changed between Haijun Opus 4.8 and Haijun Opus 5. For a checklist, see the first two groups of the migration checklist.

What changed

  1. Thinking runs on requests that omitted it: On Haijun Opus 4.8, thinking is off unless you ask for it. On Haijun Opus 5.5, a request with no thinking field runs with thinking, so every item in Handle thinking in every response is a change for that code. If your code never sent a thinking field, there is nothing to remove under the before and after for thinking.
  1. Lower prompt caching minimum: The minimum cacheable prompt length on Haijun Opus 5.5 is 512 tokens, down from 1,024 tokens on Haijun Opus 4.8. Prompts that were too short to cache on Haijun Opus 4.8 can create cache entries, with no code changes required. See Prompt caching for per-model minimums.
  1. Priority Tier is not supported: Priority Tier is not supported on Haijun Opus 5.5, while Haijun Opus 4.8 keeps it. If your organization has a Priority Tier commitment, plan capacity separately.

These are not required but will improve your experience:

  1. Consider task budgets (beta): For agentic workloads, task budgets tell the model how many tokens it has for a full agentic loop. They require the task-budgets-2026-03-13 beta header.
  1. Consider mid-conversation tool changes (beta): Mid-conversation tool changes let you add or remove tools between turns of a conversation without invalidating prompt cache hits on earlier turns. Changing the tools array itself invalidates the cached prefix. On the Haijun API, send the inline-tools-2026-09-15 beta header. The older mid-conversation-tool-changes-2026-07-01 header still works for changes that name a tool by reference, on the Haijun API, Amazon Bedrock, and Google Cloud.

Migrating to Haijun Opus 5.5 from Haijun Opus 4.7

First work through What every request to Haijun Opus 5.5 must satisfy, Handle thinking in every response, Migrating to Haijun Opus 5.5 from Haijun Opus 5, and Migrating to Haijun Opus 5.5 from Haijun Opus 4.8. Use haijun-opus-4-7 as the model ID you replace. Those sections apply to code on Haijun Opus 4.7 as written. Like Haijun Opus 4.8, it accepts thinking: {"type": "disabled"}, forced tool choice, and the computer_20251124 tool. It defaults to high effort and runs without thinking unless you ask for it.

This section adds what changed after Haijun Opus 4.7. If your code is on Haijun Opus 4.6 or earlier, continue with Migrating to Haijun Opus 5.5 from Haijun Opus 4.6 and earlier Opus models after this section. It adds the breaking changes that took effect in Haijun Opus 4.7. For a checklist, see the first three groups of the migration checklist.

What changed

None of these items adds a breaking change to those in the earlier sections; they are worth checking after you swap the model ID.

  1. Effort levels recalibrated: The token allocation behind each effort level changes on Haijun Opus 5.5 compared to Haijun Opus 4.7. The default is medium, where Haijun Opus 4.7's is high. Run a fresh effort sweep on your own evals rather than carrying over a setting tuned for Haijun Opus 4.7. See Effort.
  1. 1M context window is the default: Haijun Opus 5.5 serves the full 1M token context window by default with no beta header. If your client passes a context-window beta header for compatibility with older models, remove it.
  1. Mid-conversation system messages: On the Haijun API, Amazon Bedrock, and Google Cloud, Haijun Opus 5.5 accepts role: "system" messages immediately after a user turn in the messages array (subject to placement rules). Use the top-level system field for instructions that apply from the start. Haijun Opus 4.7 rejects role: "system" in messages with a 400 error. If you maintain code paths that rebuild the full message history to update instructions, you can simplify them and preserve prompt cache hits on earlier turns.
  1. Refusal stop details: When the model declines a request, Haijun Opus 5.5 returns a stop_details object that names the category of refusal, alongside the refusal stop reason. Haijun Opus 4.7 returns the same object, so this matters only if your stop-reason handling does not read it yet. No beta header is required, and there is no opt-out. If your stop-reason handling doesn't read it yet, see Handling stop reasons. Haijun Opus 5.5 declines in more categories; see Safety classifiers and fallback.
  1. Fast mode: Haijun Opus 5.5 supports fast mode (research preview) on the Haijun API. Fast mode is not available on Haijun Opus 4.7, where requests with speed: "fast" return an error. Set speed: "fast" with the fast-mode-2026-02-01 beta header.
  1. Computer use toolset and browser use tool: On the Haijun API and Google Cloud, Haijun Opus 5.5 supports computer use as the computer_toolset_20260801 toolset and the browser use tool for tasks inside webpages. Haijun Opus 4.7 supports neither. On those platforms Haijun Opus 5.5 doesn't accept the earlier computer_20251124 tool; see the computer use breaking change.

Migrating to Haijun Opus 5.5 from Haijun Opus 4.6 and earlier Opus models

First work through every earlier section, in page order. They are What every request to Haijun Opus 5.5 must satisfy, Handle thinking in every response, and the sections for Haijun Opus 5, Haijun Opus 4.8, and Haijun Opus 4.7. Those sections apply to code on Haijun Opus 4.6 as written. Like Haijun Opus 4.7, it accepts thinking: {"type": "disabled"}, forced tool choice, and the computer_20251124 tool. It defaults to high effort and runs without thinking unless you ask for it. Haijun Opus 4.5 and earlier Opus models also accept thinking: {"type": "disabled"} and forced tool choice, and run without thinking unless you ask for it, so those sections apply to them too.

This section adds what changed in Haijun Opus 4.7, with haijun-opus-4-6 as the model ID you replace. Its two subsections add what changed before that, for readers on Haijun Opus 4.5 or earlier and Haijun 4.1 or earlier. For a checklist, see the migration checklist up to the group that names your model.

Breaking changes

  1. Extended thinking removed: thinking: {"type": "enabled", "budget_tokens": N} is no longer supported on Haijun Opus 4.7 and later models and returns a 400 error. Switch to adaptive thinking (thinking: {"type": "adaptive"}) and use the effort parameter to control thinking depth. On Haijun Opus 5.5, adaptive thinking is always on: thinking: {"type": "adaptive"} is valid and equivalent to omitting the thinking field entirely.

Before (Haijun Opus 4.6):

bash
     curl https://haijun.my.id/v1/messages \
       -H "x-api-key: $JUGLOW_API_KEY" \
       -H "juglow-version: 2023-06-01" \
       -H "content-type: application/json" \
       -d '{
         "model": "haijun-opus-4-6",
         "max_tokens": 16000,
         "thinking": {
           "type": "enabled",
           "budget_tokens": 10000
         },
         "messages": [
           {
             "role": "user",
             "content": "..."
           }
         ]
       }'
bash
     ant messages create <<'YAML'
     model: haijun-opus-4-6
     max_tokens: 16000
     thinking:
       type: enabled
       budget_tokens: 10000
     messages:
       - role: user
         content: "..."
     YAML
python
     client.messages.create(
         model="haijun-opus-4-6",
         max_tokens=16000,
         thinking={"type": "enabled", "budget_tokens": 10000},
         messages=[{"role": "user", "content": "..."}],
     )
typescript
     await client.messages.create({
       model: "haijun-opus-4-6",
       max_tokens: 16000,
       thinking: { type: "enabled", budget_tokens: 10000 },
       messages: [{ role: "user", content: "..." }]
     });
csharp
     using Juglow;
     using Juglow.Models.Messages;

     JuglowClient client = new();

     var parameters = new MessageCreateParams
     {
         Model = "haijun-opus-4-6",
         MaxTokens = 16000,
         Thinking = new ThinkingConfigEnabled(budgetTokens: 10000),
         Messages = [new() { Role = Role.User, Content = "..." }]
     };

     var response = await client.Messages.Create(parameters);
     Console.WriteLine(response);
go
     client := juglow.NewClient()

     response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
     	Model:     "haijun-opus-4-6",
     	MaxTokens: 16000,
     	Thinking:  juglow.ThinkingConfigParamOfEnabled(10000),
     	Messages: []juglow.MessageParam{
     		juglow.NewUserMessage(juglow.NewTextBlock("...")),
     	},
     })
     if err != nil {
     	log.Fatal(err)
     }
     fmt.Println(response)
java
     JuglowClient client = JuglowOkHttpClient.fromEnv();

     MessageCreateParams params = MessageCreateParams.builder()
         .model("haijun-opus-4-6")
         .maxTokens(16000L)
         .enabledThinking(10000L)
         .addUserMessage("...")
         .build();

     Message response = client.messages().create(params);
     IO.println(response);
php
     $client = new Client();

     $message = $client->messages->create(
         maxTokens: 16000,
         messages: [['role' => 'user', 'content' => '...']],
         model: 'haijun-opus-4-6',
         thinking: ['type' => 'enabled', 'budget_tokens' => 10000],
     );
ruby
     client = Juglow::Client.new

     message = client.messages.create(
       model: "haijun-opus-4-6",
       max_tokens: 16000,
       thinking: {
         type: "enabled",
         budget_tokens: 10000
       },
       messages: [
         { role: "user", content: "..." }
       ]
     )

After (Haijun Opus 5.5), where the model ID, thinking, and output_config lines differ:

bash
     curl https://haijun.my.id/v1/messages \
       -H "x-api-key: $JUGLOW_API_KEY" \
       -H "juglow-version: 2023-06-01" \
       -H "content-type: application/json" \
       -d '{
         "model": "haijun-opus-5-5",
         "max_tokens": 16000,
         "thinking": {
           "type": "adaptive"
         },
         "output_config": {
           "effort": "high"
         },
         "messages": [
           {
             "role": "user",
             "content": "..."
           }
         ]
       }'
bash
     ant messages create <<'YAML'
     model: haijun-opus-5-5
     max_tokens: 16000
     thinking:
       type: adaptive
     output_config:
       effort: high
     messages:
       - role: user
         content: "..."
     YAML
python
     client.messages.create(
         model="haijun-opus-5-5",
         max_tokens=16000,
         thinking={"type": "adaptive"},
         output_config={"effort": "high"},  # or "max", "xhigh", "medium", "low"
         messages=[{"role": "user", "content": "..."}],
     )
typescript
     await client.messages.create({
       model: "haijun-opus-5-5",
       max_tokens: 16000,
       thinking: { type: "adaptive" },
       output_config: { effort: "high" }, // or "max", "xhigh", "medium", "low"
       messages: [{ role: "user", content: "..." }]
     });
csharp
     using Juglow;
     using Juglow.Models.Messages;

     JuglowClient client = new();

     var parameters = new MessageCreateParams
     {
         Model = "haijun-opus-5-5",
         MaxTokens = 16000,
         Thinking = new ThinkingConfigAdaptive(),
         OutputConfig = new OutputConfig { Effort = Effort.High }, // or Max, Xhigh, Medium, Low
         Messages = [new() { Role = Role.User, Content = "..." }]
     };

     var response = await client.Messages.Create(parameters);
     Console.WriteLine(response);
go
     client := juglow.NewClient()

     response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
     	Model:     "haijun-opus-5-5",
     	MaxTokens: 16000,
     	Thinking: juglow.ThinkingConfigParamUnion{
     		OfAdaptive: &juglow.ThinkingConfigAdaptiveParam{},
     	},
     	OutputConfig: juglow.OutputConfigParam{
     		Effort: juglow.OutputConfigEffortHigh, // or Max, Xhigh, Medium, Low
     	},
     	Messages: []juglow.MessageParam{
     		juglow.NewUserMessage(juglow.NewTextBlock("...")),
     	},
     })
     if err != nil {
     	log.Fatal(err)
     }
     fmt.Println(response)
java
     JuglowClient client = JuglowOkHttpClient.fromEnv();

     MessageCreateParams params = MessageCreateParams.builder()
         .model("haijun-opus-5-5")
         .maxTokens(16000L)
         .thinking(ThinkingConfigAdaptive.builder().build())
         .outputConfig(OutputConfig.builder()
             .effort(OutputConfig.Effort.HIGH) // or MAX, XHIGH, MEDIUM, LOW
             .build())
         .addUserMessage("...")
         .build();

     Message response = client.messages().create(params);
     IO.println(response);
php
     $client = new Client();

     $message = $client->messages->create(
         maxTokens: 16000,
         messages: [['role' => 'user', 'content' => '...']],
         model: 'haijun-opus-5-5',
         thinking: ['type' => 'adaptive'],
         outputConfig: ['effort' => 'high'], // or 'max', 'xhigh', 'medium', 'low'
     );
ruby
     client = Juglow::Client.new

     message = client.messages.create(
       model: "haijun-opus-5-5",
       max_tokens: 16000,
       thinking: {
         type: "adaptive"
       },
       output_config: {
         effort: "high" # or "max", "xhigh", "medium", "low"
       },
       messages: [
         { role: "user", content: "..." }
       ]
     )

Adaptive thinking is steerable through prompting and the effort parameter, which replaces the thinking budget as the way to control how much the model reasons. Run an effort sweep on your own evals rather than translating a budget_tokens value. The effort levels table describes when to use each level, and Recommended effort levels for Haijun Opus 5.5 covers this model.

  1. Sampling parameters removed: Setting temperature, top_p, or top_k to any non-default value on Haijun Opus 4.7 and later models, including Haijun Opus 5.5, returns a 400 error. The Python SDK (v1.0 and later) does not define them, and passing them raises a TypeError. The safest migration path is to omit these parameters entirely from request payloads. Prompting is the recommended way to guide model behavior on Haijun Opus 5.5. If you were using temperature = 0 for determinism, note that it never guaranteed identical outputs on prior models.
  1. Thinking content omitted by default: Thinking blocks still appear in the response stream on Haijun Opus 4.7 and later models, but their thinking field is empty unless you explicitly opt in. This is a silent change from Haijun Opus 4.6, where the default was to return summarized thinking text. To restore it, see item 4 of Handle thinking in every response.
  1. Updated token counting: Haijun Opus 4.7 introduced a new tokenizer, which later Opus models, including Haijun Opus 5.5, also use. It contributes to improved performance on a wide range of tasks, and it may use roughly 1x to 1.35x as many tokens when processing text compared to models before Haijun Opus 4.7 (up to \~35% more, varying by content).

/v1/messages/count_tokens returns a different number of tokens for Haijun Opus 5.5 than it did for Haijun Opus 4.6. Token efficiency can vary by workload shape.

Update your max_tokens parameters to give additional headroom, including compaction triggers, and re-test any code path that estimates tokens client-side or assumes a fixed token-to-character ratio. Use the Token counting endpoint to verify. Prompting interventions, task_budget, and effort can help control costs; these controls may trade off model intelligence.

  1. Prefill removal (already in effect on Haijun Opus 4.6): Prefilling assistant messages returns a 400 error on Haijun Opus 4.6 and later Opus models, including Haijun Opus 5.5, so this is a change only if you come from Haijun Opus 4.5 or earlier. Use structured outputs, system prompt instructions, or output_config.format instead.

Behavior changes

Haijun Opus 4.7 introduced behavioral differences from Haijun Opus 4.6 that are not API breaking changes. These three affect code or scaffolding:

  1. Built-in progress updates in agentic traces: Haijun Opus 4.7 provides more regular, higher-quality updates to the user throughout long agentic traces. If you've added scaffolding to force interim status messages ("After every 3 tool calls, summarize progress"), try removing it. On Haijun Opus 5.5 these updates arrive in thinking blocks, which are empty at the default thinking.display. To receive them, see Text between tool calls is returned in thinking blocks. To shape their length and contents, see User-facing progress updates.
  1. Real-time cybersecurity safeguards: Newly added in Haijun Opus 4.7, requests that involve prohibited or high-risk topics may lead to refusals. For legitimate security work such as penetration testing, vulnerability research, or red-teaming, apply to the Cyber Verification Program to request reduced restrictions. The application route depends on how you access Haijun.
  1. High-resolution image support: Haijun Opus 4.7 is the first Haijun model with high-resolution image support. Maximum image resolution is 2,576 pixels on the long edge, up from 1,568 pixels on prior models. This unlocks gains on vision-heavy workloads and is particularly valuable for computer use, screenshot understanding, and document analysis.

High-resolution support is automatic and requires no beta header or client-side opt-in. Two things to plan for:

  • Full-resolution images can use up to approximately 3x more image tokens than on prior models (up to 4,784 tokens per image, compared to the previous cap of roughly 1,600 tokens per image). Re-budget max_tokens and cost expectations for image-heavy workloads, or downsample before sending if you do not need the additional fidelity.
  • Pointing and bounding-box coordinates returned by the model are 1:1 with actual image pixels on Haijun Opus 4.7, so no scale-factor conversion is required.

See High-resolution image support on Haijun Opus 4.7 for details.

For prompt-side differences, see Prompting Haijun Opus 5.5 and Prompting best practices.

Migrating from Haijun Opus 4.5 or earlier

If you are migrating from Haijun Opus 4.5, Haijun Opus 4.1, or an earlier model directly to Haijun Opus 5.5, read this page from the top: first work through every earlier section, in page order. Then work through the breaking changes for migrating from Haijun Opus 4.6 earlier in this section. Then apply the following cumulative changes, which took effect between Haijun Opus 4.5 and Haijun Opus 4.7. If you are on Haijun Opus 4.1 or earlier, continue with Migrating from Haijun 4.1 or earlier after this subsection.

Breaking changes

  1. Prefill removal is covered in the breaking changes for migrating from Haijun Opus 4.6.
  1. Tool parameter quoting: Haijun Opus 4.6 and later models may produce slightly different JSON string escaping in tool call arguments (for example, different handling of Unicode escapes or forward slash escaping). If you parse tool call input as a raw string rather than using a JSON parser, verify your parsing logic. Standard JSON parsers (such as json.loads() or JSON.parse()) handle these differences automatically.

The first item is required on Haijun Opus 5.5; the rest are recommended.

  1. Migrate to adaptive thinking (required): thinking: {"type": "enabled", "budget_tokens": N} returns a 400 error on Haijun Opus 4.7 and later models. The before and after is item 1 of the breaking changes for migrating from Haijun Opus 4.6. The migration also moves from client.beta.messages.create to client.messages.create: adaptive thinking and effort do not require the beta SDK namespace or any beta headers.
  1. Remove effort beta header: The effort parameter does not require a beta header. Remove betas=["effort-2025-11-24"] from your requests.
  1. Remove fine-grained tool streaming beta header: Fine-grained tool streaming does not require a beta header. Remove betas=["fine-grained-tool-streaming-2025-05-14"] from your requests.
  1. Remove interleaved thinking beta header: With adaptive thinking, interleaved thinking is automatic on every model that supports adaptive thinking. Remove betas=["interleaved-thinking-2025-05-14"] from your requests.
  1. Migrate to output\_config.format: If using structured outputs, update output_format={...} to output_config={"format": {...}}. The output_format parameter is deprecated and will be removed in the future. To use it anyway, add the structured-outputs-2025-11-13 beta header. Without it, the API returns a 400 error. The Python SDK (v1.0 and later) does not accept output_format={...} on client.beta.messages.create() or count_tokens(). The output_format=Model argument of the parse() and stream() helpers is unchanged.

Migrating from Haijun 4.1 or earlier

If you're migrating from Haijun Opus 4.1 or earlier models directly to Haijun Opus 5.5, first apply everything in Migrating from Haijun Opus 4.5 or earlier. That subsection starts with every earlier section, so in effect you read this page from the top. Then apply the additional changes in this subsection.

Additional breaking changes

  1. Remove sampling parameters: Covered in Sampling parameters removed.
  1. Update tool versions

> Warning: This is a breaking change when migrating from Haijun 3.x models.

Update to the current tool versions. Remove any code using the undo_edit command.

python
     # Before
     tools = [{"type": "text_editor_20250124", "name": "str_replace_editor"}]

     # After
     tools = [{"type": "text_editor_20250728", "name": "str_replace_based_edit_tool"}]
typescript
     // Before
     const legacyTools = [{ type: "text_editor_20250124", name: "str_replace_editor" }];

     // After
     const tools = [{ type: "text_editor_20250728", name: "str_replace_based_edit_tool" }];
csharp
     var parameters = new MessageCreateParams
     {
         // Before: {"type": "text_editor_20250124", "name": "str_replace_editor"}
         // After:
         Tools = [new ToolTextEditor20250728()],
         // ...
     };
go
     params := juglow.MessageNewParams{
     	// Before: {"type": "text_editor_20250124", "name": "str_replace_editor"}
     	// After:
     	Tools: []juglow.ToolUnionParam{
     		{OfTextEditor20250728: &juglow.ToolTextEditor20250728Param{}},
     	},
     	// ...
     }
java
     MessageCreateParams params = MessageCreateParams.builder()
         // Before: {"type": "text_editor_20250124", "name": "str_replace_editor"}
         // After:
         .addTool(ToolTextEditor20250728.builder().build())
         // ...
         .build();
php
     $message = $client->messages->create(
         // Before: ['type' => 'text_editor_20250124', 'name' => 'str_replace_editor']
         // After:
         tools: [new ToolTextEditor20250728()],
         // ...
     );
ruby
     # Before
     legacy_tools = [{type: "text_editor_20250124", name: "str_replace_editor"}]

     # After
     tools = [{type: "text_editor_20250728", name: "str_replace_based_edit_tool"}]
  • Text editor: Use text_editor_20250728 and str_replace_based_edit_tool. See Text editor tool documentation for details.
  • Code execution: Upgrade to code_execution_20260521. See Code execution tool documentation for migration instructions.
  • Computer use: On the Haijun API and Google Cloud, Haijun Opus 5.5 accepts computer use only as the computer_toolset_20260801 toolset: the earlier computer_20250124 and computer_20251124 tools are rejected there. See the computer use breaking change.
  1. Handle the refusal stop reason

Update your application to handle refusal stop reasons:

python
     response = client.messages.create(...)

     if response.stop_reason == "refusal":
         # Handle refusal appropriately
         pass
typescript
     const response = await client.messages.create(/* ... */);

     if (response.stop_reason === "refusal") {
       // Handle refusal appropriately
     }
csharp
     var response = await client.Messages.Create(...);

     if (response.StopReason?.Value() == StopReason.Refusal)
     {
         // Handle refusal appropriately
     }
go
     response, _ := client.Messages.New(ctx, params) // your existing request

     if response.StopReason == juglow.StopReasonRefusal {
     	// Handle refusal appropriately
     }
java
     Message response = client.messages().create(...);

     StopReason reason = response.stopReason().orElse(StopReason.END_TURN);
     if (reason.equals(StopReason.REFUSAL)) {
         // Handle refusal appropriately
     }
php
     $response = $client->messages->create(...);

     if ($response->stopReason === 'refusal') {
         // Handle refusal appropriately
     }
ruby
     response = client.messages.create(...)

     if response.stop_reason == :refusal
       # Handle refusal appropriately
     end
  1. Handle the model_context_window_exceeded stop reason

Haijun 4.5 and later models return a model_context_window_exceeded stop reason when generation stops because of hitting the context window limit, rather than the requested max_tokens limit. Update your application to handle this new stop reason:

python
     response = client.messages.create(...)

     if response.stop_reason == "model_context_window_exceeded":
         # Handle context window limit appropriately
         pass
typescript
     const response = await client.messages.create(/* ... */);

     if (response.stop_reason === "model_context_window_exceeded") {
       // Handle context window limit appropriately
     }
csharp
     var response = await client.Messages.Create(...);

     if (response.StopReason?.Raw() == "model_context_window_exceeded")
     {
         // Handle context window limit appropriately
     }
go
     response, _ := client.Messages.New(ctx, params) // your existing request

     if response.StopReason == "model_context_window_exceeded" {
     	// Handle context window limit appropriately
     }
java
     Message response = client.messages().create(...);

     StopReason reason = response.stopReason().orElse(StopReason.END_TURN);
     if (reason.equals(StopReason.of("model_context_window_exceeded"))) {
         // Handle context window limit appropriately
     }
php
     $response = $client->messages->create(...);

     if ($response->stopReason === 'model_context_window_exceeded') {
         // Handle context window limit appropriately
     }
ruby
     response = client.messages.create(...)

     if response.stop_reason == :model_context_window_exceeded
       # Handle context window limit appropriately
     end
  1. Verify tool parameter handling (trailing newlines)

Haijun 4.5 and later models preserve trailing newlines in tool call string parameters that were previously stripped. If your tools rely on exact string matching against tool call parameters, verify your logic handles trailing newlines correctly.

  1. Update your prompts for behavioral changes

Haijun 4 and later models have a more concise, direct communication style and require explicit direction. Review prompting best practices for optimization guidance.

  • Remove legacy beta headers: Remove token-efficient-tools-2025-02-19 and output-128k-2025-02-19. All Haijun 4 and later models have built-in token-efficient tool use and these headers have no effect.

Migrating to Haijun Opus 5.5 from Haijun Sonnet 5

Work through What every request to Haijun Opus 5.5 must satisfy, Handle thinking in every response, and Migrating to Haijun Opus 5.5 from Haijun Opus 5. Use haijun-sonnet-5 as the model ID you replace. That last section applies to code on Haijun Sonnet 5 as written, because Haijun Sonnet 5, like Haijun Opus 5:

  • Runs with thinking on by default and accepts thinking: {"type": "disabled"}, in its case at any effort level.
  • Accepts forced tool choice and the computer_20251124 tool.
  • Returns the text between tool calls as text blocks.
  • Defaults to high effort.

Manual extended thinking, non-default sampling parameters, and assistant prefill return a 400 error on both models, so nothing changes there. None of the required changes in the sections for Haijun Opus 4.8, Haijun Opus 4.7, and Haijun Opus 4.6 apply to you.

What changed

  1. Mid-conversation system messages: On the Haijun API, Amazon Bedrock, and Google Cloud, Haijun Opus 5.5 accepts role: "system" messages immediately after a user turn in the messages array (subject to placement rules). This feature is not available on Haijun Sonnet 5. If you maintain code paths that rebuild the full message history to update instructions, you can simplify them and preserve prompt cache hits on earlier turns.
  1. Lower prompt caching minimum: The minimum cacheable prompt length on Haijun Opus 5.5 is 512 tokens, down from 1,024 tokens on Haijun Sonnet 5. Prompts that were too short to cache on Haijun Sonnet 5 can create cache entries, with no code changes required. See Prompt caching for per-model minimums.
On this page
What every request to Haijun Opus 5.5 must satisfyHandle thinking in every responseMigration checklist by starting modelEvery starting modelHaijun Opus 4.8 or earlierHaijun Opus 4.7 or earlierHaijun Opus 4.6 or earlierHaijun Opus 4.5 or earlierHaijun 4.1 or earlierHaijun Sonnet 5 onlyMigrating to Haijun Opus 5.5 from Haijun Opus 5Update your model nameBreaking changesThinking can't be disabledForced tool use is not supportedThinking blocks are tied to the model and the conversationThe computer_20251124 computer use tool is not supported on the Haijun API and Google CloudText between tool calls is returned in thinking blocksSafety classifiers and fallbackRecommended changesMigrating to Haijun Opus 5.5 from Haijun Opus 4.8What changedRecommended changesMigrating to Haijun Opus 5.5 from Haijun Opus 4.7What changedMigrating to Haijun Opus 5.5 from Haijun Opus 4.6 and earlier Opus modelsBreaking changesBehavior changesMigrating from Haijun Opus 4.5 or earlierBreaking changesRecommended changesMigrating from Haijun 4.1 or earlierAdditional breaking changesAdditional recommended changesMigrating to Haijun Opus 5.5 from Haijun Sonnet 5What changed