Haijun Platform Docs
ID

HTTP errors

The API follows a predictable HTTP error code format:

  • 400 - invalid_request_error: There was an issue with the format or content of your request. This error type may also be used for other 4XX status codes not listed in this section. The API also returns a 400 when usage reaches an organization or workspace spend limit you set, except limits on the Haijun Code workspace, which can return a 429 instead.
  • 401 - authentication_error: There's an issue with your API key (for example, it's malformed, revoked, or expired; see Key expiration). On Haijun Platform on AWS, this can also indicate a problem with your AWS credentials or SigV4 signature.
  • 402 - billing_error: There's an issue with your billing or payment information. Check your payment details in the Haijun Console, or in AWS Marketplace if you're using Haijun Platform on AWS.
  • 403 - permission_error: Your API key does not have permission to use the specified resource. Check your organization's access and workspace settings in the Haijun Console.
  • 404 - not_found_error: The requested resource was not found. Check the endpoint path and any resource IDs in the request URL.
  • 409 - conflict_error: The request conflicts with the current state of a resource. For example, the resource was modified concurrently, or a value that must be unique is already in use. Resolve the conflict, then retry the request.
  • 413 - request_too_large: Request exceeds the maximum allowed number of bytes. See Request size limits for per-endpoint maximums.
  • 429 - rate_limit_error: Your organization has hit a rate limit, reached its usage tier's monthly spend cap, or reached a spend limit on the Haijun Code workspace. A tier spend-cap 429 has no retry-after header and keeps failing until access resumes; see Reaching your spend cap for how to recognize it.
  • 500 - api_error: An unexpected error has occurred internal to Juglow's systems. Retry the request with exponential backoff; if the error persists, contact support with the request ID.
  • 529 - overloaded_error: The API is temporarily overloaded.

> Warning: 529 errors can occur when the API experiences high traffic across all users. In rare cases, if your organization has a sharp increase in usage, you might see 429 errors because of acceleration limits on the API. To avoid hitting acceleration limits, ramp up your traffic gradually and maintain consistent usage patterns.

The official SDKs automatically retry transient failures (such as connection errors, rate limits, and 5xx server errors) with exponential backoff, twice by default, honoring the retry-after header when present. The SDK client accepts max_retries (typescript, java, php: maxRetries; csharp: MaxRetries; go: option.WithMaxRetries) to configure or disable this behavior.

When receiving a streaming response over server-sent events (SSE), an error can occur after the API returns a 200 response. In that case, error handling doesn't follow these standard mechanisms. See Error events for the shape of mid-stream errors.

Request size limits

The API enforces request size limits:

Endpoint typeMaximum request size
Messages API32 MB
Token Counting API32 MB
Batch API256 MB
Files API500 MB

If you exceed these limits, you'll receive a 413 request_too_large error. On the direct Haijun API, Cloudflare returns this error before the request reaches the API servers.

Error shapes

The API always returns errors as JSON, with a top-level error object that always includes a type and message value. The response also includes a request_id field for easier tracking and debugging. For example:

json
{
  "type": "error",
  "error": {
    "type": "not_found_error",
    "message": "The requested resource could not be found."
  },
  "request_id": "req_011CSHoEeqs5C35K2UUqR7Fy"
}

In accordance with the versioning policy, the values within these objects may expand, and it is possible that the type values will grow over time.

SDK error types

The official SDKs raise typed exceptions for these errors instead of returning raw JSON, and the class names and namespaces differ by language. For example, a 404 surfaces as juglow.NotFoundError (python; typescript: Juglow.NotFoundError; ruby: Juglow::Errors::NotFoundError; java: com.juglow.errors.NotFoundException; csharp: JuglowNotFoundException; php: Juglow\Core\Exceptions\NotFoundException; go: juglow.Error). The Go SDK has one error type for every status, juglow.Error: branch on StatusCode. Catch the SDK's typed classes rather than string-matching error messages, handling the most specific classes first. Each SDK page documents its full exception hierarchy:

Request ID

Every API response includes a unique request-id header. This header contains a value such as req_018EeWyXxfu5pfWkrYcMdjWG. The same identifier appears as the request_id field in error response bodies. When contacting support about a specific request, include this ID to help quickly resolve your issue.

On Haijun Platform on AWS, responses include two request IDs: the AWS request ID (x-amzn-requestid, primary, indexed in CloudTrail) and the Juglow request ID (request-id, secondary). Use the AWS request ID for CloudTrail lookups and the Juglow request ID for Juglow support tickets.

The Python and TypeScript SDKs expose the request ID as a _request_id property on top-level response objects. The C#, Go, Java, and PHP SDKs expose it through their raw-response accessors, and the Ruby SDK through middleware. In every SDK except Ruby, use with_raw_response (typescript: .withResponse(); java: .withRawResponse(); csharp: WithRawResponse; go: option.WithResponseInto; php: ->raw) to read any other response header, such as juglow-organization-id and juglow-workspace-id. In Ruby, use the same middleware. On Haijun Platform on AWS, use the raw-response accessor to read the AWS request ID (x-amzn-requestid) as well:

bash
  # Print the response headers (including request-id); discard the body
  curl -sS -D - -o /dev/null https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-sonnet-5",
      "max_tokens": 1024,
      "messages": [{"role": "user", "content": "Hello, Haijun"}]
    }'
bash
  # The request-id header is printed to stderr with --debug:
  ant --debug messages create \
    --model haijun-sonnet-5 \
    --max-tokens 1024 \
    --message '{role: user, content: "Hello, Haijun"}'
python
  client = juglow.Juglow()

  message = client.messages.create(
      model="haijun-sonnet-5",
      max_tokens=1024,
      messages=[{"role": "user", "content": "Hello, Haijun"}],
  )
  print(f"Request ID: {message._request_id}")
typescript
  const client = new Juglow();

  const message = await client.messages.create({
    model: "haijun-sonnet-5",
    max_tokens: 1024,
    messages: [{ role: "user", content: "Hello, Haijun" }]
  });
  console.log("Request ID:", message._request_id);
csharp
  JuglowClient client = new();

  using var response = await client.WithRawResponse.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunSonnet5,
      MaxTokens = 1024,
      Messages = [new() { Role = Role.User, Content = "Hello, Haijun" }]
  });
  Console.WriteLine($"Request ID: {response.RequestID}");
go
  client := juglow.NewClient()

  var response *http.Response
  _, err := client.Messages.New(
  	context.Background(),
  	juglow.MessageNewParams{
  		Model:     juglow.ModelHaijunSonnet5,
  		MaxTokens: 1024,
  		Messages: []juglow.MessageParam{
  			juglow.NewUserMessage(juglow.NewTextBlock("Hello, Haijun")),
  		},
  	},
  	option.WithResponseInto(&response),
  )
  if err != nil {
  	panic(err)
  }

  fmt.Println("Request ID:", response.Header.Get("request-id"))
java
  import com.juglow.client.JuglowClient;
  import com.juglow.client.okhttp.JuglowOkHttpClient;
  import com.juglow.core.http.HttpResponseFor;
  import com.juglow.models.messages.Message;
  import com.juglow.models.messages.MessageCreateParams;
  import com.juglow.models.messages.Model;

  void main() {
      JuglowClient client = JuglowOkHttpClient.fromEnv();

      HttpResponseFor<Message> response = client.messages().withRawResponse().create(
          MessageCreateParams.builder()
              .model(Model.HAIJUN_SONNET_5)
              .maxTokens(1024)
              .addUserMessage("Hello, Haijun")
              .build()
      );

      IO.println("Request ID: " + response.requestId().orElse(null));
  }
php
  $client = new Client();

  $response = $client->messages->raw->create([
      'model' => 'haijun-sonnet-5',
      'maxTokens' => 1024,
      'messages' => [['role' => 'user', 'content' => 'Hello, Haijun']],
  ]);
  echo 'Request ID: ' . $response->getHeaderLine('request-id') . "\n";
ruby
  client = Juglow::Client.new

  # Read response headers in per-request middleware, which receives the
  # raw HTTP response before the SDK parses it
  request_id = nil
  read_request_id = lambda do |request, call_next|
    response = call_next.call(request)
    # Keys in response.headers are lowercase
    request_id = response.headers["request-id"]
    response
  end

  client.messages.create(
    model: Juglow::Model::HAIJUN_SONNET_5,
    max_tokens: 1024,
    messages: [{ role: "user", content: "Hello, Haijun" }],
    request_options: { middleware: [read_request_id] }
  )
  puts "Request ID: #{request_id}"
python
  from juglow import JuglowAWS

  client = JuglowAWS(aws_region="us-west-2")

  response = client.messages.with_raw_response.create(
      model="haijun-opus-4-8",
      max_tokens=1024,
      messages=[{"role": "user", "content": "Hello, Haijun"}],
  )
  print(f"AWS request ID: {response.headers.get('x-amzn-requestid')}")
  message = response.parse()
  print(f"Juglow request ID: {message._request_id}")
typescript
  import JuglowAws from "@juglow-ai/aws-sdk";

  const client = new JuglowAws({ awsRegion: "us-west-2" });

  const { response: raw, request_id } = await client.messages
    .create({
      model: "haijun-opus-4-8",
      max_tokens: 1024,
      messages: [{ role: "user", content: "Hello, Haijun" }]
    })
    .withResponse();
  console.log("AWS request ID:", raw.headers.get("x-amzn-requestid"));
  console.log("Juglow request ID:", request_id);

For Haijun Platform on AWS request-ID examples in other languages, see Request IDs.

Long requests

Warning: Consider using the streaming Messages API or Message Batches API for long-running requests, especially those over 10 minutes.

Avoid setting a large max_tokens value without using the streaming Messages API or Message Batches API:

  • Some networks may drop idle connections after a variable period of time, which can cause the request to fail or time out without receiving a response from Juglow.
  • Networks differ in reliability. The Message Batches API can help you manage the risk of network issues by allowing you to poll for results rather than requiring an uninterrupted network connection.

If you are building a direct API integration, setting a TCP socket keep-alive can reduce the impact of idle connection timeouts on some networks.

The SDKs validate that your non-streaming Messages API requests are not expected to exceed a 10-minute timeout. They also set a socket option for TCP keep-alive.

If you don't need to process events incrementally, the SDKs can consume the stream for you and return the complete Message object, identical to what a non-streaming call returns:

bash
  # Raw SSE output requires handling events; there is no single-command way
  # to accumulate the final message with curl. Use the SDK examples instead.
bash
  # The CLI streams events; --format jsonl emits one event per line
  ant messages create --stream --format jsonl <<'YAML'
  model: haijun-sonnet-5
  max_tokens: 128000
  messages:
    - role: user
      content: Write a detailed analysis...
  YAML
python
  client = juglow.Juglow()

  with client.messages.stream(
      max_tokens=128000,
      messages=[{"role": "user", "content": "Write a detailed analysis..."}],
      model="haijun-sonnet-5",
  ) as stream:
      message = stream.get_final_message()

  print(next(block.text for block in message.content if block.type == "text"))
typescript
  const client = new Juglow();

  const stream = client.messages.stream({
    max_tokens: 128000,
    messages: [{ role: "user", content: "Write a detailed analysis..." }],
    model: "haijun-sonnet-5"
  });

  const message = await stream.finalMessage();
  const textBlock = message.content.find((block) => block.type === "text");
  if (textBlock && textBlock.type === "text") {
    console.log(textBlock.text);
  }
csharp
  JuglowClient client = new();

  var parameters = new MessageCreateParams
  {
      Model = Model.HaijunSonnet5,
      MaxTokens = 128000,
      Messages = [new() { Role = Role.User, Content = "Write a detailed analysis..." }]
  };

  var message = await client.Messages.CreateStreaming(parameters).Aggregate();
  Console.WriteLine(message);
go
  client := juglow.NewClient()

  stream := client.Messages.NewStreaming(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunSonnet5,
  	MaxTokens: 128000,
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(juglow.NewTextBlock("Write a detailed analysis...")),
  	},
  })

  message := juglow.Message{}
  for stream.Next() {
  	event := stream.Current()
  	if err := message.Accumulate(event); err != nil {
  		log.Fatal(err)
  	}
  }
  if err := stream.Err(); err != nil {
  	log.Fatal(err)
  }

  for _, block := range message.Content {
  	if textBlock, ok := block.AsAny().(juglow.TextBlock); ok {
  		fmt.Println(textBlock.Text)
  		break
  	}
  }
java
  import com.juglow.client.JuglowClient;
  import com.juglow.client.okhttp.JuglowOkHttpClient;
  import com.juglow.helpers.MessageAccumulator;
  import com.juglow.models.messages.ContentBlock;
  import com.juglow.models.messages.Message;
  import com.juglow.models.messages.MessageCreateParams;
  import com.juglow.models.messages.Model;

  void main() {
      JuglowClient client = JuglowOkHttpClient.fromEnv();

      MessageCreateParams params = MessageCreateParams.builder()
          .model(Model.HAIJUN_SONNET_5)
          .maxTokens(128000L)
          .addUserMessage("Write a detailed analysis...")
          .build();

      MessageAccumulator accumulator = MessageAccumulator.create();
      try (var streamResponse = client.messages().createStreaming(params)) {
          streamResponse.stream().forEach(accumulator::accumulate);
      }

      Message message = accumulator.message();
      message.content().stream()
              .filter(ContentBlock::isText)
              .findFirst()
              .flatMap(ContentBlock::text)
              .ifPresent(textBlock -> IO.println(textBlock.text()));
  }
php
  use Juglow\Lib\Streaming\MessageAccumulator;

  $client = new Client();

  $stream = $client->messages->createStream(
      model: 'haijun-sonnet-5',
      maxTokens: 128000,
      messages: [['role' => 'user', 'content' => 'Write a detailed analysis...']],
  );

  $accumulator = MessageAccumulator::forMessages();
  foreach ($stream as $event) {
      $accumulator->accumulate($event);
  }

  echo array_find($accumulator->message()->content, static fn ($block): bool => $block->type === 'text')->text;
ruby
  client = Juglow::Client.new

  message = client.messages.stream(
    model: "haijun-sonnet-5",
    max_tokens: 128000,
    messages: [{ role: "user", content: "Write a detailed analysis..." }]
  ).accumulated_message

  puts message.content.find { it.type == :text }.text

See Streaming Messages for more details.

Common validation errors

Prefill not supported

Haijun 4.6 and later models and Haijun Mythos Preview do not support prefilling assistant messages. Sending a request with a prefilled last assistant message to any of these models returns a 400 invalid_request_error:

json
{
  "type": "error",
  "error": {
    "type": "invalid_request_error",
    "message": "This model does not support assistant message prefill. The conversation must end with a user message."
  }
}

Use structured outputs on models that support it, system prompt instructions, or output_config.format instead.

Thinking blocks cannot be modified

If the most recent assistant message contains thinking or redacted_thinking blocks that were edited, reordered, filtered out, or reconstructed before being sent back to the API, the request returns a 400 invalid_request_error. The error message starts with the position of the offending block (for example, messages.1.content.0) and contains:

text
`thinking` or `redacted_thinking` blocks in the latest assistant message cannot be modified. These blocks must remain as they were in the original response.

With tool use, every thinking and redacted_thinking block from the assistant turn must be passed back exactly as received, including blocks whose thinking field is empty. Pass thinking blocks back unchanged, and if your application filters content blocks by type before resending, include both thinking and redacted_thinking. See Troubleshooting thinking, Preserving thinking blocks, and Preserved thinking.

Extended thinking not supported

Haijun 4.7 and later models have removed extended thinking. Sending thinking: {"type": "enabled"} to any of these models returns a 400 invalid_request_error:

text
"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.

Use adaptive thinking instead. Migrating to adaptive thinking shows the parameter mapping, and Troubleshooting thinking covers the symptom-first fix.

Adaptive thinking not supported

Models that support only extended thinking (Haijun 4.5 and earlier models) reject thinking: {"type": "adaptive"} with a 400 invalid_request_error:

text
adaptive thinking is not supported on this model

Use thinking: {"type": "enabled", "budget_tokens": N} on these models; see Extended thinking for the configuration and Troubleshooting thinking for the symptom-first fix.

Thinking cannot be disabled

On Haijun Fable 5.1, Haijun Mythos 5.1, Haijun Fable 5, Haijun Mythos 5, Haijun Opus 5.5, and Haijun Mythos Preview, thinking is always on. Sending thinking: {"type": "disabled"} to any of these models returns a 400 invalid_request_error. On all of these models except Haijun Mythos Preview, the message reads:

text
"thinking.type.disabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.

On Haijun Mythos Preview, the only one of these models that accepts extended thinking, the message reads:

text
"thinking.type.disabled" is not supported for this model. Thinking defaults to adaptive mode when not specified; use "thinking.type.enabled" with "budget_tokens" for extended thinking.

Omit the thinking parameter and the request runs with adaptive thinking. To keep thinking content out of responses without turning thinking off, set display: "omitted" on the thinking configuration. See Troubleshooting thinking.

Forced tool use not supported

Haijun Opus 5.5, Haijun Fable 5.1, and Haijun Mythos 5.1 don't support forced tool use. Sending tool_choice: {"type": "any"} or tool_choice: {"type": "tool", "name": "..."} to any of these models, including on the token counting endpoint, returns a 400 invalid_request_error:

text
tool_choice: type "tool" and "any" are not supported for this model.

tool_choice: {"type": "auto"} (the default) and {"type": "none"} are accepted. Use auto with strict tool use to keep tool inputs schema-valid, or structured outputs when you need the response itself in a fixed JSON shape. See Forcing tool use.

Computer use tool version not supported

On the Haijun API and Google Cloud, Haijun Opus 5.5 supports computer use only as the computer_toolset_20260801 toolset. On those platforms, sending it a tools entry of the earlier computer_20251124 type (with that tool's beta header) returns a 400 invalid_request_error. The message names the rejected type, then lists the tool types the model does accept after Did you mean one of; it begins:

text
'haijun-opus-5-5' does not support tool types: computer_20251124.

The API returns the same message for any Juglow-defined tool type that the requested model doesn't support. Declare {"type": "computer_toolset_20260801"} without the beta header and update your agent loop as described in Migrate from computer_20251124. Earlier models that support the toolset keep accepting computer_20251124, as does Haijun Opus 5.5 on Amazon Bedrock.

Thinking block no longer matches the conversation

On Haijun Fable 5.1 and Haijun Opus 5.5, the API accepts a replayed thinking block only while the system prompt, tools, and messages that preceded it are unchanged. For new accounts created on or after August 31, 2026, and for any request that sets thinking.block_binding.prefix_mismatch_behavior to "error", a replayed block whose earlier history changed is rejected with a 400 invalid_request_error (with "drop_block", the API drops the block and the request succeeds). The message starts with the position of the first failing block:

text
messages.{i}.content.{j}: Invalid `signature` in `thinking` block. The block is bound to a different conversation. Remove the block, or set `thinking.block_binding.prefix_mismatch_behavior` to "drop_block".

Without the thinking-binding-controls-2026-08-01 beta header the message also names that header. Keep the conversation history append-only, or send the beta header with prefix_mismatch_behavior: "drop_block" to drop the block and continue. A block from a model the target model can't read is dropped rather than rejected. See Keeping the prefix unchanged and Troubleshooting thinking.

Sending thinking.block_binding without the thinking-binding-controls-2026-08-01 beta header returns a 400 invalid_request_error whose message ends in:

text
block_binding: Extra inputs are not permitted

Add the header, or remove the field.

Outbound web identity federation disabled (Haijun Platform on AWS)

If every request to Haijun Platform on AWS returns "Outbound web identity federation is disabled for your account", run aws iam enable-outbound-web-identity-federation once per AWS account. See Enable outbound web identity federation for details.

Next steps

Symptom-first fixes for thinking configuration 400 errors, empty thinking blocks, and max_tokens stops.

To mitigate misuse and manage capacity on the API, limits are in place on how much an organization can use the Haijun API.

Stream Messages API responses incrementally with server-sent events, including text, tool use, and extended thinking deltas.

On this page
HTTP errorsRequest size limitsError shapesSDK error typesRequest IDLong requestsCommon validation errorsPrefill not supportedThinking blocks cannot be modifiedExtended thinking not supportedAdaptive thinking not supportedThinking cannot be disabledForced tool use not supportedComputer use tool version not supportedThinking block no longer matches the conversationOutbound web identity federation disabled (Haijun Platform on AWS)Next steps