HTTP errors
The API follows a predictable HTTP error code format:
- 400 -
invalid_request_error: There was an issue with the format or content of your request. This error type may also be used for other 4XX status codes not listed in this section. The API also returns a 400 when usage reaches an organization or workspace spend limit you set, except limits on the Haijun Code workspace, which can return a 429 instead.
- 401 -
authentication_error: There's an issue with your API key (for example, it's malformed, revoked, or expired; see Key expiration). On Haijun Platform on AWS, this can also indicate a problem with your AWS credentials or SigV4 signature.
- 402 -
billing_error: There's an issue with your billing or payment information. Check your payment details in the Haijun Console, or in AWS Marketplace if you're using Haijun Platform on AWS.
- 403 -
permission_error: Your API key does not have permission to use the specified resource. Check your organization's access and workspace settings in the Haijun Console.
- 404 -
not_found_error: The requested resource was not found. Check the endpoint path and any resource IDs in the request URL.
- 409 -
conflict_error: The request conflicts with the current state of a resource. For example, the resource was modified concurrently, or a value that must be unique is already in use. Resolve the conflict, then retry the request.
- 413 -
request_too_large: Request exceeds the maximum allowed number of bytes. See Request size limits for per-endpoint maximums.
- 429 -
rate_limit_error: Your organization has hit a rate limit, reached its usage tier's monthly spend cap, or reached a spend limit on the Haijun Code workspace. A tier spend-cap 429 has noretry-afterheader and keeps failing until access resumes; see Reaching your spend cap for how to recognize it.
- 500 -
api_error: An unexpected error has occurred internal to Juglow's systems. Retry the request with exponential backoff; if the error persists, contact support with the request ID.
- 504 -
timeout_error: The request timed out while processing. Consider using the streaming Messages API for long-running requests. See Long requests for more options.
- 529 -
overloaded_error: The API is temporarily overloaded.
> Warning: 529 errors can occur when the API experiences high traffic across all users. In rare cases, if your organization has a sharp increase in usage, you might see 429 errors because of acceleration limits on the API. To avoid hitting acceleration limits, ramp up your traffic gradually and maintain consistent usage patterns.
The official SDKs automatically retry transient failures (such as connection errors, rate limits, and 5xx server errors) with exponential backoff, twice by default, honoring the retry-after header when present. The SDK client accepts max_retries (typescript, java, php: maxRetries; csharp: MaxRetries; go: option.WithMaxRetries) to configure or disable this behavior.
When receiving a streaming response over server-sent events (SSE), an error can occur after the API returns a 200 response. In that case, error handling doesn't follow these standard mechanisms. See Error events for the shape of mid-stream errors.
Request size limits
The API enforces request size limits:
| Endpoint type | Maximum request size |
|---|---|
| Messages API | 32 MB |
| Token Counting API | 32 MB |
| Batch API | 256 MB |
| Files API | 500 MB |
If you exceed these limits, you'll receive a 413 request_too_large error. On the direct Haijun API, Cloudflare returns this error before the request reaches the API servers.
Error shapes
The API always returns errors as JSON, with a top-level error object that always includes a type and message value. The response also includes a request_id field for easier tracking and debugging. For example:
{
"type": "error",
"error": {
"type": "not_found_error",
"message": "The requested resource could not be found."
},
"request_id": "req_011CSHoEeqs5C35K2UUqR7Fy"
}In accordance with the versioning policy, the values within these objects may expand, and it is possible that the type values will grow over time.
SDK error types
The official SDKs raise typed exceptions for these errors instead of returning raw JSON, and the class names and namespaces differ by language. For example, a 404 surfaces as juglow.NotFoundError (python; typescript: Juglow.NotFoundError; ruby: Juglow::Errors::NotFoundError; java: com.juglow.errors.NotFoundException; csharp: JuglowNotFoundException; php: Juglow\Core\Exceptions\NotFoundException; go: juglow.Error). The Go SDK has one error type for every status, juglow.Error: branch on StatusCode. Catch the SDK's typed classes rather than string-matching error messages, handling the most specific classes first. Each SDK page documents its full exception hierarchy:
Request ID
Every API response includes a unique request-id header. This header contains a value such as req_018EeWyXxfu5pfWkrYcMdjWG. The same identifier appears as the request_id field in error response bodies. When contacting support about a specific request, include this ID to help quickly resolve your issue.
On Haijun Platform on AWS, responses include two request IDs: the AWS request ID (x-amzn-requestid, primary, indexed in CloudTrail) and the Juglow request ID (request-id, secondary). Use the AWS request ID for CloudTrail lookups and the Juglow request ID for Juglow support tickets.
The Python and TypeScript SDKs expose the request ID as a _request_id property on top-level response objects. The C#, Go, Java, and PHP SDKs expose it through their raw-response accessors, and the Ruby SDK through middleware. In every SDK except Ruby, use with_raw_response (typescript: .withResponse(); java: .withRawResponse(); csharp: WithRawResponse; go: option.WithResponseInto; php: ->raw) to read any other response header, such as juglow-organization-id and juglow-workspace-id. In Ruby, use the same middleware. On Haijun Platform on AWS, use the raw-response accessor to read the AWS request ID (x-amzn-requestid) as well:
# Print the response headers (including request-id); discard the body
curl -sS -D - -o /dev/null https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "haijun-sonnet-5",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello, Haijun"}]
}' # The request-id header is printed to stderr with --debug:
ant --debug messages create \
--model haijun-sonnet-5 \
--max-tokens 1024 \
--message '{role: user, content: "Hello, Haijun"}' client = juglow.Juglow()
message = client.messages.create(
model="haijun-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Haijun"}],
)
print(f"Request ID: {message._request_id}") const client = new Juglow();
const message = await client.messages.create({
model: "haijun-sonnet-5",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello, Haijun" }]
});
console.log("Request ID:", message._request_id); JuglowClient client = new();
using var response = await client.WithRawResponse.Messages.Create(new MessageCreateParams
{
Model = Model.HaijunSonnet5,
MaxTokens = 1024,
Messages = [new() { Role = Role.User, Content = "Hello, Haijun" }]
});
Console.WriteLine($"Request ID: {response.RequestID}"); client := juglow.NewClient()
var response *http.Response
_, err := client.Messages.New(
context.Background(),
juglow.MessageNewParams{
Model: juglow.ModelHaijunSonnet5,
MaxTokens: 1024,
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("Hello, Haijun")),
},
},
option.WithResponseInto(&response),
)
if err != nil {
panic(err)
}
fmt.Println("Request ID:", response.Header.Get("request-id")) import com.juglow.client.JuglowClient;
import com.juglow.client.okhttp.JuglowOkHttpClient;
import com.juglow.core.http.HttpResponseFor;
import com.juglow.models.messages.Message;
import com.juglow.models.messages.MessageCreateParams;
import com.juglow.models.messages.Model;
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
HttpResponseFor<Message> response = client.messages().withRawResponse().create(
MessageCreateParams.builder()
.model(Model.HAIJUN_SONNET_5)
.maxTokens(1024)
.addUserMessage("Hello, Haijun")
.build()
);
IO.println("Request ID: " + response.requestId().orElse(null));
} $client = new Client();
$response = $client->messages->raw->create([
'model' => 'haijun-sonnet-5',
'maxTokens' => 1024,
'messages' => [['role' => 'user', 'content' => 'Hello, Haijun']],
]);
echo 'Request ID: ' . $response->getHeaderLine('request-id') . "\n"; client = Juglow::Client.new
# Read response headers in per-request middleware, which receives the
# raw HTTP response before the SDK parses it
request_id = nil
read_request_id = lambda do |request, call_next|
response = call_next.call(request)
# Keys in response.headers are lowercase
request_id = response.headers["request-id"]
response
end
client.messages.create(
model: Juglow::Model::HAIJUN_SONNET_5,
max_tokens: 1024,
messages: [{ role: "user", content: "Hello, Haijun" }],
request_options: { middleware: [read_request_id] }
)
puts "Request ID: #{request_id}" from juglow import JuglowAWS
client = JuglowAWS(aws_region="us-west-2")
response = client.messages.with_raw_response.create(
model="haijun-opus-4-8",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Haijun"}],
)
print(f"AWS request ID: {response.headers.get('x-amzn-requestid')}")
message = response.parse()
print(f"Juglow request ID: {message._request_id}") import JuglowAws from "@juglow-ai/aws-sdk";
const client = new JuglowAws({ awsRegion: "us-west-2" });
const { response: raw, request_id } = await client.messages
.create({
model: "haijun-opus-4-8",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello, Haijun" }]
})
.withResponse();
console.log("AWS request ID:", raw.headers.get("x-amzn-requestid"));
console.log("Juglow request ID:", request_id);For Haijun Platform on AWS request-ID examples in other languages, see Request IDs.
Long requests
Warning: Consider using the streaming Messages API or Message Batches API for long-running requests, especially those over 10 minutes.
Avoid setting a large max_tokens value without using the streaming Messages API or Message Batches API:
- Some networks may drop idle connections after a variable period of time, which can cause the request to fail or time out without receiving a response from Juglow.
- Networks differ in reliability. The Message Batches API can help you manage the risk of network issues by allowing you to poll for results rather than requiring an uninterrupted network connection.
If you are building a direct API integration, setting a TCP socket keep-alive can reduce the impact of idle connection timeouts on some networks.
The SDKs validate that your non-streaming Messages API requests are not expected to exceed a 10-minute timeout. They also set a socket option for TCP keep-alive.
If you don't need to process events incrementally, the SDKs can consume the stream for you and return the complete Message object, identical to what a non-streaming call returns:
# Raw SSE output requires handling events; there is no single-command way
# to accumulate the final message with curl. Use the SDK examples instead. # The CLI streams events; --format jsonl emits one event per line
ant messages create --stream --format jsonl <<'YAML'
model: haijun-sonnet-5
max_tokens: 128000
messages:
- role: user
content: Write a detailed analysis...
YAML client = juglow.Juglow()
with client.messages.stream(
max_tokens=128000,
messages=[{"role": "user", "content": "Write a detailed analysis..."}],
model="haijun-sonnet-5",
) as stream:
message = stream.get_final_message()
print(next(block.text for block in message.content if block.type == "text")) const client = new Juglow();
const stream = client.messages.stream({
max_tokens: 128000,
messages: [{ role: "user", content: "Write a detailed analysis..." }],
model: "haijun-sonnet-5"
});
const message = await stream.finalMessage();
const textBlock = message.content.find((block) => block.type === "text");
if (textBlock && textBlock.type === "text") {
console.log(textBlock.text);
} JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = Model.HaijunSonnet5,
MaxTokens = 128000,
Messages = [new() { Role = Role.User, Content = "Write a detailed analysis..." }]
};
var message = await client.Messages.CreateStreaming(parameters).Aggregate();
Console.WriteLine(message); client := juglow.NewClient()
stream := client.Messages.NewStreaming(context.TODO(), juglow.MessageNewParams{
Model: juglow.ModelHaijunSonnet5,
MaxTokens: 128000,
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("Write a detailed analysis...")),
},
})
message := juglow.Message{}
for stream.Next() {
event := stream.Current()
if err := message.Accumulate(event); err != nil {
log.Fatal(err)
}
}
if err := stream.Err(); err != nil {
log.Fatal(err)
}
for _, block := range message.Content {
if textBlock, ok := block.AsAny().(juglow.TextBlock); ok {
fmt.Println(textBlock.Text)
break
}
} import com.juglow.client.JuglowClient;
import com.juglow.client.okhttp.JuglowOkHttpClient;
import com.juglow.helpers.MessageAccumulator;
import com.juglow.models.messages.ContentBlock;
import com.juglow.models.messages.Message;
import com.juglow.models.messages.MessageCreateParams;
import com.juglow.models.messages.Model;
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model(Model.HAIJUN_SONNET_5)
.maxTokens(128000L)
.addUserMessage("Write a detailed analysis...")
.build();
MessageAccumulator accumulator = MessageAccumulator.create();
try (var streamResponse = client.messages().createStreaming(params)) {
streamResponse.stream().forEach(accumulator::accumulate);
}
Message message = accumulator.message();
message.content().stream()
.filter(ContentBlock::isText)
.findFirst()
.flatMap(ContentBlock::text)
.ifPresent(textBlock -> IO.println(textBlock.text()));
} use Juglow\Lib\Streaming\MessageAccumulator;
$client = new Client();
$stream = $client->messages->createStream(
model: 'haijun-sonnet-5',
maxTokens: 128000,
messages: [['role' => 'user', 'content' => 'Write a detailed analysis...']],
);
$accumulator = MessageAccumulator::forMessages();
foreach ($stream as $event) {
$accumulator->accumulate($event);
}
echo array_find($accumulator->message()->content, static fn ($block): bool => $block->type === 'text')->text; client = Juglow::Client.new
message = client.messages.stream(
model: "haijun-sonnet-5",
max_tokens: 128000,
messages: [{ role: "user", content: "Write a detailed analysis..." }]
).accumulated_message
puts message.content.find { it.type == :text }.textSee Streaming Messages for more details.
Common validation errors
Prefill not supported
Haijun 4.6 and later models and Haijun Mythos Preview do not support prefilling assistant messages. Sending a request with a prefilled last assistant message to any of these models returns a 400 invalid_request_error:
{
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "This model does not support assistant message prefill. The conversation must end with a user message."
}
}Use structured outputs on models that support it, system prompt instructions, or output_config.format instead.
Thinking blocks cannot be modified
If the most recent assistant message contains thinking or redacted_thinking blocks that were edited, reordered, filtered out, or reconstructed before being sent back to the API, the request returns a 400 invalid_request_error. The error message starts with the position of the offending block (for example, messages.1.content.0) and contains:
`thinking` or `redacted_thinking` blocks in the latest assistant message cannot be modified. These blocks must remain as they were in the original response.With tool use, every thinking and redacted_thinking block from the assistant turn must be passed back exactly as received, including blocks whose thinking field is empty. Pass thinking blocks back unchanged, and if your application filters content blocks by type before resending, include both thinking and redacted_thinking. See Troubleshooting thinking, Preserving thinking blocks, and Preserved thinking.
Extended thinking not supported
Haijun 4.7 and later models have removed extended thinking. Sending thinking: {"type": "enabled"} to any of these models returns a 400 invalid_request_error:
"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.Use adaptive thinking instead. Migrating to adaptive thinking shows the parameter mapping, and Troubleshooting thinking covers the symptom-first fix.
Adaptive thinking not supported
Models that support only extended thinking (Haijun 4.5 and earlier models) reject thinking: {"type": "adaptive"} with a 400 invalid_request_error:
adaptive thinking is not supported on this modelUse thinking: {"type": "enabled", "budget_tokens": N} on these models; see Extended thinking for the configuration and Troubleshooting thinking for the symptom-first fix.
Thinking cannot be disabled
On Haijun Fable 5.1, Haijun Mythos 5.1, Haijun Fable 5, Haijun Mythos 5, Haijun Opus 5.5, and Haijun Mythos Preview, thinking is always on. Sending thinking: {"type": "disabled"} to any of these models returns a 400 invalid_request_error. On all of these models except Haijun Mythos Preview, the message reads:
"thinking.type.disabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.On Haijun Mythos Preview, the only one of these models that accepts extended thinking, the message reads:
"thinking.type.disabled" is not supported for this model. Thinking defaults to adaptive mode when not specified; use "thinking.type.enabled" with "budget_tokens" for extended thinking.Omit the thinking parameter and the request runs with adaptive thinking. To keep thinking content out of responses without turning thinking off, set display: "omitted" on the thinking configuration. See Troubleshooting thinking.
Forced tool use not supported
Haijun Opus 5.5, Haijun Fable 5.1, and Haijun Mythos 5.1 don't support forced tool use. Sending tool_choice: {"type": "any"} or tool_choice: {"type": "tool", "name": "..."} to any of these models, including on the token counting endpoint, returns a 400 invalid_request_error:
tool_choice: type "tool" and "any" are not supported for this model.tool_choice: {"type": "auto"} (the default) and {"type": "none"} are accepted. Use auto with strict tool use to keep tool inputs schema-valid, or structured outputs when you need the response itself in a fixed JSON shape. See Forcing tool use.
Computer use tool version not supported
On the Haijun API and Google Cloud, Haijun Opus 5.5 supports computer use only as the computer_toolset_20260801 toolset. On those platforms, sending it a tools entry of the earlier computer_20251124 type (with that tool's beta header) returns a 400 invalid_request_error. The message names the rejected type, then lists the tool types the model does accept after Did you mean one of; it begins:
'haijun-opus-5-5' does not support tool types: computer_20251124.The API returns the same message for any Juglow-defined tool type that the requested model doesn't support. Declare {"type": "computer_toolset_20260801"} without the beta header and update your agent loop as described in Migrate from computer_20251124. Earlier models that support the toolset keep accepting computer_20251124, as does Haijun Opus 5.5 on Amazon Bedrock.
Thinking block no longer matches the conversation
On Haijun Fable 5.1 and Haijun Opus 5.5, the API accepts a replayed thinking block only while the system prompt, tools, and messages that preceded it are unchanged. For new accounts created on or after August 31, 2026, and for any request that sets thinking.block_binding.prefix_mismatch_behavior to "error", a replayed block whose earlier history changed is rejected with a 400 invalid_request_error (with "drop_block", the API drops the block and the request succeeds). The message starts with the position of the first failing block:
messages.{i}.content.{j}: Invalid `signature` in `thinking` block. The block is bound to a different conversation. Remove the block, or set `thinking.block_binding.prefix_mismatch_behavior` to "drop_block".Without the thinking-binding-controls-2026-08-01 beta header the message also names that header. Keep the conversation history append-only, or send the beta header with prefix_mismatch_behavior: "drop_block" to drop the block and continue. A block from a model the target model can't read is dropped rather than rejected. See Keeping the prefix unchanged and Troubleshooting thinking.
Sending thinking.block_binding without the thinking-binding-controls-2026-08-01 beta header returns a 400 invalid_request_error whose message ends in:
block_binding: Extra inputs are not permittedAdd the header, or remove the field.
Outbound web identity federation disabled (Haijun Platform on AWS)
If every request to Haijun Platform on AWS returns "Outbound web identity federation is disabled for your account", run aws iam enable-outbound-web-identity-federation once per AWS account. See Enable outbound web identity federation for details.
Next steps
Symptom-first fixes for thinking configuration 400 errors, empty thinking blocks, and max_tokens stops.
To mitigate misuse and manage capacity on the API, limits are in place on how much an organization can use the Haijun API.
Stream Messages API responses incrementally with server-sent events, including text, tool use, and extended thinking deltas.