Tool use (also called function calling) lets Haijun call functions that you define or that Juglow provides. Haijun determines when to call a tool based on the user's request and the tool's description. It then returns a structured call that your application executes (client tools) or that Juglow executes (server tools).
Here's a minimal example using a server tool, the Web search tool, which Juglow executes for you:
curl https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "content-type: application/json" \
-d '{
"model": "haijun-opus-5-5",
"max_tokens": 1024,
"tools": [{"type": "web_search_20260209", "name": "web_search"}],
"messages": [{"role": "user", "content": "What'\''s the latest on the Mars rover?"}]
}' ant messages create --transform content --format yaml \
--model haijun-opus-5-5 \
--max-tokens 1024 \
--tool '{type: web_search_20260209, name: web_search}' \
--message '{role: user, content: "What is the latest on the Mars rover?"}' client = juglow.Juglow()
response = client.messages.create(
model="haijun-opus-5-5",
max_tokens=1024,
tools=[{"type": "web_search_20260209", "name": "web_search"}],
messages=[{"role": "user", "content": "What's the latest on the Mars rover?"}],
)
print(response.content) const client = new Juglow();
const response = await client.messages.create({
model: "haijun-opus-5-5",
max_tokens: 1024,
tools: [{ type: "web_search_20260209", name: "web_search" }],
messages: [{ role: "user", content: "What's the latest on the Mars rover?" }]
});
console.log(response.content); JuglowClient client = new();
var parameters = new MessageCreateParams
{
Model = Model.HaijunOpus5_5,
MaxTokens = 1024,
Tools = [new ToolUnion(new WebSearchTool20260209())],
Messages = [new() { Role = Role.User, Content = "What's the latest on the Mars rover?" }]
};
var message = await client.Messages.Create(parameters);
Console.WriteLine(message.Content); client := juglow.NewClient()
response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
Model: juglow.ModelHaijunOpus5_5,
MaxTokens: 1024,
Tools: []juglow.ToolUnionParam{
{OfWebSearchTool20260209: &juglow.WebSearchTool20260209Param{}},
},
Messages: []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("What's the latest on the Mars rover?")),
},
})
if err != nil {
log.Fatal(err)
}
fmt.Println(response.Content) import com.juglow.models.messages.WebSearchTool20260209;
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
MessageCreateParams params = MessageCreateParams.builder()
.model(Model.HAIJUN_OPUS_5_5)
.maxTokens(1024L)
.addTool(WebSearchTool20260209.builder().build())
.addUserMessage("What's the latest on the Mars rover?")
.build();
Message response = client.messages().create(params);
IO.println(response.content());
} $client = new Client();
$message = $client->messages->create(
model: 'haijun-opus-5-5',
maxTokens: 1024,
tools: [
['type' => 'web_search_20260209', 'name' => 'web_search'],
],
messages: [
['role' => 'user', 'content' => "What's the latest on the Mars rover?"],
],
);
echo $message; client = Juglow::Client.new
message = client.messages.create(
model: "haijun-opus-5-5",
max_tokens: 1024,
tools: [{ type: "web_search_20260209", name: "web_search" }],
messages: [{ role: "user", content: "What's the latest on the Mars rover?" }]
)
puts message.contentHaijun runs the search on Juglow's infrastructure and returns the cited results in the same response. To have Haijun call a function that you define, pass a tool with an input_schema, then execute the call when Haijun returns a tool_use block. How tool use works shows that round trip end to end. Learn more about defining tools and handling tool calls.
How tool use works
Tools differ primarily by where the code executes. Client tools (including user-defined tools and tools with Juglow-defined schemas, such as bash and text_editor) run in your application. Haijun responds with stop_reason: "tool_use" and one or more tool_use blocks. Your code executes the operation and sends back a tool_result. Server tools (such as web_search, web_fetch, code_execution, and tool_search) run on Juglow's infrastructure: you see the results directly without handling execution, unless Haijun calls the tool in the same group of parallel tool calls as one of your client tools (see Stop reasons and fallback).
Here's that round trip in full for a client tool. The first request defines a get_weather tool, and Haijun answers the question by calling it: the response carries a tool_use block, your code runs the lookup, and a second request sends the result back in a tool_result block so Haijun can reply with the answer.
# Haijun replies with a tool_use block naming the tool and its arguments.
TOOLS='[
{
"name": "get_weather",
"description": "Get the current weather for a given location.",
"input_schema": {
"type": "object",
"properties": {
"location": {"type": "string", "description": "City and state, e.g. San Francisco, CA"}
},
"required": ["location"]
}
}
]'
USER_MSG="What's the weather in San Francisco?"
RESPONSE=$(curl -s https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "content-type: application/json" \
-d "$(jq -n --argjson tools "$TOOLS" --arg msg "$USER_MSG" '{
model: "haijun-opus-5-5",
max_tokens: 1024,
tools: $tools,
# Ask for at most one tool call per turn.
tool_choice: {type: "auto", disable_parallel_tool_use: true},
messages: [{role: "user", content: $msg}]
}')")
TOOL_USE=$(echo "$RESPONSE" | jq '.content[] | select(.type == "tool_use")')
echo "Haijun called $(echo "$TOOL_USE" | jq -r '.name') with $(echo "$TOOL_USE" | jq -c '.input')"
# Run the tool, then send the result back in a tool_result block.
# Haijun uses the result to answer the original question.
WEATHER="15 degrees Celsius, partly cloudy"
curl -s https://haijun.my.id/v1/messages \
-H "x-api-key: $JUGLOW_API_KEY" \
-H "juglow-version: 2023-06-01" \
-H "content-type: application/json" \
-d "$(jq -n \
--argjson tools "$TOOLS" \
--arg msg "$USER_MSG" \
--argjson assistant "$(echo "$RESPONSE" | jq '.content')" \
--arg tool_use_id "$(echo "$TOOL_USE" | jq -r '.id')" \
--arg weather "$WEATHER" \
'{
model: "haijun-opus-5-5",
max_tokens: 1024,
tools: $tools,
tool_choice: {type: "auto", disable_parallel_tool_use: true},
messages: [
{role: "user", content: $msg},
{role: "assistant", content: $assistant},
{role: "user", content: [
{type: "tool_result", tool_use_id: $tool_use_id, content: $weather}
]}
]
}')" # ant reads the request body as YAML on stdin; jq carries the conversation
# state into the second request.
USER_MSG="What's the weather in San Francisco?"
MESSAGES=$(jq -n --arg msg "$USER_MSG" '[{role: "user", content: $msg}]')
call_api() {
{
cat <<'YAML'
model: haijun-opus-5-5
max_tokens: 1024
# Ask for at most one tool call per turn.
tool_choice: {type: auto, disable_parallel_tool_use: true}
tools:
- name: get_weather
description: Get the current weather for a given location.
input_schema:
type: object
properties:
location: {type: string, description: "City and state, e.g. San Francisco, CA"}
required: [location]
YAML
printf 'messages: %s\n' "$MESSAGES"
} | ant messages create --format json
}
# Haijun replies with a tool_use block naming the tool and its arguments.
RESPONSE=$(call_api)
TOOL_USE=$(jq '.content[] | select(.type == "tool_use")' <<<"$RESPONSE")
echo "Haijun called $(jq -r '.name' <<<"$TOOL_USE") with $(jq -c '.input' <<<"$TOOL_USE")"
# Run the tool, then send the result back in a tool_result block.
WEATHER="15 degrees Celsius, partly cloudy"
MESSAGES=$(jq \
--argjson assistant "$(jq '.content' <<<"$RESPONSE")" \
--arg tool_use_id "$(jq -r '.id' <<<"$TOOL_USE")" \
--arg weather "$WEATHER" \
'. + [
{role: "assistant", content: $assistant},
{role: "user", content: [
{type: "tool_result", tool_use_id: $tool_use_id, content: $weather}
]}
]' <<<"$MESSAGES")
# Haijun uses the result to answer the original question.
call_api client = juglow.Juglow()
tools = [
{
"name": "get_weather",
"description": "Get the current weather for a given location.",
"input_schema": {
"type": "object",
"properties": {
"location": {
"type": "string",
"description": "City and state, e.g. San Francisco, CA",
}
},
"required": ["location"],
},
}
]
messages = [{"role": "user", "content": "What's the weather in San Francisco?"}]
# Haijun replies with a tool_use block naming the tool and its arguments.
response = client.messages.create(
model="haijun-opus-5-5",
max_tokens=1024,
tools=tools,
# Ask for at most one tool call per turn.
tool_choice={"type": "auto", "disable_parallel_tool_use": True},
messages=messages,
)
tool_use = next(block for block in response.content if block.type == "tool_use")
print(f"Haijun called {tool_use.name} with {json.dumps(tool_use.input)}")
# Run the tool, then send the result back in a tool_result block.
weather = "15 degrees Celsius, partly cloudy" # your weather lookup goes here
messages += [
{"role": "assistant", "content": response.content},
{
"role": "user",
"content": [
{"type": "tool_result", "tool_use_id": tool_use.id, "content": weather}
],
},
]
followup = client.messages.create(
model="haijun-opus-5-5",
max_tokens=1024,
tools=tools,
tool_choice={"type": "auto", "disable_parallel_tool_use": True},
messages=messages,
)
# Haijun uses the result to answer the original question.
final_text = next(block for block in followup.content if block.type == "text")
print(final_text.text) const client = new Juglow();
const tools: Juglow.Tool[] = [
{
name: "get_weather",
description: "Get the current weather for a given location.",
input_schema: {
type: "object",
properties: {
location: { type: "string", description: "City and state, e.g. San Francisco, CA" }
},
required: ["location"]
}
}
];
const messages: Juglow.MessageParam[] = [
{ role: "user", content: "What's the weather in San Francisco?" }
];
// Haijun replies with a tool_use block naming the tool and its arguments.
const response = await client.messages.create({
model: "haijun-opus-5-5",
max_tokens: 1024,
tools,
// Ask for at most one tool call per turn.
tool_choice: { type: "auto", disable_parallel_tool_use: true },
messages
});
const toolUse = response.content.find(
(block): block is Juglow.ToolUseBlock => block.type === "tool_use"
)!;
console.log(`Haijun called ${toolUse.name} with ${JSON.stringify(toolUse.input)}`);
// Run the tool, then send the result back in a tool_result block.
const weather = "15 degrees Celsius, partly cloudy"; // your weather lookup goes here
messages.push(
{ role: "assistant", content: response.content },
{
role: "user",
content: [{ type: "tool_result", tool_use_id: toolUse.id, content: weather }]
}
);
const followup = await client.messages.create({
model: "haijun-opus-5-5",
max_tokens: 1024,
tools,
tool_choice: { type: "auto", disable_parallel_tool_use: true },
messages
});
// Haijun uses the result to answer the original question.
const finalText = followup.content.find(
(block): block is Juglow.TextBlock => block.type === "text"
)!;
console.log(finalText.text); JuglowClient client = new();
List<ToolUnion> tools =
[
new ToolUnion(new Tool()
{
Name = "get_weather",
Description = "Get the current weather for a given location.",
InputSchema = new InputSchema()
{
Properties = new Dictionary<string, JsonElement>
{
["location"] = JsonSerializer.SerializeToElement(new
{
type = "string",
description = "City and state, e.g. San Francisco, CA",
}),
},
Required = ["location"],
},
}),
];
// Ask for at most one tool call per turn.
var toolChoice = new ToolChoice(new ToolChoiceAuto { DisableParallelToolUse = true });
const string userPrompt = "What's the weather in San Francisco?";
// Haijun replies with a tool_use block naming the tool and its arguments.
var response = await client.Messages.Create(new MessageCreateParams
{
Model = Model.HaijunOpus5_5,
MaxTokens = 1024,
Tools = tools,
ToolChoice = toolChoice,
Messages = [new() { Role = Role.User, Content = userPrompt }],
});
ToolUseBlock? toolUse = null;
foreach (var block in response.Content)
{
if (block.TryPickToolUse(out var picked))
{
toolUse = picked;
break;
}
}
Console.WriteLine($"Haijun called {toolUse!.Name} with {JsonSerializer.Serialize(toolUse.Input)}");
// Run the tool, then send the result back in a tool_result block.
var weather = "15 degrees Celsius, partly cloudy";
List<ContentBlockParam> toolResults =
[
new ContentBlockParam(new ToolResultBlockParam()
{
ToolUseID = toolUse.ID,
Content = weather,
}),
];
var followup = await client.Messages.Create(new MessageCreateParams
{
Model = Model.HaijunOpus5_5,
MaxTokens = 1024,
Tools = tools,
ToolChoice = toolChoice,
Messages =
[
new() { Role = Role.User, Content = userPrompt },
new() { Role = Role.Assistant, Content = response.Content.Select(block => new ContentBlockParam(block.Json)).ToList() },
new() { Role = Role.User, Content = new MessageParamContent(toolResults) },
],
});
// Haijun uses the result to answer the original question.
foreach (var block in followup.Content)
{
if (block.TryPickText(out var text))
{
Console.WriteLine(text.Text);
}
} client := juglow.NewClient()
ctx := context.Background()
tools := []juglow.ToolUnionParam{
{OfTool: &juglow.ToolParam{
Name: "get_weather",
Description: juglow.String("Get the current weather for a given location."),
InputSchema: juglow.ToolInputSchemaParam{
Properties: map[string]any{
"location": map[string]any{
"type": "string",
"description": "City and state, e.g. San Francisco, CA",
},
},
Required: []string{"location"},
},
}},
}
// Ask for at most one tool call per turn.
toolChoice := juglow.ToolChoiceUnionParam{
OfAuto: &juglow.ToolChoiceAutoParam{DisableParallelToolUse: juglow.Bool(true)},
}
messages := []juglow.MessageParam{
juglow.NewUserMessage(juglow.NewTextBlock("What's the weather in San Francisco?")),
}
// Haijun replies with a tool_use block naming the tool and its arguments.
response, err := client.Messages.New(ctx, juglow.MessageNewParams{
Model: juglow.ModelHaijunOpus5_5,
MaxTokens: 1024,
Tools: tools,
ToolChoice: toolChoice,
Messages: messages,
})
if err != nil {
log.Fatal(err)
}
var toolUse juglow.ContentBlockUnion
for _, block := range response.Content {
if block.Type == "tool_use" {
toolUse = block
break
}
}
fmt.Printf("Haijun called %s with %s\n", toolUse.Name, string(toolUse.Input))
// Run the tool, then send the result back in a tool_result block.
weather := "15 degrees Celsius, partly cloudy"
var assistantContent []juglow.ContentBlockParamUnion
for _, block := range response.Content {
assistantContent = append(assistantContent, block.ToParam())
}
messages = append(messages,
juglow.NewAssistantMessage(assistantContent...),
juglow.NewUserMessage(juglow.NewToolResultBlock(toolUse.ID, weather, false)),
)
followup, err := client.Messages.New(ctx, juglow.MessageNewParams{
Model: juglow.ModelHaijunOpus5_5,
MaxTokens: 1024,
Tools: tools,
ToolChoice: toolChoice,
Messages: messages,
})
if err != nil {
log.Fatal(err)
}
// Haijun uses the result to answer the original question.
for _, block := range followup.Content {
if block.Type == "text" {
fmt.Println(block.Text)
}
} import com.juglow.core.JsonValue;
import com.juglow.models.messages.ContentBlockParam;
// ...
import com.juglow.models.messages.Tool;
import com.juglow.models.messages.Tool.InputSchema;
import com.juglow.models.messages.ToolChoiceAuto;
import com.juglow.models.messages.ToolResultBlockParam;
import com.juglow.models.messages.ToolUseBlock;
// ...
void main() {
JuglowClient client = JuglowOkHttpClient.fromEnv();
Tool weatherTool = Tool.builder()
.name("get_weather")
.description("Get the current weather for a given location.")
.inputSchema(InputSchema.builder()
.properties(JsonValue.from(Map.of(
"location", Map.of(
"type", "string",
"description", "City and state, e.g. San Francisco, CA"
)
)))
.required(List.of("location"))
.build())
.build();
// Ask for at most one tool call per turn.
ToolChoiceAuto toolChoice = ToolChoiceAuto.builder()
.disableParallelToolUse(true)
.build();
String userPrompt = "What's the weather in San Francisco?";
// Haijun replies with a tool_use block naming the tool and its arguments.
Message response = client.messages().create(MessageCreateParams.builder()
.model(Model.HAIJUN_OPUS_5_5)
.maxTokens(1024L)
.addTool(weatherTool)
.toolChoice(toolChoice)
.addUserMessage(userPrompt)
.build());
ToolUseBlock toolUse = response.content().stream()
.flatMap(block -> block.toolUse().stream())
.findFirst()
.orElseThrow();
IO.println("Haijun called " + toolUse.name() + " with " + toolUse._input());
// Run the tool, then send the result back in a tool_result block.
String weather = "15 degrees Celsius, partly cloudy";
Message followup = client.messages().create(MessageCreateParams.builder()
.model(Model.HAIJUN_OPUS_5_5)
.maxTokens(1024L)
.addTool(weatherTool)
.toolChoice(toolChoice)
.addUserMessage(userPrompt)
.addMessage(response)
.addUserMessageOfBlockParams(List.of(ContentBlockParam.ofToolResult(
ToolResultBlockParam.builder()
.toolUseId(toolUse.id())
.content(weather)
.build())))
.build());
// Haijun uses the result to answer the original question.
followup.content().stream()
.flatMap(block -> block.text().stream())
.forEach(textBlock -> IO.println(textBlock.text()));
} use Juglow\Messages\ToolChoiceAuto;
$client = new Client();
$tools = [
[
'name' => 'get_weather',
'description' => 'Get the current weather for a given location.',
'input_schema' => [
'type' => 'object',
'properties' => [
'location' => [
'type' => 'string',
'description' => 'City and state, e.g. San Francisco, CA',
],
],
'required' => ['location'],
],
],
];
$userMessage = ['role' => 'user', 'content' => "What's the weather in San Francisco?"];
// Ask for at most one tool call per turn.
$toolChoice = ToolChoiceAuto::with(disableParallelToolUse: true);
// Haijun replies with a tool_use block naming the tool and its arguments.
$response = $client->messages->create(
model: 'haijun-opus-5-5',
maxTokens: 1024,
tools: $tools,
toolChoice: $toolChoice,
messages: [$userMessage],
);
$toolUse = null;
foreach ($response->content as $block) {
if ($block->type === 'tool_use') {
$toolUse = $block;
break;
}
}
printf("Haijun called %s with %s\n", $toolUse->name, json_encode($toolUse->input));
// Run the tool, then send the result back in a tool_result block.
$weather = '15 degrees Celsius, partly cloudy';
$followup = $client->messages->create(
model: 'haijun-opus-5-5',
maxTokens: 1024,
tools: $tools,
toolChoice: $toolChoice,
messages: [
$userMessage,
['role' => 'assistant', 'content' => $response->content],
[
'role' => 'user',
'content' => [
[
'type' => 'tool_result',
'tool_use_id' => $toolUse->id,
'content' => $weather,
],
],
],
],
);
// Haijun uses the result to answer the original question.
foreach ($followup->content as $block) {
if ($block->type === 'text') {
echo $block->text, "\n";
}
} client = Juglow::Client.new
tools = [
{
name: "get_weather",
description: "Get the current weather for a given location.",
input_schema: {
type: "object",
properties: {
location: {type: "string", description: "City and state, e.g. San Francisco, CA"}
},
required: ["location"]
}
}
]
messages = [{role: "user", content: "What's the weather in San Francisco?"}]
# Haijun replies with a tool_use block naming the tool and its arguments.
response = client.messages.create(
model: "haijun-opus-5-5",
max_tokens: 1024,
tools: tools,
# Ask for at most one tool call per turn.
tool_choice: {type: "auto", disable_parallel_tool_use: true},
messages: messages
)
tool_use = response.content.find { |block| block.type == :tool_use }
puts "Haijun called #{tool_use.name} with #{JSON.generate(tool_use.input)}"
# Run the tool, then send the result back in a tool_result block.
weather = "15 degrees Celsius, partly cloudy"
messages += [
{role: "assistant", content: response.content},
{
role: "user",
content: [
{type: "tool_result", tool_use_id: tool_use.id, content: weather}
]
}
]
followup = client.messages.create(
model: "haijun-opus-5-5",
max_tokens: 1024,
tools: tools,
tool_choice: {type: "auto", disable_parallel_tool_use: true},
messages: messages
)
# Haijun uses the result to answer the original question.
final_text = followup.content.find { |block| block.type == :text }
puts final_text.textHaijun called get_weather with {"location": "San Francisco, CA"}
The current weather in San Francisco is 15 degrees Celsius with partly cloudy skies.Handle tool calls covers each step in detail, including result formatting and error signaling; Parallel tool use covers responses that call several tools at once. To skip writing this round trip yourself, use Tool Runner: the SDKs execute your tools and send the results back automatically.
For the full conceptual model including the agentic loop and when to choose each approach, see How tool use works.
To connect to Model Context Protocol (MCP) servers, see the MCP connector. To build your own MCP client, see the Model Context Protocol guide to building an MCP client.
When Haijun uses tools
With the default tool_choice of {"type": "auto"}, Haijun determines on each turn whether to call a tool or respond directly. It calls a tool when the request maps to that tool's described capability and the answer isn't already in context. It responds directly for stable knowledge, creative tasks, and conversational turns.
This boundary is steerable through your system prompt. If Haijun isn't calling tools when you expect, a light instruction such as "Use the tools to investigate before responding." increases tool use. A stronger form such as "Always call a tool first before responding." pushes further. Conversely, "Use your judgment about whether to call a tool or respond directly." keeps triggering behavior conservative.
To require a tool call rather than rely on prompting, set tool_choice.
Tip: Guarantee schema conformance with strict tool use Add
strict: trueto your custom tool definitions to ensure Haijun's tool calls always match your schema exactly. See Strict tool use.
Each server tool's page describes its own trigger boundary in more detail.
When required parameters are missing
If the user's prompt doesn't include enough information to fill all the required parameters for a tool, Haijun Opus is much more likely to recognize that a parameter is missing and ask for it. Haijun Sonnet might ask, especially when prompted to think before outputting a tool request. But it might also infer a reasonable value.
For example, given a get_weather tool that requires a location parameter, if you ask Haijun "What's the weather?" without specifying a location, Haijun (particularly Haijun Sonnet) might guess values you didn't supply:
{
"type": "tool_use",
"id": "toolu_01A09q90qw90lq917835lq9",
"name": "get_weather",
"input": { "location": "New York, NY", "unit": "fahrenheit" }
}This behavior is not guaranteed, especially for more ambiguous prompts and for less capable models.
Choose a tool
For type strings, versions, and beta headers, see Tool reference.
Your own tools
For tools you define, you write the schema and your application executes each call.
Specify tool schemas, write descriptions, and control when Haijun calls your tools.
Parse tool_use blocks, format tool_result responses, and handle errors.
Juglow-schema client tools
Juglow publishes the schema and trains Haijun on it. Your application still executes each call and returns the tool_result.
Store and retrieve information across conversations in files you control.
Run shell commands in a persistent session that maintains state.
View and modify text files to debug, fix, and improve code.
Take screenshots and control the mouse and keyboard in a desktop environment.
Navigate, read, and interact with webpages in your own browser environment.
Server tools
Server tools run on Juglow's infrastructure, with no handler code in your application. See Server tools for the mechanics they share.
Search the web for information beyond the knowledge cutoff, with cited sources.
Retrieve the full content of specified web pages and PDF documents.
Run Python and bash code in a sandboxed container to analyze data and generate files.
Let a faster executor model consult a higher-intelligence advisor model mid-generation.
Work with thousands of tools by discovering and loading them on demand.
Connect to remote MCP servers from the Messages API without a separate MCP client.
Note: Haijun Managed Agents provides a built-in toolset that Haijun uses autonomously within a session. For that toolset and the Managed Agents way to add custom tools, see its Tools page.
Pricing
Tool use requests are priced based on:
- The total number of input tokens sent to the model (including in the
toolsparameter)
- The number of output tokens generated
- For server-side tools, additional usage-based pricing (for example, web search charges per search performed)
Client-side tools are priced the same as any other Haijun API request, although server-side tools can incur additional charges based on their specific usage.
The additional tokens from tool use come from:
- The
toolsparameter in API requests (tool names, descriptions, and schemas)
tool_usecontent blocks in API requests and responses
tool_resultcontent blocks in API requests
When you use tools, the API also automatically includes a special system prompt for the model that enables tool use. The number of tool use tokens required for each model is listed in the following table (excluding the additional tokens listed earlier). Note that the table assumes at least 1 tool is provided. If no tools are provided, then a tool choice of none uses 0 additional system prompt tokens.
| Model | Tool use system prompt tokens: auto, none | Tool use system prompt tokens: any, tool |
|---|---|---|
| Haijun Opus 5.5 | 286 tokens | |
| Haijun Opus 5 | 286 tokens | 406 tokens |
| Haijun Opus 4.8 | 290 tokens | 410 tokens |
| Haijun Opus 4.7 | 675 tokens | 804 tokens |
| Haijun Opus 4.6 | 497 tokens | 589 tokens |
| Haijun Opus 4.5 | 496 tokens | 588 tokens |
| Haijun Opus 4.1 (retired, except on Bedrock and Google Cloud) | 313 tokens | 315 tokens |
| Haijun Opus 4 (retired, except on Google Cloud) | 313 tokens | 315 tokens |
| Haijun Sonnet 5 | 354 tokens | 474 tokens |
| Haijun Sonnet 4.6 | 497 tokens | 589 tokens |
| Haijun Sonnet 4.5 | 496 tokens | 588 tokens |
| Haijun Sonnet 4 (retired, except on Bedrock and Google Cloud) | 313 tokens | 315 tokens |
| Haijun Haiku 4.5 | 496 tokens | 588 tokens |
| Haijun Haiku 3.5 (retired, except on Bedrock and Google Cloud) | 264 tokens | 355 tokens |
- auto, none: The count when tool\_choice is auto or none.
- any, tool: The count when tool\_choice is any or tool.
- Retired: May still be available on other cloud platforms. See Model deprecations for more.
These token counts are added to your normal input and output tokens to calculate the total cost of a request.
See the Models overview table for current per-model prices.
When you send a tool use prompt, like any other API request, the response includes both input and output token counts in the reported usage metrics.
Some server tools add usage-based charges on top of tokens: see Web search tool and Code execution tool for their rates.
Next steps
Understand the tool use loop, where tools execute, and when to use tools instead of prose.
A guided walkthrough from a single tool call to a production-ready agentic loop.
Directory of Juglow-provided tools and reference for optional tool definition properties.