Haijun Platform Docs
ID

Task budgets let you tell Haijun how many tokens it has for a full agentic loop, including thinking, tool calls, tool results, and output. The model sees a running countdown and uses it to prioritize work and finish gracefully as the budget is consumed.

When to use task budgets

Task budgets work best for agentic workflows where Haijun makes multiple tool calls and decisions before finalizing its output to await the next human response. Use them when:

  • You want Haijun to self-regulate token spend on long-horizon tasks.
  • You have a predictable per-task cost or latency ceiling to enforce.
  • You want the model to finish gracefully (summarize findings, report progress) as it approaches the budget rather than cutting off mid-action.

Task budgets complement the effort parameter: effort controls how thoroughly Haijun reasons about each step, while task budgets cap the total work Haijun can do across an agentic loop.

Setting a task budget

Add task_budget to output_config and include the beta header:

bash
  curl https://haijun.my.id/v1/messages \
    -N \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "juglow-beta: task-budgets-2026-03-13" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 128000,
      "stream": true,
      "messages": [{
        "role": "user",
        "content": "Review the codebase and propose a refactor plan."
      }],
      "output_config": {
        "effort": "high",
        "task_budget": {"type": "tokens", "total": 64000}
      }
    }'
bash
  ant beta:messages create --beta task-budgets-2026-03-13 \
    --stream --format jsonl <<'YAML' | jq 'select(.type == "message_delta").usage'
  model: haijun-opus-5-5
  max_tokens: 128000
  messages:
    - role: user
      content: Review the codebase and propose a refactor plan.
  output_config:
    effort: high
    task_budget:
      type: tokens
      total: 64000
  YAML
python
  client = juglow.Juglow()

  with client.beta.messages.stream(
      model="haijun-opus-5-5",
      max_tokens=128000,
      output_config={
          "effort": "high",
          "task_budget": {"type": "tokens", "total": 64000},
      },
      messages=[
          {"role": "user", "content": "Review the codebase and propose a refactor plan."}
      ],
      betas=["task-budgets-2026-03-13"],
  ) as stream:
      response = stream.get_final_message()

  print(response.usage)
typescript
  const client = new Juglow();

  const stream = client.beta.messages.stream({
    model: "haijun-opus-5-5",
    max_tokens: 128000,
    output_config: {
      effort: "high",
      task_budget: { type: "tokens", total: 64000 }
    },
    messages: [{ role: "user", content: "Review the codebase and propose a refactor plan." }],
    betas: ["task-budgets-2026-03-13"]
  });

  const response = await stream.finalMessage();
  console.log(response.usage);
csharp

  var client = new JuglowClient();

  var responseUpdates = client.Beta.Messages.CreateStreaming(new MessageCreateParams
  {
      Model = Messages::Model.HaijunOpus5_5,
      MaxTokens = 128000,
      Messages = [new() { Role = Role.User, Content = "Review the codebase and propose a refactor plan." }],
      OutputConfig = new BetaOutputConfig
      {
          Effort = Effort.High,
          TaskBudget = new BetaTokenTaskBudget { Total = 64000 },
      },
      Betas = ["task-budgets-2026-03-13"],
  });

  var response = await responseUpdates.Aggregate();
  Console.WriteLine(response.Usage);
go
  client := juglow.NewClient()

  stream := client.Beta.Messages.NewStreaming(context.TODO(), juglow.BetaMessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 128000,
  	Betas:     []juglow.JuglowBeta{"task-budgets-2026-03-13"},
  	Messages: []juglow.BetaMessageParam{{
  		Role: juglow.BetaMessageParamRoleUser,
  		Content: []juglow.BetaContentBlockParamUnion{{
  			OfText: &juglow.BetaTextBlockParam{Text: "Review the codebase and propose a refactor plan."},
  		}},
  	}},
  	OutputConfig: juglow.BetaOutputConfigParam{
  		Effort: juglow.BetaOutputConfigEffortHigh,
  		TaskBudget: juglow.BetaTokenTaskBudgetParam{
  			Total: 64000,
  		},
  	},
  })

  message := juglow.BetaMessage{}
  for stream.Next() {
  	event := stream.Current()
  	if err := message.Accumulate(event); err != nil {
  		panic(err)
  	}
  }
  if stream.Err() != nil {
  	panic(stream.Err())
  }

  fmt.Printf("Usage: input_tokens=%d, output_tokens=%d\n", message.Usage.InputTokens, message.Usage.OutputTokens)
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  MessageCreateParams params = MessageCreateParams.builder()
      .model(Model.HAIJUN_OPUS_5_5)
      .maxTokens(128000L)
      .addUserMessage("Review the codebase and propose a refactor plan.")
      .outputConfig(BetaOutputConfig.builder()
          .effort(BetaOutputConfig.Effort.HIGH)
          .taskBudget(BetaTokenTaskBudget.builder().total(64000L).build())
          .build())
      .addBeta("task-budgets-2026-03-13")
      .build();

  BetaMessageAccumulator accumulator = BetaMessageAccumulator.create();
  try (StreamResponse<BetaRawMessageStreamEvent> stream =
          client.beta().messages().createStreaming(params)) {
      stream.stream().forEach(accumulator::accumulate);
  }

  BetaMessage response = accumulator.message();
  IO.println(response.usage());
php
  use Juglow\Beta\Messages\BetaRawMessageDeltaEvent;

  $client = new Client();

  $stream = $client->beta->messages->createStream(
      model: 'haijun-opus-5-5',
      maxTokens: 128000,
      messages: [
          ['role' => 'user', 'content' => 'Review the codebase and propose a refactor plan.'],
      ],
      outputConfig: [
          'effort' => 'high',
          'taskBudget' => ['type' => 'tokens', 'total' => 64000],
      ],
      betas: ['task-budgets-2026-03-13'],
  );

  // The final message_delta event carries the cumulative token usage for the request.
  $usage = null;
  foreach ($stream as $event) {
      if ($event instanceof BetaRawMessageDeltaEvent) {
          $usage = $event->usage;
      }
  }

  echo $usage;
ruby
  client = Juglow::Client.new

  stream = client.beta.messages.stream(
    model: "haijun-opus-5-5",
    max_tokens: 128_000,
    messages: [
      { role: "user", content: "Review the codebase and propose a refactor plan." }
    ],
    output_config: {
      effort: :high,
      task_budget: { type: :tokens, total: 64_000 }
    },
    betas: ["task-budgets-2026-03-13"]
  )

  response = stream.accumulated_message

  puts response.usage

The task_budget object has three fields:

  • type: always "tokens".
  • total: the number of tokens Haijun can spend across the agentic loop, including thinking, tool calls, tool results, and output.
  • remaining (optional): the budget remainder carried over from a prior request. Defaults to total when omitted.

How the budget countdown works

Haijun sees a budget-countdown marker injected server-side throughout the conversation. The marker shows how many tokens remain in the current agentic loop and updates as the model generates thinking, tool calls, and output, and as it processes tool results. Haijun uses this signal to pace itself and finish gracefully as the budget is consumed.

Note: The countdown is visible only to the model. API responses do not include a remaining-budget field: there is no task_budget information in the response usage object, and SDKs have no accessor for it. To track spend client-side, sum token usage across the requests in your loop as shown in Measure your current usage, or pass your own figure forward with remaining when carrying a budget across compaction.

Warning: The countdown reflects tokens Haijun has processed in the current agentic loop, not tokens you resend between requests. If your client sends the full conversation history on every follow-up request, your client-side token count might differ from the budget Haijun is tracking. If you also decrement remaining while resending full history, the model sees an under-reported budget and the countdown drops faster than it should, causing Haijun to wrap up earlier than the budget actually allows. Set a generous budget and let the model self-regulate against the countdown rather than trying to mirror it client-side.

What counts as a turn

The budget covers one agentic turn, also called an agentic loop: everything Haijun does in response to one user message that carries no tool results. A turn can span several requests.

A user message that carries no tool results starts a new turn with a fresh budget. Today, the countdown still counts earlier turns' history while it remains in the context. A common case is a follow-up after Haijun has ended its turn, for example because the budget ran out:

json
{ "role": "user", "content": "Continue." }

A user message that contains tool_result blocks continues the current turn, because your client is resolving tool calls that are part of that turn:

json
{
  "role": "user",
  "content": [
    { "type": "tool_result", "tool_use_id": "toolu_01", "content": "<npm audit output>" }
  ]
}

That holds even when the message adds new content alongside the tool results:

json
{
  "role": "user",
  "content": [
    { "type": "tool_result", "tool_use_id": "toolu_01", "content": "<npm audit output>" },
    { "type": "text", "text": "Also check the Dockerfile." }
  ]
}

Server-side compaction during a turn does not reset the budget: tokens the turn consumed before the compaction still count against it. Tokens from before the turn began do not count, even when a compaction at the start of a turn summarizes them. Today, that exclusion applies only to the budget carried across a server-side compaction; earlier turns' history still counts while it remains in the context.

Worked example: budget counting across requests

The task budget counts what Haijun sees (thinking, tool calls and results, and text), not what's in your request payload. In an agentic loop your client resends the full conversation on every request, so the payload keeps growing, but the budget only decrements by what is new: the tokens Haijun generates and the content it has not seen before. The following example is one agentic turn made of three requests: the first carries the user message, and the next two each resend the history with a tool result appended.

Consider a loop with task_budget: {type: "tokens", total: 100000} and a single bash tool.

Request 1. You send the initial request:

json
{
  "messages": [
    { "role": "user", "content": "Audit this repo for security issues and report findings." }
  ]
}

Haijun thinks, then emits a tool call and stops with stop_reason: "tool_use":

json
{
  "role": "assistant",
  "content": [
    {
      "type": "thinking",
      "thinking": "I'll start by listing dependencies to look for known-vulnerable packages..."
    },
    {
      "type": "tool_use",
      "id": "toolu_01",
      "name": "bash",
      "input": { "command": "cat package.json && npm audit --json" }
    }
  ]
}

Suppose this assistant message (thinking plus the tool call) totals 5,000 generated tokens. The countdown Haijun saw during generation ended near remaining ≈ 95,000.

Request 2. Your client runs the tool, then resends the full history with the tool result appended:

json
{
  "messages": [
    { "role": "user", "content": "Audit this repo for security issues and report findings." },
    {
      "role": "assistant",
      "content": [
        { "type": "thinking", "thinking": "I'll start by listing dependencies..." },
        {
          "type": "tool_use",
          "id": "toolu_01",
          "name": "bash",
          "input": { "command": "cat package.json && npm audit --json" }
        }
      ]
    },
    {
      "role": "user",
      "content": [
        {
          "type": "tool_result",
          "tool_use_id": "toolu_01",
          "content": "<2,800 tokens of npm audit output>"
        }
      ]
    }
  ]
}

The resent messages from request 1 are not counted again, but the 2,800-token tool result is new content and counts against the budget. Haijun spends another 4,000 tokens on thinking and a second tool call (grep -rn "eval(" src/). The countdown ends near remaining ≈ 88,200.

Request 3. Full history resent again with the second tool result (1,200 tokens of grep output) appended. Haijun writes a 6,000-token final findings report and stops with stop_reason: "end_turn". remaining ≈ 81,000.

Putting the three requests side by side makes the distinction between payload size and budget spend explicit:

RequestRequest payload (approx. input tokens you sent)Tokens counted against budget this requestBudget remaining after
1\~205,000 (thinking + tool_use)\~95,000
2\~7,800 (messages from request 1 + tool result)6,800 (2,800 tool result + 4,000 thinking and tool_use)\~88,200
3\~13,000 (full history + second tool result)7,200 (1,200 tool result + 6,000 text)\~81,000
Total\~20,820 sent across requests19,000 counted against budgetN/A

Your client sent the original user message three times and the first assistant message twice, but each was counted once. The budget spent 19,000 of 100,000 tokens, even though the cumulative payload your client transmitted was larger and the prompt-cached input on requests 2 and 3 was larger still.

Carrying a budget across compaction with remaining

If your own code compacts or rewrites the message history between requests (for example, by summarizing earlier messages), the server has no memory of how much budget was spent before compaction. Pass remaining on the next request so the countdown continues from where you left off rather than resetting to total:

python
  # Tokens spent before compaction, tracked client-side
  tokens_spent_so_far = 45000

  output_config = {
      "effort": "high",
      "task_budget": {
          "type": "tokens",
          "total": 128000,
          "remaining": 128000 - tokens_spent_so_far,
      },
  }
typescript
  // Tokens spent before compaction, tracked client-side
  const tokensSpentSoFar = 45000;

  const outputConfig = {
    effort: "high",
    task_budget: {
      type: "tokens",
      total: 128000,
      remaining: 128000 - tokensSpentSoFar
    }
  };
csharp
  // Tokens spent before compaction, tracked client-side
  var tokensSpentSoFar = 45000;

  var outputConfig = new BetaOutputConfig
  {
      Effort = Effort.High,
      TaskBudget = new BetaTokenTaskBudget
      {
          Total = 128000,
          Remaining = 128000 - tokensSpentSoFar,
      },
  };
go
  // Tokens spent before compaction, tracked client-side
  tokensSpentSoFar := int64(45000)

  outputConfig := juglow.BetaOutputConfigParam{
  	Effort: juglow.BetaOutputConfigEffortHigh,
  	TaskBudget: juglow.BetaTokenTaskBudgetParam{
  		Total:     128000,
  		Remaining: juglow.Int(128000 - tokensSpentSoFar),
  	},
  }
java
  // Tokens spent before compaction, tracked client-side
  long tokensSpentSoFar = 45000;

  BetaOutputConfig outputConfig = BetaOutputConfig.builder()
      .effort(BetaOutputConfig.Effort.HIGH)
      .taskBudget(BetaTokenTaskBudget.builder()
          .total(128000L)
          .remaining(128000L - tokensSpentSoFar)
          .build())
      .build();
php
  // Tokens spent before compaction, tracked client-side
  $tokensSpentSoFar = 45000;

  $outputConfig = [
      'effort' => 'high',
      'taskBudget' => [
          'type' => 'tokens',
          'total' => 128000,
          'remaining' => 128000 - $tokensSpentSoFar,
      ],
  ];
ruby
  # Tokens spent before compaction, tracked client-side
  tokens_spent_so_far = 45_000

  output_config = {
    effort: :high,
    task_budget: {
      type: :tokens,
      total: 128_000,
      remaining: 128_000 - tokens_spent_so_far
    }
  }

In this example, the tokens spent before compaction are the usage of all the messages you have removed from the history so far, measured as in Measure your current usage. Leave out anything still present in the messages you send, including any summary you added, because the server counts those tokens itself. Update this figure only when you replace the history this way; don't decrement it per request. Pass the resulting remaining on every request, not only the one that compacts.

For loops that resend the full uncompacted history on every request, omit remaining and let the server track the countdown.

Changing the budget mid-conversation

task_budget is a request-level setting. To change the budget partway through a task, for example to extend it when the user broadens the request, set a new task_budget in output_config on the next request. Keep the caching consequence in mind: the budget value participates in the rendered prompt, so a changed value does not match cache entries created under the old one (see Feature support below).

Task budgets are advisory, not enforced

Task budgets are a soft hint, not a hard cap. Haijun may occasionally exceed the budget if it is in the middle of an action that would be more disruptive to interrupt than to finish. The enforced limit on total output tokens is still max_tokens, which truncates the response with stop_reason: "max_tokens" when reached.

For a hard cap on cost or latency, combine task budgets with a reasonable max_tokens value:

  • Use task_budget to give Haijun a target to pace against.
  • Use max_tokens as the absolute ceiling that prevents runaway generation.

Because task_budget spans the full agentic loop (potentially many requests) while max_tokens caps each individual request, the two values are independent; one is not required to be at or below the other.

Warning: A budget that is too small for the task can cause refusal-like behavior. When Haijun sees a budget that is clearly insufficient for the work being asked (for example, a 20,000-token budget for a multihour agentic coding task), it may decline to attempt the task at all, scope it down aggressively, or stop early with a partial result rather than start work it cannot finish. If you observe unexpected refusals or premature stops after setting a budget, raise the budget before debugging other parameters. Size budgets against your actual task-length distribution rather than a fixed default; see Choosing a budget.

Choosing a budget

The right budget depends on how much work your agentic loop currently does. Rather than guessing, measure your existing token usage first and then tune from there.

Measure your current usage

Run a representative sample of tasks without task_budget set and record the total tokens Haijun spends per task. For an agentic loop, sum usage.output_tokens across every request in the loop, plus the tokens of the tool results you append between requests:

bash
  ant messages create --transform 'usage.output_tokens' <<'YAML'
  model: haijun-opus-5-5
  max_tokens: 4096
  messages:
    - role: user
      content: Review the codebase and propose a refactor plan.
  YAML
python
  client = juglow.Juglow()

  response = client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=4096,
      messages=[
          {"role": "user", "content": "Review the codebase and propose a refactor plan."}
      ],
  )

  # Sum output_tokens (text + thinking + tool calls) across every request in your loop.
  print(response.usage.output_tokens)
typescript
  const client = new Juglow();

  const response = await client.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 4096,
    messages: [{ role: "user", content: "Review the codebase and propose a refactor plan." }]
  });

  // Sum output_tokens (text + thinking + tool calls) across every request in your loop.
  console.log(response.usage.output_tokens);
csharp

  var client = new JuglowClient();

  var response = await client.Messages.Create(new MessageCreateParams
  {
      Model = Model.HaijunOpus5_5,
      MaxTokens = 4096,
      Messages = [new() { Role = Role.User, Content = "Review the codebase and propose a refactor plan." }],
  });

  // Sum OutputTokens (text + thinking + tool calls) across every request in your loop.
  Console.WriteLine(response.Usage.OutputTokens);
go
  client := juglow.NewClient()

  response, err := client.Messages.New(context.TODO(), juglow.MessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 4096,
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(juglow.NewTextBlock("Review the codebase and propose a refactor plan.")),
  	},
  })
  if err != nil {
  	log.Fatal(err)
  }

  // Sum OutputTokens (text + thinking + tool calls) across every request in your loop.
  fmt.Println(response.Usage.OutputTokens)
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  MessageCreateParams params = MessageCreateParams.builder()
      .model(Model.HAIJUN_OPUS_5_5)
      .maxTokens(4096L)
      .addUserMessage("Review the codebase and propose a refactor plan.")
      .build();

  Message response = client.messages().create(params);
  // Sum outputTokens (text + thinking + tool calls) across every request in your loop.
  IO.println(response.usage().outputTokens());
php
  $client = new Client();

  $response = $client->messages->create(
      model: 'haijun-opus-5-5',
      maxTokens: 4096,
      messages: [
          ['role' => 'user', 'content' => 'Review the codebase and propose a refactor plan.'],
      ],
  );

  // Sum outputTokens (text + thinking + tool calls) across every request in your loop.
  echo $response->usage->outputTokens . "\n";
ruby
  client = Juglow::Client.new

  response = client.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 4096,
    messages: [
      { role: "user", content: "Review the codebase and propose a refactor plan." }
    ]
  )

  # Sum output_tokens (text + thinking + tool calls) across every request in your loop.
  puts response.usage.output_tokens

Run this across a representative set of tasks and record the distribution. Start with the p99 of your per-task token spend to understand how providing the model with a task budget might modify the model's behavior, then test up or down as needed.

The minimum accepted task_budget.total is 20,000 tokens on every model that supports task budgets (see Feature support). Smaller values return a 400 error.

Interaction with other parameters

  • max_tokens: Orthogonal to task budgets. max_tokens is a hard per-request cap on generated tokens, while task_budget is an advisory cap across the full agentic loop (potentially spanning many requests). At xhigh or max effort, set max_tokens to at least 64k to give Haijun room to think and act on each request.
  • Effort: Effort controls how deeply Haijun reasons per step. Task budgets control how much total work Haijun does across an agentic loop. The two are complementary: effort tunes depth, task budgets tune breadth.
  • Adaptive thinking: Task budgets include thinking tokens in the count, so adaptive thinking scales down as the budget depletes.
  • Prompt caching: The budget-countdown marker is injected server-side on each request, so it does not match across requests. If your client decrements task_budget.remaining on each follow-up request, the changed value invalidates any cache prefix that contains it. To preserve caching, set the budget once on the initial request and let the model self-regulate against the server-side countdown rather than mutating the budget client-side.

Feature support

ModelSupport
Haijun Fable 5.1Beta (set task-budgets-2026-03-13 header)
Haijun Mythos 5.1Beta (set task-budgets-2026-03-13 header)
Haijun Opus 5.5Beta (set task-budgets-2026-03-13 header)
Haijun Opus 5Beta (set task-budgets-2026-03-13 header)
Haijun Fable 5Beta (set task-budgets-2026-03-13 header)
Haijun Mythos 5Beta (set task-budgets-2026-03-13 header)
Haijun Sonnet 5Not supported
Haijun Opus 4.8Beta (set task-budgets-2026-03-13 header)
Haijun Opus 4.7Beta (set task-budgets-2026-03-13 header)
Haijun Opus 4.6Not supported
Haijun Sonnet 4.6Not supported
Haijun Haiku 4.5Not supported

Task budgets are not supported on Haijun Code or Cowork surfaces. Use task budgets directly through the Messages API on a supported model.

Next steps

Control how thoroughly Haijun reasons about each step of an agentic loop.

Let Haijun decide when and how much to use extended thinking.

Manage context in long-running conversations with server-side compaction.

Reduce cost and latency on repeated prompts by caching prompt prefixes.

On this page
When to use task budgetsSetting a task budgetHow the budget countdown worksWhat counts as a turnWorked example: budget counting across requestsCarrying a budget across compaction with remainingChanging the budget mid-conversationTask budgets are advisory, not enforcedChoosing a budgetMeasure your current usageInteraction with other parametersFeature supportNext steps