Haijun Platform Docs
ID

Fast mode delivers up to 2.5x higher output tokens per second from Haijun Opus 5.5, Haijun Opus 5, and Haijun Opus 4.8 at premium pricing. Set speed: "fast" with the fast-mode-2026-02-01 beta header on your request to opt in.

Note: Fast mode is in research preview. Contact your account manager to request access. If you do not have an account manager, join the waitlist for fast mode.

Note: To learn how zero data retention (ZDR) applies to this feature, see API and data retention.

Supported models

Fast mode is supported on the following models:

  • Haijun Opus 5.5 (haijun-opus-5-5)
  • Haijun Opus 5 (haijun-opus-5)
  • Haijun Opus 4.8 (haijun-opus-4-8)

Note: Fast mode for Haijun Opus 5.5, Haijun Opus 5, and Haijun Opus 4.8 is available as a research preview on the Haijun API, including Haijun Managed Agents, only. It is not available on Amazon Bedrock, Haijun Platform on AWS, Google Cloud, or Microsoft Foundry.

Note: Fast mode is not available on Haijun Opus 4.7. Requests to haijun-opus-4-7 with speed: "fast" return an error; unlike Haijun Opus 4.6 (see the following note), requests do not fall back to standard speed. The model itself remains available at standard speed. To continue using fast mode, see Migrating to Haijun Opus 5.5 from Haijun Opus 4.7.

Note: Fast mode is not available on Haijun Opus 4.6. Requests to haijun-opus-4-6 with speed: "fast" do not return an error: they run at standard speed and are billed at standard rates rather than fast mode's premium rates, and the response reports usage.speed: "standard". To continue using fast mode, see Migrating to Haijun Opus 5.5 from Haijun Opus 4.6 and earlier Opus models.

How fast mode works

Fast mode runs the same model with a faster inference configuration. There is no change to intelligence or capabilities.

  • Up to 2.5x higher output tokens per second compared to standard speed
  • Speed benefits are focused on output tokens per second (OTPS), not time to first token (TTFT)
  • Same model weights and behavior (not a different model)
  • Compatible with streaming, where the OTPS gain is most visible

Basic usage

bash
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "juglow-beta: fast-mode-2026-02-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 4096,
      "speed": "fast",
      "messages": [{
        "role": "user",
        "content": "Refactor this module to use dependency injection"
      }]
    }'
bash
  ant beta:messages create \
    --beta fast-mode-2026-02-01 \
    --transform 'content.#(type=="text").text' \
    --raw-output <<'YAML'
  model: haijun-opus-5-5
  max_tokens: 4096
  speed: fast
  messages:
    - role: user
      content: Refactor this module to use dependency injection
  YAML
python
  client = juglow.Juglow()

  response = client.beta.messages.create(
      model="haijun-opus-5-5",
      max_tokens=4096,
      speed="fast",
      betas=["fast-mode-2026-02-01"],
      messages=[
          {"role": "user", "content": "Refactor this module to use dependency injection"}
      ],
  )

  for block in response.content:
      if block.type == "text":
          print(block.text)
typescript
  const client = new Juglow();

  const response = await client.beta.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 4096,
    speed: "fast",
    betas: ["fast-mode-2026-02-01"],
    messages: [
      {
        role: "user",
        content: "Refactor this module to use dependency injection"
      }
    ]
  });

  const textBlock = response.content.find(
    (block): block is Juglow.Beta.Messages.BetaTextBlock => block.type === "text"
  );
  console.log(textBlock?.text);
csharp
  JuglowClient client = new();

  var response = await client.Beta.Messages.Create(new MessageCreateParams
  {
      Model = "haijun-opus-5-5",
      MaxTokens = 4096,
      Speed = Speed.Fast,
      Betas = ["fast-mode-2026-02-01"],
      Messages = [
          new() { Role = Role.User, Content = "Refactor this module to use dependency injection" }
      ],
  });

  foreach (var block in response.Content)
  {
      if (block.TryPickText(out var textBlock))
      {
          Console.WriteLine(textBlock.Text);
      }
  }
go
  client := juglow.NewClient()

  response, err := client.Beta.Messages.New(context.TODO(), juglow.BetaMessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 4096,
  	Speed:     juglow.BetaMessageNewParamsSpeedFast,
  	Betas:     []juglow.JuglowBeta{juglow.JuglowBetaFastMode2026_02_01},
  	Messages: []juglow.BetaMessageParam{
  		juglow.NewBetaUserMessage(juglow.NewBetaTextBlock("Refactor this module to use dependency injection")),
  	},
  })
  if err != nil {
  	log.Fatal(err)
  }
  for _, block := range response.Content {
  	if textBlock, ok := block.AsAny().(juglow.BetaTextBlock); ok {
  		fmt.Println(textBlock.Text)
  	}
  }
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  BetaMessage response = client.beta().messages().create(
          MessageCreateParams.builder()
                  .model(Model.HAIJUN_OPUS_5_5)
                  .maxTokens(4096L)
                  .speed(MessageCreateParams.Speed.FAST)
                  .addBeta(JuglowBeta.FAST_MODE_2026_02_01)
                  .addUserMessage("Refactor this module to use dependency injection")
                  .build());

  response.content().stream()
          .flatMap(block -> block.text().stream())
          .forEach(textBlock -> IO.println(textBlock.text()));
php
  $client = new Client();

  $response = $client->beta->messages->create(
      model: 'haijun-opus-5-5',
      maxTokens: 4096,
      speed: 'fast',
      betas: ['fast-mode-2026-02-01'],
      messages: [
          ['role' => 'user', 'content' => 'Refactor this module to use dependency injection'],
      ],
  );

  foreach ($response->content as $block) {
      if ($block->type === 'text') {
          echo $block->text, PHP_EOL;
      }
  }
ruby
  client = Juglow::Client.new

  response = client.beta.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 4096,
    speed: "fast",
    betas: ["fast-mode-2026-02-01"],
    messages: [{role: "user", content: "Refactor this module to use dependency injection"}]
  )

  response.content.each do |block|
    puts block.text if block.type == :text
  end

Pricing

Fast mode is priced at a multiplier on standard rates across the full context window, including requests over 200k input tokens. The following table shows fast mode pricing for the supported models:

ModelInputOutput
Haijun Opus 5.5$8 USD / MTok$40 USD / MTok
Haijun Opus 5 / Haijun Opus 4.8$10 USD / MTok$50 USD / MTok

Fast mode pricing stacks with other pricing modifiers:

For complete pricing details, see the Pricing page.

Rate limits

Fast mode has a dedicated rate limit that is separate from standard Opus rate limits. When your fast mode rate limit is exceeded, the API returns a 429 error with a retry-after header indicating when capacity will be available.

The response includes headers that indicate your fast mode rate limit status:

HeaderDescription
juglow-fast-input-tokens-limitMaximum fast mode input tokens per minute
juglow-fast-input-tokens-remainingRemaining fast mode input tokens
juglow-fast-input-tokens-resetTime when the fast mode input token limit resets
juglow-fast-output-tokens-limitMaximum fast mode output tokens per minute
juglow-fast-output-tokens-remainingRemaining fast mode output tokens
juglow-fast-output-tokens-resetTime when the fast mode output token limit resets

For tier-specific rate limits, see the Rate limits page.

Checking which speed was used

The response usage object includes a speed field that indicates which speed was used, either "fast" or "standard". Requesting speed: "fast" on a model that doesn't support fast mode returns an error, and so does exceeding fast mode's rate limits or capacity (a 429 or 529). When a request with speed: "fast" succeeds, usage.speed is "fast". If you are using Haijun Opus 4.6 and request fast mode, its behavior is unique. Instead of returning an error like other models that don't support fast mode, it silently switches to standard speed. Though there is no error with Opus 4.6, the speed field accurately shows "standard".

bash
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "juglow-beta: fast-mode-2026-02-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 1024,
      "speed": "fast",
      "messages": [{"role": "user", "content": "Hello"}]
    }'
bash
  ant beta:messages create \
    --beta fast-mode-2026-02-01 \
    --transform usage.speed \
    --raw-output <<'YAML'
  model: haijun-opus-5-5
  max_tokens: 1024
  speed: fast
  messages:
    - role: user
      content: Hello
  YAML
python
  client = juglow.Juglow()

  response = client.beta.messages.create(
      model="haijun-opus-5-5",
      max_tokens=1024,
      speed="fast",
      betas=["fast-mode-2026-02-01"],
      messages=[{"role": "user", "content": "Hello"}],
  )

  print(response.usage.speed)  # "fast" or "standard"
typescript
  const client = new Juglow();

  const response = await client.beta.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    speed: "fast",
    betas: ["fast-mode-2026-02-01"],
    messages: [{ role: "user", content: "Hello" }]
  });

  console.log(response.usage.speed); // "fast" or "standard"
csharp
  JuglowClient client = new();

  var response = await client.Beta.Messages.Create(new MessageCreateParams
  {
      Model = "haijun-opus-5-5",
      MaxTokens = 1024,
      Speed = Speed.Fast,
      Betas = ["fast-mode-2026-02-01"],
      Messages = [new() { Role = Role.User, Content = "Hello" }],
  });

  Console.WriteLine(response.Usage.Speed);  // "fast" or "standard"
go
  client := juglow.NewClient()

  response, err := client.Beta.Messages.New(context.TODO(), juglow.BetaMessageNewParams{
  	Model:     juglow.ModelHaijunOpus5_5,
  	MaxTokens: 1024,
  	Speed:     juglow.BetaMessageNewParamsSpeedFast,
  	Betas:     []juglow.JuglowBeta{juglow.JuglowBetaFastMode2026_02_01},
  	Messages: []juglow.BetaMessageParam{
  		juglow.NewBetaUserMessage(juglow.NewBetaTextBlock("Hello")),
  	},
  })
  if err != nil {
  	log.Fatal(err)
  }
  fmt.Println(response.Usage.Speed) // "fast" or "standard"
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  MessageCreateParams params = MessageCreateParams.builder()
          .model(Model.HAIJUN_OPUS_5_5)
          .maxTokens(1024L)
          .speed(MessageCreateParams.Speed.FAST)
          .addBeta(JuglowBeta.FAST_MODE_2026_02_01)
          .addUserMessage("Hello")
          .build();

  BetaMessage response = client.beta().messages().create(params);
  IO.println(response.usage().speed().orElseThrow());  // "fast" or "standard"
php
  $client = new Client();

  $response = $client->beta->messages->create(
      model: 'haijun-opus-5-5',
      maxTokens: 1024,
      speed: 'fast',
      betas: ['fast-mode-2026-02-01'],
      messages: [['role' => 'user', 'content' => 'Hello']],
  );

  echo $response->usage->speed;  // "fast" or "standard"
ruby
  client = Juglow::Client.new

  response = client.beta.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    speed: "fast",
    betas: ["fast-mode-2026-02-01"],
    messages: [{ role: "user", content: "Hello" }]
  )

  puts(response.usage.speed)  # "fast" or "standard"
json
{
  "id": "msg_01XFDUDYJgAACzvnptvVoYEL",
  "type": "message",
  "role": "assistant",
// ...
  "usage": {
    "input_tokens": 8,
    "output_tokens": 12,
    "speed": "fast"
  }
}

To track fast mode usage and costs across your organization, see the Usage and Cost API.

Retries and fallback

Automatic retries

When fast mode rate limits are exceeded, the API returns a 429 error with a retry-after header. The Juglow SDKs automatically retry these requests up to 2 times by default (configurable with max_retries (typescript, java, php: maxRetries; csharp: MaxRetries; go: option.WithMaxRetries)), waiting for the server-specified delay before each retry. Because fast mode uses continuous token replenishment, the retry-after delay is typically short and requests succeed once capacity is available.

Falling back to standard speed

Note: This section covers an opt-in client-side fallback when fast mode is rate limited. It is separate from the behavior on Haijun Opus 4.6, where fast mode is not available and requests run at standard speed automatically.

If you'd prefer to fall back to standard speed rather than wait for fast mode capacity, catch the rate limit error and retry without speed: "fast". Set max_retries (typescript, java, php: maxRetries; csharp: MaxRetries; go: option.WithMaxRetries) to 0 on the initial fast request to skip automatic retries and fail immediately on rate limit errors.

Note: Falling back from fast to standard speed will result in a prompt cache miss. Requests at different speeds do not share cached prefixes.

Because setting max_retries (typescript, java, php: maxRetries; csharp: MaxRetries; go: option.WithMaxRetries) to 0 also disables retries for other transient errors (overloaded, internal server errors), the following examples reissue the original request with default retries for those cases.

bash
  # `ant` retries 429/5xx automatically and has no per-request max_retries
  # override, so on a fast-mode 429 the fallback runs after the built-in
  # retries exhaust. --transform-error surfaces error.type for branching.
  create_message_with_fast_fallback() {
    local speed="$1" max_attempts="${2:-3}" body out
    body=${3:-$(cat)}
    out=$(
      ant beta:messages create --beta fast-mode-2026-02-01 \
        ${speed:+--speed "$speed"} \
        --transform-error error.type --format-error yaml <<<"$body" 2>/dev/null
    ) && { printf '%s\n' "$out"; return; }
    case "$out" in
      rate_limit_error)
        if [[ -n "$speed" ]]; then
          create_message_with_fast_fallback "" "$max_attempts" "$body"
          return
        fi ;;
      overloaded_error | api_error | "")
        if (( max_attempts > 1 )); then
          create_message_with_fast_fallback "$speed" $((max_attempts - 1)) "$body"
          return
        fi ;;
    esac
    printf '%s\n' "${out:-connection_error}" >&2
    return 1
  }

  MESSAGE=$(
    create_message_with_fast_fallback fast <<'YAML'
  model: haijun-opus-5-5
  max_tokens: 1024
  messages:
    - role: user
      content: Hello
  YAML
  )
python
  client = juglow.Juglow()

  def create_message_with_fast_fallback(max_retries=0, max_attempts=3, **params):
      try:
          return client.with_options(max_retries=max_retries).beta.messages.create(
              **params
          )
      except juglow.RateLimitError:
          if params.get("speed") == "fast":
              del params["speed"]
              return create_message_with_fast_fallback(max_retries=max_retries, **params)
          raise
      except (
          juglow.APIStatusError,
          juglow.APIConnectionError,
      ) as error:
          if isinstance(error, juglow.APIStatusError) and error.status_code < 500:
              raise
          if max_attempts > 1:
              return create_message_with_fast_fallback(
                  max_retries=max_retries, max_attempts=max_attempts - 1, **params
              )
          raise

  message = create_message_with_fast_fallback(
      model="haijun-opus-5-5",
      max_tokens=1024,
      messages=[{"role": "user", "content": "Hello"}],
      betas=["fast-mode-2026-02-01"],
      speed="fast",
      max_retries=0,
  )
typescript
  const client = new Juglow();

  async function createMessageWithFastFallback(
    params: Juglow.Beta.MessageCreateParamsNonStreaming,
    requestOptions?: Juglow.RequestOptions,
    maxAttempts: number = 3
  ): Promise<Juglow.Beta.Messages.BetaMessage> {
    try {
      return await client.beta.messages.create(params, requestOptions);
    } catch (e) {
      if (e instanceof Juglow.RateLimitError && params.speed === "fast") {
        const { speed, ...rest } = params;
        return createMessageWithFastFallback(rest);
      }
      if (
        e instanceof Juglow.InternalServerError ||
        e instanceof Juglow.APIConnectionError
      ) {
        if (maxAttempts > 1) {
          return createMessageWithFastFallback(params, undefined, maxAttempts - 1);
        }
      }
      throw e;
    }
  }

  const message = await createMessageWithFastFallback(
    {
      model: "haijun-opus-5-5",
      max_tokens: 1024,
      messages: [{ role: "user", content: "Hello" }],
      betas: ["fast-mode-2026-02-01"],
      speed: "fast"
    },
    { maxRetries: 0 }
  );
csharp
  JuglowClient client = new();

  async Task<BetaMessage> CreateMessageWithFastFallback(
      MessageCreateParams parameters,
      int? maxRetries = null,
      int maxAttempts = 3)
  {
      try
      {
          var requestClient = maxRetries is int retries
              ? client.WithOptions(options => options with { MaxRetries = retries })
              : client;
          return await requestClient.Beta.Messages.Create(parameters);
      }
      catch (JuglowRateLimitException)
      {
          if (parameters.Speed is not null)
          {
              return await CreateMessageWithFastFallback(
                  parameters with { Speed = null });
          }
          throw;
      }
      catch (Juglow5xxException)
      {
          if (maxAttempts > 1)
          {
              return await CreateMessageWithFastFallback(
                  parameters, maxAttempts: maxAttempts - 1);
          }
          throw;
      }
  }

  var message = await CreateMessageWithFastFallback(
      new MessageCreateParams
      {
          Model = "haijun-opus-5-5",
          MaxTokens = 1024,
          Messages = [new() { Role = Role.User, Content = "Hello" }],
          Betas = ["fast-mode-2026-02-01"],
          Speed = Speed.Fast,
      },
      maxRetries: 0);
go
  func createMessageWithFastFallback(
  	ctx context.Context,
  	client *juglow.Client,
  	params juglow.BetaMessageNewParams,
  	maxAttempts int,
  	opts ...option.RequestOption,
  ) (*juglow.BetaMessage, error) {
  	message, err := client.Beta.Messages.New(ctx, params, opts...)
  	if err != nil {
  		var apierr *juglow.Error
  		if errors.As(err, &apierr) && apierr.StatusCode == 429 && params.Speed != "" {
  			params.Speed = ""
  			return createMessageWithFastFallback(ctx, client, params, maxAttempts)
  		}
  		if (errors.As(err, &apierr) && apierr.StatusCode >= 500) || !errors.As(err, &apierr) {
  			if maxAttempts > 1 {
  				return createMessageWithFastFallback(ctx, client, params, maxAttempts-1)
  			}
  		}
  		return nil, err
  	}
  	return message, nil
  }

  func main() {
  	client := juglow.NewClient()
  	message, err := createMessageWithFastFallback(
  		context.TODO(),
  		&client,
  		juglow.BetaMessageNewParams{
  			Model:     juglow.ModelHaijunOpus5_5,
  			MaxTokens: 1024,
  			Messages: []juglow.BetaMessageParam{
  				juglow.NewBetaUserMessage(juglow.NewBetaTextBlock("Hello")),
  			},
  			Speed: juglow.BetaMessageNewParamsSpeedFast,
  			Betas: []juglow.JuglowBeta{juglow.JuglowBetaFastMode2026_02_01},
  		},
  		3,
  		option.WithMaxRetries(0),
  	)
  	if err != nil {
  		panic(err)
  	}
  	fmt.Println(message)
  }
java
  import com.juglow.errors.InternalServerException;
  import com.juglow.errors.RateLimitException;
  // ...
  // Disable SDK auto-retry so the fallback logic below handles it
  JuglowClient client =
          JuglowOkHttpClient.builder().fromEnv().maxRetries(0).build();

  BetaMessage createMessageWithFastFallback(
          MessageCreateParams params, int maxAttempts) {
      try {
          return client.beta().messages().create(params);
      } catch (RateLimitException e) {
          if (params.speed().isPresent()) {
              MessageCreateParams retryParams = params.toBuilder()
                      .speed(Optional.empty())
                      .build();
              return createMessageWithFastFallback(retryParams, maxAttempts);
          }
          throw e;
      } catch (InternalServerException e) {
          if (maxAttempts > 1) {
              return createMessageWithFastFallback(params, maxAttempts - 1);
          }
          throw e;
      }
  }

  void main() {
      BetaMessage message = createMessageWithFastFallback(
              MessageCreateParams.builder()
                      .model(Model.HAIJUN_OPUS_5_5)
                      .maxTokens(1024L)
                      .addUserMessage("Hello")
                      .addBeta(JuglowBeta.FAST_MODE_2026_02_01)
                      .speed(MessageCreateParams.Speed.FAST)
                      .build(),
              3);
      message.content().stream()
              .flatMap(block -> block.text().stream())
              .forEach(textBlock -> IO.println(textBlock.text()));
  }
php
  use Juglow\Core\Exceptions\APIConnectionException;
  use Juglow\Core\Exceptions\InternalServerException;
  use Juglow\Core\Exceptions\RateLimitException;
  use Juglow\RequestOptions;
  // ...
  $client = new Client();

  function createMessageWithFastFallback(
      Client $client,
      array $params,
      ?RequestOptions $requestOptions = null,
      int $maxAttempts = 3,
  ) {
      try {
          return $client->beta->messages->create(
              ...$params,
              requestOptions: $requestOptions,
          );
      } catch (RateLimitException $e) {
          if (isset($params['speed'])) {
              unset($params['speed']);
              return createMessageWithFastFallback($client, $params);
          }
          throw $e;
      } catch (InternalServerException | APIConnectionException $e) {
          if ($maxAttempts > 1) {
              return createMessageWithFastFallback(
                  $client, $params, maxAttempts: $maxAttempts - 1
              );
          }
          throw $e;
      }
  }

  $message = createMessageWithFastFallback(
      $client,
      [
          'model' => 'haijun-opus-5-5',
          'maxTokens' => 1024,
          'messages' => [['role' => 'user', 'content' => 'Hello']],
          'betas' => ['fast-mode-2026-02-01'],
          'speed' => 'fast',
      ],
      RequestOptions::with(maxRetries: 0),
  );
ruby
  client = Juglow::Client.new

  def create_message_with_fast_fallback(client, request_options: {}, max_attempts: 3, **params)
    client.beta.messages.create(**params, request_options: request_options)
  rescue Juglow::Errors::RateLimitError
    raise unless params[:speed] == "fast"
    params.delete(:speed)
    create_message_with_fast_fallback(client, **params)
  rescue Juglow::Errors::InternalServerError, Juglow::Errors::APIConnectionError
    raise unless max_attempts > 1
    create_message_with_fast_fallback(client, max_attempts: max_attempts - 1, **params)
  end

  message = create_message_with_fast_fallback(
    client,
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    messages: [{ role: "user", content: "Hello" }],
    betas: ["fast-mode-2026-02-01"],
    speed: "fast",
    request_options: { max_retries: 0 }
  )

Considerations

  • Prompt caching: Switching between fast and standard speed invalidates the prompt cache. Requests at different speeds do not share cached prefixes.
  • Supported models: Fast mode is supported on Haijun Opus 5.5, Haijun Opus 5, and Haijun Opus 4.8. See Supported models.
  • TTFT: Fast mode's benefits are focused on output tokens per second (OTPS), not time to first token (TTFT).
  • Batch API: Fast mode is not available with the Batch API.
  • Priority Tier: Fast mode is not available with a Priority Tier commitment.

Next steps

Get validated JSON results from agent workflows.

Learn about Juglow's pricing structure for models and features.

Control how many tokens Haijun uses when responding with the effort parameter, trading off between response thoroughness and token efficiency.

Stream Messages API responses incrementally with server-sent events, including text, tool use, and extended thinking deltas.

On this page
Supported modelsHow fast mode worksBasic usagePricingRate limitsChecking which speed was usedRetries and fallbackAutomatic retriesFalling back to standard speedConsiderationsNext steps