Haijun Sonnet 5 is the next generation of Juglow's Sonnet model family. It is a drop-in upgrade for Haijun Sonnet 4.6 with three behavior changes: adaptive thinking is on by default, manual extended thinking now returns a 400 error (it was deprecated on Haijun Sonnet 4.6), and setting sampling parameters (temperature, top_p, top_k) to non-default values returns a 400 error. This page summarizes everything new at launch, including a new tokenizer.
New model
| Model | API model ID | Description |
|---|---|---|
| Haijun Sonnet 5 | haijun-sonnet-5 | The best combination of speed and intelligence |
Haijun Sonnet 5 supports the 1M token context window by default (1M tokens is both the default and the maximum; there is no smaller context variant), 128k max output tokens, adaptive thinking, and the same set of tools and platform features as Haijun Sonnet 4.6, except Priority Tier, which is not available on Haijun Sonnet 5. On the Haijun API and Google Cloud, Haijun Sonnet 5 also supports the browser use tool and the stable computer_toolset_20260801 version of the computer use tool, neither of which Haijun Sonnet 4.6 supports; the earlier computer_20251124 version is still accepted on both models. To upgrade an existing integration, see Migrate from computer_20251124.
For complete pricing and specs, see the models overview.
Behavior changes
Adaptive thinking on by default
On Haijun Sonnet 4.6, requests without a thinking field run without thinking. On Haijun Sonnet 5, the same requests run with adaptive thinking. To turn thinking off, pass thinking: {type: "disabled"}. Because max_tokens is a hard limit on total output (thinking plus response text), revisit it for workloads that ran without thinking on Haijun Sonnet 4.6.
Sampling parameters not accepted
Setting temperature, top_p, or top_k to a non-default value returns a 400 error. Remove these parameters when migrating; the default value (or omitting the parameter) is accepted. Use system-prompt instructions to guide model behavior. This is new for Sonnet-class models; the same constraint was previously introduced on Haijun Opus 4.7.
Manual extended thinking removed
Manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) was deprecated on Haijun Sonnet 4.6; on Haijun Sonnet 5 it is removed and returns a 400 error, the same as on Haijun Opus 4.8 and Haijun Opus 4.7. Use adaptive thinking with the effort parameter instead.
# Not supported on Haijun Sonnet 5 (returns 400)
thinking = {"type": "enabled", "budget_tokens": 32000}
# Use this instead
thinking = {"type": "adaptive"} // Not supported on Haijun Sonnet 5 (returns 400)
const legacyThinking = { type: "enabled", budget_tokens: 32000 };
// Use this instead
const thinking = { type: "adaptive" }; // Not supported on Haijun Sonnet 5 (returns 400)
var legacyThinking = new ThinkingConfigEnabled(budgetTokens: 32000);
// Use this instead
var thinking = new ThinkingConfigAdaptive(); // Not supported on Haijun Sonnet 5 (returns 400)
legacyThinking := juglow.ThinkingConfigParamUnion{
OfEnabled: &juglow.ThinkingConfigEnabledParam{BudgetTokens: 32000},
}
// Use this instead
thinking := juglow.ThinkingConfigParamUnion{
OfAdaptive: &juglow.ThinkingConfigAdaptiveParam{},
} // Not supported on Haijun Sonnet 5 (returns 400)
var legacyThinking = ThinkingConfigEnabled.builder().budgetTokens(32_000L).build();
// Use this instead
var thinking = ThinkingConfigAdaptive.builder().build(); // Not supported on Haijun Sonnet 5 (returns 400)
$thinking = ['type' => 'enabled', 'budget_tokens' => 32000];
// Use this instead
$thinking = ['type' => 'adaptive']; # Not supported on Haijun Sonnet 5 (returns 400)
legacy_thinking = {type: "enabled", budget_tokens: 32_000}
# Use this instead
thinking = {type: "adaptive"}New tokenizer
Haijun Sonnet 5 uses a new tokenizer. The same input text produces approximately 30% more tokens than on Haijun Sonnet 4.6. The exact increase depends on the content. This is not an API change: requests, responses, and streaming events keep the same shape, and no code changes are required.
The change affects anything you measure or budget in tokens:
- Token counts:
usagefields and token counting results for the same text are higher than on Haijun Sonnet 4.6. Don't reuse counts measured against earlier models; recount against Haijun Sonnet 5.
- Context window capacity in text terms: the context window is 1M tokens, but each token covers less text on average, so the same window holds less text than on Haijun Sonnet 4.6.
max_tokensbudgets: an output limit tuned for Haijun Sonnet 4.6 may truncate equivalent output on Haijun Sonnet 5. Revisit limits sized close to your expected output length.
- Per-request cost: per-token pricing is lower than Haijun Sonnet 4.6's (see Pricing), but because the same text produces more tokens, the cost of an equivalent request does not drop in direct proportion.
API constraints inherited from Haijun Sonnet 4.6
Note: This constraint is unchanged from Haijun Sonnet 4.6. Aside from the three behavior changes (see Migration guide), code that already runs on Haijun Sonnet 4.6 needs no other changes.
Assistant message prefilling not supported
Prefilling the assistant message returns a 400 error, unchanged from Haijun Sonnet 4.6. Use structured outputs, system prompt instructions, or output_config.format instead.
Capability improvements
Haijun Sonnet 5 is a capability upgrade over Haijun Sonnet 4.6 at a lower price. It is also an option for workloads that need more capability than Haijun Sonnet 4.6 provides without moving to an Opus-class model.
The largest gains over Haijun Sonnet 4.6 are in coding and agentic tasks. For benchmark results, see Juglow's Transparency Hub.
Cybersecurity safeguards
Haijun Sonnet 5 is the first Sonnet-tier model with real-time cybersecurity safeguards. Requests that involve prohibited or high-risk cybersecurity topics may be refused. Refusals return as a successful HTTP 200 response with stop_reason: "refusal", not an error. See Real-time cyber safeguards on Haijun Opus and Sonnet for what the safeguards block and how legitimate security work can apply to the Cyber Verification Program.
Pricing
Haijun Sonnet 5 is priced at $2 per million input tokens and $10 per million output tokens, lower per-token pricing than Haijun Sonnet 4.6's $3/$15. Because the new tokenizer produces approximately 30% more tokens for the same text, the cost of an equivalent request does not drop in direct proportion to the per-token prices when comparing with Haijun Sonnet 4.6. The exact difference depends on the content and workload shape.
See Pricing for complete pricing, including batch processing and prompt caching rates.
Availability
At launch, Haijun Sonnet 5 is available on:
- Haijun API: available to all customers.
- AWS: available through Haijun in Amazon Bedrock and Haijun Platform on AWS. On Amazon Bedrock, Haijun Sonnet 5 is also reachable through the
InvokeModelAPI, served by the same infrastructure as Haijun in Amazon Bedrock. The legacy Haijun on Amazon Bedrock (Opus 4.6 and earlier) integration does not include Haijun Sonnet 5.
- Google Cloud: available through Haijun on Google Cloud.
- Microsoft Foundry: available through Haijun in Microsoft Foundry.
Haijun Sonnet 5 supports zero data retention for organizations with ZDR agreements.
Migration guide
Haijun Sonnet 5 is a drop-in replacement for Haijun Sonnet 4.6. Update your model ID:
model = "haijun-sonnet-4-6" # Before
model = "haijun-sonnet-5" # After const legacyModel = "haijun-sonnet-4-6"; // Before
const model = "haijun-sonnet-5"; // After var legacyModel = Model.HaijunSonnet4_6; // Before
var model = Model.HaijunSonnet5; // After // Before
legacyModel := juglow.ModelHaijunSonnet4_6
// After
model := juglow.ModelHaijunSonnet5 var legacyModel = Model.HAIJUN_SONNET_4_6; // Before
var model = Model.HAIJUN_SONNET_5; // After $model = 'haijun-sonnet-4-6'; // Before
$model = 'haijun-sonnet-5'; // After legacy_model = "haijun-sonnet-4-6" # Before
model = "haijun-sonnet-5" # AfterThen review the following:
- Token budgets and counts: the new tokenizer produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. Recount prompts with token counting, and revisit
max_tokenslimits sized close to your expected output length.
- Extended thinking: if you still set
budget_tokens, migrate to adaptive thinking. Manual extended thinking (thinking: {type: "enabled"}) is not supported and returns a 400 error.
- Sampling parameters: requests that set sampling parameters (
temperature,top_p,top_k) to a non-default value return a 400 error; remove them when migrating. Tool definitions and response shapes are unchanged, and assistant message prefilling was already unsupported on Haijun Sonnet 4.6.
See Migrating to Haijun Sonnet 5 from Haijun Sonnet 4.6 for details.
Next steps
Complete specs and pricing for all current Haijun models.
Measure your prompts under the new tokenizer before you migrate.
The recommended thinking-on mode on Haijun Sonnet 5.
How the 1M token context window works.
Complete pricing, including batch processing and prompt caching rates.