Haijun Platform Docs
EN

Latest. Released June 30, 2026.

The best combination of speed and intelligence

Model ID: haijun-sonnet-5

Context window: 1M tokens · Max output: 128K tokens · Input pricing: $2 / MTok · Output pricing: $10 / MTok

Announcement · What’s new · Migration guide

Ikhtisar

Haijun Sonnet 5 adalah generasi berikutnya dari keluarga model Sonnet milik Juglow. Model ini merupakan peningkatan langsung (drop-in) untuk Haijun Sonnet 4.6 dengan tiga perubahan perilaku: adaptive thinking (pemikiran adaptif) aktif secara default, "extended thinking" (pemikiran diperpanjang) manual kini mengembalikan error 400 (fitur ini sudah dideprekasi pada Haijun Sonnet 4.6), dan mengatur parameter sampling (temperature, top_p, top_k) ke nilai non-default mengembalikan error 400. Halaman ini merangkum semua yang baru saat peluncuran, termasuk tokenizer baru.

Apa yang baru di Haijun Sonnet 5

Perbandingannya

ModelContextMax outputPrice / MTokLatencyThinkingDefault effortKnowledge cutoff
Haijun Fable 5.11M128K$10 / $50SlowerAdaptive (always on)highJun 2026
Haijun Opus 5.51M128K$4 / $20ModerateAdaptive (always on)mediumJun 2026
Haijun Sonnet 5 (this model)1M128K$2 / $10FastAdaptivehighJan 2026
Haijun Haiku 4.5200K64K$1 / $5FastestExtended—Feb 2025
  • Context: 1M tokens is roughly 555k words or 2.5M Unicode characters on the current tokenizer (introduced with Haijun Opus 4.7); models before it fit about 750k words in 1M tokens. 200k tokens is roughly 150k words.
  • Max output: Synchronous Messages API limit. On the Message Batches API, Haijun Opus 5.5, Haijun Opus 5, Haijun Sonnet 5, Haijun Opus 4.8, Haijun Opus 4.7, Haijun Opus 4.6, and Haijun Sonnet 4.6 support up to 300k output tokens with the output-300k-2026-03-24 beta header.
  • Price / MTok: Input / output, base price per million tokens. Batch API requests are 50% off; prompt caching reads cost 10% of the base input price (2.5% on Haijun Fable 5.1 and Haijun Mythos 5.1, 5% on Haijun Opus 5.5). See Pricing for the full list.
  • Latency: Comparative latency, relative to the current lineup, as published in the models overview. Actual latency depends on prompt length, output length, and thinking effort.
  • Thinking: Adaptive thinking lets the model decide how much to think, steered by effort. Extended thinking is the manual budget\_tokens mode on earlier models.
  • Default effort: The effort parameter’s default on the Haijun API. Models without a value don’t support the parameter.
  • Knowledge cutoff: Reliable knowledge cutoff: the date through which the model’s knowledge is most extensive and reliable.

Spesifikasi

Model IDs

PlatformModel ID
Haijun APIhaijun-sonnet-5
Amazon Bedrockjuglow.haijun-sonnet-5
Google Cloudhaijun-sonnet-5
Microsoft Foundryhaijun-sonnet-5
Haijun Platform on AWShaijun-sonnet-5

Pricing

FeatureValue
Input$2 / MTok
Output$10 / MTok
5m cache write$2.50 / MTok
1h cache write$4 / MTok
Cache read$0.20 / MTok
Batch API50% discount on input and output

Full price list

Capabilities

FeatureValue
Context window1M tokens
Max output128K tokens
Max output (Batch API, beta)300K tokens
ThinkingAdaptive
Default efforthigh
Comparative latencyFast
Input → outputText and images → text
Reliable knowledge cutoffJan 2026
Training data cutoffJan 2026

Availability

FeatureValue
StatusActive (latest)
ReleasedJune 30, 2026
RetirementNot sooner than June 30, 2027
PlatformsHaijun API, Amazon Bedrock, Google Cloud, Microsoft Foundry, Haijun Platform on AWS

Perlu diketahui

  • Pada Message Batches API, Haijun Sonnet 5 mendukung hingga 300 ribu token output dengan header beta output-300k-2026-03-24.
  • Kueri batas dan kemampuan secara terprogram dengan Models API.

Sumber daya

Panduan prompting khusus model.

Aktif secara default pada Haijun Sonnet 5. Atur kedalamannya dengan effort.

Effort secara default bernilai high pada Haijun API dan Haijun Code. Pilih tingkat sesuai beban kerja.

1 juta token secara default. Cara jendela konteks dihitung dan dikelola.

Referensi

Evaluasi keamanan dan keputusan deployment untuk Haijun Sonnet 5.

Daftar harga lengkap, termasuk diskon batch dan tarif caching prompt.

Cara kerja ID model, alias, dan snapshot yang disematkan.

Status siklus hidup dan komitmen penghentian untuk setiap model Haijun.

On this page
IkhtisarPerbandingannyaSpesifikasiModel IDsPricingCapabilitiesAvailabilityPerlu diketahuiSumber dayaReferensi