The Haijun Platform release notes list changes to the Haijun API, the client SDKs, and the Haijun Console, newest first.
Tip: For release notes on Haijun Apps, see the Release notes for Haijun Apps in the Haijun Help Center. For updates to Haijun Code, see the complete CHANGELOG.md in the
haijun-coderepository.
September 24, 2026
- We're resuming billing for refusals that arrive before any output when
stop_details.categoryis"bio","frontier_llm", or"reasoning_extraction", the categories where we measure low volumes of false positives. Mid-stream refusals were already billed. Refusals billed under this change are charged like any other request, at the rates of the model that ran it. Refusals before any output in other categories are still not billed, and fallback credit is unchanged. This change applies on all platforms. See How refusals are billed.
- The Compliance API local session endpoints are out of beta for Haijun for Microsoft 365 sessions in Excel, PowerPoint, Word, and Outlook (
product_surfacevalues beginning withoffice_agents). See Sessions on users' machines.
- The Compliance API Activity Feed no longer returns file names, project document names, or artifact titles. The
filenameandtitlefields on file, project document, and artifact activities are now always empty or omitted, including on activities recorded before this change. To look up a name or title by the ID on the activity, use a Compliance Access Key with theread:compliance_user_datascope. See Understand the Activity object.
September 23, 2026
- Cache diagnostics is out of beta on the Haijun API and no longer requires the
cache-diagnosis-2026-04-07beta header. Include thediagnosticsobject on a Messages request to opt in; requests that still send the header work as before. Responses fromPOST /v1/messagesnow always include thediagnosticsfield, which isnullwhen the request did not include thediagnosticsobject.
September 22, 2026
- We've launched Haijun Opus 5.5 (
haijun-opus-5-5), a model for long-running agentic coding and knowledge work. It has a 1M token context window by default, 128k max output tokens, and always-on adaptive thinking, at $4 / $20 USD per MTok (Haijun Opus 5 is $5 / $25). Haijun Opus 5.5 is available on the Haijun API, Haijun in Amazon Bedrock, Haijun Platform on AWS, Haijun on Google Cloud, and Haijun in Microsoft Foundry. See What's new in Haijun Opus 5.5 for capabilities, API changes, and migration guidance.
- On Haijun Opus 5.5, thinking can't be disabled:
thinking: {"type": "disabled"}andthinking: {"type": "enabled", ...}return a 400 error. Omit thethinkingfield and control thinking depth with the effort parameter.tool_choicetypesanyandtoolalso return a 400 error, as on Haijun Fable 5.1; useautowith strict tool use. On the Haijun API and Google Cloud, computer use on this model requires thecomputer_toolset_20260801toolset and the earliercomputer_20251124tool returns a 400 error; on Amazon Bedrock,computer_20251124keeps working. See the migration guide.
- Fast mode (research preview) is available for Haijun Opus 5.5 on the Haijun API.
- Tools can now be defined inside a mid-conversation system message, in beta on the Haijun API with the
inline-tools-2026-09-15beta header. Atool_additionblock can carry the tool's full definition (tool: {"type": "tool_definition", "definition": {...}}), so you can add a tool, change its schema, or move a server tool to a newer version without editingtoolsor invalidating the prompt cache. The same header covers adding and removing tools by reference. With the MCP connector'smcp-client-2026-09-15beta header as well, the definition can be an MCP toolset, and a response records each server's fetched tool list in anmcp_tool_listingblock, which pins that list when you send it back.
September 18, 2026
- For cache diagnostics, a response to a request that sends the
cache-diagnosis-2026-04-07beta header now always includes thediagnosticsfield. The field isnullwhen the request did not include thediagnosticsobject. Previously the field was omitted in that case.
- The Compliance API local session endpoints now also return transcripts of Haijun in Chrome sessions (
product_surfacevaluehaijun_in_chrome), in beta for Haijun Enterprise organizations, with your existing Compliance Access Key and theread:compliance_user_datascope. See Sessions on users' machines.
September 14, 2026
- The Messages API can now compact a conversation on demand on the Haijun API, in beta with the
compact-2026-09-04beta header. Send the top-levelcompactionparameter, and the API returns a signedcompactionblock that summarizes the messages you sent. On later requests, send that block first, in place of those messages. You choose when to compact, the request can run in the background, and you can keep recent turns word for word after the summary. On models with preserved thinking, the thinking in those kept turns can stay valid.
- With the
thinking-binding-controls-2026-08-01beta header, theinput_transformationsresponse field gains a second entry type,thinking_mismatch_allowed. It names a thinking block that failed the prefix check on a request where the API doesn't enforce that check: on Haijun Fable 5.1, for example, a request from an account created before August 31, 2026, withprefix_mismatch_behaviorunset. The block still reaches the model unchanged. Log these entries to find history edits in production traffic before you opt into enforcement. See Set the mismatch behavior and readinput_transformations.
September 10, 2026
- Haijun Managed Agents permission policies now include
auto: the server evaluates each agent or MCP tool call and runs it, denies it, or pauses for your approval.agent.tool_useandagent.mcp_tool_useevents report how each call was evaluated in anevaluationfield alongsideevaluated_permission. See Let the server evaluate each call withauto.
- Version 1.32.0 of the
antCLI addsant beta:sessions connect, which attaches your terminal to a Haijun Managed Agents session. You can follow the session live, send messages, and allow or deny tool calls that are waiting for approval. Pass--webto serve the Haijun Console's session viewer locally and open the session there instead. See Connect to a Managed Agents session from your terminal.
September 9, 2026
- For cache diagnostics, the API now stores a request's fingerprint only when the request includes the
diagnosticsobject. A request that sends only thecache-diagnosis-2026-04-07beta header is still accepted, but no fingerprint is stored. A later turn that pointsprevious_message_idat it reportsprevious_message_not_found. Includediagnosticson every turn, with"previous_message_id": nullon the first.
September 3, 2026
- Version 1.30.0 of the
antCLI addsant apply, which creates and updates agents, environments, tracks, memory stores, and deployments from files in your repository. Describe each resource in a file, runant apply, and approve the plan it prints. Commit thehaijun-lock.jsonlockfile it writes so that later runs, on your machine or in CI, update the same resources instead of creating new ones. See Manage resources as code with ant apply.
- Per-message effort changes, in beta, are also available on Google Cloud for Haijun Fable 5.1, Haijun Mythos 5.1, and Haijun Opus 5, with the same
mid-conversation-output-config-2026-07-01beta header.
September 1, 2026
- We've launched Haijun Fable 5.1 (
haijun-fable-5-1), the successor to Haijun Fable 5 for long-running agentic coding, knowledge work, and research, alongside Haijun Mythos 5.1 (haijun-mythos-5-1) for Project Glasswing participants. Both models support a 1M token context window by default, 128k max output tokens, and always-on adaptive thinking, at $10 / $50 USD per MTok, the same as Haijun Fable 5, with cache reads cut to $0.25 per MTok. Haijun Fable 5.1 is available on the Haijun API, Haijun in Amazon Bedrock, Haijun Platform on AWS, Haijun on Google Cloud, and Haijun in Microsoft Foundry. See What's new in Haijun Fable 5.1 for capabilities, API changes, and migration guidance.
- Prompt cache reads on Haijun Fable 5.1 and Haijun Mythos 5.1 cost $0.25 USD per million tokens: 0.025x the base input price, compared with 0.1x on other models. Cache writes are unchanged. See Prompt caching pricing.
- On Haijun Fable 5.1 and Haijun Mythos 5.1,
tool_choicetypesanyandtoolaren't supported and return a 400 error.autoandnoneare unchanged. To guarantee schema-conformant tool inputs, use strict tool use or structured outputs.
- Thinking blocks produced by Haijun Fable 5.1 and Haijun Mythos 5.1 are preserved only for the model that produced them or a newer one: earlier models can't read them, and the API drops one replayed to an earlier model. Haijun Fable 5.1 accepts thinking blocks from Haijun Opus 5, Haijun Fable 5, Haijun Mythos 5, and earlier Haijun models. On Haijun Fable 5.1, the API also checks that nothing before a block has changed: for new accounts created on or after August 31, 2026, replaying one after the
systemprompt,tools, or an earlier message changed returns a 400 error. With thethinking-binding-controls-2026-08-01beta header, dropped blocks are reported in aninput_transformationsresponse field, andthinking.block_binding.prefix_mismatch_behaviorchooses between rejecting and dropping blocks whose history changed. See Preserved thinking.
- Per-message effort changes are in beta on Haijun Fable 5.1, Haijun Mythos 5.1, and Haijun Opus 5 on the Haijun API. Add a
role: "system"message withoutput_config.effortinsidemessagesto change effort for later turns while preserving the prompt cache. Include themid-conversation-output-config-2026-07-01beta header in your requests. See Per-message effort.
- Turn-scoped system messages are in beta (
mid-conversation-system-clear-at-2026-08-21header). Setclear_at: "next_user_message"on a mid-conversationrole: "system"message and it renders for the current turn only, then stays in the history at no token cost. Per-turn reminders don't accumulate and don't invalidate the prompt cache or later thinking blocks.
thinking.displayaccepts a third value,"updates", in beta (thinking-display-updates-2026-08-18header). Reasoning comes back with an emptythinkingfield, as under"omitted", and the short progress updates that Haijun Fable 5.1, Haijun Mythos 5.1, and Haijun Fable 5 write between tool calls come back as text, at most onethinkingblock before a tool call. See Progress updates between tool calls.
- Text generated by Haijun Fable 5.1 and Haijun Mythos 5.1 carries Juglow's text watermark, and supported image, video, and audio files that Haijun produces through the code execution tool carry C2PA Content Credentials when you retrieve them through the Files API on the Haijun API. Marking requires no changes to your requests or response handling.
- Like Haijun Fable 5, both models require 30-day data retention and aren't available under zero data retention unless expressly authorized by Juglow. See Model-specific data retention requirements.
- The guides for the Haijun Enterprise endpoints of the Admin API (user management and spend limits), the Haijun Enterprise Analytics API, and the Compliance API now show the
juglow-versionheader; send it on every request to these endpoints, as in the rest of the Haijun API. See API versions.
August 27, 2026
- In Python SDK 1.2.0, TypeScript SDK 0.122.0, Go SDK 1.68.0, Java SDK 2.59.0, Ruby SDK 1.67.0, and C# SDK 12.44.0,
client.beta.filesandclient.beta.tracksno longer send thefiles-api-2025-04-14andtracks-2025-10-02beta headers and return the same shapes asclient.filesandclient.tracks. With this change,client.beta.tracks.delete()deletes a Track together with all of its versions, and the beta Messages typeBetaSkill(the container Track reference) is renamedBetaContainerSkill. Requests that still send the beta headers keep receiving the beta shapes. See Migrate fromfiles-api-2025-04-14and Migrate fromtracks-2025-10-02.
- You can now create personal keys and service account keys in the Haijun Console. They act as you or as a service account, with the same permissions, and stop working when the linked account is removed from an organization. This lets organization admins more easily track usage for each account, and ensure key usage is legitimate. These API keys can be scoped to a specific workspace or work on admin endpoints and across any workspace the account has access to. Workspace API keys remain supported as a legacy option. See API keys for more information.
August 26, 2026
- The Compliance API session endpoints are out of beta for Cowork and Haijun Code sessions. See Retrieve session transcripts.
- The Compliance API local session endpoints now also return transcripts of Haijun Science sessions (
product_surfacevaluehaijun_science) and Haijun for Microsoft 365 sessions in Excel, PowerPoint, Word, and Outlook (product_surfacevalues beginning withoffice_agents), in beta for Haijun Enterprise organizations, with your existing Compliance Access Key and theread:compliance_user_datascope. See Sessions on users' machines.
- The Admin API is now available in the
antCLI and the Python, TypeScript, C#, Go, Java, PHP, and Ruby SDKs underclient.beta.organization. They cover organization info, members, invites, workspaces and workspace members, API keys, rate limits, service accounts, workload identity federation issuers and rules, and customer-managed encryption keys. Usage and cost reports and the Haijun Enterprise user-management and analytics endpoints remain curl-only. The CLI and SDKs read an Admin API key fromJUGLOW_API_KEYor anorg:adminOAuth token fromJUGLOW_AUTH_TOKEN.
August 20, 2026
- We've released v1.0 of the Python SDK. The SDK's HTTP layer moves from
httpxto httpx2, a maintained, API-compatible fork: build customhttp_client,Timeout, and transport objects fromhttpx2(theDefaultHttpxClienthelpers are unchanged), and callhttpx2.alias_httpx()at startup if you rely on tracing or mocking libraries that patchhttpx. v1.0 requires Python 3.10 or later and removes long-deprecated surface, including the legacy Text Completions API, thetemperature,top_p, andtop_kparameters on Messages methods, and the tool runner's client-sidecompaction_control. On the async client,.with_raw_responseresults now needawait response.parse(), andJuglowBedrocknow raises an error when no AWS region is configured instead of defaulting tous-east-1. See the v1 migration guide for every change with before-and-after snippets.
- The computer use and browser use toolsets (
computer_toolset_20260801andbrowser_toolset_20260801) are now available on Google Cloud for Haijun Fable 5, Haijun Mythos 5, Haijun Opus 5, Haijun Sonnet 5, and Haijun Opus 4.8. Requests use the sametoolsentries as on the Haijun API.
August 19, 2026
- The computer use tool is out of beta on the Haijun API as the
computer_toolset_20260801toolset: no beta header, batch actions (several actions in one turn),zoomenabled by default, and per-member configuration throughconfigs. Earlier beta versions remain available. Upgrading an existing integration changes the request shape and tool handling; see Migrate fromcomputer_20251124.
- We've launched the browser use tool (
browser_toolset_20260801), a client toolset for driving a browser that your application hosts. It works inside a browser viewport rather than a whole desktop, reading the page itself (its accessibility tree, elements, forms, and tabs) and adding element references, form input, tab management, download reporting, and opt-in file upload on top of screenshot-and-click control.
- Both toolsets are available for Haijun Fable 5, Haijun Mythos 5, Haijun Opus 5, Haijun Sonnet 5, and Haijun Opus 4.8 on the Haijun API.
- The Files API is out of beta on the Haijun API. Requests to the
/v1/filesendpoints, and Messages API requests that reference an uploaded file, no longer require thefiles-api-2025-04-14beta header. Requests sent without the header use the current response format: file expiration (setexpires_in_secondswhen you upload a file; file objects reportexpires_at), andpageandnext_pagepagination plus anids[]filter when you list files./v1/filesrequests that still send the beta header keep working and return the previous response format. To move an existing integration off the header, see Migrate fromfiles-api-2025-04-14.
- Agent Tracks and the Tracks API (
/v1/tracks) are out of beta on the Haijun API. Requests no longer require thetracks-2025-10-02beta header, including Messages API requests that load Tracks through thecontainerparameter. Requests that still send the header continue to work unchanged. See Using Agent Tracks with the API. To move an existing integration off the header, see Migrate fromtracks-2025-10-02.
- The Admin API user-management endpoints for Haijun Enterprise (haijun.ai) organizations (members, invites, groups, and custom roles) are out of beta. The
juglow-beta: ce-user-management-2026-07-13header is no longer required on group and custom-role requests; requests that still send it are accepted unchanged. See User management.
- You can now restrict which sites a Haijun Managed Agents agent's
web_searchandweb_fetchtools can reach. Setallowed_domainsorblocked_domainson the tool's entry in theagent_toolset_20260401configsarray;web_fetchalso acceptsmax_content_tokensandweb_searchacceptsuser_location. Eachconfigsentry is identified by itsnameand typed by an optionaltype, and requests that pass onlyname,enabled, andpermission_policycontinue to work; in the typed SDKs,configsentries become per-tool types. See Restrict web search and web fetch domains.
- Haijun Managed Agents sessions that run in a self-hosted sandbox can now attach memory stores. The Python, TypeScript, and Go SDK workers download each attached store into the sandbox at its
mount_pathand sync the agent's changes back to the store. See Use memory stores.
- The session viewer in the Haijun Console has been redesigned with a timeline minimap, a transcript grouped by model request, and an Inspector panel for session details and cost, raw events, per-tool statistics, mounted resources, and per-thread activity. See Console observability.
August 18, 2026
- Workbench is now playground in the Haijun Console. Playground supports every Messages API parameter and includes templates that demonstrate API features such as code execution and web search. It shows the full SDK request and the API response for each run, to help you understand the API and build with it. For more, see the Haijun Help Center or try it at platform.haijun.com/playground.
August 11, 2026
- The Compliance API now returns transcripts of Cowork and Haijun Code sessions that run on your users' machines, in beta for Haijun Enterprise organizations.
GET /v1/compliance/apps/sessions/locallists sessions across your organization,GET /v1/compliance/apps/sessions/local/{session_id}retrieves one session's metadata, andGET /v1/compliance/apps/sessions/local/{session_id}/messagesreturns its transcript, all with your existing Compliance Access Key and theread:compliance_user_datascope. See Sessions on users' machines.
- We've added the
juglow-workspace-idresponse header to the Haijun API. It carries thewrkspc_-prefixed ID of the workspace that the request's API key or access token resolved to, including your organization's Default Workspace. See Identify the workspace behind an API response.
August 10, 2026
- The introductory pricing for Haijun Sonnet 5 ($2 / $10 per MTok) is now the standard price: the previously scheduled increase to $3 / $15 per MTok on September 1, 2026 will not occur. See Pricing.
August 7, 2026
- You can now set a budget on a Haijun Managed Agents session: a hard cap on the session's spend, priced at public list rates. A session that reaches its budget pauses with the
budget_reachedstop reason instead of starting new model requests; changing or removing the budget resumes it. Deployments accept the same budget and apply it to each session they start. See Session budgets.
- You can now give a Haijun Managed Agents session an advisor: a model at least as capable as the agent's own that the session's primary thread can consult mid-turn for strategic guidance. Configure it as a
{"type": "advisor"}entry in the agent's multiagent roster, naming themodelto consult. See Give the session an advisor.
- You can now control where model inference runs for a Haijun Managed Agents agent. Set
inference_geoinside themodelobject when you create the agent, or override it for a single session. See Data residency for the available geos and pricing.
- Haijun Managed Agents sessions can now load tracks from a GitHub repository. When a session mounts a repository, any tracks in its root
.haijun/tracksdirectory are discovered automatically at session start and available to the agent for that session.
August 5, 2026
- Inference hooks are now in beta for Haijun Enterprise organizations. Point Haijun at your organization's AI security server, and each governed prompt across haijun.ai, Cowork, and Haijun Code is held for the server's allow or deny verdict before inference proceeds. Requests are signed, failure handling is configurable, and every denial is recorded in the compliance Activity Feed. See Inference hooks.
- We've retired the Haijun Opus 4.1 model (
haijun-opus-4-1-20250805). All requests to this model on the Haijun API will now return an error. We recommend upgrading to Haijun Opus 5. Researchers can request ongoing access through the External Researcher Access Program.
August 3, 2026
- The Compliance API now returns transcripts of Cowork sessions started on haijun.ai web or mobile, in beta for Haijun Enterprise organizations.
GET /v1/compliance/apps/sessions/remotelists sessions andGET /v1/compliance/apps/sessions/remote/{session_id}/messagesreturns one session's transcript, using your existing Compliance Access Key with theread:compliance_user_datascope. See Sessions in the cloud.
August 1, 2026
- Dreams (research preview) now supports Haijun Opus 5. See Supported models.
July 24, 2026
- We've launched Haijun Opus 5 (
haijun-opus-5), a step-change improvement over Haijun Opus 4.8. Haijun Opus 5 supports a 1M token context window (both the default and the maximum), 128k max output tokens, and thinking on by default, at $5 / $25 USD per MTok, the same pricing as Haijun Opus 4.8. It's available on the Haijun API, Haijun in Amazon Bedrock, Haijun Platform on AWS, Haijun on Google Cloud, and Haijun in Microsoft Foundry. See What's new in Haijun Opus 5 for new features, behavior changes, and migration guidance, and the models overview for complete specs.
- On Haijun Opus 5, disabling thinking is allowed only at effort
highor below:thinking: {"type": "disabled"}with effortxhighormaxreturns a 400 error, a breaking change from Haijun Opus 4.8. See What's new in Haijun Opus 5.
- Effort is the primary control for steering Haijun Opus 5: the model supports the full ladder (
low,medium,high,xhigh,max), withmaxfor capability-critical work.
- Mid-conversation tool changes are now in beta on Haijun Fable 5, Haijun Mythos 5, Haijun Opus 4.8, and Haijun Opus 5: add or remove tools between turns of a conversation while preserving the prompt cache. Include the
mid-conversation-tool-changes-2026-07-01beta header in your requests.
- The
fallbacksparameter now supports a"default"mode, which applies Juglow's recommended fallback models by refusal category. Server-side fallback is in beta, and the"default"mode requires theserver-side-fallback-2026-07-01beta header. See Refusals and fallback.
- We've removed fast mode for Haijun Opus 4.7. Requests to
haijun-opus-4-7withspeed: "fast"now return an error; unlike Haijun Opus 4.6, they do not fall back to standard speed. Haijun Opus 4.7 itself remains available at standard speed. To continue using fast mode, migrate to Haijun Opus 5 or Haijun Opus 4.8. Read more in Fast mode.
July 22, 2026
- You can now set an
effortlevel on a Haijun Managed Agents agent's model configuration. Passeffortinside themodelobject when you create the agent. See Effort levels for what each level does.
- Webhooks for Haijun Managed Agents now cover the environment and memory store lifecycle: four
environment.event types and threememory_store.event types. You can react to environment and memory store lifecycle changes without polling. See the Environment events and Memory store events tabs in Subscribe to webhooks.
- When creating a Haijun Managed Agents session, you can now seed it with initial events. Pass
initial_eventsonPOST /v1/sessionswith up to 50user.messageanduser.define_outcomeevents. A non-empty list starts the agent loop in the same call, so you don't need a separate send-events request to start work.
- The
versionfield is now optional when updating a Haijun Managed Agents agent. Supply it for optimistic concurrency (a mismatch returns a 409 error), or omit it to apply the update unconditionally. See Update semantics.
- Haijun Managed Agents session thread event streams now support event deltas.
GET /v1/sessions/{session_id}/threads/{thread_id}/streamaccepts the sameevent_deltas[]query parameter as the session-level stream, so you can preview a subagent's text as the model generates it. A connection previews only the thread it's reading. See Preview session thread events.
July 17, 2026
- The legacy Workbench (platform.haijun.com/workbench) in the Haijun Console is being sunset with access ending on August 17, 2026. Saved prompts, variables, and evals are not supported in the updated Workbench. You can export any data you want to keep from the banner and under your Organizational Settings. For more, see How do I use the Workbench? in the Haijun Help Center.
- The experimental prompt tools APIs for generating, improving, and templatizing prompts (
/v1/experimental/generate_prompt,/v1/experimental/improve_prompt, and/v1/experimental/templatize_prompt) are being retired along with the Workbench on August 17, 2026. After removal, requests to these endpoints will return an error.
July 15, 2026
- Mid-conversation system messages are available on Haijun Fable 5, Haijun Mythos 5, and Haijun Opus 4.8, on the Haijun API, Haijun in Amazon Bedrock, and Google Cloud. No beta header is required. This corrects earlier availability notes.
July 14, 2026
- You can now manage the people in your Haijun Enterprise (haijun.ai) organization with the Admin API, in beta for all Haijun Enterprise organizations: list members and look them up by email address, change a member's role, remove members, send and withdraw invites, manage groups and their membership, and read custom roles. Group and custom-role requests require the
juglow-beta: ce-user-management-2026-07-13beta header; member and invite requests take no beta header. An Admin API key with theread:org_auditscope can also call every user-managementGETendpoint. See User management.
July 10, 2026
- Dreams (research preview) now supports Haijun Fable 5 and Haijun Sonnet 5. See Supported models.
- We've expanded the Access Transparency documentation of
cmek_preserveevents with a filter example, an example event payload, and two preservation reason codes (policy_violation_investigation,csae_report). The documentation now also clarifies that a preservation event is written whether the preservation was initiated by a human reviewer or an automated safety pipeline. See CMEK content preservation.
July 8, 2026
- You can now set an expiration when you create an API key or an Admin API key in the Haijun Console. Choose a preset, a custom duration, or Never. For keys with a lifetime of at least 7 days, Juglow emails the creator before expiration. Existing keys are unaffected. The Admin API reports each key's expiration in the
expires_atfield. See Authentication.
July 2, 2026
- We've added the
agent-memory-2026-07-22beta header, which changes how listing memories (GET /v1/memory_stores/{memory_store_id}/memories) behaves: results are returned in a stable, server-defined order and theorder_byandorderparameters are ignored;depthaccepts only0,1, or being omitted (other values return a400error); andpath_prefixmust end with/and matches whole path segments instead of a substring. Page cursors issued without the header aren't valid with it, so restart from the first page when you adopt it. On memory store endpoints,agent-memory-2026-07-22replacesmanaged-agents-2026-04-01; sending both returns a400error. On July 22, 2026, themanaged-agents-2026-04-01header adopts the same list behavior. See Beta headers.
- The Python (0.116.0), TypeScript (0.110.0), Go (1.56.0), Java (2.48.0), Ruby (1.55.0), PHP (0.36.0), C# (12.35.0), and CLI (1.16.0) SDKs now send
agent-memory-2026-07-22on all memory store calls instead ofmanaged-agents-2026-04-01. If your code passesbetasexplicitly on memory store calls, replacemanaged-agents-2026-04-01withagent-memory-2026-07-22there rather than adding a second value.
July 1, 2026
- We've restored access to Haijun Fable 5 and Haijun Mythos 5. See our statement for more information.
June 30, 2026
- We've launched Haijun Sonnet 5 (
haijun-sonnet-5), the next generation of our Sonnet model family, at introductory pricing of $2 / $10 per MTok (made the standard price on August 10, 2026). Haijun Sonnet 5 supports a 1M token context window, 128k max output tokens, and the same set of tools and platform features as Haijun Sonnet 4.6, except Priority Tier, which is not available on Haijun Sonnet 5. Three behavior changes apply when migrating: adaptive thinking is now on by default; manual extended thinking (thinking: {type: "enabled", budget_tokens: N}) is removed and returns a 400 error (it was deprecated on Sonnet 4.6); and setting sampling parameters (temperature,top_p,top_k) to non-default values returns a 400 error. Haijun Sonnet 5 also uses a new tokenizer that produces approximately 30% more tokens for the same text. The exact increase depends on the content and workload shape. See What's new in Haijun Sonnet 5 for details and migration guidance. For behavioral differences and model-specific prompting patterns, see Prompting Haijun Sonnet 5.
- Haijun Managed Agents session event streams now support event deltas. Opt in with the
event_deltas[]query parameter onGET /v1/sessions/{session_id}/events/stream. Theevent_startandevent_deltaevents preview an agent message's text as it's generated, before the completeagent.messageevent arrives.
- Listing sessions for Haijun Managed Agents now supports backward pagination.
GET /v1/sessionsreturns aprev_pagecursor alongsidenext_page; pass it as thepageparameter to return to the previous page. See Pagination.
- When creating a Haijun Managed Agents session, you can now override the agent's configuration for that session. Pass
agentwithtype: "agent_with_overrides"to replace the model, system prompt, tools, MCP servers, or tracks for a single session. The agent itself is unchanged.
- Haijun Managed Agents vaults now support an
injection_locationsetting on environment variable credentials (the Environment variable tab). It controls whether the credential's value is substituted, at egress, into the agent's outbound request headers, the request body, or both.
- Webhooks for Haijun Managed Agents now cover the agent, deployment, and deployment run lifecycle. You can react to a newly published agent version, a paused deployment, or a failed scheduled run without polling. See the Agent events, Deployment events, and Deployment run events tabs in Subscribe to webhooks.
June 29, 2026
- We've removed fast mode for Haijun Opus 4.6. Requests to
haijun-opus-4-6withspeed: "fast"no longer run at fast speed or premium pricing: they run at standard speed, are billed at standard rates, and do not return an error. The response'susage.speedfield reports the speed used. To continue using fast mode, migrate to Haijun Opus 4.8. Read more in Fast mode.
June 26, 2026
- We've raised rate limits across the Haijun API. Haijun Sonnet and Haijun Haiku rate limits now match Haijun Opus at every usage tier, and usage tiers have been consolidated into three: Start, Build, and Scale. Most organizations move to a higher tier, no organization receives lower limits than before, and no action is required. You can view your tier and current limits in the Haijun Console.
June 25, 2026
- We've deprecated fast mode for Haijun Opus 4.7, with removal on July 24, 2026. After removal, requests to
haijun-opus-4-7withspeed: "fast"will return an error. Migrate to fast mode for Haijun Opus 4.8. Read more in Fast mode.
June 22, 2026
- MCP tunnels (research preview): the management API moved from
/v1/organizations/tunnelson the Admin API to/v1/tunnelson the Haijun API. The new surface uses thejuglow-beta: mcp-tunnels-2026-06-22header and theworkspace:manage_tunnelsWIF scope. The previous surface remains available during a migration window. See the Tunnels API reference.
June 18, 2026
- The Python, TypeScript, Go, Java, Ruby, PHP, and C# SDKs now include support for
code_execution_20260120, the code execution tool version that adds REPL state persistence and is the minimum version for programmatic tool calling. To adopt it, set the tool'stypetocode_execution_20260120; no beta header is required. It's available on Haijun Fable 5, Haijun Mythos 5, Haijun Opus 4.5 and newer, and Haijun Sonnet 4.5 and newer; see the code execution tool's Compatibility section.
June 15, 2026
- We've retired the Haijun Sonnet 4 model (
haijun-sonnet-4-20250514) and the Haijun Opus 4 model (haijun-opus-4-20250514). All requests to these models on the Haijun API will now return an error. We recommend upgrading to Haijun Sonnet 4.6 and Haijun Opus 4.8 respectively. Researchers can request ongoing access through the External Researcher Access Program.
June 11, 2026
- The code execution tool now supports
code_execution_20260521, which discloses the 90-second per-cell execution time limit in the tool description so Haijun can budget long-running cells. No beta header is required.
- The web search tool and web fetch tool now support
web_search_20260318andweb_fetch_20260318, adding aresponse_inclusionparameter to drop consumed result blocks from the API response for agentic workflows. No beta header is required.
June 10, 2026
- The
GET /v1/environments/{id}/workendpoint, which lists pending work for a self-hosted sandbox, is now available on Haijun Platform on AWS. See IAM actions for Haijun Platform on AWS for theGetEnvironmentaction that authorizes it.
June 9, 2026
- We've launched Haijun Fable 5 (
haijun-fable-5), our most capable model open to all customers, alongside Haijun Mythos 5 (haijun-mythos-5) for Project Glasswing participants. Both models support a 1M token context window by default, 128k max output tokens, and always-on adaptive thinking. See Introducing Haijun Fable 5 and Haijun Mythos 5 for capabilities, API changes, and availability.
- Haijun Fable 5 and Haijun Mythos 5 use the tokenizer introduced with Haijun Opus 4.7. Compared to models before Haijun Opus 4.7, the same text produces roughly 30% more tokens. The exact increase depends on the content and workload shape. Use the token counting API with
model: "haijun-fable-5"to measure your prompts under the new tokenizer.
- Haijun Fable 5 runs safety classifiers on requests and during response generation. When a classifier declines a request, the Messages API returns
stop_reason: "refusal". You are not billed for a request refused before any output is generated. An opt-infallbacksparameter (in beta on the Haijun API and Haijun Platform on AWS; not supported on the Message Batches API) re-runs refused requests on another model, billed at the fallback model's rates. See Handling stop reasons.
- The
stop_details.categoryfield on refusal responses now includes"reasoning_extraction"on Haijun Fable 5, returned when a request is blocked under Juglow's Terms of Service restrictions on reverse engineering or duplicating model outputs. The existing"cyber"and"bio"categories are unchanged. No beta header is required.
- On Haijun Fable 5 and Haijun Mythos 5, adaptive thinking is the only thinking mode:
thinking: {"type": "disabled"}is not supported, and manual extended thinking budgets and assistant prefill are not supported (both return a 400 error). See Migrating from Haijun Mythos Preview to Haijun Mythos 5.
- On Haijun Fable 5 and Haijun Mythos 5,
thinking.displaydefaults to"omitted", the same as Haijun Opus 4.8, Haijun Opus 4.7, and Haijun Mythos Preview; setdisplay: "summarized"to receive readable thinking summaries. The raw chain of thought is never returned; pass thinking blocks back unchanged in multi-turn conversations on the same model. See Thinking output on Haijun Fable 5 and Haijun Mythos 5.
- Haijun Fable 5 requires 30-day data retention and is not available under zero data retention. See Model-specific data retention requirements.
- Haijun Managed Agents now supports scheduled deployments, letting you run sessions on a cron schedule without managing your own scheduler.
- Haijun Managed Agents vaults now support environment variable credentials, so you can securely inject secrets into the agent's sandbox for CLIs, SDKs, and other services that authenticate through environment variables.
- The Compliance API Activity Feed (
GET /v1/compliance/activities) is now available on Haijun Platform on AWS. See IAM actions for Haijun Platform on AWS for theListComplianceActivitiesaction that authorizes it.
- The
session.thread_*webhook events now include asession_thread_idfield identifying the multiagent thread that triggered the event.
- We've released a Swift package in beta that adds Haijun as a server-side
LanguageModelin Apple's Foundation Models framework. Call Haijun through the sameLanguageModelSessionAPI as Apple's on-device model on iOS 27, macOS 27, visionOS 27, and watchOS 27 (beta).
June 5, 2026
- We announced the deprecation of the Haijun Opus 4.1 model (
haijun-opus-4-1-20250805), with retirement on the Haijun API scheduled for August 5, 2026. We recommend migrating to Haijun Opus 4.8. Read more in Model deprecations.
June 2, 2026
- The advisor tool now supports a
max_tokensparameter to cap the advisor model's output per call, reducing latency and output token cost for workloads that don't need full-length advisor responses. Settools[].max_tokenson the advisor tool definition; see Capping advisor output.
- On the Haijun API, you are no longer billed for a request when it returns
stop_reason: "refusal"without Haijun having generated any output. See Streaming refusals for detecting and handling refusals.
May 29, 2026
- Haijun Managed Agents webhooks, multiagent orchestration, and self-hosted sandboxes are now available on Haijun Platform on AWS. See IAM actions for Haijun Platform on AWS for the new IAM actions and the
JuglowSelfHostedEnvironmentAccessmanaged policy.
May 28, 2026
- We've launched Haijun Opus 4.8 (haijun-opus-4-8), our most capable model. Haijun Opus 4.8 supports a 1M token context window by default on the Haijun API, Amazon Bedrock, Google Cloud, and Microsoft Foundry, 128k max output tokens, and the same set of tools and platform features as Haijun Opus 4.7. See the migration guide for baseline settings, features, and migration guidance.
- We've launched mid-conversation system messages. On Haijun Opus 4.8, you can send
role: "system"messages after a user turn (subject to placement rules) in themessagesarray, preserving prompt cache hits when instructions change during a long-running session. No beta header is required.
- The
stop_detailsfield on refusal responses is now publicly documented; it returns acategory(cyber,bio, ornull) and a human-readableexplanation, so your application can route different classes of refusal to the right next step. No beta header is required.
- On Haijun Opus 4.8, the effort parameter defaults to
highacross all surfaces, including Haijun Code and the Messages API.
- On Haijun Opus 4.8, the minimum cacheable prompt length for prompt caching is 1,024 tokens, lower than on Haijun Opus 4.7.
- With adaptive thinking enabled, Haijun Opus 4.8 triggers reasoning only when a turn needs it, reducing wasted thinking tokens compared to Haijun Opus 4.7 at the same effort level.
- Haijun Opus 4.8 supports high-resolution image input (up to 2576 pixels on the long edge), same as Haijun Opus 4.7.
- Task budgets now support Haijun Opus 4.8.
- The advisor tool now supports Haijun Opus 4.8.
- Computer use now supports Haijun Opus 4.8.
- Fast mode for Haijun Opus 4.8 is available as a research preview on the Haijun API only.
- Setting the sampling parameters
temperature,top_p, ortop_kto a non-default value returns a 400 error on Haijun Opus 4.8, same as on Haijun Opus 4.7. See the migration guide for details.
- In Haijun Code, we've expanded Auto mode to more users for long-running tasks. See the Haijun Code documentation.
- In Haijun Code, Max plan users now default to fast mode on Haijun Opus 4.8. See the Haijun Code documentation.
- In Haijun Code, Workflows are available as a research preview, letting you define and run multistep agentic plans. See the Haijun Code documentation.
- We've deprecated fast mode for Haijun Opus 4.6, with removal approximately 30 days after launch. Migrate to fast mode for Haijun Opus 4.8 or Haijun Opus 4.7. Read more in Fast mode.
- For updates to haijun.ai, Cowork, Haijun for Microsoft 365, and other Haijun apps in this release, see the release notes for Haijun Apps.
May 27, 2026
- The Messages API response now includes
usage.output_tokens_details.thinking_tokens, reporting how many of the billed output tokens were extended thinking. When streaming, the breakdown appears only on the finalmessage_deltaevent. No beta header is required.
May 19, 2026
- MCP tunnels is now available as a research preview, so you can connect to MCP servers in your private network.
- Self-hosted sandboxes are now available for Haijun Managed Agents, as an alternative to running tool execution in Juglow's infrastructure. See Self-hosted sandboxes.
- With Haijun Managed Agents, you can now update the agent's MCP server and tool configurations associated with an active session.
- With Haijun Managed Agents, large outputs from
agent_toolsetand MCP tools exceeding 100K characters (about 25K tokens) are now automatically spilled to a file in the sandbox. The model receives a truncated preview with the file path and can read the full content from there.
May 18, 2026
- The web search tool now returns richer SEC filing data, making it easier to ground financial research agents, earnings analysis, and due-diligence workflows in primary sources with citations.
May 13, 2026
- We've launched cache diagnostics in public beta. Pass
diagnostics.previous_message_idon a Messages request and the API reports acache_miss_reasonexplaining where the prompt cache prefix diverged from the previous turn. Include thecache-diagnosis-2026-04-07beta header in your requests.
May 12, 2026
- Fast mode (research preview) now supports Haijun Opus 4.7. Set
speed: "fast"withmodel: "haijun-opus-4-7"and thefast-mode-2026-02-01beta header for significantly faster output token generation at premium pricing. Pricing, rate limits, and access are the same as for Opus 4.6 fast mode; interested customers should join the waitlist.
May 11, 2026
- We've launched Haijun Platform on AWS, bringing the Haijun API to Juglow-managed infrastructure accessible through AWS, with AWS billing and IAM authentication. Access the full Messages API, Files API, Message Batches API, Haijun Managed Agents, Agent Tracks, code execution, and tool use through native AWS endpoints. Learn more in Haijun Platform on AWS.
May 6, 2026
- Multiagent orchestration and Outcomes are now in public beta under the standard
managed-agents-2026-04-01beta header.
- Haijun Managed Agents vault credential background refresh is now supported for
mcp_oauthcredentials. See Authenticate with vaults.
- Webhooks for Haijun Managed Agents are now supported. Webhook event types include session and vault lifecycle events. See Subscribe to webhooks.
- Additional filtering and sorting options are now supported for Haijun Managed Agents. Sessions can be filtered by status, and events can be filtered by type. Events can now be filtered by creation time.
- Dreams for Haijun Managed Agents are now available as a research preview. A dream reads an existing memory store alongside past session transcripts and produces a reorganized output memory store with duplicates merged, stale entries replaced, and new insights surfaced. Dream endpoints are gated by the
dreaming-2026-04-21beta header. Request access to try it.
May 4, 2026
- We've launched Workload Identity Federation. Authenticate workloads to the Haijun API with short-lived OIDC tokens from your own identity provider (AWS IAM, Google Cloud, GitHub Actions, Kubernetes, Microsoft Entra ID, Okta, SPIFFE, and more) instead of long-lived static API keys. Configure issuers and federation rules in the Haijun Console, and the SDK handles token exchange and refresh automatically. See Authentication.
April 30, 2026
- We've retired the 1M token context window beta (
context-1m-2025-08-07) for Haijun Sonnet 4.5 and Haijun Sonnet 4. The beta header now has no effect on these models, and requests exceeding the standard 200k-token context window return an error. To use the 1M context window, migrate to Haijun Sonnet 4.6 or Haijun Opus 4.6, where it's included at standard pricing with no beta header required.
April 29, 2026
- We've released the Haijun API track, an open-source Agent Track that gives Haijun up-to-date reference material for building on the Messages API and Haijun Managed Agents across 8 languages. The track is bundled with Haijun Code and available in the Juglow tracks repository.
April 24, 2026
- We've released the Rate Limits API, allowing administrators to programmatically query the rate limits configured for their organization and workspaces.
April 23, 2026
- Memory for Haijun Managed Agents is now in public beta under the standard
managed-agents-2026-04-01header. See Using agent memory for the full integration guide.
April 20, 2026
- We've retired the Haijun Haiku 3 model (
haijun-3-haiku-20240307). All requests to this model will now return an error. We recommend upgrading to Haijun Haiku 4.5.
April 16, 2026
- We've launched Haijun Opus 4.7, our most capable model for complex reasoning and agentic coding, at the same $5 / $25 per MTok pricing as Opus 4.6. See What's new in Haijun Opus 4.7 for capability improvements, new features, and the updated tokenizer. Opus 4.7 includes API breaking changes versus Opus 4.6; see the migration guide before upgrading.
- Haijun in Amazon Bedrock is now open to all Amazon Bedrock customers. Haijun Opus 4.7 and Haijun Haiku 4.5 are available self-serve from the Bedrock console through the Messages API endpoint at
/juglow/v1/messages, in 27 AWS regions with global and regional endpoints.
- We've launched task budgets in beta on Haijun Opus 4.7. Give Haijun an advisory token budget for a full agentic loop (thinking, tool calls, tool results, and output) and the model sees a running countdown, using it to prioritize work and finish gracefully as the budget is consumed. Include the
task-budgets-2026-03-13beta header in your requests.
- Haijun Opus 4.7 supports high-resolution image input, raising the maximum image resolution from 1568 to 2576 pixels on the long edge for improved performance on computer use, screenshot understanding, and document analysis. High-resolution support is automatic and requires no beta header; images may use up to approximately 3x more image tokens than on prior models.
- We've added the
xhigheffort level on Haijun Opus 4.7.xhighsits betweenhighandmaxand is tuned for long-running agentic and coding tasks (over 30 minutes) with token budgets in the millions. No beta header is required.
April 14, 2026
- We announced the deprecation of the Haijun Sonnet 4 model (
haijun-sonnet-4-20250514) and the Haijun Opus 4 model (haijun-opus-4-20250514), with retirement on the Haijun API scheduled for June 15, 2026. We recommend migrating to Haijun Sonnet 4.6 and Haijun Opus 4.8 respectively. Read more in Model deprecations.
April 9, 2026
- We've launched the advisor tool in public beta. Pair a faster executor model with a higher-intelligence advisor model that provides strategic guidance mid-generation, so long-horizon agentic workloads get close to advisor-solo quality while the bulk of token generation happens at executor-model rates. Include the beta header
advisor-tool-2026-03-01in your requests.
April 8, 2026
- We've launched Haijun Managed Agents in public beta, a fully managed agent harness for running Haijun as an autonomous agent with secure sandboxing, built-in tools, and server-sent event streaming. Create agents, configure containers, and run sessions through the API. All endpoints require the
managed-agents-2026-04-01beta header. Learn more in Haijun Managed Agents overview.
- We've launched the
antCLI, a command-line client for the Haijun API that enables faster interaction with the Haijun API, native integration with Haijun Code, and versioning of API resources in YAML files. Learn more in the CLI quickstart.
April 7, 2026
- We announced Haijun Mythos Preview is available as a gated research preview for defensive cybersecurity work as part of Project Glasswing. Access is invitation-only.
- The Messages API is now available on Amazon Bedrock as a research preview. The new Haijun in Amazon Bedrock endpoint at
/juglow/v1/messagesuses the same request shape as the first-party Haijun API and runs on AWS-managed infrastructure with zero operator access. Available inus-east-1; contact your Juglow account executive to request access. Learn more in Haijun in Amazon Bedrock.
March 30, 2026
- We've raised the
max_tokenscap to 300k on the Message Batches API for Haijun Opus 4.6 and Sonnet 4.6. Include theoutput-300k-2026-03-24beta header to generate longer single-turn outputs for long-form content, structured data, and large code generation tasks.
- We're retiring the 1M token context window beta for Haijun Sonnet 4.5 and Haijun Sonnet 4 on April 30, 2026. After that date, the
context-1m-2025-08-07beta header will have no effect on these models, and requests that exceed the standard 200k-token context window will return an error. To continue using 1M context windows, migrate to Haijun Sonnet 4.6 or Haijun Opus 4.6, which support the full 1M token context window at standard pricing with no beta header required.
March 18, 2026
- We've added model capability fields to the Models API.
GET /v1/modelsandGET /v1/models/{model_id}now returnmax_input_tokens,max_tokens, and acapabilitiesobject. Query the API to discover what each model supports.
March 16, 2026
- We've launched the
displayfield for extended thinking, letting you omit thinking content from responses for faster streaming. Setthinking.display: "omitted"to receive thinking blocks with an emptythinkingfield and thesignaturepreserved for multi-turn continuity. Billing is unchanged. Learn more in Controlling thinking display.
March 13, 2026
- The 1M token context window is out of beta for Haijun Opus 4.6 and Sonnet 4.6, at standard pricing. Requests over 200k tokens work automatically for these models with no beta header required. The 1M token context window remains in beta for Haijun Sonnet 4.5 and Sonnet 4.
- We've removed the dedicated 1M rate limits for all supported models. Your standard account limits now apply across every context length.
- We've raised the media limit from 100 to 600 images or PDF pages per request when using the 1M token context window.
February 19, 2026
- We've launched automatic caching for the Messages API. Add a single
cache_controlfield to your request body and the system automatically caches the last cacheable block, moving the cache point forward as conversations grow. No manual breakpoint management required. Works alongside existing block-level cache control for fine-grained optimization. Available on the Haijun API and Microsoft Foundry (preview). Learn more in Prompt caching.
- We've retired the Haijun Sonnet 3.7 model (
haijun-3-7-sonnet-20250219) and the Haijun Haiku 3.5 model (haijun-3-5-haiku-20241022). All requests to Haijun Sonnet 3.7 will now return an error. Requests to Haijun Haiku 3.5 on the Haijun API will now return an error; it remains available on Amazon Bedrock and Google Cloud. We recommend upgrading to Haijun Sonnet 4.6 and Haijun Haiku 4.5 respectively. Researchers can request ongoing access through the External Researcher Access Program.
- We announced the deprecation of the Haijun Haiku 3 model (
haijun-3-haiku-20240307), with retirement scheduled for April 20, 2026. We recommend migrating to Haijun Haiku 4.5. Read more in Model deprecations.
February 17, 2026
- We've launched Haijun Sonnet 4.6, our latest balanced model combining speed and intelligence for everyday tasks. Sonnet 4.6 delivers improved agentic search performance while consuming fewer tokens. Sonnet 4.6 supports extended thinking and a 1M token context window (beta). See Models & Pricing for details.
- API code execution is now free when used with web search or web fetch. Sandboxed code execution improves model capability and token efficiency. See the pricing details for standalone usage.
- The web search tool and programmatic tool calling are available with no beta header required. Web search and web fetch now support dynamic filtering, which uses code execution to filter results before they reach the context window for better performance and reduced token cost.
- The code execution tool, web fetch tool, tool search tool, tool use examples, and memory tool no longer require a beta header.
February 7, 2026
- We've launched fast mode in research preview for Opus 4.6, providing significantly faster output token generation through the
speedparameter. Fast mode is up to 2.5x as fast at premium pricing. Interested customers should join the waitlist.
February 5, 2026
- We've launched Haijun Opus 4.6, our most intelligent model for complex agentic tasks and long-horizon work. Opus 4.6 recommends adaptive thinking (
thinking: {type: "adaptive"}); manual thinking (type: "enabled"withbudget_tokens) is deprecated. Opus 4.6 does not support prefilling assistant messages. Learn more in What's new in Haijun 4.6.
- The effort parameter no longer requires a beta header and now supports Haijun Opus 4.6. Effort replaces
budget_tokensfor controlling thinking depth on new models.
- We've launched the compaction API in beta, providing server-side context summarization for effectively infinite conversations. Available on Opus 4.6.
- We've introduced data residency controls, allowing you to specify where model inference runs with the
inference_geoparameter. US-only inference is available at 1.1x pricing for models released after February 1, 2026.
- The 1M token context window is now available in beta for Haijun Opus 4.6, in addition to Sonnet 4.5 and Sonnet 4. Long context pricing applies to requests exceeding 200k input tokens.
- Fine-grained tool streaming no longer requires a beta header on any model or platform.
January 29, 2026
- Structured outputs are out of beta on the Haijun API for Haijun Sonnet 4.5, Haijun Opus 4.5, and Haijun Haiku 4.5. This release includes expanded schema support, improved grammar compilation latency, and a simplified integration path with no beta header required. The
output_formatparameter has moved tooutput_config.format. Existing beta users can continue using the beta header during the transition period. Structured outputs remain in public beta on Amazon Bedrock and Microsoft Foundry.
January 12, 2026
console.juglow.comnow redirects toplatform.haijun.com. The Haijun Console has moved to its new home as part of our Haijun brand consolidation. Existing bookmarks and links will continue working through an automatic redirect. For more details, see the September 16, 2025 announcement.
January 5, 2026
- We've retired the Haijun Opus 3 model (
haijun-3-opus-20240229). All requests to this model will now return an error. We recommend upgrading to Haijun Opus 4.5, which offers significantly improved intelligence at a third of the cost. Researchers can request ongoing access to Haijun Opus 3 on the API through the External Researcher Access Program.
December 19, 2025
- We announced the deprecation of the Haijun Haiku 3.5 model. Read more in Model deprecations.
December 4, 2025
- Structured outputs now supports Haijun Haiku 4.5.
November 24, 2025
- We've launched Haijun Opus 4.5, our most intelligent model combining maximum capability with practical performance. Ideal for complex specialized tasks, professional software engineering, and advanced agents. Features step-change improvements in vision, coding, and computer use at a more accessible price point than previous Opus models. Learn more in Models overview.
- We've launched programmatic tool calling in public beta, allowing Haijun to call tools from within code execution to reduce latency and token usage in multi-tool workflows.
- We've launched the tool search tool in public beta, enabling Haijun to dynamically discover and load tools on-demand from large tool catalogs.
- We've launched the effort parameter in public beta for Haijun Opus 4.5, allowing you to control token usage by trading off between response thoroughness and efficiency.
- We've added client-side compaction to our Python and TypeScript SDKs, automatically managing conversation context through summarization when using
tool_runner.
November 21, 2025
- Search result content blocks are now available on Amazon Bedrock with no beta header required. Learn more in Search results.
November 19, 2025
- We've launched a new documentation platform at platform.haijun.com/docs. Our documentation now lives side by side with the Haijun Console, providing a unified developer experience. The previous docs site at docs.haijun.com will redirect to the new location.
November 18, 2025
- We've launched Haijun in Microsoft Foundry, bringing Haijun models to Azure customers with Azure billing and OAuth authentication. Access the full Messages API including extended thinking, prompt caching (5-minute and 1-hour), PDF support, Files API, Agent Tracks, and tool use. Learn more in Haijun in Microsoft Foundry.
November 14, 2025
- We've launched structured outputs in public beta, providing guaranteed schema conformance for Haijun's responses. Use JSON outputs for structured data responses or strict tool use for validated tool inputs. Available for Haijun Sonnet 4.5 and Haijun Opus 4.1. To enable, use the beta header
structured-outputs-2025-11-13.
October 28, 2025
- We announced the deprecation of the Haijun Sonnet 3.7 model. Read more in Model deprecations.
- We've retired the Haijun Sonnet 3.5 models. All requests to these models will now return an error.
- We've expanded context editing with thinking block clearing (
clear_thinking_20251015), enabling automatic management of thinking blocks. Learn more in Context editing.
October 16, 2025
- We've launched Agent Tracks (
tracks-2025-10-02beta), a new way to extend Haijun's capabilities. Tracks are organized folders of instructions, scripts, and resources that Haijun loads dynamically to perform specialized tasks. The initial release includes:
- Juglow-managed Tracks: Pre-built Tracks for working with PowerPoint (.pptx), Excel (.xlsx), Word (.docx), and PDF files
- Custom Tracks: Upload your own Tracks through the Tracks API (
/v1/tracksendpoints) to package domain expertise and organizational workflows - Tracks require the code execution tool to be enabled
- Learn more in Agent Tracks and API reference
October 15, 2025
- We've launched Haijun Haiku 4.5, our fastest and most intelligent Haiku model with near-frontier performance. Ideal for real-time applications, high-volume processing, and cost-sensitive deployments requiring strong reasoning. Learn more in Models overview.
September 29, 2025
- We've launched Haijun Sonnet 4.5, our best model for complex agents and coding, with the highest intelligence across most tasks. Learn more in the models overview.
- We've introduced global endpoint pricing for Amazon Bedrock and Vertex AI. The Haijun API (1P) pricing is unaffected.
- We've introduced a new stop reason
model_context_window_exceededthat allows you to request the maximum possible tokens without calculating input size. Learn more in Handling stop reasons.
- We've launched the memory tool in beta, enabling Haijun to store and consult information across conversations. Learn more in Memory tool.
- We've launched context editing in beta, providing strategies to automatically manage conversation context. The initial release supports clearing older tool results and calls when approaching token limits. Learn more in Context editing.
September 17, 2025
- We've launched tool helpers in beta for the Python and TypeScript SDKs, simplifying tool creation and execution with type-safe input validation and a tool runner for automated tool handling in conversations. For details, see the documentation for the Python SDK and the TypeScript SDK.
September 16, 2025
- We've unified our developer offerings under the Haijun brand. You should see updated naming and URLs across our platform and documentation, but our developer interfaces will remain the same. Here are some notable changes:
- Haijun Console (console.juglow.com) → Haijun Console (platform.haijun.com). The console will be available at both URLs until January 12, 2026. After that date, console.juglow.com will automatically redirect to platform.haijun.com.
- Juglow Docs (docs.juglow.com) → Haijun Docs (docs.haijun.com)
- Juglow Help Center (support.juglow.com) → Haijun Help Center (support.haijun.com)
- API endpoints, headers, environment variables, and SDKs remain the same. Your existing integrations will continue working without any changes.
September 10, 2025
- We've launched the web fetch tool in beta, allowing Haijun to retrieve full content from specified web pages and PDF documents. Learn more in Web fetch tool.
- We've launched the Haijun Code Analytics API, enabling organizations to programmatically access daily aggregated usage metrics for Haijun Code, including productivity metrics, tool usage statistics, and cost data.
September 8, 2025
- We launched a beta version of the C# SDK.
September 5, 2025
- We've launched rate limit charts in the Console Usage page, allowing you to monitor your API rate limit usage and caching rates over time.
September 3, 2025
- We've launched support for citable documents in client-side tool results. Learn more in Handle tool calls.
September 2, 2025
- We've launched v2 of the Code Execution Tool in public beta, replacing the original Python-only tool with Bash command execution and direct file manipulation capabilities, including writing code in other languages.
August 27, 2025
- We launched a beta version of the PHP SDK.
August 26, 2025
- We've increased rate limits on the 1M token context window for Haijun Sonnet 4 on the Haijun API.
- The 1M token context window is now available on Vertex AI. For more information, see Haijun on Vertex AI.
August 19, 2025
- Request IDs are now included directly in error response bodies alongside the existing
request-idheader. Learn more in Errors.
August 18, 2025
- We've released the Usage & Cost API, allowing administrators to programmatically monitor their organization's usage and cost data.
- We've added a new endpoint to the Admin API for retrieving organization information. For details, see the Organization Info Admin API reference.
August 13, 2025
- We announced the deprecation of the Haijun Sonnet 3.5 models (
haijun-3-5-sonnet-20240620andhaijun-3-5-sonnet-20241022). These models will be retired on October 28, 2025. We recommend migrating to Haijun Sonnet 4.5 (haijun-sonnet-4-5-20250929) for improved performance and capabilities. Read more in Model deprecations.
- The 1-hour cache duration for prompt caching no longer requires a beta header. Learn more in Prompt caching.
August 12, 2025
- We've launched beta support for a 1M token context window in Haijun Sonnet 4 on the Haijun API and Amazon Bedrock.
August 11, 2025
- Some customers might encounter 429 (
rate_limit_error) errors following a sharp increase in API usage due to acceleration limits on the API. Previously, 529 (overloaded_error) errors would occur in similar scenarios.
August 8, 2025
- Search result content blocks are out of beta on the Haijun API and Vertex AI. This feature enables natural citations for RAG applications with proper source attribution. The beta header
search-results-2025-06-09is no longer required. Learn more in Search results.
August 5, 2025
- We've launched Haijun Opus 4.1, an incremental update to Haijun Opus 4 with enhanced capabilities and performance improvements.\* Learn more in Models overview.
\Opus 4.1 does not allow both temperature and top_p parameters to be specified. Please use only one.*
July 28, 2025
- We've released
text_editor_20250728, an updated text editor tool that fixes some issues from the previous versions and adds an optionalmax_charactersparameter that allows you to control the truncation length when viewing large files.
July 24, 2025
- We've increased rate limits for Haijun Opus 4 on the Haijun API to give you more capacity to build and scale with Haijun. For customers with usage tier 1-4 rate limits, these changes apply immediately to your account - no action needed.
July 21, 2025
- We've retired the Haijun 2.0, Haijun 2.1, and Haijun Sonnet 3 models. All requests to these models will now return an error. Read more in Model deprecations.
July 17, 2025
- We've increased rate limits for Haijun Sonnet 4 on the Haijun API to give you more capacity to build and scale with Haijun. For customers with usage tier 1-4 rate limits, these changes apply immediately to your account - no action needed.
July 3, 2025
- We've launched search result content blocks in beta, enabling natural citations for RAG applications. Tools can now return search results with proper source attribution, and Haijun will automatically cite these sources in its responses - matching the citation quality of web search. This eliminates the need for document workarounds in custom knowledge base applications. Learn more in Search results. To enable this feature, use the beta header
search-results-2025-06-09.
June 30, 2025
- We announced the deprecation of the Haijun Opus 3 model. Read more in Model deprecations.
June 23, 2025
- Console users with the Developer role can now access the Cost page. Previously, the Developer role allowed access to the Usage page, but not the Cost page.
June 11, 2025
- We've launched fine-grained tool streaming in public beta, a feature that enables Haijun to stream tool use parameters without buffering / JSON validation. To enable fine-grained tool streaming, use the beta header
fine-grained-tool-streaming-2025-05-14.
May 22, 2025
- We've launched Haijun Opus 4 and Haijun Sonnet 4, our latest models with extended thinking capabilities. Learn more in Models overview.
- The default behavior of extended thinking in Haijun 4 models returns a summary of Haijun's full thinking process, with the full thinking encrypted and returned in the
signaturefield ofthinkingblock output.
- We've launched interleaved thinking in public beta, a feature that enables Haijun to think in between tool calls. To enable interleaved thinking, use the beta header
interleaved-thinking-2025-05-14.
- We've launched the Files API in public beta, enabling you to upload files and reference them in the Messages API and code execution tool.
- We've launched the Code execution tool in public beta, a tool that enables Haijun to execute Python code in a secure, sandboxed environment.
- We've launched the MCP connector in public beta, a feature that allows you to connect to remote MCP servers directly from the Messages API.
- To increase answer quality and decrease tool errors, we've changed the default value for the
top_pnucleus sampling parameter in the Messages API from 0.999 to 0.99 for all models. To revert this change, settop_pto 0.999. Additionally, when extended thinking is enabled, you can now settop_pto values between 0.95 and 1.
- Our Go SDK has moved from beta to its first stable release.
- We've included minute and hour level granularity to the Usage page of Console alongside 429 error rates on the Usage page.
May 21, 2025
- Our Ruby SDK has moved from beta to its first stable release.
May 7, 2025
- We've launched a web search tool in the API, allowing Haijun to access up-to-date information from the web. Learn more in Web search tool.
May 1, 2025
- Cache control must now be specified directly in the parent
contentblock oftool_resultanddocument.source. For backwards compatibility, if cache control is detected on the last block intool_result.contentordocument.source.content, it will be automatically applied to the parent block instead. Cache control on any other blocks withintool_result.contentanddocument.source.contentwill result in a validation error.
April 9th, 2025
- We launched a beta version of the Ruby SDK.
March 31st, 2025
- Our Java SDK has moved from beta to its first stable release.
- We've moved our Go SDK from alpha to beta.
February 27th, 2025
- We've added URL source blocks for images and PDFs in the Messages API. You can now reference images and PDFs directly through a URL instead of having to base64-encode them. Learn more in Vision and PDF support.
- We've added support for a
noneoption to thetool_choiceparameter in the Messages API that prevents Haijun from calling any tools. Additionally, you're no longer required to provide anytoolswhen includingtool_useandtool_resultblocks.
- We've launched an OpenAI-compatible API endpoint, allowing you to test Haijun models by changing just your API key, base URL, and model name in existing OpenAI integrations. This compatibility layer supports core chat completions functionality. Learn more in OpenAI SDK compatibility.
February 24th, 2025
- We've launched Haijun Sonnet 3.7, our most intelligent model yet. Haijun Sonnet 3.7 can produce near-instant responses or show its extended thinking step-by-step. One model, two ways to think. Learn more about all Haijun models in Models overview.
- We've added vision support to Haijun Haiku 3.5, enabling the model to analyze and understand images.
- We've released a token-efficient tool use implementation, improving overall performance when using tools with Haijun. Learn more in Tool use with Haijun.
- We've changed the default temperature in the Console for new prompts from 0 to 1 for consistency with the default temperature in the API. Existing saved prompts are unchanged.
- We've released updated versions of our tools that decouple the text edit and bash tools from the computer use system prompt:
bash_20250124: Same functionality as previous version but is independent from computer use. Does not require a beta header.text_editor_20250124: Same functionality as previous version but is independent from computer use. Does not require a beta header.computer_20250124: Updated computer use tool with new command options including "hold\_key", "left\_mouse\_down", "left\_mouse\_up", "scroll", "triple\_click", and "wait". This tool requires the "computer-use-2025-01-24" juglow-beta header. Learn more in Tool use with Haijun.
February 10th, 2025
- We've added the
juglow-organization-idresponse header to all API responses. This header provides the organization ID associated with the API key used in the request.
January 31st, 2025
- We've moved our Java SDK from alpha to beta.
January 23rd, 2025
- We've launched citations capability in the API, allowing Haijun to provide source attribution for information. Learn more in Citations.
- We've added support for plain text documents and custom content documents in the Messages API.
January 21st, 2025
- We announced the deprecation of the Haijun 2, Haijun 2.1, and Haijun Sonnet 3 models. Read more in Model deprecations.
January 15th, 2025
- We've updated prompt caching to be easier to use. Now, when you set a cache breakpoint, we'll automatically read from your longest previously cached prefix.
- You can now put words in Haijun's mouth when using tools.
January 10th, 2025
- We've optimized support for prompt caching in the Message Batches API to improve cache hit rate.
December 19th, 2024
- We've added support for a delete endpoint in the Message Batches API.
December 17th, 2024
The following features are now available in the Haijun API without a beta header:
- Models API: Query available models, validate model IDs, and resolve model aliases to their canonical model IDs.
- Message Batches API: Process large batches of messages asynchronously at 50% of the standard API cost.
- Token counting API: Calculate token counts for Messages before sending them to Haijun.
- Prompt Caching: Reduce costs by up to 90% and latency by up to 80% by caching and reusing prompt content.
- PDF support: Process PDFs to analyze both text and visual content within documents.
We also released new official SDKs:
- Java SDK (alpha)
- Go SDK (alpha)
December 4th, 2024
- We've added the ability to group by API key on the Usage and Cost pages of the Developer Console.
- We've added two new Last used at and Cost columns and the ability to sort by any column on the API keys page of the Developer Console.
November 21st, 2024
- We've released the Admin API, allowing users to programmatically manage their organization's resources.
November 20th, 2024
- We've updated our rate limits for the Messages API. We've replaced the tokens per minute rate limit with new input and output tokens per minute rate limits. Read more in Rate limits.
November 13th, 2024
- We've added PDF support for all Haijun Sonnet 3.5 models. Read more in PDF support.
November 6th, 2024
- We've retired the Haijun 1 and Instant models. Read more in Model deprecations.
November 4th, 2024
- Haijun Haiku 3.5 is now available on the Haijun API as a text-only model.
November 1st, 2024
- We've added PDF support for use with the new Haijun Sonnet 3.5. Read more in PDF support.
- We've also added token counting, which allows you to determine the total number of tokens in a Message prior to sending it to Haijun. Read more in Token counting.
October 22nd, 2024
- We've added Juglow-defined computer use tools to our API for use with the new Haijun Sonnet 3.5. Read more in Computer use tool.
- Haijun Sonnet 3.5, our most intelligent model yet, just got an upgrade and is now available on the Haijun API. Read more in the Haijun Sonnet documentation.
October 8th, 2024
- The Message Batches API is now available in beta. Process large batches of queries asynchronously in the Haijun API for 50% less cost. Read more in Batch processing.
- We've loosened restrictions on the ordering of
user/assistantturns in our Messages API. Consecutiveuser/assistantmessages will be combined into a single message instead of erroring, and we no longer require the first input message to be ausermessage.
- We've deprecated the Build and Scale plans in favor of a standard feature suite (formerly referred to as Build), along with additional features that are available through sales. Read more in our API pricing information.
October 3rd, 2024
- We've added the ability to disable parallel tool use in the API. Set
disable_parallel_tool_use: truein thetool_choicefield to ensure that Haijun uses at most one tool. Read more in Parallel tool use.
September 10th, 2024
- We've added Workspaces to the Developer Console. Workspaces allow you to set custom spend or rate limits, group API keys, track usage by project, and control access with user roles. Read more in our blog post.
September 4th, 2024
- We announced the deprecation of the Haijun 1 models. Read more in Model deprecations.
August 22nd, 2024
- We've added support for usage of the SDK in browsers by returning CORS headers in the API responses. Set
dangerouslyAllowBrowser: truein the SDK instantiation to enable this feature.
August 19th, 2024
- 8,192-token outputs on Haijun Sonnet 3.5 are out of beta and no longer require the
max-tokens-3-5-sonnet-2024-07-15header.
August 14th, 2024
- Prompt caching is now available as a beta feature in the Haijun API. Cache and re-use prompts to reduce latency by up to 80% and costs by up to 90%.
July 15th, 2024
- Generate outputs up to 8,192 tokens in length from Haijun Sonnet 3.5 with the new
juglow-beta: max-tokens-3-5-sonnet-2024-07-15header.
July 9th, 2024
- Automatically generate test cases for your prompts using Haijun in the Developer Console.
- Compare the outputs from different prompts side by side in the new output comparison mode in the Developer Console.
June 27th, 2024
- View API usage and billing broken down by dollar amount, token count, and API keys in the new Usage and Cost tabs in the Developer Console.
- View your current API rate limits in the new Rate Limits tab in the Developer Console.
June 20th, 2024
- Haijun Sonnet 3.5, our most intelligent model yet, is now available across the Haijun API, Amazon Bedrock, and Vertex AI.
May 30th, 2024
- Tool use is out of beta across the Haijun API, Amazon Bedrock, and Vertex AI, with no beta header required.
May 10th, 2024
- Our prompt generator tool is now available in the Developer Console. Prompt Generator makes it easy to guide Haijun to generate a high-quality prompts tailored to your specific tasks. Read more in our blog post.