Haijun Platform Docs
ID

Haijun's API surface is organized into five areas:

  • Model capabilities: Control how Haijun reasons and formats responses.
  • Tools: Let Haijun take actions on the web or in your environment.
  • Tool infrastructure: Handles discovery and orchestration at scale.
  • Context management: Keeps long-running sessions efficient.
  • Files and assets: Manage the documents and data you provide to Haijun.

If you're new, start with model capabilities and tools. Return to the other sections when you're ready to optimize cost, latency, or scale.

For administration and governance, see the Admin API, the Usage and Cost API, and the Compliance API.

Feature availability

The Availability column in each of the following tables lists the platforms that offer a feature. A platform listed without a label offers the feature as stable, fully supported, and recommended for production use, with no beta header and with standard API versioning guarantees. A label after a platform name marks one of the following classifications on that platform. Not all features pass through every stage, and a feature may enter at any stage or skip stages.

ClassificationDescription
BetaPreview features used for gathering feedback and iterating on a less mature use case. Availability may be limited, including through sign-up requirements or waitlists, and may not be publicly announced. Features may change significantly or be discontinued based on feedback. Not guaranteed for ongoing production use. Breaking changes are possible with notice, and some platform-specific limitations may apply. Most beta features on the Haijun API and Haijun Platform on AWS have a beta header.
DeprecatedFeature is still functional but no longer recommended. A migration path and removal timeline are provided.
RetiredFeature is no longer available.

Platform labels: Haijun API (Juglow first-party) · Bedrock (AWS-operated) · Haijun Platform on AWS (Juglow-operated on AWS) · Google Cloud (Google-operated) · Microsoft Foundry (Juglow-operated on Azure)

Model capabilities

Ways to steer Haijun and Haijun's direct outputs, including response format, reasoning depth, and input modalities.

Tip: You can discover which capabilities a model supports programmatically. The Models API returns max_input_tokens, max_tokens, and a capabilities object for every available model.

The ZDR column indicates whether a feature is available under a Zero Data Retention arrangement. For most features this depends only on what the feature mechanism retains; for features tied to specific models, model-level ZDR availability also applies. See Model-specific data retention requirements.

FeatureDescriptionZero Data Retention (ZDR)Availability
Context windowsUp to 1M tokens for processing large documents, extensive code bases, and long conversations.ZDR eligible
Adaptive thinkingLet Haijun dynamically decide when and how much to think. The only thinking mode on Haijun 4.7 and later models. Use the effort parameter to control thinking depth.ZDR eligible
Batch processingProcess large volumes of requests asynchronously for cost savings. Send batches with a large number of queries per batch. Batch API calls cost 50% less than standard API calls.Not ZDR eligible
CitationsGround Haijun's responses in source documents. With Citations, Haijun can provide detailed references to the exact sentences and passages it uses to generate responses, leading to more verifiable, trustworthy outputs.ZDR eligible
Data residencyControl where model inference runs using geographic controls. Specify "global" or "us" routing per request through the inference_geo parameter.ZDR eligible
EffortControl how many tokens Haijun uses when responding with the effort parameter, trading off between response thoroughness and token efficiency.ZDR eligible
Fallback creditAvoid paying the prompt-cache cost twice when you retry a refused request on another model. The refusal carries a credit token, and echoing it on the retry bills the retry as though the conversation had been on the new model all along. Message Batches results do not include fallback credit tokens.Not ZDR eligible\*
PDF supportProcess and analyze text and visual content from PDF documents.ZDR eligible
Search resultsEnable natural citations for RAG applications by providing search results with proper source attribution. Achieve web search-quality citations for custom knowledge bases and tools.ZDR eligible
Server-side fallbackRetry a refused request inside a single API call. Use the "default" mode to apply Juglow's recommended fallback models, or name up to three models of your own; when the requested model declines, the API runs the next model in the chain on the same request. The fallbacks parameter is not available in the Message Batches API.Not ZDR eligible\*
Structured outputsGuarantee schema conformance with two approaches: JSON outputs for structured data responses, and strict tool use for validated tool inputs.ZDR eligible (qualified)\*
ThinkingEnhanced reasoning capabilities for complex tasks, providing transparency into Haijun's step-by-step thought process before delivering its final answer.ZDR eligible

Tools

Built-in tools that Haijun invokes through tool_use. Server-side tools are run by the platform; client-side tools are implemented and executed by you.

Server-side tools

FeatureDescriptionZDRAvailability
Advisor toolPair a faster executor model with a higher-intelligence advisor model that provides strategic guidance mid-generation for long-horizon agentic workloads.ZDR eligible
Code executionRun code in a sandboxed environment for advanced data analysis, calculations, and file processing. Free when used with web search or web fetch.Not ZDR eligible†
Web fetchRetrieve full content from specified web pages and PDF documents for in-depth analysis.ZDR eligible\*
Web searchAugment Haijun's comprehensive knowledge with current, real-world data from across the web.ZDR eligible\*

Client-side tools

FeatureDescriptionZDRAvailability
BashExecute bash commands and scripts to interact with the system shell and perform command-line operations.ZDR eligible
Browser useNavigate, read, and interact with webpages in your own browser environment.ZDR eligible
Computer useControl computer interfaces by taking screenshots and issuing mouse and keyboard commands.ZDR eligible
MemoryEnable Haijun to store and retrieve information across conversations. Build knowledge bases over time, maintain project context, and learn from past interactions.ZDR eligible
Text editorCreate and edit text files with a built-in text editor interface for file manipulation tasks.ZDR eligible

Tool infrastructure

Infrastructure that supports discovering, orchestrating, and scaling tool use.

FeatureDescriptionZDRAvailability
Agent TracksExtend Haijun's capabilities with Tracks. Use pre-built Tracks (PowerPoint, Excel, Word, PDF) or create custom Tracks with instructions and scripts. Tracks use progressive disclosure to efficiently manage context.Not ZDR eligible†‡
Fine-grained tool streamingStream tool use parameters without buffering/JSON validation, reducing latency for receiving large parameters.ZDR eligible
MCP connectorConnect to remote MCP servers directly from the Messages API without a separate MCP client.Not ZDR eligible
Programmatic tool callingEnable Haijun to call your tools programmatically from within code execution containers, reducing latency and token consumption for multi-tool workflows.Not ZDR eligible†
Tool searchScale to thousands of tools by dynamically discovering and loading tools on-demand using regex- and BM25-based search, optimizing context usage and improving tool selection accuracy.ZDR eligible

Context management

Infrastructure for controlling and optimizing Haijun's context window.

FeatureDescriptionZDRAvailability
Compaction at a token thresholdServer-side context summarization for long-running conversations. When input tokens reach the trigger threshold, the API automatically summarizes earlier parts of the conversation.ZDR eligible
Context editingAutomatically manage conversation context with configurable strategies. Supports clearing tool results when approaching token limits and managing thinking blocks in extended thinking conversations.ZDR eligible
Automatic prompt cachingSimplify prompt caching to a single API parameter. The system automatically caches the last cacheable block in your request, moving the cache point forward as conversations grow.ZDR eligible
Prompt caching (5m)Provide Haijun with more background knowledge and example outputs to reduce costs and latency.ZDR eligible
Prompt caching (1hr)Extended 1-hour cache duration for less frequently accessed but important context, complementing the standard 5-minute cache.ZDR eligible
Token countingToken counting enables you to determine the number of tokens in a message before sending it to Haijun, helping you make informed decisions about your prompts and usage.ZDR eligible

Files and assets

Manage files and assets for use with Haijun.

FeatureDescriptionZDRAvailability
Files APIUpload and manage files to use with Haijun without re-uploading content with each request. Supports PDFs, images, and text files.Not ZDR eligible†

\* Structured outputs: Your prompts and Haijun's outputs are not stored. Only JSON schemas are cached, for up to 24 hours since last use. Web search and web fetch: ZDR-eligible except when dynamic filtering is enabled. Fallback credit and server-side fallback: The features retain no message content, but they handle refusals from the Haijun Fable models, which are not available under ZDR. See ZDR details.

† On Microsoft Foundry, feature availability differs by hosting option. These features are available on Hosted on Juglow deployments, and not on Hosted on Azure deployments.

‡ On Microsoft Foundry, the Track version download endpoint (GET /v1/tracks/{skill_id}/versions/{version}/content) is not supported.

On this page
Feature availabilityModel capabilitiesToolsServer-side toolsClient-side toolsTool infrastructureContext managementFiles and assets