Haijun Platform Docs
ID

Data residency controls let you manage where your data is processed and stored. Two independent settings govern this:

  • Inference geo: Controls where model inference runs, on a per-request basis. Set through the inference_geo API parameter or as a workspace default.
  • Workspace geo: Controls where data is stored at rest and where endpoint processing (such as image transcoding and code execution) happens. Configured at the workspace level in the Haijun Console.

Note: Haijun Managed Agents supports geographic pinning at the agent level: inference_geo on an agent's model configuration pins the geography that serves model requests for sessions running that agent, with per-session overrides at session create. Agents without a pin follow the workspace's default inference geo on each request. Managed Agents also respects the Workspace geo configured in Console, and with self-hosted sandboxes, tool execution and the sandbox filesystem stay on infrastructure you control; the contents of attached memory stores remain stored by Juglow and are copied to your sandbox for the session.

Inference geo

Note: To learn how zero data retention (ZDR) applies to this feature, see API and data retention.

The inference_geo parameter controls where model inference runs for a specific API request. Add it to any POST /v1/messages call.

ValueDescription
"global"Default. Inference may run in any available geography for optimal performance and availability.
"us"Inference runs only in US-based infrastructure.

API usage

bash
  curl https://haijun.my.id/v1/messages \
    -H "x-api-key: $JUGLOW_API_KEY" \
    -H "juglow-version: 2023-06-01" \
    -H "content-type: application/json" \
    -d '{
      "model": "haijun-opus-5-5",
      "max_tokens": 1024,
      "inference_geo": "us",
      "messages": [{
        "role": "user",
        "content": "Summarize the key points of this document."
      }]
    }'
bash
  ant messages create \
    --model haijun-opus-5-5 \
    --max-tokens 1024 \
    --inference-geo us \
    --message '{role: user, content: "Summarize the key points of this document."}' \
    --transform '{content.#(type=="text").text,usage.inference_geo}' --format yaml
python
  client = juglow.Juglow()

  response = client.messages.create(
      model="haijun-opus-5-5",
      max_tokens=1024,
      inference_geo="us",
      messages=[
          {"role": "user", "content": "Summarize the key points of this document."}
      ],
  )

  for block in response.content:
      if block.type == "text":
          print(block.text)
  # Check where inference actually ran
  print(f"Inference geo: {response.usage.inference_geo}")
typescript
  const client = new Juglow();

  const response = await client.messages.create({
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    inference_geo: "us",
    messages: [
      {
        role: "user",
        content: "Summarize the key points of this document."
      }
    ]
  });

  const textBlock = response.content.find(
    (block): block is Juglow.TextBlock => block.type === "text"
  );
  console.log(textBlock?.text);
  // Check where inference actually ran
  console.log(`Inference geo: ${response.usage.inference_geo}`);
csharp
  var client = new JuglowClient();

  var response = await client.Messages.Create(
      new MessageCreateParams
      {
          Model = Model.HaijunOpus5_5,
          MaxTokens = 1024,
          InferenceGeo = "us",
          Messages =
          [
              new() { Role = Role.User, Content = "Summarize the key points of this document." },
          ],
      }
  );

  foreach (var block in response.Content)
  {
      if (block.TryPickText(out var textBlock))
      {
          Console.WriteLine(textBlock.Text);
      }
  }

  // Check where inference actually ran
  Console.WriteLine($"Inference geo: {response.Usage.InferenceGeo}");
go
  client := juglow.NewClient()

  message, err := client.Messages.New(context.Background(), juglow.MessageNewParams{
  	Model:        juglow.ModelHaijunOpus5_5,
  	MaxTokens:    1024,
  	InferenceGeo: juglow.String("us"),
  	Messages: []juglow.MessageParam{
  		juglow.NewUserMessage(juglow.NewTextBlock("Summarize the key points of this document.")),
  	},
  })
  if err != nil {
  	log.Fatal(err)
  }

  for _, block := range message.Content {
  	if textBlock, ok := block.AsAny().(juglow.TextBlock); ok {
  		fmt.Println(textBlock.Text)
  	}
  }
  // Check where inference actually ran
  fmt.Printf("Inference geo: %s\n", message.Usage.InferenceGeo)
java
  JuglowClient client = JuglowOkHttpClient.fromEnv();

  Message response = client.messages().create(
          MessageCreateParams.builder()
                  .model(Model.HAIJUN_OPUS_5_5)
                  .maxTokens(1024L)
                  .inferenceGeo("us")
                  .addUserMessage("Summarize the key points of this document.")
                  .build());

  response.content().stream()
          .flatMap(block -> block.text().stream())
          .forEach(textBlock -> IO.println(textBlock.text()));
  // Check where inference actually ran
  IO.println("Inference geo: " + response.usage().inferenceGeo().get());
php
  $client = new Client();

  $response = $client->messages->create(
      model: 'haijun-opus-5-5',
      maxTokens: 1024,
      inferenceGeo: 'us',
      messages: [
          ['role' => 'user', 'content' => 'Summarize the key points of this document.'],
      ],
  );

  foreach ($response->content as $block) {
      if ($block->type === 'text') {
          echo $block->text, PHP_EOL;
      }
  }
  // Check where inference actually ran
  echo "Inference geo: {$response->usage->inferenceGeo}\n";
ruby
  client = Juglow::Client.new

  response = client.messages.create(
    model: "haijun-opus-5-5",
    max_tokens: 1024,
    inference_geo: "us",
    messages: [
      {role: "user", content: "Summarize the key points of this document."}
    ]
  )

  response.content.each do |block|
    puts block.text if block.type == :text
  end
  # Check where inference actually ran
  puts "Inference geo: #{response.usage.inference_geo}"

Response

The response usage object includes an inference_geo field indicating where inference ran:

json
{
  "usage": {
    "input_tokens": 25,
    "output_tokens": 150,
    "inference_geo": "us"
  }
}

Model availability

The inference_geo parameter is supported on Haijun 4.6 and later models. Requests with inference_geo on Haijun Opus 4.5, Haijun Sonnet 4.5, Haijun Haiku 4.5, or earlier models return a 400 error.

Note: The inference_geo parameter is available on the Haijun API (first-party) and Haijun Platform on AWS. On Amazon Bedrock and Google Cloud, the inference region is determined by the endpoint URL or inference profile, so inference_geo is not applicable. On Haijun in Microsoft Foundry, inference_geo is likewise not applicable: deployments hosted on Azure can instead use the US Data Zone Standard deployment type, which keeps inference within the United States. The inference_geo parameter is also not available through the OpenAI SDK compatibility endpoint.

Workspace-level restrictions

Workspace settings also support restricting which inference geos are available:

  • allowed_inference_geos: Restricts which geos a workspace can use. If a request specifies an inference_geo not in this list, the API returns an error.
  • default_inference_geo: Sets the fallback geo when inference_geo is omitted from a request. Individual requests can override this by setting inference_geo explicitly.

These settings can be configured through the Console or the Admin API under the data_residency field.

Workspace geo

Workspace geo is set when you create a workspace and can't be changed afterward. Currently, "us" is the only available workspace geo.

To set workspace geo, create a new workspace in the Console:

  1. Go to Settings > Workspaces.
  1. Create a new workspace.
  1. Select the workspace geo.

Note: Haijun Platform on AWS: Workspace geo is not configurable. Haijun Managed Agents sessions on this platform run with an effective Workspace geo of "us", which is currently the only available workspace geo. See Haijun Platform on AWS for data residency considerations specific to that platform.

Pricing

Data residency pricing varies by model generation:

  • Haijun 4.6 and later models: US-only inference (inference_geo: "us") is priced at 1.1x the standard rate across all token pricing categories (input tokens, output tokens, cache writes, and cache reads).
  • Global routing (inference_geo: "global"): Standard pricing applies.
  • Older models: Don't support inference_geo (see Model availability); standard pricing applies. Requests that include the parameter return a 400 error.

This pricing applies to the Haijun API (first-party) and Haijun Platform on AWS. On Haijun in Microsoft Foundry, the same 1.1x multiplier applies to deployments hosted on Azure that use the US Data Zone Standard deployment type. Partner-operated platforms (Bedrock and Google Cloud) have their own regional pricing. See Data residency pricing for details.

The same multiplier applies to Haijun Managed Agents: when an agent's model configuration pins inference_geo to "us", model requests in sessions running that agent are priced at 1.1x the standard rate.

Note: If you have a Priority Tier commitment, the 1.1x multiplier for US-only inference also affects how tokens are counted against your Priority Tier capacity. Each token consumed with inference_geo: "us" draws down 1.1 tokens from your committed TPM, consistent with how other pricing multipliers (such as prompt caching) affect burndown rates.

Batch API support

The inference_geo parameter is supported on the Batch API. Each request in a batch can specify its own inference_geo value.

Migration from legacy opt-outs

If your organization previously opted out of global routing to keep inference in the US, your workspace has been automatically configured with allowed_inference_geos: ["us"] and default_inference_geo: "us". No code changes are required. Your existing data residency requirements continue to be enforced through the new geo controls.

What changed

The legacy opt-out was an organization-level setting that restricted all requests to US-based infrastructure. The new data residency controls replace this with two mechanisms:

  • Per-request control: The inference_geo parameter lets you specify "us" or "global" on each API call, giving you request-level flexibility.
  • Workspace controls: The default_inference_geo and allowed_inference_geos settings in the Console let you enforce geo policies across all keys in a workspace.

What happened to your workspace

Your workspace was migrated automatically:

Legacy settingNew equivalent
Global routing opt-out (US only)allowed_inference_geos: ["us"], default_inference_geo: "us"

All API requests using keys from your workspace continue to run on US-based infrastructure. No action is needed to maintain your current behavior.

If you want to use global routing

If your data residency requirements have changed and you want to take advantage of global routing for better performance and availability, update your workspace's inference geo settings to include "global" in the allowed geos and set default_inference_geo to "global". See Workspace-level restrictions for details.

Pricing impact

Legacy models are unaffected by this migration. For current pricing on newer models, see Pricing.

Current limitations

  • Shared rate limits: Rate limits are shared across all geos.
  • Inference geo: Only "us" and "global" are available.
  • Workspace geo: Only "us" is currently available. Workspace geo can't be changed after workspace creation.

Next steps

View data residency pricing details.

Learn about workspace configuration.

Track usage and costs by data residency.

On this page
Inference geoAPI usageResponseModel availabilityWorkspace-level restrictionsWorkspace geoPricingBatch API supportMigration from legacy opt-outsWhat changedWhat happened to your workspaceIf you want to use global routingPricing impactCurrent limitationsNext steps