Haijun Platform Docs
ID

Prompt caching with the Haijun API

Prompt caching lets you store and reuse context within your prompts, reducing latency by >2x and costs by up to 90% for repetitive tasks.

There are two ways to enable prompt caching:

trol

field at the top level of your request. The system automatically manages cache breakpoints for you.

Explicit cache breakpoints

: Place

cache_control

on individual content blocks for fine-grained control over exactly what gets cached.

This cookbook demonstrates both approaches, starting with the simpler automatic method.

import time

import juglow

import requests

from bs4 import BeautifulSoup

from dotenv import load_dotenv

load_dotenv()

client = juglow.Juglow()

MODEL_NAME = "haijun-sonnet-4-6"

Unique prefix to ensure we don't hit a stale cache from a previous run

TIMESTAMP = int(time.time())

Let's fetch the full text of Pride and Prejudice (~187k tokens) to use as our large context.

= response.usage

cache_create = getattr(usage, "cache_creation_input_tokens", 0)

cache_read = getattr(usage, "cache_read_input_tokens", 0)

print(f" Time: {elapsed:.2f}s")

print(f" Input tokens: {usage.input_tokens}")

print(f" Output tokens: {usage.output_tokens}")

if cache_create:

print(f" Cache write tokens: {cache_create}")

if cache_read:

print(f" Cache read tokens: {cache_read}")


Example 1: Automatic caching (single turn)

+ ""

+ "\n\nWhat is the title of this book? Only output the title.",

}

],

)

hit_time = time.time() - start

print(f"Response: {hit_response.content[0].text}")

print_usage(hit_response, hit_time)

print("\n" + "=" * 50)

print("COMPARISON")

print("=" * 50)

print(f"No caching: {baseline_time:.2f}s")

print(f"Cache write: {write_time:.2f}s")

print(f"Cache hit: {hit_time:.2f}s")

print(f"Speedup: {baseline_time / hit_time:.1f}x")

Response: Pride and Prejudice Time: 1.48s Input tokens: 3 Output tokens: 8 Cache read tokens: 187361 ================================================== COMPARISON ================================================== No caching: 4.89s Cache write: 4.28s Cache hit: 1.48s Speedup: 3.3x Example 2: Automatic caching in a multi-turn conversation Automatic caching really shines in multi-turn conversations. The cache breakpoint automatically moves forward as the conversation grows — you don't need to manage any markers yourself.

ed on each hit). A 1-hour TTL is available at 2x base input price.

Pricing:

Cache writes cost 1.25x base input price. Cache reads cost 0.1x base input price.

Breakpoint limit:

Up to 4 explicit breakpoints per request. Automatic caching uses one slot.

For full details, see the prompt caching documentation.

On this page
Example 1: Automatic caching (single turn)