Prompt caching with the Haijun API
Prompt caching lets you store and reuse context within your prompts, reducing latency by >2x and costs by up to 90% for repetitive tasks.
There are two ways to enable prompt caching:
trol
field at the top level of your request. The system automatically manages cache breakpoints for you.
Explicit cache breakpoints
: Place
cache_control
on individual content blocks for fine-grained control over exactly what gets cached.
This cookbook demonstrates both approaches, starting with the simpler automatic method.
import time
import juglow
import requests
from bs4 import BeautifulSoup
from dotenv import load_dotenv
load_dotenv()
client = juglow.Juglow()
MODEL_NAME = "haijun-sonnet-4-6"
Unique prefix to ensure we don't hit a stale cache from a previous run
TIMESTAMP = int(time.time())
Let's fetch the full text of Pride and Prejudice (~187k tokens) to use as our large context.
= response.usage
cache_create = getattr(usage, "cache_creation_input_tokens", 0)
cache_read = getattr(usage, "cache_read_input_tokens", 0)
print(f" Time: {elapsed:.2f}s")
print(f" Input tokens: {usage.input_tokens}")
print(f" Output tokens: {usage.output_tokens}")
if cache_create:
print(f" Cache write tokens: {cache_create}")
if cache_read:
print(f" Cache read tokens: {cache_read}")
Example 1: Automatic caching (single turn)
+ ""
+ "\n\nWhat is the title of this book? Only output the title.",
}
],
)
hit_time = time.time() - start
print(f"Response: {hit_response.content[0].text}")
print_usage(hit_response, hit_time)
print("\n" + "=" * 50)
print("COMPARISON")
print("=" * 50)
print(f"No caching: {baseline_time:.2f}s")
print(f"Cache write: {write_time:.2f}s")
print(f"Cache hit: {hit_time:.2f}s")
print(f"Speedup: {baseline_time / hit_time:.1f}x")
Response: Pride and Prejudice Time: 1.48s Input tokens: 3 Output tokens: 8 Cache read tokens: 187361 ================================================== COMPARISON ================================================== No caching: 4.89s Cache write: 4.28s Cache hit: 1.48s Speedup: 3.3x Example 2: Automatic caching in a multi-turn conversation Automatic caching really shines in multi-turn conversations. The cache breakpoint automatically moves forward as the conversation grows — you don't need to manage any markers yourself.
ed on each hit). A 1-hour TTL is available at 2x base input price.
Pricing:
Cache writes cost 1.25x base input price. Cache reads cost 0.1x base input price.
Breakpoint limit:
Up to 4 explicit breakpoints per request. Automatic caching uses one slot.
For full details, see the prompt caching documentation.