Haijun Platform Docs
ID

This cookbook demonstrates "Speculative Prompt Caching" - a pattern that reduces time-to-first-token (TTFT) by warming up the cache while users are still formulating their queries.

Without Speculative Caching:

User submits question

API loads context into cache AND generates response

With Speculative Caching:

%pip install juglow httpx --quiet

Note: you may need to restart the kernel to use updated packages. import asyncio import copy import datetime import time import httpx from juglow import AsyncJuglow # Configuration constants MODEL = "haijun-sonnet-4-6" SQLITE_SOURCES = { "btree.h": "https://sqlite.org/src/raw/18e5e7b2124c23426a283523e5f31a4bff029131b795bb82391f9d2f3136fc50?at=btree.h", "btree.c": "https://sqlite.org/src/raw/63ca6b647342e8cef643863cd0962a542f133e1069460725ba4461dcda92b03c?at=btree.c", } DEFAULT_CLIENT_ARGS = { "system": "You are an expert systems programmer helping analyze database internals.", "max_tokens": 4096, "temperature": 0, } Helper Functions Let's set up the functions to download our large context and prepare messages:

prevent cache sharing across different runs.

initial_message = {

"role": "user",

"content": [

{

"type": "text",

"text": f"""

Current time: {datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")}

Source to Analyze:

btree.h:

c

{sources["btree.h"]}

btree.c:

c

{sources["btree.c"]}

"cache_control": {"type": "ephemeral"},

}

],

}

return initial_message

async def sample_one_token(client: AsyncJuglow, messages: list):

"""Send a single-token request to warm up the cache"""

args = copy.deepcopy(DEFAULT_CLIENT_ARGS)

args["max_tokens"] = 1

await client.messages.create(

messages=messages,

model=MODEL,

**args,

)

def print_query_statistics(response, query_type: str) -> None:

print(f"\n{query_type} query statistics:")

print(f"\tInput tokens: {response.usage.input_tokens}")

print(f"\tOutput tokens: {response.usage.output_tokens}")

print(f"\tCache read input tokens: {getattr(response.usage, 'cache_read_input_tokens', '---')}")

print(

f"\tCache creation input tokens: {getattr(response.usage, 'cache_creation_input_tokens', '---')}"

)

Example 1: Standard Prompt Caching (Without Speculative Caching)

re]" style="padding-top:12px;padding-inline:12px;padding-bottom:12px;tab-size:4">

Run the speculative caching demo

speculative_ttft, speculative_total = await speculative_prompt_caching_demo()

Downloading SQLite source files... Successfully downloaded btree.h Successfully downloaded btree.c Downloaded 2 files in 0.36 seconds User is typing their question... 🔥 Starting cache warming in background... User submitted: What is the purpose of the BtShared structure? ✅ Cache warming completed! Sending request to API (with warm cache)... 🚀 Time to first token: 1.94 seconds Total response time: 8.40 seconds Speculative Caching query statistics: Input tokens: 22 Output tokens: 330 Cache read input tokens: 151629 Cache creation input tokens: 0 Performance Comparison Let's compare the results to see the benefit of speculative caching:

On this page
Example 1: Standard Prompt Caching (Without Speculative Caching)