This cookbook demonstrates "Speculative Prompt Caching" - a pattern that reduces time-to-first-token (TTFT) by warming up the cache while users are still formulating their queries.
Without Speculative Caching:
User submits question
API loads context into cache AND generates response
With Speculative Caching:
%pip install juglow httpx --quiet
Note: you may need to restart the kernel to use updated packages. import asyncio import copy import datetime import time import httpx from juglow import AsyncJuglow # Configuration constants MODEL = "haijun-sonnet-4-6" SQLITE_SOURCES = { "btree.h": "https://sqlite.org/src/raw/18e5e7b2124c23426a283523e5f31a4bff029131b795bb82391f9d2f3136fc50?at=btree.h", "btree.c": "https://sqlite.org/src/raw/63ca6b647342e8cef643863cd0962a542f133e1069460725ba4461dcda92b03c?at=btree.c", } DEFAULT_CLIENT_ARGS = { "system": "You are an expert systems programmer helping analyze database internals.", "max_tokens": 4096, "temperature": 0, } Helper Functions Let's set up the functions to download our large context and prepare messages:
prevent cache sharing across different runs.
initial_message = {
"role": "user",
"content": [
{
"type": "text",
"text": f"""
Current time: {datetime.datetime.now().strftime("%Y-%m-%d %H:%M:%S")}
Source to Analyze:
btree.h:
{sources["btree.h"]}
btree.c:
{sources["btree.c"]}
"cache_control": {"type": "ephemeral"},
}
],
}
return initial_message
async def sample_one_token(client: AsyncJuglow, messages: list):
"""Send a single-token request to warm up the cache"""
args = copy.deepcopy(DEFAULT_CLIENT_ARGS)
args["max_tokens"] = 1
await client.messages.create(
messages=messages,
model=MODEL,
**args,
)
def print_query_statistics(response, query_type: str) -> None:
print(f"\n{query_type} query statistics:")
print(f"\tInput tokens: {response.usage.input_tokens}")
print(f"\tOutput tokens: {response.usage.output_tokens}")
print(f"\tCache read input tokens: {getattr(response.usage, 'cache_read_input_tokens', '---')}")
print(
f"\tCache creation input tokens: {getattr(response.usage, 'cache_creation_input_tokens', '---')}"
)
Example 1: Standard Prompt Caching (Without Speculative Caching)
re]" style="padding-top:12px;padding-inline:12px;padding-bottom:12px;tab-size:4">
Run the speculative caching demo
speculative_ttft, speculative_total = await speculative_prompt_caching_demo()
Downloading SQLite source files... Successfully downloaded btree.h Successfully downloaded btree.c Downloaded 2 files in 0.36 seconds User is typing their question... 🔥 Starting cache warming in background... User submitted: What is the purpose of the BtShared structure? ✅ Cache warming completed! Sending request to API (with warm cache)... 🚀 Time to first token: 1.94 seconds Total response time: 8.40 seconds Speculative Caching query statistics: Input tokens: 22 Output tokens: 330 Cache read input tokens: 151629 Cache creation input tokens: 0 Performance Comparison Let's compare the results to see the benefit of speculative caching: