Long-running conversations with Haijun can exceed context limits, causing loss of important information. Whether you're building a coding assistant, creative writing tool, or customer service agent, managing session memory is critical for maintaining continuity and quality.
This cookbook teaches you how to proactively manage session memory to avoid jarring context limit interruptions. Unlike reactive approaches that wait until the context is full, you'll learn to build session memory in the background so compaction is instant when needed.
Learning Objectives
By the end of this cookbook, you will be able to:
Required Knowledge
- Basic understanding of Haijun API usage and message formatting
- Familiarity with Python threading concepts (helpful but not required)
Required Tools
xt.strip().split("\n")
if len(lines) <= max_lines:
return text
return "\n".join(lines[:max_lines]) + f"\n... ({len(lines) - max_lines} more lines)"
def remove_thinking_blocks(text: str) -> tuple[str, str]:
"""Remove
import re
matches = re.findall(r"
cleaned = re.sub(r"
return cleaned, "".join(matches)
def add_cache_control(messages: list[dict]) -> list[MessageParam]:
"""Add cache_control to the last user message for prompt caching.
For prompt caching to work, the message prefix structure must be identical between requests.
All messages are converted to list format for consistency, and cache_control is placed on
the last user message to match the standard API call pattern.
"""
cached_messages: list[MessageParam] = []
last_user_idx = None
Find last user message index
for i, msg in enumerate(messages):
if msg["role"] == "user":
last_user_idx = i
for i, msg in enumerate(messages):
content = msg["content"]
text = content if isinstance(content, str) else content[0]["text"]
content_block: TextBlockParam = {"type": "text", "text": text}
if i == last_user_idx:
content_block["cache_control"] = {"type": "ephemeral"}
cached_messages.append({"role": msg["role"], "content": [content_block]})
return cached_messages
def estimate_tokens(text: str) -> int:
"""Rudimentary token estimation: 1 token per 4 characters."""
return len(text) // 4
/root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:676: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior return Regex(regex, options) /root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:457: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior result = add_action(grammar, unpack).parseWithTabs().transformString(text) SESSION_MEMORY_PROMPT = """ Compress the conversation into a structured summary that preserves all information needed to continue work seamlessly. Optimize for the assistant's ability to continue working, not human readability.
print(f"{'-' * 60}")
Update token count to reflect compacted state
self.current_context_window_tokens = approximate_summary_tokens
/root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:403: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior grammar.streamline() /root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:457: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior result = add_action(grammar, unpack).parseWithTabs().transformString(text) Below we simulate a conversation between an author and an LLM that helps write stories.
luded in the session memory
)
self.tokens_at_last_update = 0 # To track tokens at last memory update and see if enough new tokens have been added to trigger another update
Background update tracking
self._update_thread: threading.Thread | None = None
self.last_update_time = None
self._lock = threading.Lock()
def chat(self, user_message: str) -> tuple[str, juglow.types.Usage, str | None]:
"""Process a chat turn with background session memory updates."""
if self.current_context_window_tokens + estimate_tokens(user_message) >= self.context_limit:
self.compact() # note that when this is triggered, the compaction has already been created and is just swapped in instantly
self.messages.append({"role": "user", "content": user_message})
response = client.messages.create(
model=MODEL,
max_tokens=3500,
system=self.system_message,
messages=add_cache_control(self.messages),
)
assistant_message = response.content[0].text
self.messages.append({"role": "assistant", "content": assistant_message})
Calculate token usage including cache
cache_read = getattr(response.usage, "cache_read_input_tokens", 0) or 0
total_input = response.usage.input_tokens + cache_read
Update context window tokens (includes cached tokens since they still count toward context)
self.current_context_window_tokens = total_input + response.usage.output_tokens
KEY DIFFERENCE: Trigger background memory update if needed proactively, before compaction is needed
background_status = None
if self._should_init_memory() or self._should_update_memory():
self._trigger_background_update()
background_status = "initializing" if self.session_memory is None else "updating"
Return usage info with cache stats
return assistant_message, response.usage, background_status
Helper methods to determine when to init session memory
def _should_init_memory(self) -> bool:
return (
self.session_memory is None
and self.current_context_window_tokens >= self.min_tokens_to_init
)
Helper method to determine if memory should be updated
def _should_update_memory(self) -> bool:
if self.session_memory is None:
return False
tokens_since = self.current_context_window_tokens - self.tokens_at_last_update
return tokens_since >= self.min_tokens_between_updates
Methods to create initial session memory
def _create_session_memory(self, messages: list[dict]) -> str:
"""Generate initial session memory from messages."""
Put compaction instructions in user message to share cache with main chat
compaction_messages = [{"role": "user", "content": SESSION_MEMORY_PROMPT}]
response = client.messages.create(
model=MODEL,
max_tokens=5000,
system=self.system_message, # Same as main chat for cache sharing
messages=add_cache_control(messages) + compaction_messages,
)
summary, _ = remove_thinking_blocks(
response.content[0].text
) # clean up any
print(
f" [Background] Initial session memory created. Cache hit={getattr(response.usage, 'cache_read_input_tokens', 0) > 0}"
)
return summary
def _update_session_memory(self, new_messages: list[dict]) -> str:
"""Update existing session memory with new messages. In practice, you may want to do this via file edit rather than full re-generation. But for demo purposes we do full regeneration here."""
Put compaction instructions in user message to share cache with main chat
compaction_update_messages = [
{
"role": "user",
"content": SESSION_MEMORY_PROMPT
+ f"""There is an existing session memory: {self.session_memory}. Return the entire session memory with updates to reflect new messages.""",
}
]
response = client.messages.create(
model=MODEL,
max_tokens=5000,
system=self.system_message,
messages=new_messages
+ compaction_update_messages, # you may want to use prompt caching instead, in which case you'd use add_cache_control(self.messages) here
)
updated_summary, _ = remove_thinking_blocks(
response.content[0].text
) # clean up any
print(" [Background] Session memory updated.")
return updated_summary
Background memory update methods
def _background_memory_update(
self, messages_snapshot: list[dict], snapshot_index: int, current_tokens: int
) -> None:
"""Run session memory update in a background thread."""
try:
with self._lock:
current_session_memory = self.session_memory
last_index = self.last_summarized_index
if current_session_memory is None:
new_memory = self._create_session_memory(messages_snapshot)
else:
Get new messages since last summary
new_messages = messages_snapshot[last_index:]
if not new_messages:
return
new_memory = self._update_session_memory(new_messages)
Update state (thread-safe)
with self._lock:
self.session_memory = new_memory
self.last_summarized_index = snapshot_index
self.tokens_at_last_update = current_tokens
self.last_update_time = time.time()
except Exception as e:
print(f" [Background] Error updating memory: {e}")
This makes sure only one background update runs at a time. If one is already running, we skip starting another. If not, we start a new thread to do the update.
def _trigger_background_update(self):
"""Trigger a background session memory update."""
if self._update_thread is not None and self._update_thread.is_alive():
return
messages_snapshot = self.messages.copy()
snapshot_index = len(messages_snapshot)
current_tokens = self.current_context_window_tokens
self._update_thread = threading.Thread(
target=self._background_memory_update,
args=(messages_snapshot, snapshot_index, current_tokens),
daemon=True,
)
self._update_thread.start()
Function to compact
def compact(self) -> None:
"""INSTANT compaction using pre-built session memory."""
prev_msg_count = len(self.messages)
Ensure session memory is ready. Shouldn't be an issue normally, but here for safety.
if self.session_memory is None:
if self._update_thread is not None and self._update_thread.is_alive():
print(" ⏳ Waiting for background memory update...")
self._update_thread.join(timeout=30.0)
if self.session_memory is None:
print(" ⚠️ No pre-built memory, creating synchronously...")
start = time.perf_counter()
self.session_memory = self._create_session_memory(self.messages)
elapsed = time.perf_counter() - start
print(f" ⏱️ Took {elapsed:.2f}s (but should be instant normally!)")
self.last_summarized_index = len(self.messages)
with self._lock:
unsummarized = self.messages[self.last_summarized_index :]
summary_message = [
{
"role": "user",
"content": f"""This session is being continued from a previous conversation. Here is the session memory: {self.session_memory}.Continue from where we left off.""",
}
]
self.messages = summary_message + unsummarized
self.last_summarized_index = 1
print(f"\n{'=' * 60}")
print(f"⚡ INSTANT COMPACTION! Messages: {prev_msg_count} → {len(self.messages)}")
print(" Session memory was pre-built (no wait time!)")
print(f"{'=' * 60}")
/root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:403: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior grammar.streamline() /root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:457: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior result = add_action(grammar, unpack).parseWithTabs().transformString(text) Example use of Instant Compaction # Low thresholds for demo - in production you'd use higher values session = InstantCompactingChatSession( system_message=SYSTEM_PROMPT, ) messages = [ "I want to create a story about a young detective solving a mysterious case in a small town. Generate 3 well thought out plot ideas for me to consider.", "I don't like those ideas, can you think of one plot something more unique and unexpected?", "Ok I like it. Can you help me develop the main character's backstory and motivations?", "Can you draft a detailed outline for the story, breaking it down into chapters and key events?", "Can you draft me a first chapter based on the plot and character ideas we've discussed so far? Make it around 2,000 words.", "Can you draft a second chapter that builds on the first one?", ] print("Starting conversation with instant compacting chat session...\n") turn_count = 0 for message in messages: response, usage, background_status = session.chat(message) turn_count += 1 # Calculate cache stats cache_read = getattr(usage, "cache_read_input_tokens", 0) or 0 cache_created = getattr(usage, "cache_creation_input_tokens", 0) or 0 total_input = usage.input_tokens + cache_read print(f"{'=' * 60}") print(f"Turn {turn_count}:") print(f"\nUser: {message}") print(f"\nAssistant: \n{truncate_response(response, max_lines=3)}") print("\nToken Usage:") print(f" Input: {total_input:,} (new: {usage.input_tokens:,}, cached: {cache_read:,})") print(f" Output: {usage.output_tokens:,}") print( f" Messages: {len(session.messages)} | Memory: {'ready' if session.session_memory else 'not yet'}" ) if cache_read > 0: cache_pct = (cache_read / total_input) * 100 print(f" ✓ Cache hit! {cache_pct:.0f}% of input from cache") if background_status: print(f"\n [Background] Proactively {background_status} session memory...") print(f" Context window: {session.current_context_window_tokens:,} tokens") print() Starting conversation with instant compacting chat session... ============================================================ Turn 1: User: I want to create a story about a young detective solving a mysterious case in a small town. Generate 3 well thought out plot ideas for me to consider. Assistant: # Three Mystery Plot Ideas ## 1. The Vanishing Choir ... (36 more lines) Token Usage: Input: 317 (new: 317, cached: 0) Output: 902 Messages: 2 | Memory: not yet ============================================================ Turn 2: User: I don't like those ideas, can you think of one plot something more unique and unexpected? Assistant: # The Forgetting House Setup: Your young detective arrives in Ember Falls to investigate a string of burglaries—except the victims don't realize they've been robbed until weeks later. A woman discovers her wedding ring gone and insists she lost it yesterday, but security footage shows she hasn't worn it in a month. A man reports his grandfather's watch stolen, then his sister shows him photos proving he sold it himself at a pawn shop—which he has no memory of doing. ... (16 more lines) Token Usage: Input: 1,241 (new: 1,241, cached: 0) Output: 592 Messages: 4 | Memory: not yet ============================================================ Turn 3: User: Ok I like it. Can you help me develop the main character's backstory and motivations? Assistant: # Your Detective: Building From The Inside Out ## Core Identity ... (79 more lines) Token Usage: Input: 1,856 (new: 1,856, cached: 0) Output: 1,329 Messages: 6 | Memory: not yet ============================================================ Turn 4: User: Can you draft a detailed outline for the story, breaking it down into chapters and key events? Assistant: # The Forgetting House: Chapter Outline --- ... (272 more lines) Token Usage: Input: 3,207 (new: 3,207, cached: 0) Output: 3,500 Messages: 8 | Memory: not yet ============================================================ Turn 5: User: Can you draft me a first chapter based on the plot and character ideas we've discussed so far? Make it around 2,000 words. Assistant: # Chapter One: The Impossible Theft The apartment smelled like burnt coffee and old paper. ... (196 more lines) Token Usage: Input: 6,743 (new: 6,743, cached: 0) Output: 3,155 Messages: 10 | Memory: not yet [Background] Proactively initializing session memory... Context window: 9,898 tokens [Background] Initial session memory created. Cache hit=True ============================================================ Turn 6: User: Can you draft a second chapter that builds on the first one? Assistant: # Chapter Two: Rosemont Manor The house appeared through the trees like something from a postcard. ... (190 more lines) Token Usage: Input: 9,914 (new: 5,818, cached: 4,096) Output: 3,500 Messages: 12 | Memory: ready ✓ Cache hit! 41% of input from cache [Background] Proactively updating session memory... Context window: 13,414 tokens message = "What did we just talk about? Give me one sentence" response, usage, background_status = session.chat(message) # Calculate cache stats cache_read = getattr(usage, "cache_read_input_tokens", 0) or 0 total_input = usage.input_tokens + cache_read print(f"\nUser: {message}") print(f"\nAssistant: \n{truncate_response(response, max_lines=3)}") print("\nToken Usage:") print(f" Input: {total_input:,} (new: {usage.input_tokens:,}, cached: {cache_read:,})") print(f" Output: {usage.output_tokens:,}") print( f" Messages: {len(session.messages)} | Memory: {'ready' if session.session_memory else 'not yet'}" ) if cache_read > 0: cache_pct = (cache_read / total_input) * 100 print(f" ✓ Cache hit! {cache_pct:.0f}% of input from cache") ============================================================ ⚡ INSTANT COMPACTION! Messages: 12 → 3 Session memory was pre-built (no wait time!) ============================================================ User: What did we just talk about? Give me one sentence Assistant: I drafted Chapter 2 where Casey arrives at Rosemont Manor, interviews Iris (who deflects questions about her past and shows moments of disorientation), and realizes through comparing photos that Iris Hale is definitely their missing grandmother Iris Whitmore. Token Usage: Input: 5,490 (new: 5,490, cached: 0) Output: 60 Messages: 5 | Memory: ready You'll notice here that once we hit the context limit, the session memory was instantaly swapped in, meaning the user had zero waiting time for a response!