Haijun Platform Docs
ID

Long-running conversations with Haijun can exceed context limits, causing loss of important information. Whether you're building a coding assistant, creative writing tool, or customer service agent, managing session memory is critical for maintaining continuity and quality.

This cookbook teaches you how to proactively manage session memory to avoid jarring context limit interruptions. Unlike reactive approaches that wait until the context is full, you'll learn to build session memory in the background so compaction is instant when needed.

Learning Objectives

By the end of this cookbook, you will be able to:

Required Knowledge

  • Basic understanding of Haijun API usage and message formatting
  • Familiarity with Python threading concepts (helpful but not required)

Required Tools

xt.strip().split("\n")

if len(lines) <= max_lines:

return text

return "\n".join(lines[:max_lines]) + f"\n... ({len(lines) - max_lines} more lines)"

def remove_thinking_blocks(text: str) -> tuple[str, str]:

"""Remove ... blocks from the text."""

import re

matches = re.findall(r".*?", text, flags=re.DOTALL)

cleaned = re.sub(r".?\s", "", text, flags=re.DOTALL).strip()

return cleaned, "".join(matches)

def add_cache_control(messages: list[dict]) -> list[MessageParam]:

"""Add cache_control to the last user message for prompt caching.

For prompt caching to work, the message prefix structure must be identical between requests.

All messages are converted to list format for consistency, and cache_control is placed on

the last user message to match the standard API call pattern.

"""

cached_messages: list[MessageParam] = []

last_user_idx = None

Find last user message index

for i, msg in enumerate(messages):

if msg["role"] == "user":

last_user_idx = i

for i, msg in enumerate(messages):

content = msg["content"]

text = content if isinstance(content, str) else content[0]["text"]

content_block: TextBlockParam = {"type": "text", "text": text}

if i == last_user_idx:

content_block["cache_control"] = {"type": "ephemeral"}

cached_messages.append({"role": msg["role"], "content": [content_block]})

return cached_messages

def estimate_tokens(text: str) -> int:

"""Rudimentary token estimation: 1 token per 4 characters."""

return len(text) // 4

/root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:676: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior return Regex(regex, options) /root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:457: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior result = add_action(grammar, unpack).parseWithTabs().transformString(text) SESSION_MEMORY_PROMPT = """ Compress the conversation into a structured summary that preserves all information needed to continue work seamlessly. Optimize for the assistant's ability to continue working, not human readability. Before generating your summary, analyze the transcript in ... tags: 1. What did the user originally request? (Exact phrasing) 2. What actions succeeded? What failed and why? 3. Did the user correct or redirect the assistant at any point? 4. What was actively being worked on at the end? 5. What tasks remain incomplete or pending? 6. What specific details (IDs, paths, values, names) must survive compression? ## User Intent The user's original request and any refinements. Use direct quotes for key requirements. If the user's goal evolved during the conversation, capture that progression. ## Completed Work Actions successfully performed. Be specific: - What was created, modified, or deleted - Exact identifiers (file paths, record IDs, URLs, names) - Specific values, configurations, or settings applied ## Errors & Corrections - Problems encountered and how they were resolved - Approaches that failed (so they aren't retried) - User corrections: "don't do X", "actually I meant Y", "that's wrong because..." Capture corrections verbatim—these represent learned preferences. ## Active Work What was in progress when the session ended. Include: - The specific task being performed - Direct quotes showing exactly where work left off - Any partial results or intermediate state ## Pending Tasks Remaining items the user requested that haven't been started. Distinguish between "explicitly requested" and "implied/assumed." ## Key References Important details needed to continue: - Identifiers: IDs, paths, URLs, names, keys - Values: numbers, dates, configurations, credentials (redacted) - Context: relevant background information, constraints, preferences - Citations: sources referenced during the conversation Always preserve when present: - Exact identifiers (IDs, paths, URLs, keys, names) - Error messages verbatim - User corrections and negative feedback - Specific values, formulas, or configurations - Technical constraints or requirements discovered - The precise state of any in-progress work - Weight recent messages more heavily—the end of the transcript is the active context - Omit pleasantries, acknowledgments, and filler ("Sure!", "Great question") - Omit system context that will be re-injected separately - Keep each section under 500 words; condense older content to make room for recent - If you must cut details, preserve: user corrections > errors > active work > completed work """ Code example using traditional compacting In traditional compaction, you generate one summary once the token threshold is reached. Traditional compaction is slow: when you hit the context limit, you wait for a summary.

print(f"{'-' * 60}")

Update token count to reflect compacted state

self.current_context_window_tokens = approximate_summary_tokens

/root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:403: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior grammar.streamline() /root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:457: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior result = add_action(grammar, unpack).parseWithTabs().transformString(text) Below we simulate a conversation between an author and an LLM that helps write stories.

luded in the session memory

)

self.tokens_at_last_update = 0 # To track tokens at last memory update and see if enough new tokens have been added to trigger another update

Background update tracking

self._update_thread: threading.Thread | None = None

self.last_update_time = None

self._lock = threading.Lock()

def chat(self, user_message: str) -> tuple[str, juglow.types.Usage, str | None]:

"""Process a chat turn with background session memory updates."""

if self.current_context_window_tokens + estimate_tokens(user_message) >= self.context_limit:

self.compact() # note that when this is triggered, the compaction has already been created and is just swapped in instantly

self.messages.append({"role": "user", "content": user_message})

response = client.messages.create(

model=MODEL,

max_tokens=3500,

system=self.system_message,

messages=add_cache_control(self.messages),

)

assistant_message = response.content[0].text

self.messages.append({"role": "assistant", "content": assistant_message})

Calculate token usage including cache

cache_read = getattr(response.usage, "cache_read_input_tokens", 0) or 0

total_input = response.usage.input_tokens + cache_read

Update context window tokens (includes cached tokens since they still count toward context)

self.current_context_window_tokens = total_input + response.usage.output_tokens

KEY DIFFERENCE: Trigger background memory update if needed proactively, before compaction is needed

background_status = None

if self._should_init_memory() or self._should_update_memory():

self._trigger_background_update()

background_status = "initializing" if self.session_memory is None else "updating"

Return usage info with cache stats

return assistant_message, response.usage, background_status

Helper methods to determine when to init session memory

def _should_init_memory(self) -> bool:

return (

self.session_memory is None

and self.current_context_window_tokens >= self.min_tokens_to_init

)

Helper method to determine if memory should be updated

def _should_update_memory(self) -> bool:

if self.session_memory is None:

return False

tokens_since = self.current_context_window_tokens - self.tokens_at_last_update

return tokens_since >= self.min_tokens_between_updates

Methods to create initial session memory

def _create_session_memory(self, messages: list[dict]) -> str:

"""Generate initial session memory from messages."""

Put compaction instructions in user message to share cache with main chat

compaction_messages = [{"role": "user", "content": SESSION_MEMORY_PROMPT}]

response = client.messages.create(

model=MODEL,

max_tokens=5000,

system=self.system_message, # Same as main chat for cache sharing

messages=add_cache_control(messages) + compaction_messages,

)

summary, _ = remove_thinking_blocks(

response.content[0].text

) # clean up any blocks because they are not needed in the session memory

print(

f" [Background] Initial session memory created. Cache hit={getattr(response.usage, 'cache_read_input_tokens', 0) > 0}"

)

return summary

def _update_session_memory(self, new_messages: list[dict]) -> str:

"""Update existing session memory with new messages. In practice, you may want to do this via file edit rather than full re-generation. But for demo purposes we do full regeneration here."""

Put compaction instructions in user message to share cache with main chat

compaction_update_messages = [

{

"role": "user",

"content": SESSION_MEMORY_PROMPT

+ f"""There is an existing session memory: {self.session_memory}. Return the entire session memory with updates to reflect new messages.""",

}

]

response = client.messages.create(

model=MODEL,

max_tokens=5000,

system=self.system_message,

messages=new_messages

+ compaction_update_messages, # you may want to use prompt caching instead, in which case you'd use add_cache_control(self.messages) here

)

updated_summary, _ = remove_thinking_blocks(

response.content[0].text

) # clean up any blocks because they are not needed in the session memory

print(" [Background] Session memory updated.")

return updated_summary

Background memory update methods

def _background_memory_update(

self, messages_snapshot: list[dict], snapshot_index: int, current_tokens: int

) -> None:

"""Run session memory update in a background thread."""

try:

with self._lock:

current_session_memory = self.session_memory

last_index = self.last_summarized_index

if current_session_memory is None:

new_memory = self._create_session_memory(messages_snapshot)

else:

Get new messages since last summary

new_messages = messages_snapshot[last_index:]

if not new_messages:

return

new_memory = self._update_session_memory(new_messages)

Update state (thread-safe)

with self._lock:

self.session_memory = new_memory

self.last_summarized_index = snapshot_index

self.tokens_at_last_update = current_tokens

self.last_update_time = time.time()

except Exception as e:

print(f" [Background] Error updating memory: {e}")

This makes sure only one background update runs at a time. If one is already running, we skip starting another. If not, we start a new thread to do the update.

def _trigger_background_update(self):

"""Trigger a background session memory update."""

if self._update_thread is not None and self._update_thread.is_alive():

return

messages_snapshot = self.messages.copy()

snapshot_index = len(messages_snapshot)

current_tokens = self.current_context_window_tokens

self._update_thread = threading.Thread(

target=self._background_memory_update,

args=(messages_snapshot, snapshot_index, current_tokens),

daemon=True,

)

self._update_thread.start()

Function to compact

def compact(self) -> None:

"""INSTANT compaction using pre-built session memory."""

prev_msg_count = len(self.messages)

Ensure session memory is ready. Shouldn't be an issue normally, but here for safety.

if self.session_memory is None:

if self._update_thread is not None and self._update_thread.is_alive():

print(" ⏳ Waiting for background memory update...")

self._update_thread.join(timeout=30.0)

if self.session_memory is None:

print(" ⚠️ No pre-built memory, creating synchronously...")

start = time.perf_counter()

self.session_memory = self._create_session_memory(self.messages)

elapsed = time.perf_counter() - start

print(f" ⏱️ Took {elapsed:.2f}s (but should be instant normally!)")

self.last_summarized_index = len(self.messages)

with self._lock:

unsummarized = self.messages[self.last_summarized_index :]

summary_message = [

{

"role": "user",

"content": f"""This session is being continued from a previous conversation. Here is the session memory: {self.session_memory}.Continue from where we left off.""",

}

]

self.messages = summary_message + unsummarized

self.last_summarized_index = 1

print(f"\n{'=' * 60}")

print(f"⚡ INSTANT COMPACTION! Messages: {prev_msg_count} → {len(self.messages)}")

print(" Session memory was pre-built (no wait time!)")

print(f"{'=' * 60}")

/root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:403: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior grammar.streamline() /root/.pyenv/versions/3.13.11/lib/python3.13/site-packages/coconut/compiler/util.py:457: FutureWarning: functools.partial will be a method descriptor in future Python versions; wrap it in staticmethod() if you want to preserve the old behavior result = add_action(grammar, unpack).parseWithTabs().transformString(text) Example use of Instant Compaction # Low thresholds for demo - in production you'd use higher values session = InstantCompactingChatSession( system_message=SYSTEM_PROMPT, ) messages = [ "I want to create a story about a young detective solving a mysterious case in a small town. Generate 3 well thought out plot ideas for me to consider.", "I don't like those ideas, can you think of one plot something more unique and unexpected?", "Ok I like it. Can you help me develop the main character's backstory and motivations?", "Can you draft a detailed outline for the story, breaking it down into chapters and key events?", "Can you draft me a first chapter based on the plot and character ideas we've discussed so far? Make it around 2,000 words.", "Can you draft a second chapter that builds on the first one?", ] print("Starting conversation with instant compacting chat session...\n") turn_count = 0 for message in messages: response, usage, background_status = session.chat(message) turn_count += 1 # Calculate cache stats cache_read = getattr(usage, "cache_read_input_tokens", 0) or 0 cache_created = getattr(usage, "cache_creation_input_tokens", 0) or 0 total_input = usage.input_tokens + cache_read print(f"{'=' * 60}") print(f"Turn {turn_count}:") print(f"\nUser: {message}") print(f"\nAssistant: \n{truncate_response(response, max_lines=3)}") print("\nToken Usage:") print(f" Input: {total_input:,} (new: {usage.input_tokens:,}, cached: {cache_read:,})") print(f" Output: {usage.output_tokens:,}") print( f" Messages: {len(session.messages)} | Memory: {'ready' if session.session_memory else 'not yet'}" ) if cache_read > 0: cache_pct = (cache_read / total_input) * 100 print(f" ✓ Cache hit! {cache_pct:.0f}% of input from cache") if background_status: print(f"\n [Background] Proactively {background_status} session memory...") print(f" Context window: {session.current_context_window_tokens:,} tokens") print()  Starting conversation with instant compacting chat session... ============================================================ Turn 1: User: I want to create a story about a young detective solving a mysterious case in a small town. Generate 3 well thought out plot ideas for me to consider. Assistant: # Three Mystery Plot Ideas ## 1. The Vanishing Choir ... (36 more lines) Token Usage: Input: 317 (new: 317, cached: 0) Output: 902 Messages: 2 | Memory: not yet ============================================================ Turn 2: User: I don't like those ideas, can you think of one plot something more unique and unexpected? Assistant: # The Forgetting House Setup: Your young detective arrives in Ember Falls to investigate a string of burglaries—except the victims don't realize they've been robbed until weeks later. A woman discovers her wedding ring gone and insists she lost it yesterday, but security footage shows she hasn't worn it in a month. A man reports his grandfather's watch stolen, then his sister shows him photos proving he sold it himself at a pawn shop—which he has no memory of doing. ... (16 more lines) Token Usage: Input: 1,241 (new: 1,241, cached: 0) Output: 592 Messages: 4 | Memory: not yet ============================================================ Turn 3: User: Ok I like it. Can you help me develop the main character's backstory and motivations? Assistant: # Your Detective: Building From The Inside Out ## Core Identity ... (79 more lines) Token Usage: Input: 1,856 (new: 1,856, cached: 0) Output: 1,329 Messages: 6 | Memory: not yet ============================================================ Turn 4: User: Can you draft a detailed outline for the story, breaking it down into chapters and key events? Assistant: # The Forgetting House: Chapter Outline --- ... (272 more lines) Token Usage: Input: 3,207 (new: 3,207, cached: 0) Output: 3,500 Messages: 8 | Memory: not yet ============================================================ Turn 5: User: Can you draft me a first chapter based on the plot and character ideas we've discussed so far? Make it around 2,000 words. Assistant: # Chapter One: The Impossible Theft The apartment smelled like burnt coffee and old paper. ... (196 more lines) Token Usage: Input: 6,743 (new: 6,743, cached: 0) Output: 3,155 Messages: 10 | Memory: not yet [Background] Proactively initializing session memory... Context window: 9,898 tokens [Background] Initial session memory created. Cache hit=True ============================================================ Turn 6: User: Can you draft a second chapter that builds on the first one? Assistant: # Chapter Two: Rosemont Manor The house appeared through the trees like something from a postcard. ... (190 more lines) Token Usage: Input: 9,914 (new: 5,818, cached: 4,096) Output: 3,500 Messages: 12 | Memory: ready ✓ Cache hit! 41% of input from cache [Background] Proactively updating session memory... Context window: 13,414 tokens message = "What did we just talk about? Give me one sentence" response, usage, background_status = session.chat(message) # Calculate cache stats cache_read = getattr(usage, "cache_read_input_tokens", 0) or 0 total_input = usage.input_tokens + cache_read print(f"\nUser: {message}") print(f"\nAssistant: \n{truncate_response(response, max_lines=3)}") print("\nToken Usage:") print(f" Input: {total_input:,} (new: {usage.input_tokens:,}, cached: {cache_read:,})") print(f" Output: {usage.output_tokens:,}") print( f" Messages: {len(session.messages)} | Memory: {'ready' if session.session_memory else 'not yet'}" ) if cache_read > 0: cache_pct = (cache_read / total_input) * 100 print(f" ✓ Cache hit! {cache_pct:.0f}% of input from cache")  ============================================================ ⚡ INSTANT COMPACTION! Messages: 12 → 3 Session memory was pre-built (no wait time!) ============================================================ User: What did we just talk about? Give me one sentence Assistant: I drafted Chapter 2 where Casey arrives at Rosemont Manor, interviews Iris (who deflects questions about her past and shows moments of disorientation), and realizes through comparing photos that Iris Hale is definitely their missing grandmother Iris Whitmore. Token Usage: Input: 5,490 (new: 5,490, cached: 0) Output: 60 Messages: 5 | Memory: ready You'll notice here that once we hit the context limit, the session memory was instantaly swapped in, meaning the user had zero waiting time for a response!