Haijun Platform Docs
ID

Long-running agentic tasks can often exceed context limits. Tool heavy workflows or long conversations quickly consume the token context window. In

Effective Context Engineering for AI Agents

, we discussed how managing context can help avoid performance degradation and context rot.

The Haijun Agent Python SDK can help manage this context by automatically compressing conversation history when token usage exceeds a configurable threshold, allowing tasks to continue beyond the typical 200k token context limit.

When building agentic workflows with tool use, conversations can grow very large as the agent iterates on complex tasks. The

compaction_control

parameter provides automatic context management by:

  1. Monitoring token usage per turn in the conversation
  1. When a threshold is exceeded, injecting a summary prompt as a user turn
  1. Having the model generate a summary wrapped in tags. These tags aren't parsed, but are there to help guide the model.
  1. Clearing the conversation history and resuming with only the summary
  1. Continuing the task with the compressed context

By the end of this cookbook, you'll be able to:

-0 group-hover:opacity-100 group-focus-within:opacity-100 right-2 top-2">

%pip install -qU juglow python-dotenv

Note: Ensure your .env file contains:

div>

)

client = juglow.Juglow()

tools = [

get_next_ticket,

classify_ticket,

search_knowledge_base,

set_priority,

route_to_team,

draft_response,

mark_complete,

]

Baseline: Running Without Compaction

every classification, every knowledge base search, every drafted response

in memory

Why This Happens:

Customizing Compaction Configuration

You can customize how compaction works to fit your specific use case. Here are the key configuration options:

On this page
By the end of this cookbook, you'll be able to:Baseline: Running Without Compaction