Context Engineering for AI Agents: Memory vs. Compaction vs. Tool Clearing
Introduction
A common challenge when building long-horizon agents is managing context. Tool results, the model's own reasoning, and user messages all accumulate, and eventually you either hit the token limit or start paying for context that isn't helping anymore. Studies on needle-in-a-haystack style benchmarking have uncovered the concept of context rot: as the number of tokens in the context window increases, the model's ability to accurately recall information from that context decreases. So, even before the hard context limit is reached, the agent may be getting less out of each token.
re several levers for this: subagents that isolate work in their own context, programmatic tool calling that keeps large results out of the window entirely, and others.
This cookbook focuses on three: compaction, tool-result clearing, and memory. All three are effective strategies for context engineering, but since they all operate to make the context window more efficient in different ways, they can be hard to distinguish. Understanding those distinctions is what lets you map each tool to the part of your workload it actually helps with. Alongside other core context management strategies like utilizing subagents, these three are crucial for teams building long-running agents to understand. They also all have first-party API support, so you can adopt them without building orchestration infrastructure.
server-side compaction
,
context editing
(which includes tool-result clearing), and the
memory tool
. This cookbook works through how to think about designing with them: when each one applies, how to configure them, what changes when you use them independently vs. together, and sample use-cases where different combinations make sense.
The examples center on a long-running research agent: one that reads a corpus of documents, takes notes, and builds on its findings across multiple sessions. It's a useful test case because it naturally hits all three problems: bulky document reads (clearing), long analytical conversations (compaction), and knowledge that needs to survive between sessions (memory).
lockquote>
Step 0: Environment Setup
ent, the agent's context grows to hundreds of thousands of tokens mid-task. And since the work spans sessions, even a completed run starts the next session with no memory of what was learned.
The research task
content = CORPUS.get(path)
if content is None:
return f"Error: '{path}' not found. Available: {', '.join(CORPUS.keys())}"
return content
if name == "record_finding":
finding = tool_input.get("finding", "")
notes.append(finding)
return f"Finding #{len(notes)} recorded (session-local)."
return f"Error: unknown tool '{name}'"
── Session result container ─────────────────────────────────────────────
SessionResult = namedtuple(
"SessionResult",
[
"messages", # final message list
"notes", # notes taken this session
"token_trajectory", # list of (turn, total_context_tokens)
"events", # list of dicts describing compaction / clearing events
"tool_counts", # dict of tool_name -> call count
"file_reads", # list of (turn, path) for each read_file call
"hit_limit", # True if the session stopped because it hit the context window
"final_text", # final assistant text (convenience)
],
)
def _format_tool_arg(name: str, tool_input: dict) -> str:
"""Human-readable one-line summary of a tool call's arguments for verbose output."""
if name == "search_files":
return repr(tool_input.get("query", ""))
if name == "read_file":
return tool_input.get("path", "?")
if name == "record_finding":
finding = tool_input.get("finding", "")
preview = finding[:50].replace("\n", " ")
return f'"{preview}{"..." if len(finding) > 50 else ""}"'
if name == "memory":
cmd = tool_input.get("command", "?")
path = tool_input.get("path", tool_input.get("old_path", ""))
return f"{cmd} {path}"
return str(tool_input)[:60]
── Agent loop ───────────────────────────────────────────────────────────
def run_research_session(
initial_prompt: str,
*,
context_management: dict | None = None,
betas: list[str] | None = None,
memory_handler=None,
max_turns: int = 12,
label: str = "session",
verbose: bool = True,
) -> SessionResult:
"""Run the research agent. Catches context-window overflow gracefully."""
tools = list(BASE_TOOLS)
if memory_handler is not None:
tools.append(MEMORY_TOOL_SPEC)
messages: list[dict] = [{"role": "user", "content": initial_prompt}]
notes: list[str] = []
token_trajectory: list[tuple[int, int]] = []
events: list[dict] = []
tool_counts: dict[str, int] = {}
file_reads: list[tuple[int, str]] = [] # (turn, path) for each read_file
hit_limit = False
final_text = ""
if verbose:
print(f"┌─ [{label}]")
for turn in range(1, max_turns + 1):
kwargs: dict = dict(
model=MODEL,
max_tokens=4096,
system=SYSTEM_PROMPT,
tools=tools,
messages=messages,
)
if context_management:
kwargs["context_management"] = context_management
if betas:
kwargs["betas"] = betas
Call the API; catch context-window overflow
try:
if betas:
response = client.beta.messages.create(**kwargs)
else:
response = client.messages.create(**kwargs)
except juglow.BadRequestError as e:
Context window exceeded (or similar input-too-large error)
hit_limit = True
if verbose:
print(f"│ ⚠ CONTEXT WINDOW LIMIT REACHED at turn {turn} (API rejected)")
print(f"│ {str(e)[:200]}")
break
Track TOTAL context size: uncached + cache-read + cache-created.
usage.input_tokens alone excludes cached tokens, which makes the
plot show dips that are just cache hits, not context management.
u = response.usage
total_in = (
u.input_tokens
+ (getattr(u, "cache_read_input_tokens", None) or 0)
+ (getattr(u, "cache_creation_input_tokens", None) or 0)
)
token_trajectory.append((turn, total_in))
Surface context-management events on their own prominent lines
cm = getattr(response, "context_management", None)
if cm is not None and getattr(cm, "applied_edits", None):
for edit in cm.applied_edits:
cleared_uses = getattr(edit, "cleared_tool_uses", None)
cleared_toks = getattr(edit, "cleared_input_tokens", None)
events.append(
{
"turn": turn,
"kind": "clearing",
"cleared_tool_uses": cleared_uses,
"cleared_input_tokens": cleared_toks,
}
)
if verbose:
print(
f"│ ✂ CLEARING (turn {turn}): {cleared_uses or '?'} tool results cleared, "
f"~{cleared_toks:,} tokens freed"
if cleared_toks
else f"│ ✂ CLEARING (turn {turn}): applied"
)
Serialize assistant content, track tool calls and compaction
serialized: list[dict] = []
tool_calls: list[dict] = []
turn_tool_calls: list[tuple[str, dict]] = [] # (name, input) for verbose display
compaction_this_turn = False
for block in response.content:
if block.type == "text":
serialized.append({"type": "text", "text": block.text})
if block.text.strip():
final_text = block.text
elif block.type == "tool_use":
serialized.append(
{
"type": "tool_use",
"id": block.id,
"name": block.name,
"input": block.input,
}
)
tool_calls.append({"id": block.id, "name": block.name, "input": block.input})
turn_tool_calls.append((block.name, block.input))
tool_counts[block.name] = tool_counts.get(block.name, 0) + 1
Track file reads specifically so we can show what clearing drops
if block.name == "read_file":
file_reads.append((turn, block.input.get("path", "?")))
elif block.type == "thinking":
serialized.append(
{"type": "thinking", "thinking": block.thinking, "signature": block.signature}
)
elif block.type == "compaction":
serialized.append({"type": "compaction", "content": block.content})
events.append({"turn": turn, "kind": "compaction", "summary": block.content})
compaction_this_turn = True
if verbose:
print(
f"│ ⊟ COMPACTION (turn {turn}): "
f"~{count_tokens(block.content):,}-token summary replaces prior turns"
)
messages.append({"role": "assistant", "content": serialized})
if not tool_calls:
if verbose and not compaction_this_turn:
print(f"│ turn {turn:2d} ctx={total_in:>7,} (final answer)")
break
Execute tools, collecting result sizes so the verbose print
can show how much each call added to context
tool_results: list[dict] = []
result_sizes: list[int] = []
for call in tool_calls:
if call["name"] == "memory" and memory_handler is not None:
result = memory_handler.handle(call["input"])
else:
result = execute_research_tool(call["name"], call["input"], notes)
tool_results.append(
{
"type": "tool_result",
"tool_use_id": call["id"],
"content": result,
}
)
result_sizes.append(len(result) if isinstance(result, str) else 0)
messages.append({"role": "user", "content": tool_results})
if verbose and not compaction_this_turn:
Header with context size, then one line per tool call with
its argument and approximate result size (so you can see
which calls are responsible for the next turn's ctx jump)
print(f"│ turn {turn:2d} ctx={total_in:>7,}")
for (name, tinput), rsize in zip(turn_tool_calls, result_sizes, strict=False):
size_note = f" → ~{rsize // 4:,} tok" if rsize > 200 else ""
print(f"│ {name:<16} {_format_tool_arg(name, tinput)}{size_note}")
if verbose:
if token_trajectory:
peak = max(t for _, t in token_trajectory)
status = "⚠ HIT CONTEXT LIMIT" if hit_limit else "completed"
print(
f"└─ {status}: {len(token_trajectory)} turns, peak ctx {peak:,}, "
f"final ctx {token_trajectory[-1][1]:,}, {len(events)} context event(s)\n"
)
else:
print("└─ HIT CONTEXT LIMIT on first turn: 0 turns completed\n")
return SessionResult(
messages, notes, token_trajectory, events, tool_counts, file_reads, hit_limit, final_text
)
def show_cleared_reads(result: SessionResult, keep: int):
"""Show which file reads are no longer in context after clearing.
Clearing replaces tool results older than the last keep tool uses
with placeholders. We reconstruct the tool-use order and mark any
read that falls outside the surviving keep-window as cleared.
Note: if clearing fires multiple times, only the last event's
boundary is considered and earlier cleared-then-re-read files may
be misclassified.
"""
if not result.file_reads:
print("No file reads in this session.")
return
clearing_events = [e for e in result.events if e["kind"] == "clearing"]
if not clearing_events:
print("Clearing never fired; all file reads remain in context.")
return
Walk messages in order to reconstruct the sequence of tool_use blocks
and which turn each one came from. The last keep of these survive
the most recent clearing; earlier ones are cleared.
tool_use_seq: list[tuple[int, str, str]] = [] # (turn, name, path-if-read)
turn = 0
for msg in result.messages:
if msg["role"] == "assistant" and isinstance(msg["content"], list):
turn += 1
for block in msg["content"]:
if block.get("type") == "tool_use":
name = block.get("name", "?")
path = block.get("input", {}).get("path", "") if name == "read_file" else ""
tool_use_seq.append((turn, name, path))
last_clear_turn = clearing_events[-1]["turn"]
seq_before = [t for t in tool_use_seq if t[0] < last_clear_turn]
cleared_reads = [(t, p) for (t, n, p) in seq_before[:-keep] if n == "read_file"]
surviving_reads = [
(t, p) for tu in seq_before[-keep:] for (t, n, p) in [tu] if n == "read_file"
]
Reads at or after the last clearing turn are untouched by it.
surviving_reads += [
(t, p) for (t, n, p) in tool_use_seq if t >= last_clear_turn and n == "read_file"
]
total = len(result.file_reads)
print(f"Total file reads across session: {total}")
print(f"Last clearing event fired at turn {last_clear_turn} (keep={keep})")
print(f"\nReads cleared from context: {len(cleared_reads)}")
for turn, path in cleared_reads[:12]:
print(f" ✗ turn {turn:2d}: {path}")
if len(cleared_reads) > 12:
print(f" ... and {len(cleared_reads) - 12} more")
print(
f"\nReads still in context (within the keep={keep} window or after "
f"the last clearing): {len(surviving_reads)}"
)
for turn, path in surviving_reads[:6]:
print(f" ✓ turn {turn:2d}: {path}")
if len(surviving_reads) > 6:
print(f" ... and {len(surviving_reads) - 6} more")
── Plot helpers ─────────────────────────────────────────────────────────
def plot_trajectories(
results: dict[str, SessionResult],
title: str = "Context size per turn",
triggers: dict[str, int] | None = None,
project_growth_for: str | None = None,
):
"""Line plot of context tokens per turn.
Vertical dashed lines mark turns where clearing fired; dash-dot lines
mark compaction. Horizontal dotted lines mark configured trigger
thresholds (pass triggers={"clearing": 20000, "compaction": 50000} etc).
project_growth_for: label of a run to extrapolate. Fits a line to the
last 5 points and draws a dotted segment forward ~8 turns to show where
unmanaged growth is heading.
"""
fig, ax = plt.subplots(figsize=(10, 4.5))
colors = plt.rcParams["axes.prop_cycle"].by_key()["color"]
for i, (label, res) in enumerate(results.items()):
turns = [t for t, _ in res.token_trajectory]
tokens = [tok for _, tok in res.token_trajectory]
color = colors[i % len(colors)]
ax.plot(turns, tokens, marker="o", label=label, markersize=4, color=color)
Mark events on the x-axis
for ev in res.events:
style = "--" if ev["kind"] == "clearing" else "-."
ax.axvline(x=ev["turn"], color=color, linestyle=style, alpha=0.25, linewidth=1)
Dotted growth projection for the named run, capped at 1M
if label == project_growth_for and len(turns) >= 3:
HARD_LIMIT = 1_000_000
fit_n = min(5, len(turns))
xs, ys = turns[-fit_n:], tokens[-fit_n:]
n = len(xs)
sx, sy = sum(xs), sum(ys)
slope = (n * sum(x * y for x, y in zip(xs, ys, strict=False)) - sx * sy) / (
n * sum(x * x for x in xs) - sx * sx
)
intercept = (sy - slope * sx) / n
proj_x, proj_y = [], []
for x in range(turns[-1], turns[-1] + 9):
y = slope * x + intercept
if y > HARD_LIMIT:
Clip the last segment to the 1M ceiling and stop
if proj_y and slope > 0:
proj_x.append(proj_x[-1] + (HARD_LIMIT - proj_y[-1]) / slope)
proj_y.append(HARD_LIMIT)
break
proj_x.append(x)
proj_y.append(y)
if proj_y:
ax.plot(proj_x, proj_y, linestyle=":", color=color, alpha=0.6, linewidth=1.5)
Horizontal reference lines for trigger thresholds
if triggers:
for name, value in triggers.items():
ax.axhline(y=value, color="gray", linestyle=":", alpha=0.6, linewidth=1)
ax.annotate(
f"{name} trigger: {value:,}",
xy=(ax.get_xlim()[1], value),
xytext=(-5, 3),
textcoords="offset points",
ha="right",
va="bottom",
fontsize=8,
color="gray",
)
200K reference: earlier models cap here and would hard-stop
if ax.get_ylim()[1] > 30_000:
ax.axhline(y=200_000, color="gray", linestyle="--", alpha=0.5, linewidth=1)
ax.annotate(
"200K: earlier models stop here",
xy=(ax.get_xlim()[0], 200_000),
xytext=(5, -12),
textcoords="offset points",
ha="left",
va="top",
fontsize=8,
color="gray",
alpha=0.8,
)
ax.set_xlabel("Turn")
ax.set_ylabel("Context tokens (incl. cached)")
ax.set_title(title)
ax.legend(loc="best")
ax.grid(alpha=0.3)
plt.tight_layout()
plt.show()
def plot_summary_bars(results: dict[str, SessionResult], title: str = "Session outcomes"):
"""Side-by-side bars: final context, file reads."""
labels = list(results.keys())
final_ctx = [r.token_trajectory[-1][1] for r in results.values()]
reads = [r.tool_counts.get("read_file", 0) for r in results.values()]
fig, axes = plt.subplots(1, 2, figsize=(9, 3.5))
fig.suptitle(title, y=1.02)
for ax, values, ylabel in zip(
axes,
[final_ctx, reads],
["Final context (tokens)", "File reads"],
strict=False,
):
bars = ax.bar(range(len(labels)), values, color=plt.cm.Set2(range(len(labels))))
ax.set_ylabel(ylabel)
ax.set_xticks(range(len(labels)))
ax.set_xticklabels(labels, rotation=30, ha="right", fontsize=8)
ax.grid(axis="y", alpha=0.3)
for bar, val in zip(bars, values, strict=False):
ax.text(
bar.get_x() + bar.get_width() / 2,
bar.get_height(),
f"{val:,}" if val > 1000 else str(val),
ha="center",
va="bottom",
fontsize=7,
)
plt.tight_layout()
plt.show()
Baseline: no context management
like compaction and memory, clearing has no prompt to tune, and the knobs are all numeric (
trigger
,
keep
,
clear_at_least
) or list-based (
exclude_tools
). One trade-off to understand: clearing invalidates cached prompt prefixes. To account for this, clear enough tokens to make the cache invalidation worthwhile; the
clear_at_least
parameter ensures a minimum number of tokens is cleared each time. You'll incur cache write costs each time clearing fires, but subsequent requests can reuse the newly cached prefix.
The right values for trigger and keep depend on how your agent uses tool results: how large they are, how often the agent revisits them, whether re-fetching is cheap. The clearing run above used trigger=30K and keep=4; the all-three run later uses a higher trigger and keep=6 so clearing and compaction split the work. Test a few configurations against your own agent's workload: the context_management.applied_edits field in each response shows how many tool uses and tokens were cleared, which makes the effect of each config directly observable.