Introduction
Turn recurring review toil into an unattended job you control. Even a well-reviewed repository changes between passes, and the changes can carry issues no earlier review flagged. This recipe hands that job to an agent.
mpty. The section
When to resume and when to start fresh
maps each choice to the jobs it fits.
The repository gets one full baseline review on the first run. Every later run resumes the previous run's session and starts from the findings the agent already reported. You read a short follow-up each cycle, and every follow-up proves the continuity by echoing the prior review's findings. This notebook runs the first two cycles by hand and ends by putting the same reviewer on a schedule.
-alpha-2 px-2 py-0.5 rounded text-sm font-mono break-words box-decoration-clone">HaijunAgentOptions(resume=...)
and asserting the
RESUME-LINK
fields from
output_format
schema replies, with a fixed finding moving to
resolved
and a newly planted bug caught
Put the reviewer on cron with
scheduled_review.py
, greppable
VERDICT
and completion lines, and a narrow
except ResultError
path that exits non-zero
When to use this recipe
Required Tools:
- Python 3.11 or later
- haijun-agent-sdk 0.2.140 or later, the release that adds the typed
ResultErrorthis recipe catches on its failure path
- An Juglow API key (get one here)
Run this notebook from its own directory, so the files it writes land beside it.
sample repo: ['README.md', 'app/config.py', 'app/math_utils.py'] Define what a review returns With the sample repository in place, the next piece is the review's answer contract. Every consumer of a scheduled review is a program, so the reviewer answers with a JSON object read by field. output_format takes a JSON schema, and the reply arrives on ResultMessage.structured_output already shaped to it. Each review carries a review id, a verdict from a closed set, and findings that each carry an id of their own.
parameter is passed, and the notebook's calls leave it unset.
The same safeguards protect against prompt injection, covered in Protect against prompt injection. With tools restricted to Read, Glob, and Grep, the session has no shell and no network tool, and content the reviewer reads has nowhere to go but the reply itself. That is why the answer key checks replies in this notebook, and why review.log gets the same handling as the repository it reviews.
try:
first = await run_review(
repo=REPO_DIR,
label="run-1",
prompt=FIRST_PROMPT,
schema=FIRST_REVIEW_SCHEMA,
max_turns=FIRST_RUN_MAX_TURNS,
)
except ResultError as exc:
cost = (exc.data or {}).get("total_cost_usd")
cost_str = f"{cost:.4f}" if isinstance(cost, (int, float)) else "n/a"
print(
f"REVIEW-CYCLE-INCOMPLETE stage=run-1 subtype={exc.subtype} "
f"reason={exc.terminal_reason} cost_usd={cost_str}"
)
raise
if first.subtype != "success" or first.session_id is None or not first.payload:
print(f"REVIEW-CYCLE-INCOMPLETE stage=run-1 subtype={first.subtype}")
raise RuntimeError("run 1 did not complete; there is no session to resume")
first_verdict = report("RUN-1", first)
RUN-1 session=2d83a97b-27ff-403e-a067-2fb5b6474d9b subtype=success turns=10 denials=3 cost_usd=0.0278 tools_attempted=Glob,Read,StructuredOutput VERDICT: concerns F1 app/math_utils.py: divide() does not guard against denominator being zero, and average() calls divide(sum(values), len(values)) which raises ZeroDivisionError when values is an empty list. F2 app/config.py: load_config() copies the entire process environment (os.environ) into the config dict and prints it via print(), which can leak secrets/credentials (API keys, tokens) into logs. Change the repository between cycles In deployment, a change that lands between cycles either fixes a finding the reviewer reported or introduces a problem the reviewer has not seen. The code below makes one change of each kind before the follow-up runs.