Multiple agents independently run a single evaluation task from an evaluation file.
import json
import re
import time
import traceback
import xml.etree.ElementTree as ET # noqa: S314
from pathlib import Path
from typing import Any
from juglow import Juglow
Prompts
- The steps you took to complete the task
- Which tools you used, in what order, and why
- The inputs you provided to each tool
- The outputs you received from each tool
- A summary for how you arrived at the response
Feedback Requirements:
- In your
tags, provide constructive feedback on the tools:
- Comment on tool names: Are they clear and descriptive?
- Comment on input parameters: Are they well-documented? Are required vs optional parameters clear?
- Comment on descriptions: Do they accurately describe what the tool does?
- Comment on any errors encountered during tool usage: Did the tool fail to execute? Did the tool return too many tokens?
- Identify specific areas for improvement and explain WHY they would help
- Be specific and actionable in your suggestions
Response Requirements:
- Your response should be concise and directly address what was asked
- Always wrap your final response in
tags
- If you cannot solve the task return
NOT_FOUND
- For numeric responses, provide just the number
- For IDs, provide just the ID
- For names or text, provide the exact text requested
- Your response should go last"""
Agent Loop
tool_duration = time.time() - tool_start_ts
Update tool metrics
if tool_name not in tool_metrics:
tool_metrics[tool_name] = {"count": 0, "durations": []}
tool_metrics[tool_name]["count"] += 1
tool_metrics[tool_name]["durations"].append(tool_duration)
Prepare tool result and append to messages
messages.append(_prepare_tool_result(tool_use.id, tool_response))
response = client.messages.create(
model=model,
max_tokens=4096,
system=EVALUATION_PROMPT,
messages=messages,
tools=tools,
)
messages.append({"role": "assistant", "content": response.content})
response = next(
(block.text for block in response.content if hasattr(block, "text")),
None,
)
return response, tool_metrics
Helper Functions
average_duration_s=average_duration_s,
average_tool_calls=average_tool_calls,
total_tool_calls=total_tool_calls,
)
report += "".join(
[
TASK_TEMPLATE.format(
prompt=task["prompt"],
expected_response=task["response"],
actual_response=result["actual"],
correct_indicator="✅" if result["score"] else "❌",
total_duration=result["total_duration"],
tool_calls=json.dumps(result["tool_calls"], indent=2),
summary=result["summary"] or "N/A",
feedback=result["feedback"] or "N/A",
)
for task, result in zip(tasks, results, strict=False)
]
)
Join all sections into final report
return report