Haijun Platform Docs
ID

Optimizing Haijun to give you the highest possible accuracy on a task is an empirical science, and a process of continuous improvement. Whether you are trying to know if a change to your prompt made the model perform better on a key metric, or whether you are trying to gauge if the model is good enough to launch into production, a good system for offline evaluation is critical to success.

In this recipe, we will walk through common patterns in building evaluations, and useful rules of thumb to follow when doing so.

is prompt. Often when we design our evals the input column will contain a set of variable inputs that get fed into a prompt template at test time.

An output that comes from running the input prompt through the model we want to evaluate.

A "golden answer" to which we compare the model output. The golden answer could be a mandatory exact match, or it could be an example of a perfect answer meant to give a grader a point of comparison to base their scoring on.

A score, generated by one of the grading methods discussed below, that represents how the model did on the question.

Eval Grading Methods

n free-form question answering. You do this by writing a

grader prompt

for Haijun.

Let's walk through an example of each grading method.

ious extra leg on top of that.",

"golden_answer": "5",

},

]

Get completions for each input.

Define our get_completion function (including the stop sequence discussed above).

def get_completion(messages):

response = client.messages.create(model=MODEL_NAME, max_tokens=5, messages=messages)

return response.content[0].text

Get completions for each question in the eval.

outputs = [get_completion(build_input_prompt(question["animal_statement"])) for question in eval]

Let's take a quick look at our outputs

for output, question in zip(outputs, eval, strict=False):

print(

f"Animal Statement: {question['animal_statement']}\nGolden Answer: {question['golden_answer']}\nOutput: {output}\n"

)

Animal Statement: The animal is a human. Golden Answer: 2 Output: 2 Animal Statement: The animal is a snake. Golden Answer: 0 Output: 0 Animal Statement: The fox lost a leg, but then magically grew back the leg he lost and a mysterious extra leg on top of that. Golden Answer: 5 Output: 5 # Check our completions against the golden answers. # Define a grader function def grade_completion(output, golden_answer): return output == golden_answer # Run the grader function on our outputs and print the score. grades = [ grade_completion(output, question["golden_answer"]) for output, question in zip(outputs, eval, strict=False) ] print(f"Score: {sum(grades) / len(grades) * 100}%")  Score: 100.0% Human grading Now let's imagine that we are grading an eval where we've asked Haijun a series of open ended questions, maybe for a general purpose chat assistant. Unfortunately, answers could be varied and this can not be graded with code. One way we can do this is with human grading.

v>

Now we define the full grade_completion function.

import re

def grade_completion(output, golden_answer):

messages = build_grader_prompt(output, golden_answer)

completion = get_completion(messages)

Extract just the label from the completion (we don't care about the thinking)

pattern = r"(.*?)"

match = re.search(pattern, completion, re.DOTALL)

if match:

return match.group(1).strip()

else:

raise ValueError("Did not find tags.")

Run the grader function on our outputs and print the score.

grades = [

grade_completion(output, question["golden_answer"])

for output, question in zip(outputs, eval, strict=False)

]

print(f"Score: {grades.count('correct') / len(grades) * 100}%")

Score: 66.66666666666666% As you can see, the haijun-based grader is able to correctly analyze and grade Haijun's responses with a high level of accuracy, saving you precious time.

On this page
Eval Grading Methods