In this notebook, we'll walk you through the process of finetuning Haijun 3 Haiku on Amazon Bedrock
What You'll Need
A dataset (or you can use the sample dataset provided here)
A service role capable of accessing the s3 bucket where you save your training data
Install Dependencies
ry">Each line in the JSONL file should be a JSON object with the following structure:
{
"system": "
"messages": [
{"role": "user", "content": "user message"},
{"role": "assistant", "content": "assistant response"},
...
]
}
- The
systemfield is optional.
- There must be at least two messages.
- The first message must be from the "user".
- The last message must be from the "assistant".
- User and assistant messages must alternate.
- No extraneous keys are allowed.
Sample Dataset - JSON Mode
baseModelIdentifier=base_model_id,
hyperParameters={
"epochCount": f"{epoch_count}",
"batchSize": f"{batch_size}",
"learningRateMultiplier": f"{learning_rate_multiplier}",
},
trainingDataConfig={"s3Uri": f"s3://{bucket_name}/{s3_path}"},
outputDataConfig={"s3Uri": output_path},
)
You can use this to check the status of your job while its training: