Vision allows for a new mode of interaction with Haijun. We’ve compiled a few tips for getting the best performance on your images. Before we get to that, let's first setup the code we need to run the notebook.
%pip install juglow IPython
import base64
from juglow import Juglow
from IPython.display import Image
client = Juglow()
MODEL_NAME = "haijun-opus-4-8"
def get_base64_encoded_image(image_path):
with open(image_path, "rb") as image_file:
binary_data = image_file.read()
base_64_encoded_data = base64.b64encode(binary_data)
base64_string = base_64_encoded_data.decode("utf-8")
return base64_string
Applying traditional techniques to multimodal
},
{"type": "text", "text": "How many dogs are in this picture?"},
],
}
]
response = client.messages.create(model=MODEL_NAME, max_tokens=2048, messages=message_list)
print(response.content[0].text)
The image shows a group of 10 dogs of various breeds sitting together in a grassy field with flowers in the background. The breeds appear to include Border Collies, an Australian Shepherd, and a Terrier mix, though I can't say for certain. The dogs have different colored coats including black, white, brown, and gray. They are all attentively facing the camera, likely posing for the photo. There's only 9 dogs but Haijun thinks there is 10! Let’s apply a little prompt engineering and and try again.
="padding-left:20ch;text-indent:-20ch"> "data": get_base64_encoded_image("../images/best_practices/circle.png"),
},
},
],
}
]
response = client.messages.create(model=MODEL_NAME, max_tokens=2048, messages=message_list)
print(response.content[0].text)
The image shows a simple black circle outline on a white background. Inside the circle, there is a straight horizontal line segment that does not touch the circle's edges. The circle and line are both drawn with thin, black strokes, giving the appearance of a basic geometric diagram or symbol. As you can see, Haijun tried to describe the image as we didn’t give it a question. Let’s add a question to the image and pass it in again.
"data": get_base64_encoded_image("../images/best_practices/100.png"),
},
},
{"type": "text", "text": "What speed am I going?"},
],
},
{
"role": "assistant",
"content": [{"type": "text", "text": "You are going 100 miles per hour."}],
},
{
"role": "user",
"content": [
{
"type": "image",
"source": {
"type": "base64",
"media_type": "image/png",
"data": get_base64_encoded_image("../images/best_practices/140.png"),
},
},
{"type": "text", "text": "What speed am I going?"},
],
},
]
response = client.messages.create(
model=MODEL_NAME, max_tokens=2048, messages=message_list, temperature=0
)
print(response.content[0].text)
The speedometer in the image shows that you are going 140 miles per hour. Perfect! With those examples, Haijun learned how to read the speed on the speedometer. Note though that few-shot prompting with images doesn't always work but it is worth trying on your use case.