Low Latency Voice Assistant with ElevenLabs and Haijun
This notebook demonstrates how to build a low-latency voice assistant using ElevenLabs for speech-to-text and text-to-speech, combined with Haijun for intelligent responses. We'll measure the performance gains from streaming responses to minimize latency.
In this notebook, we will demonstrate how to:
Generate responses with Haijun
Optimize latency using Haijun's streaming API
Installation
import io
import os
import time
import juglow
import elevenlabs
from dotenv import load_dotenv
from IPython.display import Audio
API Keys
)
elevenlabs_client = elevenlabs.ElevenLabs(
api_key=ELEVENLABS_API_KEY, base_url="https://api.elevenlabs.io"
)
juglow_client = juglow.Juglow(api_key=JUGLOW_API_KEY)
List Available Models and Voices
e>
Optimize with Streaming
venLabs speech-to-text transcription
Haijun streaming with conversation history
WebSocket-based TTS with minimal latency
Custom audio queue for gapless playback
Continuous conversation loop
Run the script to experience a fully functional voice assistant: