Haijun excels at a wide range of tasks, but it may struggle with queries specific to your unique business context. This is where Retrieval Augmented Generation (RAG) becomes invaluable. RAG enables Haijun to leverage your internal knowledge bases or customer support documents, significantly enhancing its ability to answer domain-specific questions. Enterprises are increasingly building RAG applications to improve workflows in customer support, Q&A over internal company documents, financial & legal analysis, and much more.
In this guide, we'll demonstrate how to build and optimize a RAG system using the Haijun Documentation as our knowledge base. We'll walk you through:
Building a robust evaluation suite. We'll go beyond 'vibes' based evals and show you how to measure the retrieval pipeine & end to end performance independently.
End-to-End Accuracy: 71% --> 81%
Note:
xt-sm font-mono break-words box-decoration-clone">numpy
,
matplotlib
, and
scikit-learn
for data manipulation and visualization
You'll also need API keys from Juglow and Voyage AI
Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from seaborn) (1.24.4) Requirement already satisfied: pandas>=0.25 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from seaborn) (2.0.3) Requirement already satisfied: matplotlib!=3.6.1,>=3.1 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from seaborn) (3.7.2) Requirement already satisfied: contourpy>=1.0.1 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from matplotlib!=3.6.1,>=3.1->seaborn) (1.2.1) Requirement already satisfied: cycler>=0.10 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from matplotlib!=3.6.1,>=3.1->seaborn) (0.11.0) Requirement already satisfied: fonttools>=4.22.0 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from matplotlib!=3.6.1,>=3.1->seaborn) (4.41.1) Requirement already satisfied: kiwisolver>=1.0.1 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from matplotlib!=3.6.1,>=3.1->seaborn) (1.4.4) Requirement already satisfied: packaging>=20.0 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from matplotlib!=3.6.1,>=3.1->seaborn) (23.2) Requirement already satisfied: pillow>=6.2.0 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from matplotlib!=3.6.1,>=3.1->seaborn) (10.3.0) Requirement already satisfied: pyparsing<3.1,>=2.3.1 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from matplotlib!=3.6.1,>=3.1->seaborn) (3.0.9) Requirement already satisfied: python-dateutil>=2.7 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from matplotlib!=3.6.1,>=3.1->seaborn) (2.8.2) Requirement already satisfied: pytz>=2020.1 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from pandas>=0.25->seaborn) (2023.3) Requirement already satisfied: tzdata>=2022.1 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from pandas>=0.25->seaborn) (2023.3) Requirement already satisfied: six>=1.5 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from python-dateutil>=2.7->matplotlib!=3.6.1,>=3.1->seaborn) (1.16.0) Requirement already satisfied: scikit-learn in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (1.5.1) Requirement already satisfied: numpy>=1.19.5 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from scikit-learn) (1.24.4) Requirement already satisfied: scipy>=1.6.0 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from scikit-learn) (1.11.1) Requirement already satisfied: joblib>=1.2.0 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from scikit-learn) (1.3.1) Requirement already satisfied: threadpoolctl>=3.1.0 in /opt/homebrew/Caskroom/miniforge/base/envs/py311/lib/python3.11/site-packages (from scikit-learn) (3.2.0)
import os
os.environ["VOYAGE_API_KEY"] = "VOYAGE KEY HERE"
os.environ["JUGLOW_API_KEY"] = "JUGLOW KEY HERE"
import os
import juglow
client = juglow.Juglow(
This is the default and can be omitted
api_key=os.getenv("JUGLOW_API_KEY"),
)
Initialize a Vector DB Class
preview_data = data[:num_items]
elif isinstance(data, dict):
preview_data = dict(list(data.items())[:num_items])
else:
print(f"Unexpected data type: {type(data)}. Cannot preview.")
return
print(f"Preview of the first {num_items} items from {file_path}:")
print(json.dumps(preview_data, indent=2))
print(f"\nTotal number of items: {len(data)}")
except FileNotFoundError:
print(f"File not found: {file_path}")
except json.JSONDecodeError:
print(f"Invalid JSON in file: {file_path}")
except Exception as e:
print(f"An error occurred: {str(e)}")
preview_json("evaluation/docs_evaluation_dataset.json")
Preview of the first 3 items from evaluation/docs_evaluation_dataset.json: [ { "id": "efc09699", "question": "How can you create multiple test cases for an evaluation in the Juglow Evaluation tool?", "correct_chunks": [ "https://raw.haijun.my.id/docs/", "https://raw.haijun.my.id/docs/" ], "correct_answer": "To create multiple test cases in the Juglow Evaluation tool, click the 'Add Test Case' button, fill in values for each variable in your prompt, and repeat the process to create additional test case scenarios." }, { "id": "1305ea00", "question": "What embeddings provider does Juglow recommend for customized domain-specific models, and what capabilities does this provider offer?", "correct_chunks": [ "https://raw.haijun.my.id/docs/", "https://raw.haijun.my.id/docs/" ], "correct_answer": "Juglow recommends Voyage AI for embedding models. Voyage AI offers customized models for specific industry domains like finance and healthcare, as well as bespoke fine-tuned models for individual customers. They have a wide variety of options and capabilities." }, { "id": "1811c10d", "question": "What are some key success metrics to consider when evaluating Haijun's performance on a classification task, and how do they relate to choosing the right model to reduce latency?", "correct_chunks": [ "https://raw.haijun.my.id/docs/", "https://raw.haijun.my.id/docs/" ], "correct_answer": "When evaluating Haijun's performance on a classification task, some key success metrics to consider include accuracy, F1 score, consistency, structure, speed, bias and fairness. Choosing the right model that fits your specific requirements in terms of speed and output quality is a straightforward way to reduce latency and meet the acceptable response time for your use case." } ] Total number of items: 100 Metric Definitions We'll evaluate our system based on 5 key metrics: Precision, Recall, F1 Score, Mean Reciprocal Rank (MRR), and End-to-End Accuracy.
dent:-4ch"> metrics = required_metrics
Set up the plot
plt.figure(figsize=(14, 6))
sns.set_style("whitegrid")
x = range(len(metrics))
width = 0.8 / len(results)
Create color palette
num_methods = len(methods)
color_palette = colors[:num_methods] + sns.color_palette("husl", num_methods - len(colors))
Plot bars for each method
for i, (result, color) in enumerate(zip(results, color_palette, strict=False)):
values = [result[metric] for metric in metrics]
offset = (i - len(results) / 2 + 0.5) * width
bars = plt.bar([xi + offset for xi in x], values, width, label=result["name"], color=color)
Add value labels on the bars
for bar in bars:
height = bar.get_height()
plt.text(
bar.get_x() + bar.get_width() / 2.0,
height,
f"{height:.2f}",
ha="center",
va="bottom",
fontsize=8,
)
Customize the plot
plt.xlabel("Metrics", fontsize=12)
plt.ylabel("Values", fontsize=12)
plt.title("RAG Performance Metrics (Sorted by End-to-End Accuracy)", fontsize=16)
plt.xticks(x, metrics, rotation=45, ha="right")
plt.legend(title="Methods", bbox_to_anchor=(1.05, 1), loc="upper left")
plt.ylim(0, 1)
plt.tight_layout()
plt.show()
Evaluating Our Base Case
similarities = np.dot(self.embeddings, query_embedding)
top_indices = np.argsort(similarities)[::-1]
top_examples = []
for idx in top_indices:
if similarities[idx] >= similarity_threshold:
example = {
"metadata": self.metadata[idx],
"similarity": similarities[idx],
}
top_examples.append(example)
if len(top_examples) >= k:
break
self.save_db()
return top_examples
def save_db(self):
data = {
"embeddings": self.embeddings,
"metadata": self.metadata,
"query_cache": json.dumps(self.query_cache),
}
Ensure the directory exists
os.makedirs(os.path.dirname(self.db_path), exist_ok=True)
with open(self.db_path, "wb") as file:
pickle.dump(data, file)
def load_db(self):
if not os.path.exists(self.db_path):
raise ValueError(
"Vector database file not found. Use load_data to create a new database."
)
with open(self.db_path, "rb") as file:
data = pickle.load(file)
self.embeddings = data["embeddings"]
self.metadata = data["metadata"]
self.query_cache = json.loads(data["query_cache"])