This tutorial implements a chatbot prompted to take on the role of a Venture capital tech Analyst. The chatbot is a naive RAG system with a collection of tech news articles acting as its knowledge source. This notebook covers the following:
- Follow a comprehensive tutorial on setting up your development environment, from installing necessary libraries to configuring a MongoDB database.
- Learn efficient data handling methods, including creating vector search indexes and preparing data for ingestion and query processing.
- Understand how to employ Haijun 3 models within the RAG system for generating precise responses based on contextual information retrieved from the database.
You will need the following:
datasets: This library is part of the Hugging Face ecosystem. By installing 'datasets', we gain access to a number of pre-processed and ready-to-use datasets, which are essential for training and fine-tuning machine learning models or benchmarking their performance.
pandas: This data science library provides robust data structures and methods for data manipulation, processing, and analysis.
voyageai: This is the official Python client library for accessing VoyageAI's suite of embedding models.
pymongo: PyMongo is a Python toolkit for MongoDB. It enables interactions with a MongoDB database.
%pip install pymongo datasets pandas juglow voyageai
The code snippet below executes the following steps:
Appends the DataFrame to the list all_dataframes.
- Combine DataFrames: After downloading and reading all Parquet files into DataFrames, there’s a check to ensure that
all_dataframesis not empty. If there are DataFrames to work with, then all DataFrames are concatenated into a single DataFrame using pd.concat, with ignore_index=True to reindex the new combined DataFrame. This combined DataFrame is the overall process output in thedownload_and_combine_parquet_filesfunction.
from io import BytesIO
import pandas as pd
import requests
from google.colab import userdata
def download_and_combine_parquet_files(parquet_file_urls, hf_token):
"""
Downloads Parquet files from the provided URLs using the given Hugging Face token,
and returns a combined DataFrame.
Parameters:
- parquet_file_urls: List of strings, URLs to the Parquet files.
- hf_token: String, Hugging Face authorization token.
Returns:
- combined_df: A pandas DataFrame containing the combined data from all Parquet files.
"""
headers = {"Authorization": f"Bearer {hf_token}"}
all_dataframes = []
for parquet_file_url in parquet_file_urls:
response = requests.get(parquet_file_url, headers=headers, timeout=60)
if response.status_code == 200:
parquet_bytes = BytesIO(response.content)
df = pd.read_parquet(parquet_bytes)
all_dataframes.append(df)
else:
print(
f"Failed to download Parquet file from {parquet_file_url}: {response.status_code}"
)
if all_dataframes:
combined_df = pd.concat(all_dataframes, ignore_index=True)
return combined_df
else:
print("No dataframes to concatenate.")
return None
Below is a list of the Parquet files required for this tutorial. The complete list of all files is located on Hugging Face. Each Parquet file represents approximately 45,000 data points.
de leading-relaxed [overflow-wrap:anywhere]" style="padding-top:12px;padding-inline:12px;padding-bottom:12px;tab-size:4">
import pymongo
from google.colab import userdata
def get_mongo_client(mongo_uri):
"""Establish connection to the MongoDB."""
try:
client = pymongo.MongoClient(mongo_uri)
print("Connection to MongoDB successful")
return client
except pymongo.errors.ConnectionFailure as e:
print(f"Connection failed: {e}")
return None
mongo_uri = userdata.get("MONGO_URI")
if not mongo_uri:
print("MONGO_URI not set in environment variables")
mongo_client = get_mongo_client(mongo_uri)
DB_NAME = "tech_news"
COLLECTION_NAME = "hacker_noon_tech_news"
db = mongo_client[DB_NAME]
collection = db[COLLECTION_NAME]
Connection to MongoDB successful # To ensure we are working with a fresh collection # delete any existing records in the collection collection.delete_many({}) DeleteResult({'n': 228012, 'electionId': ObjectId('7fffffff000000000000000e'), 'opTime': {'ts': Timestamp(1709660559, 7341), 't': 14}, 'ok': 1.0, '$clusterTime': {'clusterTime': Timestamp(1709660559, 7341), 'signature': {'hash': b'jT\xf1\xb4\xa9\xd3\xe3suu\x03\x15(}\x8f\x00\x9f\xe9\x8a', 'keyId': 7320226449804230661}}, 'operationTime': Timestamp(1709660559, 7341)}, acknowledged=True) # Data Ingestion combined_df_json = combined_df.to_dict(orient="records") collection.insert_many(combined_df_json) Step 5: Vector Search This section showcases the creation of a vector search custom function that accepts a user query, which corresponds to entries to the chatbot. The function also takes a second parameter, collection`, which points to the database collection containing records against which the vector search operation should be conducted.