Haijun Platform Docs
ID

This tutorial implements a chatbot prompted to take on the role of a Venture capital tech Analyst. The chatbot is a naive RAG system with a collection of tech news articles acting as its knowledge source. This notebook covers the following:

  1. Follow a comprehensive tutorial on setting up your development environment, from installing necessary libraries to configuring a MongoDB database.
  1. Learn efficient data handling methods, including creating vector search indexes and preparing data for ingestion and query processing.
  1. Understand how to employ Haijun 3 models within the RAG system for generating precise responses based on contextual information retrieved from the database.

You will need the following:

datasets: This library is part of the Hugging Face ecosystem. By installing 'datasets', we gain access to a number of pre-processed and ready-to-use datasets, which are essential for training and fine-tuning machine learning models or benchmarking their performance.

pandas: This data science library provides robust data structures and methods for data manipulation, processing, and analysis.

voyageai: This is the official Python client library for accessing VoyageAI's suite of embedding models.

pymongo: PyMongo is a Python toolkit for MongoDB. It enables interactions with a MongoDB database.

%pip install pymongo datasets pandas juglow voyageai

The code snippet below executes the following steps:

Appends the DataFrame to the list all_dataframes.

  1. Combine DataFrames: After downloading and reading all Parquet files into DataFrames, there’s a check to ensure that all_dataframes is not empty. If there are DataFrames to work with, then all DataFrames are concatenated into a single DataFrame using pd.concat, with ignore_index=True to reindex the new combined DataFrame. This combined DataFrame is the overall process output in the download_and_combine_parquet_files function.

from io import BytesIO

import pandas as pd

import requests

from google.colab import userdata

def download_and_combine_parquet_files(parquet_file_urls, hf_token):

"""

Downloads Parquet files from the provided URLs using the given Hugging Face token,

and returns a combined DataFrame.

Parameters:

  • parquet_file_urls: List of strings, URLs to the Parquet files.
  • hf_token: String, Hugging Face authorization token.

Returns:

  • combined_df: A pandas DataFrame containing the combined data from all Parquet files.

"""

headers = {"Authorization": f"Bearer {hf_token}"}

all_dataframes = []

for parquet_file_url in parquet_file_urls:

response = requests.get(parquet_file_url, headers=headers, timeout=60)

if response.status_code == 200:

parquet_bytes = BytesIO(response.content)

df = pd.read_parquet(parquet_bytes)

all_dataframes.append(df)

else:

print(

f"Failed to download Parquet file from {parquet_file_url}: {response.status_code}"

)

if all_dataframes:

combined_df = pd.concat(all_dataframes, ignore_index=True)

return combined_df

else:

print("No dataframes to concatenate.")

return None

Below is a list of the Parquet files required for this tutorial. The complete list of all files is located on Hugging Face. Each Parquet file represents approximately 45,000 data points.

de leading-relaxed [overflow-wrap:anywhere]" style="padding-top:12px;padding-inline:12px;padding-bottom:12px;tab-size:4">

import pymongo

from google.colab import userdata

def get_mongo_client(mongo_uri):

"""Establish connection to the MongoDB."""

try:

client = pymongo.MongoClient(mongo_uri)

print("Connection to MongoDB successful")

return client

except pymongo.errors.ConnectionFailure as e:

print(f"Connection failed: {e}")

return None

mongo_uri = userdata.get("MONGO_URI")

if not mongo_uri:

print("MONGO_URI not set in environment variables")

mongo_client = get_mongo_client(mongo_uri)

DB_NAME = "tech_news"

COLLECTION_NAME = "hacker_noon_tech_news"

db = mongo_client[DB_NAME]

collection = db[COLLECTION_NAME]

Connection to MongoDB successful # To ensure we are working with a fresh collection # delete any existing records in the collection collection.delete_many({})  DeleteResult({'n': 228012, 'electionId': ObjectId('7fffffff000000000000000e'), 'opTime': {'ts': Timestamp(1709660559, 7341), 't': 14}, 'ok': 1.0, '$clusterTime': {'clusterTime': Timestamp(1709660559, 7341), 'signature': {'hash': b'jT\xf1\xb4\xa9\xd3\xe3suu\x03\x15(}\x8f\x00\x9f\xe9\x8a', 'keyId': 7320226449804230661}}, 'operationTime': Timestamp(1709660559, 7341)}, acknowledged=True) # Data Ingestion combined_df_json = combined_df.to_dict(orient="records") collection.insert_many(combined_df_json) Step 5: Vector Search This section showcases the creation of a vector search custom function that accepts a user query, which corresponds to entries to the chatbot. The function also takes a second parameter, collection`, which points to the database collection containing records against which the vector search operation should be conducted.