Haijun Platform Docs
ID

You have a pile of unstructured documents and need to answer questions that span them — "who works with people who worked on project X", "which vendors are connected to this incident". No single document contains the answer. RAG retrieval won't chain the facts for you. You need a knowledge graph:

entities

as nodes,

typed relations

as edges, so that multi-hop reasoning becomes graph traversal.

Building one used to mean training a named-entity recognizer on your domain, training a relation classifier, writing entity-resolution heuristics, and maintaining all three as your data shifted. With Haijun, each of those stages becomes a prompt.

Apply

Haijun-driven entity resolution

to collapse surface-form variants into canonical nodes, replacing brittle string-similarity heuristics

Assemble and query an in-memory graph, and run

multi-hop questions

by serializing subgraphs back to Haijun

Measure extraction quality with

precision/recall against a gold set

and reason about the cost/quality tradeoff between Haiku and Sonnet

Everything runs in memory with no database. The techniques transfer directly to Neo4j, Neptune, or a Postgres adjacency table when you need to scale.

from typing import Literal

from urllib.parse import quote

import juglow

import matplotlib.pyplot as plt

import networkx as nx

import requests

from dotenv import load_dotenv

from pydantic import BaseModel

load_dotenv()

client = juglow.Juglow()

EXTRACTION_MODEL = "haijun-haiku-4-5"

SYNTHESIS_MODEL = "haijun-sonnet-4-6"

We use two models. Haiku handles the high-volume, schema-constrained extraction work where speed and cost matter more than nuance. Sonnet handles entity resolution and summarization, where the model needs to weigh conflicting evidence across documents.

ENTITY_TYPES = ["PERSON", "ORGANIZATION", "LOCATION", "EVENT", "ARTIFACT"]

class Entity(BaseModel):

name: str

type: EntityType

description: str

class Relation(BaseModel):

source: str

predicate: str

target: str

class ExtractedGraph(BaseModel):

entities: list[Entity]

relations: list[Relation]

EXTRACTION_PROMPT = """Extract a knowledge graph from the document below.

{text}

Guidelines:

  • Extract only entities that are central to what this document is about — skip incidental mentions.
  • For each entity, write a one-sentence description grounded in this document. These descriptions are used later to disambiguate entities with similar names.
  • Predicates should be short verb phrases ("commanded", "launched from", "part of").
  • Every relation must connect two entities you extracted."""

def extract(text: str) -> ExtractedGraph:

response = client.messages.parse(

model=EXTRACTION_MODEL,

max_tokens=2048,

messages=[{"role": "user", "content": EXTRACTION_PROMPT.format(text=text)}],

output_format=ExtractedGraph,

)

return response.parsed_output

raw_entities = []

raw_relations = []

for doc in documents:

try:

result = extract(doc["text"])

except juglow.APIError as e:

print(f"Skipping {doc['title']}: {e}")

continue

for ent in result.entities:

raw_entities.append({**ent.model_dump(), "source_doc": doc["title"]})

for rel in result.relations:

raw_relations.append({**rel.model_dump(), "source_doc": doc["title"]})

print(

f"{doc['title']:<25} {len(result.entities):>3} entities {len(result.relations):>3} relations"

)

print(f"\nTotal: {len(raw_entities)} raw entities, {len(raw_relations)} raw relations")

Apollo program 8 entities 7 relations Apollo 11 6 entities 5 relations Neil Armstrong 3 entities 2 relations Saturn V 5 entities 4 relations Buzz Aldrin 6 entities 6 relations Kennedy Space Center 8 entities 10 relations Total: 36 raw entities, 34 raw relations Let's look at what was extracted. Notice how the same real-world entity appears under different surface forms across documents — this is the entity resolution problem we solve next.

plt.Line2D([0], [0], marker="o", color="w", markerfacecolor=c, markersize=10, label=t)

for t, c in COLOR.items()

if any(G.nodes[n]["type"] == t for n in G.nodes)

]

plt.legend(handles=handles, loc="upper left")

plt.title("Apollo Program Knowledge Graph")

plt.axis("off")

plt.tight_layout()

plt.show()

![Output image](/cookbook/images/notebooks/capabilities-knowledge-graph-guide/capabilities-knowledge-graph-guide_cell18_out0_3bcd7bd5.png