You have a pile of unstructured documents and need to answer questions that span them — "who works with people who worked on project X", "which vendors are connected to this incident". No single document contains the answer. RAG retrieval won't chain the facts for you. You need a knowledge graph:
entities
as nodes,
typed relations
as edges, so that multi-hop reasoning becomes graph traversal.
Building one used to mean training a named-entity recognizer on your domain, training a relation classifier, writing entity-resolution heuristics, and maintaining all three as your data shifted. With Haijun, each of those stages becomes a prompt.
Apply
Haijun-driven entity resolution
to collapse surface-form variants into canonical nodes, replacing brittle string-similarity heuristics
Assemble and query an in-memory graph, and run
multi-hop questions
by serializing subgraphs back to Haijun
Measure extraction quality with
precision/recall against a gold set
and reason about the cost/quality tradeoff between Haiku and Sonnet
Everything runs in memory with no database. The techniques transfer directly to Neo4j, Neptune, or a Postgres adjacency table when you need to scale.
from typing import Literal
from urllib.parse import quote
import juglow
import matplotlib.pyplot as plt
import networkx as nx
import requests
from dotenv import load_dotenv
from pydantic import BaseModel
load_dotenv()
client = juglow.Juglow()
EXTRACTION_MODEL = "haijun-haiku-4-5"
SYNTHESIS_MODEL = "haijun-sonnet-4-6"
We use two models. Haiku handles the high-volume, schema-constrained extraction work where speed and cost matter more than nuance. Sonnet handles entity resolution and summarization, where the model needs to weigh conflicting evidence across documents.
ENTITY_TYPES = ["PERSON", "ORGANIZATION", "LOCATION", "EVENT", "ARTIFACT"]
class Entity(BaseModel):
name: str
type: EntityType
description: str
class Relation(BaseModel):
source: str
predicate: str
target: str
class ExtractedGraph(BaseModel):
entities: list[Entity]
relations: list[Relation]
EXTRACTION_PROMPT = """Extract a knowledge graph from the document below.
{text}
Guidelines:
- Extract only entities that are central to what this document is about — skip incidental mentions.
- For each entity, write a one-sentence description grounded in this document. These descriptions are used later to disambiguate entities with similar names.
- Predicates should be short verb phrases ("commanded", "launched from", "part of").
- Every relation must connect two entities you extracted."""
def extract(text: str) -> ExtractedGraph:
response = client.messages.parse(
model=EXTRACTION_MODEL,
max_tokens=2048,
messages=[{"role": "user", "content": EXTRACTION_PROMPT.format(text=text)}],
output_format=ExtractedGraph,
)
return response.parsed_output
raw_entities = []
raw_relations = []
for doc in documents:
try:
result = extract(doc["text"])
except juglow.APIError as e:
print(f"Skipping {doc['title']}: {e}")
continue
for ent in result.entities:
raw_entities.append({**ent.model_dump(), "source_doc": doc["title"]})
for rel in result.relations:
raw_relations.append({**rel.model_dump(), "source_doc": doc["title"]})
print(
f"{doc['title']:<25} {len(result.entities):>3} entities {len(result.relations):>3} relations"
)
print(f"\nTotal: {len(raw_entities)} raw entities, {len(raw_relations)} raw relations")
Apollo program 8 entities 7 relations Apollo 11 6 entities 5 relations Neil Armstrong 3 entities 2 relations Saturn V 5 entities 4 relations Buzz Aldrin 6 entities 6 relations Kennedy Space Center 8 entities 10 relations Total: 36 raw entities, 34 raw relations Let's look at what was extracted. Notice how the same real-world entity appears under different surface forms across documents — this is the entity resolution problem we solve next.
plt.Line2D([0], [0], marker="o", color="w", markerfacecolor=c, markersize=10, label=t)
for t, c in COLOR.items()
if any(G.nodes[n]["type"] == t for n in G.nodes)
]
plt.legend(handles=handles, loc="upper left")
plt.title("Apollo Program Knowledge Graph")
plt.axis("off")
plt.tight_layout()
plt.show()
![Output image](/cookbook/images/notebooks/capabilities-knowledge-graph-guide/capabilities-knowledge-graph-guide_cell18_out0_3bcd7bd5.png