This cookbook shows how to bring
MongoDB Atlas
to an agent running on
Haijun Managed Agents
(CMA) — as its retrieval engine, its graph store, and its system of record — using only the standard CMA patterns (custom tools and MCP toolsets), with
no platform-level MongoDB integration required
. The runnable helpers live beside this notebook in
mongodb_on_cma/
. Each section links the full file so the narrative stays focused on
where MongoDB plugs in
and
where the agent adds value
.
Where MongoDB Plugs In: Most agent stacks bolt together three or four systems: a vector database for semantic search, a separate engine for keywords, a graph store for relationships, and an operational database for the records themselves. Every seam is another integration, another credential, another place the agent's view of the world can drift. MongoDB collapses that into one engine: the same documents are searchable by vector ($vectorSearch), by full-text ($search), as a hybrid of the two fused with reciprocal rank fusion, or RRF ($rankFusion), and traversable as a graph ($graphLookup) — and they are the same documents the agent reads, writes, and persists its decisions to. One query language, one cluster, one connection (pymongo).
Lift vector, full-text, hybrid, and graph retrieval into an agent's custom tools.
Gate risky agent decisions behind CMA's native human-in-the-loop pause.
Make one MongoDB Atlas cluster the agent's system of record and audit backbone.
The worked example is a human-in-the-loop fraud-review agent, but the patterns are vertical-agnostic: swap the collection and the tools, and the same shape serves a support, research, or operations agent.
ry">
Setup
Required: MONGO_URI (an Atlas SRV connection string) plus Juglow API access. A free M0 cluster runs everything in this cookbook — vector, full-text, hybrid $rankFusion, and graph traversal — so there is no paid-tier requirement. Hybrid search uses the native $rankFusion stage, which needs MongoDB 8.0+; every current Atlas cluster (M0 included) is on 8.0 or later, so this holds by default.
"relying on the SDK / ant CLI to resolve credentials."
)
MODEL = resolve_model()
ENABLE_RERANK = os.getenv("ENABLE_RERANK", "").lower() in ("1", "true")
AUTO_APPROVE = os.getenv("AUTO_APPROVE", "").lower() in ("1", "true")
Quiet the SDK's one-shot notice that JUGLOW_API_KEY shadows profile/federation
auto-discovery — expected when an API key is set; harmless under ant / WIF auth.
logging.getLogger("juglow.lib.credentials._auth").setLevel(logging.ERROR)
client = Juglow()
mongo = MongoClient(os.environ["MONGO_URI"])
db = mongo["fraud_review_demo"]
coll = db["transactions"]
ai_client = make_embedding_client() # Atlas / Voyage / None (seed ships precomputed vectors)
print(
f"model={MODEL} rerank={'on' if ENABLE_RERANK else 'off'} provider={'yes' if ai_client else 'none'}"
)
warning: no JUGLOW_API_KEY / JUGLOW_AUTH_TOKEN / JUGLOW_PROFILE found; relying on the SDK / ant CLI to resolve credentials. model=haijun-haiku-4-5 rerank=on provider=yes 1. Connect MongoDB to a managed agent The agent loop and its sandbox are Juglow-hosted — that is what "Managed" means. The only thing you host is the data path to MongoDB. On every path, the MongoDB credential lives on your side of the boundary: it never enters the agent context, a cloud sandbox's environment, or a file the agent can read.
/td>
$vectorSearch
Full-text search
build_lexical_pipeline
$search
Hybrid (reciprocal rank fusion)
build_rank_fusion_pipeline
$rankFusion
(8.0+)
Graph traversal
build_graph_pipeline
$graphLookup
The builders are defined inline in the next cell — each returns a list of aggregation stages you can lift into your own collection. Index names and the projected fields come from config.py; EMBED_DIM and the index constants are imported in the setup cell above.
d)]
if enable_rerank and reranker is not None and candidates:
results = reranker(query, [c.get("text", "") for c in candidates], top_k=k)
candidates = merge_rerank(candidates, results, top_k=k)
else:
candidates = candidates[:k]
return {"similar": [_pick_fields(c) for c in candidates]}
def tool_detect_fraud_ring(coll, account_id, *, max_depth=4) -> dict:
pipeline = build_graph_pipeline(account_id, max_depth=max_depth, collection=coll.name)
docs = list(coll.aggregate(pipeline))
return _jsonable(summarize_ring(docs[0] if docs else {"chain": []}, seed_account=account_id))
def tool_record_decision(
db,
transaction_id,
decision,
*,
confidence,
risk_factors,
reasoning,
reviewed_by,
escalated=False,
recommended_decision=None,
) -> dict:
decision_doc = build_decision_doc(
transaction_id,
decision,
confidence=confidence,
risk_factors=risk_factors,
reasoning=reasoning,
reviewed_by=reviewed_by,
)
db["transaction_decisions"].insert_one(decision_doc)
if escalated:
audit = build_audit_event(
"escalated_to_human",
transaction_id,
decision_id=decision_doc["decision_id"],
severity="warning",
event_data={"human_decision": decision, "recommended_decision": recommended_decision},
)
else:
audit = build_audit_event(
"decision_stored", transaction_id, decision_id=decision_doc["decision_id"]
)
db["audit_events"].insert_one(audit)
Advance the lifecycle status to past tense (approve -> approved) so a decided case matches
the seed's vocabulary and DECIDED_STATUSES, making it eligible as precedent in later reviews.
status = {"approve": "approved", "reject": "rejected"}.get(decision, decision)
db["transactions"].update_one({"transaction_id": transaction_id}, {"$set": {"status": status}})
return {"recorded": True, "decision_id": decision_doc["decision_id"]}
TOOLS = [
{
"type": "custom",
"name": "verify_mandates",
"description": "Validate the AP2 Checkout and Payment Mandate JWTs (signature, constraints, "
"double-spend). Run this FIRST. If valid=false, constraints_satisfied=false, or "
"double_spend_detected=true, reject immediately.",
"input_schema": {
"type": "object",
"properties": {"transaction_id": {"type": "string"}},
"required": ["transaction_id"],
},
},
{
"type": "custom",
"name": "get_transaction",
"description": "Fetch the full transaction record under review.",
"input_schema": {
"type": "object",
"properties": {"transaction_id": {"type": "string"}},
"required": ["transaction_id"],
},
},
{
"type": "custom",
"name": "hybrid_search_similar_frauds",
"description": "Retrieve the most similar prior (already-decided) cases as precedent, using "
"hybrid vector + full-text search.",
"input_schema": {
"type": "object",
"properties": {
"transaction_id": {"type": "string"},
"k": {"type": "integer", "description": "how many precedents (default 5)"},
},
"required": ["transaction_id"],
},
},
{
"type": "custom",
"name": "detect_fraud_ring",
"description": "Trace the account's sender->recipient chain for circular-flow / mule / layering patterns.",
"input_schema": {
"type": "object",
"properties": {"account_id": {"type": "string"}},
"required": ["account_id"],
},
},
{
"type": "custom",
"name": "record_decision",
"description": "Persist the final approve/reject decision with reasoning and an audit event.",
"input_schema": {
"type": "object",
"properties": {
"transaction_id": {"type": "string"},
"decision": {"type": "string", "enum": ["approve", "reject"]},
"confidence": {"type": "number"},
"risk_factors": {"type": "array", "items": {"type": "string"}},
"reasoning": {"type": "string"},
"escalated": {"type": "boolean"},
"recommended_decision": {"type": "string", "enum": ["approve", "reject"]},
},
"required": ["transaction_id", "decision", "reasoning"],
},
},
{
"type": "custom",
"name": "escalate",
"description": "Send a risky case to a human reviewer for the final decision. Use for "
"medium-confidence, high-value, structuring, or fraud-ring cases.",
"input_schema": {
"type": "object",
"properties": {
"transaction_id": {"type": "string"},
"recommended_decision": {"type": "string", "enum": ["approve", "reject"]},
"confidence": {"type": "number"},
"reason": {"type": "string"},
},
"required": ["transaction_id", "recommended_decision", "reason"],
},
},
]
Host-side handlers — map each data tool to its handler function.
reranker = (
(lambda q, ds, top_k: rerank(q, ds, client=ai_client, top_k=top_k))
if (ENABLE_RERANK and ai_client)
else None
)
def _verify_mandates(inp):
result = tool_verify_mandates(db, inp["transaction_id"], ts_public_key)
ok = result["valid"] and result["constraints_satisfied"] and not result["double_spend_detected"]
print(f"\n [verify_mandates] {inp['transaction_id']}: {'pass' if ok else 'FAIL'}")
return result
def _hybrid(inp):
return tool_hybrid_search_similar_frauds(
coll,
inp["transaction_id"],
inp.get("k", 5),
enable_rerank=ENABLE_RERANK,
reranker=reranker,
)
def _record(inp):
result = tool_record_decision(
db,
inp["transaction_id"],
inp["decision"],
confidence=inp.get("confidence", 0),
risk_factors=inp.get("risk_factors", []),
reasoning=inp.get("reasoning", ""),
reviewed_by="human" if inp.get("escalated") else "agent",
escalated=inp.get("escalated", False),
recommended_decision=inp.get("recommended_decision"),
)
if (
inp["decision"] == "approve"
): # store the AP2 receipt that later powers double-spend detection
txn_doc = coll.find_one(
{"transaction_id": inp["transaction_id"]},
{"checkout_mandate_jwt": 1, "mandate_id": 1, "agent_pk": 1},
)
if txn_doc and txn_doc.get("mandate_id"):
checkout_hash = hashlib.sha256(txn_doc["checkout_mandate_jwt"].encode()).hexdigest()
store_mandate_receipt(
db, txn_doc["mandate_id"], txn_doc["agent_pk"], checkout_hash, "approve"
)
return result
HANDLERS = {
"verify_mandates": _verify_mandates,
"get_transaction": lambda inp: tool_get_transaction(coll, inp["transaction_id"]),
"hybrid_search_similar_frauds": _hybrid,
"detect_fraud_ring": lambda inp: tool_detect_fraud_ring(coll, inp["account_id"]),
"record_decision": _record,
}