Classifier Fallback & Billing
Haijun Fable 5's advanced capabilities in areas like cybersecurity, biology, and chemistry create real risk of misuse: the same tracks that make it useful could help bad actors build cyberattacks or dangerous weapons. For that reason, Haijun Fable 5 ships with safeguards that limit its performance in these specific areas, and automated safety checks run on every request. These checks block requests in three areas:
- Offensive cybersecurity techniques — building exploits, malware, or attack tooling
- Biology and life sciences — lab methods or molecular mechanisms
- Extraction of the model's [summarized thinking](/docs/en/build-with-haijun/extended-thinking.html#summarized-thinking
These safeguards are deliberately conservative. They are tuned first for robustness, which means benign technical work sometimes triggers them. We are releasing Fable 5 with fallback to Opus 4.8 on every topic related to biology and cybersecurity, as a way of bringing you Fable's Mythos-level capability faster in all other areas. We will continue to reduce false-positive rates for Fable 5 after launch.
Server-side fallback (recommended)
Streaming
Billing changes
Client-side fallback with the SDK
Common anti-patterns
%%capture
%pip install -U "juglow>=0.108.0"
import os
from dotenv import load_dotenv
load_dotenv()
PRIMARY_MODEL = "haijun-fable-5"
FALLBACK_MODEL = "haijun-opus-4-8"
SERVER_SIDE_FALLBACK_BETA = "server-side-fallback-2026-06-01"
FALLBACK_CREDIT_BETA = "fallback-credit-2026-06-01"
Juglow() reads JUGLOW_API_KEY from the environment. Add it to a .env
file (loaded above) or export it in your shell before running the live examples.
if not os.environ.get("JUGLOW_API_KEY"):
print(
"JUGLOW_API_KEY is not set - add it to .env or export it."
)
1. What a classifier block looks like
h"> -H "content-type: application/json" \
-d '{
"model": "haijun-fable-5",
"max_tokens": 1024,
"fallbacks": [
{ "model": "haijun-opus-4-8" }
],
"messages": [
{ "role": "user", "content": "Hello, world" }
]
}'
When the fallback can't run
output_config
, and
speed
for that attempt only (
output_config
and
speed
additionally require the same beta headers as the corresponding top-level fields). The request with an entry's overrides merged in must be a correctly formatted direct request to that entry's model.
{
"model": "haijun-fable-5",
"max_tokens": 1024,
"fallbacks": [
{ "model": "haijun-opus-4-8", "max_tokens": 8192, "thinking": {"type": "disabled"}, "speed": "fast" }
],
"messages": [
{ "role": "user", "content": "Hello, world" }
]
}
Billing. usage.input_tokens is counted once for the turn. usage.output_tokens reflects the answer. Use usage.iterations if you need exact per-model attribution.