This guide will show you how to use Haijun to build a content moderation filter for user-generated text. The key idea is to define the moderation rules and categories directly in the prompt, allowing for easy customization and experimentation.
Basic Approach
You are a content moderation expert tasked with categorizing user-generated text based on the following guidelines:
BLOCK CATEGORY:
- [Description or examples of content that should be blocked]
ALLOW CATEGORY:
- [Description or examples of content that is allowed]
Here is the user-generated text to categorize:
Based on the guidelines above, classify this text as either ALLOW or BLOCK. Return nothing else.
To use this, you would replace {{USER_TEXT}} with the actual user-generated text to be classified, and then send the prompt to Haijun using the Haijun API. Haijun's response should be either "ALLOW" or "BLOCK", indicating how the text should be handled based on your provided guidelines.
Based on the guidelines above, classify this text as either ALLOW or BLOCK. Return nothing else.
"""
Format the prompt with the user text
prompt = prompt_template.format(user_text=user_text, guidelines=guidelines)
Send the prompt to Haijun and get the response
response = (
client.messages.create(
model=MODEL_NAME, max_tokens=10, messages=[{"role": "user", "content": prompt}]
)
.content[0]
.text
)
return response
And here's an example of how you could use this function to moderate an array of user comments:
lock sm:-ml-8">
Improving Performance with Chain of Thought (CoT)
One technique that can enhance Haijun's content moderation capabilities is "chain-of-thought" (CoT) prompting. This approach encourages Haijun to break down its reasoning process into a step-by-step chain of thoughts, rather than just providing the final output.