Mistral AI: Mistral AI offers a Moderation API that classifies text inputs — including user prompts and AI-generated outputs — into nine harmful content categories, powering the safety guardrails inside Le Chat and available to third-party developers for their own applications. | AI Trace
Content ModerationVerified
Mistral AI offers a Moderation API that classifies text inputs — including user prompts and AI-generated outputs — into nine harmful content categories, powering the safety guardrails inside Le Chat and available to third-party developers for their own applications.
Details
The Moderation API is powered by a classifier model based on Ministral 8B and provides two endpoints: one for raw text and one for conversational content. It classifies text into categories including sexual content, hate and discrimination, violence and threats, self-harm, and PII (personally identifiable information). Developers can integrate the API into their own pipelines with customizable sensitivity thresholds. The API also powers the moderation system inside Le Chat itself.