AI TraceTrace Foundation, Inc.
Content ModerationVerified

Reviewed and published by trentmaziarz, July 23, 2026. Discovered and drafted by our automated research pipeline.

Mistral AI offers a Moderation API that classifies text inputs — including user prompts and AI-generated outputs — into nine harmful content categories, powering the safety guardrails inside Le Chat and available to third-party developers for their own applications.

Details

The Moderation API is powered by a classifier model based on Ministral 8B and provides two endpoints: one for raw text and one for conversational content. It classifies text into categories including sexual content, hate and discrimination, violence and threats, self-harm, and PII (personally identifiable information). Developers can integrate the API into their own pipelines with customizable sensitivity thresholds. The API also powers the moderation system inside Le Chat itself.

Products affected

Mistral Moderation APILe ChatMistral AI Studio

Sources & Evidence

Cite this record

Trace Foundation. (2026). Mistral AI: Mistral AI offers a Moderation API that classifies text inputs — including user prompts and AI-generated outputs — into nine harmful content categories, powering the safety guardrails inside Le Chat and available to third-party developers for their own applications (data as of 2026-07-23) [Data set record]. AI Trace. https://www.aitrace.org/r/practice/aa01625b-6b51-4416-abff-f6ff6f6eeb0e. Accessed September 7, 2026.

Stable link
https://www.aitrace.org/r/practice/aa01625b-6b51-4416-abff-f6ff6f6eeb0e
Data as of
July 23, 2026
Last verified
Not recorded

How to cite AI Trace

Other practices by Mistral AI

Have evidence about Mistral AI's AI practices? Submit a report.

Report a Sighting →