Content ModerationConsumer FacingVerified
Reviewed and published by trentmaziarz, April 8, 2026. Discovered and drafted by our automated research pipeline.
Anthropic trains Claude to refuse harmful requests using a method called Constitutional AI, which teaches the model to evaluate its own responses against a written set of principles — like giving an AI a rulebook it must check before answering. A separate layer of automated filters screens conversations in real time, and a trust and safety team enforces usage policies that banned 1.45 million accounts in the second half of 2025.
Details
Constitutional AI, published as a research paper in December 2022, is the foundational method Anthropic uses to align Claude's behavior. Instead of relying solely on human reviewers to label harmful outputs, the technique uses AI feedback guided by an explicit set of principles to shape how the model responds. In January 2025, Anthropic published Constitutional Classifiers — a separate filtering layer that withstood more than 3,000 hours of adversarial testing without a universal bypass being found. Anthropic's Transparency Hub reports that in July through December 2025, the company banned 1.45 million accounts, processed 52,000 appeals (overturning 1,700), and reported 5,005 pieces of content to the National Center for Missing and Exploited Children. Usage policies in effect since September 2025 prohibit 14 categories of harmful use and require human-in-the-loop oversight for high-risk applications in healthcare, law, finance, and journalism. Developers using the Claude API receive access to real-time moderation tools at no additional cost.
Products affected
All Claude products
Sources & Evidence
Company Disclosure
Other practices by Anthropic
OtherAnthropic has partnered with scientific and government institutions to deploy Claude in research settings. In January 2026, Claude helped guide NASA's Perseverance rover to travel 400 meters across Mars — the first time an AI assistant helped navigate a spacecraft on another planet. Partnerships with major biomedical research institutions are also underway to use Claude in laboratory and computational research.OtherAnthropic developed a specialized AI model called Claude Mythos to help find dangerous security flaws in software before malicious actors do. Because the model is powerful enough to create its own exploits, it is not available to the public — access is restricted to approximately 40 vetted organizations, including major technology and financial companies, for defensive use only.OtherAnthropic operates a formal safety framework called the Responsible Scaling Policy that sets rules for when and how it can train and release more powerful AI models. Under this policy, each new Claude model is assigned a safety level, and passing specific safety tests is required before the model can be deployed. The framework is now in its third version and has been updated as Claude's capabilities have grown.
Have evidence about Anthropic's AI practices? Submit a report.
Report a Sighting →