OtherInternal OnlyVerified
Reviewed and published by trentmaziarz, April 8, 2026. Discovered and drafted by our automated research pipeline.
Anthropic operates a formal safety framework called the Responsible Scaling Policy that sets rules for when and how it can train and release more powerful AI models. Under this policy, each new Claude model is assigned a safety level, and passing specific safety tests is required before the model can be deployed. The framework is now in its third version and has been updated as Claude's capabilities have grown.
Details
The Responsible Scaling Policy was first published in September 2023 and has been revised to Version 3.0, released in February 2026. It classifies AI models on a scale from ASL-1 (no meaningful risk) through ASL-4 (requiring the most stringent safeguards). Claude Opus 4 became the first model classified as ASL-3 in May 2025, triggering a set of enhanced deployment safeguards including Constitutional Classifiers and hardened infrastructure security. Anthropic publishes detailed model system cards — technical documents summarizing safety evaluations — for each major release, running up to 244 pages. Version 3.0 introduced public Frontier Safety Roadmaps with stated safety goals and Risk Reports, but also drew criticism for removing a prior commitment to pause model training if adequate safety measures could not be confirmed. Approximately 8% of all Anthropic employees work on security-related areas, and a named Responsible Scaling Officer (Jared Kaplan) holds accountability for the policy.
Products affected
All Claude models
Sources & Evidence
Other practices by Anthropic
OtherAnthropic has partnered with scientific and government institutions to deploy Claude in research settings. In January 2026, Claude helped guide NASA's Perseverance rover to travel 400 meters across Mars — the first time an AI assistant helped navigate a spacecraft on another planet. Partnerships with major biomedical research institutions are also underway to use Claude in laboratory and computational research.OtherAnthropic developed a specialized AI model called Claude Mythos to help find dangerous security flaws in software before malicious actors do. Because the model is powerful enough to create its own exploits, it is not available to the public — access is restricted to approximately 40 vetted organizations, including major technology and financial companies, for defensive use only.OtherAnthropic created and open-sourced the Model Context Protocol (MCP), a technical standard that gives AI models a consistent way to connect to external tools, databases, and applications — similar to how USB-C gives devices a universal port for charging and data transfer. MCP has been adopted by major technology companies including Google, Microsoft, and OpenAI, and reaches over 100 million monthly downloads.
Have evidence about Anthropic's AI practices? Submit a report.
Report a Sighting →