Productivity AutomationVerified
Reviewed and published by trentmaziarz, April 22, 2026. Discovered and drafted by our automated research pipeline.
ElevenLabs offers Scribe, an AI speech-to-text model that converts audio and video recordings into accurate, structured text transcripts in 99 languages, with features including speaker labeling, word-level timestamps, and audio-event tagging. Scribe launched in December 2024, with Scribe v2 following in January 2026.
Details
Scribe takes audio or video files as input and produces structured JSON transcripts as output, including speaker diarization (identifying who said what), character-level timestamps, and tagging of non-speech audio events such as laughter or applause. A real-time version, Scribe v2 Realtime, processes live speech with approximately 150 milliseconds of latency and is designed for use in conversational AI agents and meeting assistants. The tool is available via the web dashboard and API.
Products affected
ScribeScribe v2Scribe v2 RealtimeElevenLabs StudioElevenLabs API
Sources & Evidence
Cite this record
Trace Foundation. (2026). ElevenLabs: ElevenLabs offers Scribe, an AI speech-to-text model that converts audio and video recordings into accurate, structured text transcripts in 99 languages, with features including speaker labeling, word-level timestamps, and audio-event tagging. Scribe launched in December 2024, with Scribe v2 following in January 2026 (data as of 2026-04-22) [Data set record]. AI Trace. https://www.aitrace.org/r/practice/8c0319e4-a5eb-458e-8f31-bc6e1b6e5f05. Accessed September 11, 2026.
- Stable link
- https://www.aitrace.org/r/practice/8c0319e4-a5eb-458e-8f31-bc6e1b6e5f05
- Data as of
- April 22, 2026
- Last verified
- August 4, 2026
Other practices by ElevenLabs
Creative GenElevenLabs offers Voice Design, a tool launched in February 2023 that generates entirely new AI voices from a text description, allowing users to create a voice that has no real-world equivalent without uploading any audio recordings. The tool is available to all registered users, including those on the free tier, and produces three unique voice options per prompt for preview and selection.Creative GenElevenLabs offers a text-to-sound-effects generator that produces custom sound effects, soundscapes, and ambient audio from natural language text prompts. The tool was first released in May 2024, with a second-generation model (SFX V2) launched in September 2025.Creative GenElevenLabs offers an AI text-to-speech platform that converts written text into natural-sounding spoken audio across 70+ languages using deep learning voice models. Users can generate speech using pre-built voices, cloned voices, or custom-designed voices via a web platform or API.
Have evidence about ElevenLabs's AI practices? Submit a report.
Report a Sighting →