Details
The Voxtral family includes Voxtral Small (24B parameters) for production-scale applications, Voxtral Mini (3–5B parameters) for edge and local deployment, and Voxtral Transcribe 2 (released February 2026) adding speaker diarization, context biasing, and word-level timestamps. Voxtral Realtime supports live transcription with latency as low as 200ms. Models support audio up to 40 minutes in length, are available via Mistral's API and on Hugging Face, and are released under Apache 2.0. The models also support question-answering and summarization directly from audio without chaining to a separate language model.
Have evidence about Mistral AI's AI practices? Submit a report.
Report a Sighting →