Mistral AI: Mistral AI deployed Voxtral TTS, a text-to-speech model that converts written text into spoken audio in nine languages, including support for voice cloning from a short audio reference clip, released in March 2026. | AI Trace
Creative GenerationVerified
Mistral AI deployed Voxtral TTS, a text-to-speech model that converts written text into spoken audio in nine languages, including support for voice cloning from a short audio reference clip, released in March 2026.
Details
Voxtral TTS is a 4-billion-parameter model available via Mistral's API and Mistral Studio, as well as as open weights on Hugging Face. It supports 20 built-in preset voices and can clone a specific speaker's voice from a 3–25 second audio reference clip. The model is designed for low-latency streaming inference, with a reported 70ms latency, making it suitable for real-time voice agents and translation workflows.