AI TraceTrace Foundation, Inc.
Creative GenerationVerified

Reviewed and published by trentmaziarz, July 23, 2026. Discovered and drafted by our automated research pipeline.

Mistral AI deployed Voxtral TTS, a text-to-speech model that converts written text into spoken audio in nine languages, including support for voice cloning from a short audio reference clip, released in March 2026.

Details

Voxtral TTS is a 4-billion-parameter model available via Mistral's API and Mistral Studio, as well as as open weights on Hugging Face. It supports 20 built-in preset voices and can clone a specific speaker's voice from a 3–25 second audio reference clip. The model is designed for low-latency streaming inference, with a reported 70ms latency, making it suitable for real-time voice agents and translation workflows.

Products affected

Voxtral TTSMistral AI Studio

Sources & Evidence

Cite this record

Trace Foundation. (2026). Mistral AI: Mistral AI deployed Voxtral TTS, a text-to-speech model that converts written text into spoken audio in nine languages, including support for voice cloning from a short audio reference clip, released in March 2026 (data as of 2026-07-23) [Data set record]. AI Trace. https://www.aitrace.org/r/practice/316a4dad-c90f-4e27-9a0e-4892c778a280. Accessed September 7, 2026.

Stable link
https://www.aitrace.org/r/practice/316a4dad-c90f-4e27-9a0e-4892c778a280
Data as of
July 23, 2026
Last verified
Not recorded

How to cite AI Trace

Other practices by Mistral AI

Have evidence about Mistral AI's AI practices? Submit a report.

Report a Sighting →