All models
flagship

Voxtral TTS

CC-BY-NC-4.0

Mistral's zero-shot cloning model, ~70ms latency across 9 languages

Voxtral TTS is Mistral AI's Ministral-3B-based speech model, cloning a voice from just 3 seconds of reference audio and streaming it back in roughly 70ms — fast enough for live agents — across nine languages. The weights are non-commercial; commercial use requires a separate agreement with Mistral.

Hear Voxtral TTS

Standardized samples are pending for this release. See the official sources below for the author’s demo.

Voxtral TTS · Reference-based · Same three scripts across the directory.

Best for: Mistral's zero-shot cloning model, ~70ms latency across 9 languages

Before you choose: Compare the samples with your own use case. Check the linked documentation for code, weight, and voice license terms.

Details awaiting a fresh review. Hardware figures are estimates unless a benchmark is linked.

No samples yet

Every model in this directory is read against the same three scripts so voices can be compared honestly — Voxtral TTS’s samples just haven’t been generated yet.

Contribute samples on GitHub →

Run Voxtral TTS

pip install mistral-inference

Sources & setup details

Voxtral TTS FAQ

Do I need a GPU to run Voxtral TTS?

The directory lists roughly 16 GB of VRAM as a historical estimate. Check the model documentation for your checkpoint and runtime.

Can Voxtral TTS clone voices?

Yes — Voxtral TTS supports voice cloning from reference audio.

What license is Voxtral TTS released under?

Voxtral TTS is released under the CC-BY-NC-4.0 license. Always check the repository for the exact terms — some models license code and weights separately.

What languages does Voxtral TTS support?

Voxtral TTS supports English, French, Spanish, German, Italian, Portuguese, Dutch, Arabic, Hindi.

Is there a hosted Voxtral TTS API?

OpenSpeech Cloud is accepting production-access interest for Voxtral TTS, but this model is not hosted in the beta. Use the official repository for current deployment options.