Muyan-TTS
Apache-2.0Full training code released, trained for ~$50K on podcast audio
Muyan-TTS from MYZY-AI is a LLaMA-3.2-3B-based pipeline trained on more than 100,000 hours of podcast audio for a total training cost of around $50,000. Unusually for open TTS, MYZY-AI released the full training code, not just inference — letting anyone reproduce or fine-tune the pipeline from scratch.
Hear Muyan-TTS
Standardized samples are pending for this release. See the official sources below for the author’s demo.
Muyan-TTS · Reference-based · Same three scripts across the directory.
Best for: Full training code released, trained for ~$50K on podcast audio
Before you choose: Compare the samples with your own use case. Check the linked documentation for code, weight, and voice license terms.
Details awaiting a fresh review. Hardware figures are estimates unless a benchmark is linked.
Compare with similar
Compare all →Choose an alternative to hear both models together.
No samples yet
Every model in this directory is read against the same three scripts so voices can be compared honestly — Muyan-TTS’s samples just haven’t been generated yet.
Contribute samples on GitHub →Run Muyan-TTS
git clone https://github.com/MYZY-AI/Muyan-TTSSources & setup details
Muyan-TTS FAQ
Do I need a GPU to run Muyan-TTS?
The directory lists roughly 10 GB of VRAM as a historical estimate. Check the model documentation for your checkpoint and runtime.
Can Muyan-TTS clone voices?
Yes — Muyan-TTS supports voice cloning from reference audio.
What license is Muyan-TTS released under?
Muyan-TTS is released under the Apache-2.0 license. Always check the repository for the exact terms — some models license code and weights separately.
What languages does Muyan-TTS support?
Muyan-TTS supports English.
Is there a hosted Muyan-TTS API?
OpenSpeech Cloud is accepting production-access interest for Muyan-TTS, but this model is not hosted in the beta. Use the official repository for current deployment options.