All models
long-form

Muyan-TTS

Apache-2.0

Full training code released, trained for ~$50K on podcast audio

Muyan-TTS from MYZY-AI is a LLaMA-3.2-3B-based pipeline trained on more than 100,000 hours of podcast audio for a total training cost of around $50,000. Unusually for open TTS, MYZY-AI released the full training code, not just inference — letting anyone reproduce or fine-tune the pipeline from scratch.

Hear Muyan-TTS

Standardized samples are pending for this release. See the official sources below for the author’s demo.

Muyan-TTS · Reference-based · Same three scripts across the directory.

Best for: Full training code released, trained for ~$50K on podcast audio

Before you choose: Compare the samples with your own use case. Check the linked documentation for code, weight, and voice license terms.

Details awaiting a fresh review. Hardware figures are estimates unless a benchmark is linked.

No samples yet

Every model in this directory is read against the same three scripts so voices can be compared honestly — Muyan-TTS’s samples just haven’t been generated yet.

Contribute samples on GitHub →

Run Muyan-TTS

git clone https://github.com/MYZY-AI/Muyan-TTS

Sources & setup details

Muyan-TTS FAQ

Do I need a GPU to run Muyan-TTS?

The directory lists roughly 10 GB of VRAM as a historical estimate. Check the model documentation for your checkpoint and runtime.

Can Muyan-TTS clone voices?

Yes — Muyan-TTS supports voice cloning from reference audio.

What license is Muyan-TTS released under?

Muyan-TTS is released under the Apache-2.0 license. Always check the repository for the exact terms — some models license code and weights separately.

What languages does Muyan-TTS support?

Muyan-TTS supports English.

Is there a hosted Muyan-TTS API?

OpenSpeech Cloud is accepting production-access interest for Muyan-TTS, but this model is not hosted in the beta. Use the official repository for current deployment options.