SoulX-Podcast
Apache-2.0Multi-speaker podcast dialogue with dialect support
SoulX-Podcast, built on Qwen3 by Soul AI Lab, generates multi-speaker, multi-turn podcast dialogue in a single pass, complete with paralinguistics like laughter. It supports cross-dialect voice cloning across Sichuanese, Henanese, and Cantonese alongside standard Mandarin and English.
Hear SoulX-Podcast
Standardized samples are pending for this release. See the official sources below for the author’s demo.
SoulX-Podcast · Speaker 1 · Same three scripts across the directory.
Best for: Multi-speaker podcast dialogue with dialect support
Before you choose: Compare the samples with your own use case. Check the linked documentation for code, weight, and voice license terms.
Details awaiting a fresh review. Hardware figures are estimates unless a benchmark is linked.
Compare with similar
Compare all →Choose an alternative to hear both models together.
No samples yet
Every model in this directory is read against the same three scripts so voices can be compared honestly — SoulX-Podcast’s samples just haven’t been generated yet.
Contribute samples on GitHub →Run SoulX-Podcast
git clone https://github.com/Soul-AILab/SoulX-PodcastSources & setup details
SoulX-Podcast FAQ
Do I need a GPU to run SoulX-Podcast?
The directory lists roughly 8 GB of VRAM as a historical estimate. Check the model documentation for your checkpoint and runtime.
Can SoulX-Podcast clone voices?
Yes — SoulX-Podcast supports voice cloning from reference audio.
What license is SoulX-Podcast released under?
SoulX-Podcast is released under the Apache-2.0 license. Always check the repository for the exact terms — some models license code and weights separately.
What languages does SoulX-Podcast support?
SoulX-Podcast supports Chinese, English.
Is there a hosted SoulX-Podcast API?
OpenSpeech Cloud is accepting production-access interest for SoulX-Podcast, but this model is not hosted in the beta. Use the official repository for current deployment options.