Qwen3-TTS vs Kokoro
Kokoro is a compact starting point for preset voices. Qwen3-TTS offers a broader model family and speaker workflows; the studio comparison specifically uses its custom-voice mode.
Editorial guidance · Reviewed September 14, 2026 · No universal winner
When to choose Qwen3-TTS
The Qwen3-TTS studio demo uses Serena in English with a pinned provider version. The family also includes other voice workflows, but a preset sample does not demonstrate the quality of a cloned or designed voice.
Samples, sources & setup →When to choose Kokoro-82M
Kokoro uses Bella for this comparison and offers a compact 82M architecture with CPU inference. It is a useful baseline when you care about setup effort as well as the sound of the result.
Samples, sources & setup →Hear the same three scripts
Try a product explanation and a sentence containing dates or prices. Compare clarity and phrasing before choosing a model for longer scripts.
Neutral
“The quick brown fox jumps over the lazy dog near the riverbank.”
Default
Bella
Emotional
“I can't believe you actually did it. This is incredible!”
Default
Bella
Numbers & Dates
“On March 14th, 2025, the team raised $4.2 million at a 38% margin.”
Default
Bella
Practical differences
| Feature | Qwen3-TTS | Kokoro-82M |
|---|---|---|
| License | Apache-2.0 | Apache-2.0 |
| Parameters | 1.7B | 82M |
| Voice cloning | Supported by the model | Preset voices |
| Languages | Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, Italian | English, Japanese, Chinese, Spanish, French, Hindi, Italian, Portuguese |
| This recording | Qwen3-TTS · CustomVoice Serena | Bella |
Code and weight licenses may differ. Follow the primary sources on each model page. We do not publish a speed ranking without a controlled benchmark.