Kokoro vs XTTS v2
Choose Kokoro for preset-voice narration on modest hardware. Choose XTTS v2 when cloning a reference voice is central to your project, and check its model license before commercial use.
Editorial guidance · Reviewed September 14, 2026 · No universal winner
When to choose Kokoro-82M
Kokoro is a compact 82M model with preset voices and CPU inference. It is a practical starting point for narration, accessibility, and apps that do not need a custom speaker.
Samples, sources & setup →When to choose XTTS v2
XTTS v2 focuses on multilingual speech and voice cloning from reference audio. That makes it useful for speaker consistency and localization, but its Coqui Public Model License needs separate consideration from the code license.
Samples, sources & setup →Hear the same three scripts
Listen for how each voice handles the sentence ending and the numbers script. These recordings use different speakers, so they compare the complete output experience rather than isolating model architecture.
Neutral
“The quick brown fox jumps over the lazy dog near the riverbank.”
Bella
Reference clone
Emotional
“I can't believe you actually did it. This is incredible!”
Bella
Reference clone
Numbers & Dates
“On March 14th, 2025, the team raised $4.2 million at a 38% margin.”
Bella
Reference clone
Practical differences
| Feature | Kokoro-82M | XTTS v2 |
|---|---|---|
| License | Apache-2.0 | CPML (non-commercial) |
| Parameters | 82M | 750M |
| Voice cloning | Preset voices | Supported by the model |
| Languages | English, Japanese, Chinese, Spanish, French, Hindi, Italian, Portuguese | English, Spanish, French, German, Italian, Portuguese, Polish, Turkish, Russian, Dutch, Czech, Arabic, Chinese, Japanese, Hungarian, Korean, Hindi |
| This recording | Bella | Reference clone |
Code and weight licenses may differ. Follow the primary sources on each model page. We do not publish a speed ranking without a controlled benchmark.