Kokoro vs XTTS v2

Choose Kokoro for preset-voice narration on modest hardware. Choose XTTS v2 when cloning a reference voice is central to your project, and check its model license before commercial use.

Editorial guidance · Reviewed September 14, 2026 · No universal winner

When to choose Kokoro-82M

Kokoro is a compact 82M model with preset voices and CPU inference. It is a practical starting point for narration, accessibility, and apps that do not need a custom speaker.

Samples, sources & setup →

When to choose XTTS v2

XTTS v2 focuses on multilingual speech and voice cloning from reference audio. That makes it useful for speaker consistency and localization, but its Coqui Public Model License needs separate consideration from the code license.

Samples, sources & setup →

Hear the same three scripts

Listen for how each voice handles the sentence ending and the numbers script. These recordings use different speakers, so they compare the complete output experience rather than isolating model architecture.

Neutral

The quick brown fox jumps over the lazy dog near the riverbank.

Kokoro-82M

Bella

XTTS v2

Reference clone

Emotional

I can't believe you actually did it. This is incredible!

Kokoro-82M

Bella

XTTS v2

Reference clone

Numbers & Dates

On March 14th, 2025, the team raised $4.2 million at a 38% margin.

Kokoro-82M

Bella

XTTS v2

Reference clone

Try Kokoro with your own text →

Practical differences

FeatureKokoro-82MXTTS v2
LicenseApache-2.0CPML (non-commercial)
Parameters82M750M
Voice cloningPreset voicesSupported by the model
LanguagesEnglish, Japanese, Chinese, Spanish, French, Hindi, Italian, PortugueseEnglish, Spanish, French, German, Italian, Portuguese, Polish, Turkish, Russian, Dutch, Czech, Arabic, Chinese, Japanese, Hungarian, Korean, Hindi
This recordingBellaReference clone

Code and weight licenses may differ. Follow the primary sources on each model page. We do not publish a speed ranking without a controlled benchmark.

More listening comparisons