Measured September 20, 2026 · Technical evaluation

Open speech: listen to the evidence

Six engines, the same original English scripts. These recordings show what our tested settings produced. Human listening responses and actual invoice reconciliation remain pending.

We completed 12 hosted runs and nine local runs. This small sample cannot establish a failure rate or an overall quality winner.

Help evaluate the recordings

Score six clips with model names hidden until submission. Team and independent participant responses are kept separate; independence is self-reported.

Start the listening review →

Compare the same 238-character paragraph

The morning train leaves at nine fifteen. Bring your ticket and a light jacket. Our guide, Elena, will meet you beside the station clock. After a short walk, we will visit the old market and stop for lunch. This is the end of the passage.

Kokoro

Replicate · 15.50 seconds of audio

Observed request-to-result time: 1.19 seconds.

Estimated inference: $0.000115 for this run. This is the provider estimate, not the retail recording price.

Model and license details

Chatterbox Turbo

Replicate · 14.72 seconds of audio

Observed request-to-result time: 12.29 seconds.

Estimated inference: $0.005950 for this run. This is the provider estimate, not the retail recording price.

Model and license details

Qwen3-TTS

Replicate · 17.44 seconds of audio

Observed request-to-result time: 16.40 seconds.

Estimated inference: $0.004760 for this run. This is the provider estimate, not the retail recording price.

Model and license details

Soprano 1.1

Local CPU · 12.86 seconds of audio

Generation after loading: 2.75 seconds. No cloud cost measured.

ekwek/Soprano-1.1-80M · macOS ARM64 · four PyTorch threads. These timings use a different environment and timing boundary from the hosted runs.

Pocket TTS

Local CPU · 13.60 seconds of audio

Generation after loading: 2.52 seconds. No cloud cost measured.

pocket-tts english_2026-01 · macOS ARM64 · four PyTorch threads. These timings use a different environment and timing boundary from the hosted runs.

VoxCPM2

Local MPS · 17.92 seconds of audio

Generation after loading: 30.31 seconds. No cloud cost measured.

openbmb/VoxCPM2 · macOS ARM64 · four PyTorch threads. These timings use a different environment and timing boundary from the hosted runs.

How these measurements were made

The hosted tests pin the adapter version and record every input setting, output duration, byte count and SHA-256 hash. Three scripts contain 105, 238 and 478 characters; Kokoro has additional repeat and 1,911-character checks. Audio decoding confirms a playable file, not that every word was pronounced correctly.

Kokoro estimates multiply observed provider prediction time by the published $0.000225-per-second T4 rate. Qwen3 uses $0.02 per 1,000 input characters and Chatterbox Turbo uses $0.025. These estimates exclude other operating costs and have not been matched to an invoice. Sources: Kokoro adapter, Qwen3 pricing, Chatterbox pricing.

Local tests use PyTorch 2.14.0 on macOS ARM64, four CPU threads and seed 42. VoxCPM2 uses Apple MPS with float32, guidance 2.0 and ten inference steps, without a reference voice, denoiser or compilation. Its first measured generation includes initial device warm-up. Soprano uses its Transformers backend; Pocket uses the original english_2026-01 checkpoint and Alba prompt. Loading and downloads are recorded separately. None of these local candidates is currently offered in the paid studio.

Pocket TTS: Kyutai weights, CC-BY-4.0, MIT code and voice documentation. Voice reference: Alba MacKenna, CC-BY-4.0. These are new synthesized recordings of our original scripts, not the original voice recording or an endorsement.

Download all settings, measurements, hashes and sample paths

Automated speech screening

We transcribed all 21 recordings with faster-whisper small.en and compared normalized words against the scripts. All decoded samples contain finite audio values. Two short clips triggered our transcript-difference threshold: Kokoro and VoxCPM2. Time formatting explains part of the difference; VoxCPM2’s transcript also renders “appointment” as “point.” Listening is needed to distinguish synthesis errors from recognition errors.

This is a screening aid, not a human pronunciation score or a quality ranking. Review the original audio before drawing conclusions.

Download automatic transcripts, signal checks and limitations

Try the economical hosted workflow

Kokoro costs $0.02 per 1,000 characters with a $0.01 minimum per paragraph and whole-cent rounding. Choose one of eight voices, review your script’s total and export an MP3 or WAV. Credit packs start at $5.

Review a voiceover script →

Compare ElevenLabs alternatives · Hear all eight Kokoro voices