Step-Audio-EditX
Apache-2.0Re-edit the emotion or style of audio you already generated
Step-Audio-EditX from StepFun AI is built around iterative editing rather than one-shot generation: you can take audio you already made and regenerate just the emotion, style, or a paralinguistic detail (breathing, laughter) without touching the words or re-cloning the voice. It tops recent open-model Elo rankings for emotion control, and also handles Cantonese and Sichuanese dialects alongside standard Mandarin and English.
Hear Step-Audio-EditX
Standardized samples are pending for this release. See the official sources below for the author’s demo.
Step-Audio-EditX · Reference-based · Same three scripts across the directory.
Best for: Re-edit the emotion or style of audio you already generated
Before you choose: Compare the samples with your own use case. Check the linked documentation for code, weight, and voice license terms.
Details awaiting a fresh review. Hardware figures are estimates unless a benchmark is linked.
Compare with similar
Compare all →Choose an alternative to hear both models together.
No samples yet
Every model in this directory is read against the same three scripts so voices can be compared honestly — Step-Audio-EditX’s samples just haven’t been generated yet.
Contribute samples on GitHub →Run Step-Audio-EditX
git clone https://github.com/stepfun-ai/Step-Audio-EditXSources & setup details
Step-Audio-EditX FAQ
Do I need a GPU to run Step-Audio-EditX?
The directory lists roughly 12 GB of VRAM as a historical estimate. Check the model documentation for your checkpoint and runtime.
Can Step-Audio-EditX clone voices?
Yes — Step-Audio-EditX supports voice cloning from reference audio.
What license is Step-Audio-EditX released under?
Step-Audio-EditX is released under the Apache-2.0 license. Always check the repository for the exact terms — some models license code and weights separately.
What languages does Step-Audio-EditX support?
Step-Audio-EditX supports English, Chinese.
Is there a hosted Step-Audio-EditX API?
OpenSpeech Cloud is accepting production-access interest for Step-Audio-EditX, but this model is not hosted in the beta. Use the official repository for current deployment options.