All models
expressive

Step-Audio-EditX

Apache-2.0

Re-edit the emotion or style of audio you already generated

Step-Audio-EditX from StepFun AI is built around iterative editing rather than one-shot generation: you can take audio you already made and regenerate just the emotion, style, or a paralinguistic detail (breathing, laughter) without touching the words or re-cloning the voice. It tops recent open-model Elo rankings for emotion control, and also handles Cantonese and Sichuanese dialects alongside standard Mandarin and English.

Hear Step-Audio-EditX

Standardized samples are pending for this release. See the official sources below for the author’s demo.

Step-Audio-EditX · Reference-based · Same three scripts across the directory.

Best for: Re-edit the emotion or style of audio you already generated

Before you choose: Compare the samples with your own use case. Check the linked documentation for code, weight, and voice license terms.

Details awaiting a fresh review. Hardware figures are estimates unless a benchmark is linked.

Compare with similar

Compare all →

Choose an alternative to hear both models together.

No samples yet

Every model in this directory is read against the same three scripts so voices can be compared honestly — Step-Audio-EditX’s samples just haven’t been generated yet.

Contribute samples on GitHub →

Run Step-Audio-EditX

git clone https://github.com/stepfun-ai/Step-Audio-EditX

Sources & setup details

Step-Audio-EditX FAQ

Do I need a GPU to run Step-Audio-EditX?

The directory lists roughly 12 GB of VRAM as a historical estimate. Check the model documentation for your checkpoint and runtime.

Can Step-Audio-EditX clone voices?

Yes — Step-Audio-EditX supports voice cloning from reference audio.

What license is Step-Audio-EditX released under?

Step-Audio-EditX is released under the Apache-2.0 license. Always check the repository for the exact terms — some models license code and weights separately.

What languages does Step-Audio-EditX support?

Step-Audio-EditX supports English, Chinese.

Is there a hosted Step-Audio-EditX API?

OpenSpeech Cloud is accepting production-access interest for Step-Audio-EditX, but this model is not hosted in the beta. Use the official repository for current deployment options.