Overview
StepAudio 3 is StepFun's next-generation speech model series, released on September 15, 2026. It comprises five models: StepAudio 3 Realtime, StepAudio 3 ASR, StepAudio 3 TTS, StepAudio 3 Gen, and StepAudio 3 Music, covering real-time voice interaction, speech understanding in complex scenarios, human-level speech generation, full audio content generation, and music creation, respectively. All five models are now available on the StepFun open platform.
Realtime is the model that best illustrates the direction of this series. It supports native full-duplex real-time conversation, simultaneously understanding semantics, tone, emotion, paralinguistic features, and ambient sound to judge when to respond and when to wait, and it can handle user interruptions and continuous feedback. For complex questions, it supports parallel reasoning and speech generation, while Tool Call and long tasks can run asynchronously without blocking the current session. On the Artificial Analysis Conversational Dynamics leaderboard, it ranks first globally with a composite score of 98.9%; on the Speech Reasoning leaderboard, it also ranks first with 99.7% speech reasoning accuracy.
The other four models address different stages of audio production. ASR combines high-precision speech recognition with the contextual understanding of large models, covering Chinese, English, dialects, mixed Chinese-English speech, and long audio, as well as specialized domains such as healthcare, law, finance, automotive, and programming. It also handles complex inputs including low volume, fast speech, unclear articulation, singing, and background music, achieving a word error rate (WER) of 1.7% for non-streaming speech recognition, tied for first globally. TTS targets real-time interaction scenarios, closely matching natural speech in timbre, intonation layers, rhythm, and pauses. It can reproduce paralinguistic behaviors such as laughter, hesitation, stuttering, repetition, and self-correction, and can autonomously adjust emotional tone based on semantic content, using a streaming generation architecture that plays audio as it is generated.
Gen and Music target content production. Gen compresses the traditional audio production chain of recording, editing, and mixing into a single step: a natural language description plus reference audio is enough to uniformly generate character voices, sound effects, ambient sound, and background music, with support for specifying the timing and order in which dialogue, sound effects, and background music appear. Music supports zero-shot generation and multi-turn interactive creation based on ABC notation. Inputs can be lyrics, a cappella vocals, reference songs, or sheet music, with natural language used to control genre, vocals, melody, rhythm, instruments, and emotion; in the a cappella scoring scenario, a single a cappella passage is enough to automatically fill in harmonies, rhythm, instruments, and a complete arrangement.
Key Features
- Native full-duplex real-time conversation: StepAudio 3 Realtime supports simultaneous two-way speech, handles interruptions and continuous feedback, and judges when to respond and when to wait, ranking first globally on the Artificial Analysis Conversational Dynamics leaderboard with a composite score of 98.9%.
- Paralinguistic and ambient sound understanding: Beyond semantics, it simultaneously understands tone, emotion, paralinguistic features, and ambient sound, bringing the pacing of voice interaction closer to that of a real conversation.
- Parallel reasoning and speech, asynchronous long tasks: For complex questions, it supports parallel reasoning and speech generation, with Tool Call and long tasks executing asynchronously without blocking the current session. It achieves 99.7% speech reasoning accuracy, ranking first globally on the Speech Reasoning leaderboard.
- Speech recognition for specialized scenarios: ASR combines speech recognition with large-model contextual understanding, covering Chinese and English, dialects, mixed Chinese-English speech, and long audio, and adapting to fields such as healthcare, law, finance, automotive, and programming, with a non-streaming word error rate of 1.7%, tied for first globally.
- Recognition robustness in complex input environments: It handles low volume, fast speech, unclear articulation, and complex audio inputs such as singing and background music.
- Speech synthesis with reproducible paralinguistic cues: TTS reproduces paralinguistic behaviors such as laughter, hesitation, stuttering, repetition, and self-correction, can autonomously adjust emotional tone based on semantics, and uses a streaming architecture that plays audio as it is generated.
- Unified audio generation and voice timing control: Gen uses natural language plus reference audio to uniformly generate vocals, sound effects, ambient sound, and background music, supports multiple characters, and can specify the timing and order in which each part appears.
- Multi-turn music creation based on ABC notation: Music accepts lyrics, a cappella vocals, reference songs or sheet music as input, and lets you control genre, melody, rhythm, instrumentation and mood in natural language; in the a cappella case it can automatically add harmonies, rhythm and a full arrangement.
Use Cases
- Real-time voice assistant: full-duplex conversation with interruption handling for a voice interface that feels close to talking to a real person
- Intelligent customer service and agent assist: long tasks run asynchronously, so queries and transactions don't block the call
- Meeting and long-audio transcription: speech recognition covering dialects, mixed Chinese-English speech and specialized domains
- Audio content production: generate vocals, sound effects and ambient audio at once from a single natural-language description
- Short video and ad voiceover: relies on paralinguistic reproduction to convey emotion and tone
- Music creation: expand a snippet of a cappella or a set of lyrics into a fully arranged song
Pros
- Realtime ranks first globally on the Artificial Analysis Conversational Dynamics leaderboard with a composite score of 98.9%
- Speech reasoning accuracy of 99.7%, ranking first globally on the Speech Reasoning leaderboard
- ASR non-streaming word error rate of 1.7%, tied for first globally, with coverage of dialects and specialized domains
- Full-duplex conversation supports interruption handling and asynchronous execution of long tasks, for a voice experience that keeps pace with real human speech
- Five models span recognition, synthesis, real-time interaction, audio generation and music creation, and can be integrated individually as needed
- TTS supports streaming generation and paralinguistic reproduction, making it suited to real-time interaction scenarios
- All are now live on the StepFun open platform, with documentation available
Pricing
All five models are now live on the StepFun open platform (platform.stepfun.com), called through separate endpoints per model. ASR and TTS are relatively easy to integrate, while Realtime requires handling bidirectional audio streams and session state. Specific pricing and free quotas are subject to announcements on the StepFun open platform.
Summary
StepAudio 3 expands speech models from single-purpose recognition or generation into full audio content production and real-time interaction capabilities, with each of the five models handling one part: Realtime handles real-time conversation, ASR makes sure it hears accurately, TTS makes sure it sounds human, Gen generates a complete set of sounds from a single instruction, and Music turns a cappella into a full arrangement. All three global first-place rankings come from Artificial Analysis leaderboards, and Realtime's 98.9% conversational dynamics score and 99.7% speech reasoning accuracy are the hardest numbers of this generation. For products built around voice assistants, customer service, audio content or music creation, you can start by integrating ASR or TTS as needed and move to real-time conversation as a next step; when evaluating real-time voice experiences, this series is worth testing side by side with Gemini 3.8 Live.
Version History
- StepAudio 3 series launches: Realtime, ASR, TTS, Gen and Music now live on the open platform (2026-09-15): StepFun has released the five-model StepAudio 3 series and launched it on its open platform at the same time. Realtime supports native full-duplex real-time conversation and ranks first globally on the Artificial Analysis Conversational Dynamics leaderboard with a composite score of 98.9%, while also tying for first on the Speech Reasoning leaderboard with 99.7% speech reasoning accuracy; it supports parallel reasoning and speech generation, Tool Call and asynchronous execution of long tasks. ASR ties for first globally with a non-streaming word error rate of 1.7%, covering dialects, mixed Chinese-English speech and specialized fields such as healthcare, law and finance. TTS targets real-time interaction, reproducing paralinguistic cues such as laughter, hesitation, stuttering and self-correction, and supports streaming generation. Gen unifies the generation of vocals, sound effects, ambient audio and background music, with support for multiple characters and control over sound timing. Music supports multi-turn music creation based on ABC notation and automatic arrangement of a cappella vocals.