ChengRang

StepAudio 3

AI Audio & Music Freemium
This page covers a version or sub-product of Step. View Step overview →

A speech model series released by StepFun on 2026/9/15, comprising five models — Realtime, ASR, TTS, Gen and Music — launched simultaneously on its open platform. Realtime is natively full-duplex, ranking first worldwide on Conversational Dynamics with a composite score of 98.9% and first in speech reasoning accuracy at 99.7%. ASR ties for first worldwide with a 1.7% non-streaming word error rate. TTS reproduces paralinguistic cues such as laughter, hesitation and stuttering. Gen generates vocals, sound effects, ambient sound and BGM from a single instruction, while Music supports multi-turn composition using ABC notation.

StepFunSpeech ModelFull-DuplexASRTTSMusic Generation
Visit StepAudio 3

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

StepAudio 3 is StepFun's next-generation speech model series, released on September 15, 2026. It comprises five models: StepAudio 3 Realtime, StepAudio 3 ASR, StepAudio 3 TTS, StepAudio 3 Gen, and StepAudio 3 Music, covering real-time voice interaction, speech understanding in complex scenarios, human-level speech generation, full audio content generation, and music creation, respectively. All five models are now available on the StepFun open platform.

Realtime is the model that best illustrates the direction of this series. It supports native full-duplex real-time conversation, simultaneously understanding semantics, tone, emotion, paralinguistic features, and ambient sound to judge when to respond and when to wait, and it can handle user interruptions and continuous feedback. For complex questions, it supports parallel reasoning and speech generation, while Tool Call and long tasks can run asynchronously without blocking the current session. On the Artificial Analysis Conversational Dynamics leaderboard, it ranks first globally with a composite score of 98.9%; on the Speech Reasoning leaderboard, it also ranks first with 99.7% speech reasoning accuracy.

The other four models address different stages of audio production. ASR combines high-precision speech recognition with the contextual understanding of large models, covering Chinese, English, dialects, mixed Chinese-English speech, and long audio, as well as specialized domains such as healthcare, law, finance, automotive, and programming. It also handles complex inputs including low volume, fast speech, unclear articulation, singing, and background music, achieving a word error rate (WER) of 1.7% for non-streaming speech recognition, tied for first globally. TTS targets real-time interaction scenarios, closely matching natural speech in timbre, intonation layers, rhythm, and pauses. It can reproduce paralinguistic behaviors such as laughter, hesitation, stuttering, repetition, and self-correction, and can autonomously adjust emotional tone based on semantic content, using a streaming generation architecture that plays audio as it is generated.

Gen and Music target content production. Gen compresses the traditional audio production chain of recording, editing, and mixing into a single step: a natural language description plus reference audio is enough to uniformly generate character voices, sound effects, ambient sound, and background music, with support for specifying the timing and order in which dialogue, sound effects, and background music appear. Music supports zero-shot generation and multi-turn interactive creation based on ABC notation. Inputs can be lyrics, a cappella vocals, reference songs, or sheet music, with natural language used to control genre, vocals, melody, rhythm, instruments, and emotion; in the a cappella scoring scenario, a single a cappella passage is enough to automatically fill in harmonies, rhythm, instruments, and a complete arrangement.

Key Features

Use Cases

Pros

Pricing

All five models are now live on the StepFun open platform (platform.stepfun.com), called through separate endpoints per model. ASR and TTS are relatively easy to integrate, while Realtime requires handling bidirectional audio streams and session state. Specific pricing and free quotas are subject to announcements on the StepFun open platform.

Summary

StepAudio 3 expands speech models from single-purpose recognition or generation into full audio content production and real-time interaction capabilities, with each of the five models handling one part: Realtime handles real-time conversation, ASR makes sure it hears accurately, TTS makes sure it sounds human, Gen generates a complete set of sounds from a single instruction, and Music turns a cappella into a full arrangement. All three global first-place rankings come from Artificial Analysis leaderboards, and Realtime's 98.9% conversational dynamics score and 99.7% speech reasoning accuracy are the hardest numbers of this generation. For products built around voice assistants, customer service, audio content or music creation, you can start by integrating ASR or TTS as needed and move to real-time conversation as a next step; when evaluating real-time voice experiences, this series is worth testing side by side with Gemini 3.8 Live.

Version History

Category
AI Audio & Music
Pricing
Freemium
Tags
StepFun · Speech Model · Full-Duplex

Related Tools