Overview
Doubao Audio 1.0 is ByteDance's first independent audio large model, released at the **FORCE 2026** conference on June 23, 2026, by ByteDance Volcano Engine. It forms a **multimodal family bucket** with Seedance 2.5 (video) + Seedream 5.0 Pro (image) + Doubao 2.1 Pro (multimodal dialogue)—completing the last piece of ByteDance's 2026 "text/image/video/audio" strategy.
Core capabilities cover: **TTS speech synthesis + voice cloning + music generation + audio understanding**—four tasks in one. The most imaginative aspect is its integration with **Doubao 2.1 Pro's ability to watch 2-hour videos** to create an **end-to-end dubbing workflow**—an automated pipeline where "AI watches video → AI writes dubbing script → AI generates dubbing → AI composes music." This is the first product in the domestic model ecosystem to truly connect the entire chain from "video content to dubbing/music."
This aligns with Doubao's overall strategy: **providing capabilities comparable to overseas closed-source models at lower cost**—ElevenLabs Pro costs $99/month, while Doubao Audio 1.0 is billed per token via Volcano Engine, expected to be significantly cheaper for Chinese-language scenarios.
Key Features
- TTS Speech Synthesis: High-quality TTS covering multiple voices, emotions, and tones, optimized for Chinese scenarios
- Voice Cloning: Clone a voice by uploading a short sample; can be combined with Doubao's video-watching capability for character dubbing
- Music Generation: Supports text-to-music, image-to-music, and stylized BGM generation
- Audio Understanding: Audio transcription + content understanding + emotion analysis, usable as an input tool for agents
- End-to-End Dubbing Workflow: Integrated with Doubao 2.1 Pro's ability to watch 2-hour videos: AI watches video → writes dubbing script → generates dubbing → composes music, fully automated pipeline
- Doubao Family Bucket Synergy: Forms a complete multimodal matrix with Seedance 2.5 / Seedream 5.0 / Doubao 2.1 Pro
- Chinese Scenario Optimization: Targeted optimization for Chinese pronunciation, emotions, accents, dialects, etc.
Use Cases
- Batch dubbing for short-form/long-form video content creators
- Voice generation for AI digital humans/virtual anchors
- Batch production of audiobooks/podcasts/radio dramas
- TTS for customer service/education/training scenarios
- Multi-version dubbing for marketing videos (different voices/languages)
- End-to-end dubbing of video content in combination with Doubao 2.1 Pro
Pros
- TTS/cloning/music/understanding all in one
- Integration with Doubao 2.1 Pro's video capability is the first end-to-end dubbing workflow among domestic models
- Deep optimization for Chinese scenarios
- Completes ByteDance's multimodal family bucket
- Expected cost significantly lower than overseas competitors like ElevenLabs (for Chinese scenarios)
Pricing
**Called via Volcano Engine/Doubao platform**, billed based on a combination of audio duration + number of calls + features (TTS/cloning/music/understanding). Specific unit prices are subject to the official announcement on the Volcano Engine website. **Doubao plan** users can use it in conjunction.
Summary
Doubao Audio 1.0 is the final piece of ByteDance's multimodal family bucket released on June 23, 2026—**TTS + cloning + music + understanding in one, plus end-to-end dubbing integrated with Doubao 2.1 Pro's video capability** gives it unique competitiveness in the Chinese audio AI track. If you create Chinese video/podcast/digital human content, Doubao Audio 1.0 is worth trying on the domestic side; if you pursue the highest TTS naturalness and multilingual support (70+ languages), ElevenLabs v3 remains the primary choice; if you do Chinese short video dubbing and want to save costs, Doubao Audio 1.0 is a more cost-effective option.