ChengRang

Doubao Audio 1.0

AI Audio & Music Paid

ByteDance Doubao audio model released at 2026/6/23 FORCE, supporting reference generation and long-duration multi-role voice consistency

ByteDanceDoubaoTTSVoice CloningMusic
Visit Doubao Audio 1.0

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

Doubao Audio 1.0 is ByteDance's first independent audio large model, released at the **FORCE 2026** conference on June 23, 2026, by ByteDance Volcano Engine. It forms a **multimodal family bucket** with Seedance 2.5 (video) + Seedream 5.0 Pro (image) + Doubao 2.1 Pro (multimodal dialogue)—completing the last piece of ByteDance's 2026 "text/image/video/audio" strategy.

Core capabilities cover: **TTS speech synthesis + voice cloning + music generation + audio understanding**—four tasks in one. The most imaginative aspect is its integration with **Doubao 2.1 Pro's ability to watch 2-hour videos** to create an **end-to-end dubbing workflow**—an automated pipeline where "AI watches video → AI writes dubbing script → AI generates dubbing → AI composes music." This is the first product in the domestic model ecosystem to truly connect the entire chain from "video content to dubbing/music."

This aligns with Doubao's overall strategy: **providing capabilities comparable to overseas closed-source models at lower cost**—ElevenLabs Pro costs $99/month, while Doubao Audio 1.0 is billed per token via Volcano Engine, expected to be significantly cheaper for Chinese-language scenarios.

Key Features

Use Cases

Pros

Pricing

**Called via Volcano Engine/Doubao platform**, billed based on a combination of audio duration + number of calls + features (TTS/cloning/music/understanding). Specific unit prices are subject to the official announcement on the Volcano Engine website. **Doubao plan** users can use it in conjunction.

Summary

Doubao Audio 1.0 is the final piece of ByteDance's multimodal family bucket released on June 23, 2026—**TTS + cloning + music + understanding in one, plus end-to-end dubbing integrated with Doubao 2.1 Pro's video capability** gives it unique competitiveness in the Chinese audio AI track. If you create Chinese video/podcast/digital human content, Doubao Audio 1.0 is worth trying on the domestic side; if you pursue the highest TTS naturalness and multilingual support (70+ languages), ElevenLabs v3 remains the primary choice; if you do Chinese short video dubbing and want to save costs, Doubao Audio 1.0 is a more cost-effective option.

Category
AI Audio & Music
Pricing
Paid
Tags
ByteDance · Doubao · TTS

Related Tools