ChengRang

Vidu S2

AI Video Tools Freemium
This page covers a version or sub-product of Vidu. View Vidu overview →

A dual real-time video model released by Shengshu Technology on 2026/9/15, just 69 days after Vidu S1, and made fully available at launch. S2-Avatar raises real-time digital character output from 540P to 720P at 25-42 FPS, and lets users feed in reference images of objects, clothing or scenes at any point during generation while keeping the character holding the object. S2-Editing applies real-time style transfer, virtual try-on, character swaps and background swaps to an incoming video stream. The company is also exploring spatial video for VR headsets.

Shengshu TechnologyReal-Time InteractionReal-Time EditingDigital HumanSpatial Video
Visit Vidu S2

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

Vidu S2 is a real-time video model released by Shengshu Technology on September 15, 2026, just 69 days after Vidu S1. It comprises two models: Vidu S2-Avatar for continuous interactive digital character generation, and Vidu S2-Editing for real-time editing of input video streams. Both the web experience and the API opened to everyone on launch day, with no need to apply for early access.

S2-Avatar's improvements over the previous generation center on three areas: real-time output rises from 540P to 720P at 25 to 42 FPS; new reference images can be fed in at any point during video stream generation, letting the character pick up a specified object, change into a specified outfit, or enter a new scene; and motion capabilities expand from mostly talking to full-body, large-amplitude movements such as solo dancing and 2D and 3D animation. This kind of interaction also involves maintaining state after an action is completed—for example, "pick up the cup and then smile" requires the character to keep holding the cup while smiling. The model uses a VLM Agent to read user instructions and check output frames in temporal order, judging whether an action is complete, partially complete, or off-target, then adjusting subsequent prompts.

S2-Editing takes a continuous input video stream and performs four types of real-time editing with a single text instruction plus an optional reference image: style transfer, virtual try-on, character replacement, and background replacement. When a person raises a hand, turns, or the camera moves, the edited result must follow those changes. This relies on frame-aligned sparse attention, which lets a target frame read only the source video's frame at the same moment, preventing motion information from different moments from interfering with each other.

Beyond the two models, Shengshu Technology has extended its exploration into spatial video: converting real-time generated or real-time edited footage into synchronized left- and right-eye views with parallax, then streaming them to a VR headset for character interaction and image editing in spatial video. Currently this is mainly fixed-viewpoint binocular video, with panoramic spatial video still in the exploratory stage. The technical roadmap for the entire real-time pipeline is led by Zhang Jintao (a PhD student of Professor Zhu Jun), head of streaming video generation and inference at Shengshu Technology.

Key Features

Use Cases

Pros

Pricing

The web experience is available at vidu.com/vidu-stream, and the API is accessible through the Vidu S2 documentation at platform.vidu.com. Pricing varies by model and usage mode: for real-time digital humans you can choose a full service bundle covering speech recognition, language model, speech synthesis and real-time communication, or use the component mode to plug into an existing voice system; offline generation runs through an asynchronous interface. For specific unit prices and new-user credits, refer to the official website and API platform announcements.

Summary

Vidu S2 moves video generation from generating a clip to controlling a continuously running video stream in real time, with each of the two lines targeting a clear use case: S2-Avatar for real-time digital human interaction and S2-Editing for real-time video editing. Compared with Vidu S1, this generation's changes go beyond resolution moving from 540P to 720P: the control target expands from a fixed digital human to dynamic reference images and full video streams, motion capability extends to full-body performance, and there is an added step of spatial video exploration for VR headsets. Teams working on livestreaming, virtual try-on, real-time virtual production or batch derivation of short-drama assets should start testing from Day 1; if the need is simply to generate a finished clip offline, a mature finished-video model like Vidu Q3 remains the easier choice.

Version History

Category
AI Video Tools
Pricing
Freemium
Tags
Shengshu Technology · Real-Time Interaction · Real-Time Editing
Website

Related Tools