ByteDance's Seedance 2.5: AI Video Now Tells a 30-Second Story
ByteDance has officially launched Seedance 2.5, its latest video generation model, doubling the maximum single-video length from 15 seconds to 30 seconds. The model is being gradually introduced on Jimeng AI and Doubao Professional Edition, with API services soon to be integrated into Volcano Engine, opening doors to applications in film, advertising, education, industrial manufacturing, and even autonomous driving.

According to the company, Seedance 2.5 retains the unified multimodal audio-visual joint generation architecture, but the core breakthroughs lie in enhanced long-narrative capabilities, multimodal reference, and editing. In simpler terms, AI videos are no longer just scattered clips—they can now deliver a complete creation with a beginning, middle, and end.
One Take, Full Story
In an official demonstration, a single continuous shot of a singer's performance unfolds with clear narrative logic. The camera starts behind a red curtain, moves to a warm backstage dressing room where the young female singer adjusts her earphones, and a staff member reminds her it's time to go on stage. She walks through a passage, interacts with a partner, takes the microphone, and finally steps onto the stage. The camera pulls back to reveal the entire stadium—audience, lights, fluorescent sticks, and cheers all captured in one seamless take. The model can organize multiple logically connected shots—setup, development, turning point, and conclusion—within 30 seconds.
Extending the Story Seamlessly
The multi-round extension capability is equally impressive. The model can continue generating another 30 seconds based on an existing video, while maintaining consistency in main characters, scenes, visual style, and even sound effects. In another example, a boy runs out of a subway car holding a soccer ball, and the male lead chases and finally catches him. The continuous action is seamless, with no sense of detachment.
Handling Complex Scenes with Ease
Seedance 2.5 supports up to 30 images, 10 video clips, and 10 audio clips as reference materials in a single input. The model comprehensively understands elements like composition, scene, style, characters, and props across different materials, and accurately applies them according to instructions. In scenes with multiple people, it can simultaneously restore multiple characters' appearances and voices while keeping the main subjects stable. The official 30-second music concert clip used a 16:9 horizontal screen with a cinematic realistic style, and the reference materials covered a pianist, a cellist, a violinist, a vocalist, an orchestra, a choir, and the audience.
When AI video evolves from a "special effect toy" to a tool capable of handling complete narratives and group scheduling, the toolkit for short video creation may need to be rewritten again.
Key Points
- 30-second single-take videos with coherent storylines
- Multi-round extension maintains consistency across additional 30-second segments
- Supports up to 30 images, 10 videos, and 10 audio clips as references
- Handles complex multi-character scenes with stable main subjects
- Rolling out on Jimeng AI and Doubao Pro, with API via Volcano Engine