ByteDance · Video models
Seedance 2.0
Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.
200 credits on signup, then weekly refills through your first month. No credit card required.
Seedance 2.0 on Scenema
Seedance 2.0 is ByteDance’s second-generation multimodal video model, built on a joint audio-video architecture that accepts text, image, and audio inputs and generates video and sound together in a single pass. On Scenema, Seedance 2.0 runs as one of the shot generators inside the long-form video pipeline. Seedance 2.0 delivers the take at up to 720p, with a single generation running as long as 15 seconds, the longest of any Scenema video model. Scenema packs those takes into scenes, and scenes into full-length videos with character consistency, voiceover, and music running across every cut.
Specifications
Seedance 2.0 specifications on Scenema
- Supports
- Text to video, Image to video, Audio to video
- Duration
- 4 to 15 seconds
- Resolution
- Up to 720p
- Aspect ratios
- 16:9, 9:16, and 1:1
- Quality tier
- High
Distinctive
What makes Seedance 2.0 distinctive
Seedance 2.0 uses a unified audio-video architecture.
Video and audio are generated jointly rather than in separate passes, so effects, ambience, and lip movement are aligned with what appears on screen.
Seedance 2.0 accepts text, image, and audio inputs.
A single shot can be driven by a written prompt, a reference image, or an audio clip. That flexibility is useful when the input material already exists.
Seedance 2.0 has the longest single-generation length on Scenema.
A single Seedance 2.0 shot can run up to 15 seconds. That is the longest take available from any of Scenema’s video models, useful when a scene calls for a sustained beat before the next cut.
Scenema assembles Seedance 2.0 shots into long-form video.
Every Seedance 2.0 generation is one shot inside a Scenema scene. Scenema plans the shot list, runs each shot, and joins them into scenes and full videos with continuous character, voiceover, and music tracks.
Use cases
Use cases for Seedance 2.0
Seedance 2.0 fits sustained single-take scenes.
Use Seedance 2.0 when a scene needs a longer, uninterrupted shot before cutting, up to 15 seconds per generation.
Seedance 2.0 works for audio-driven shots.
Feed Seedance 2.0 an audio input to align motion or ambience with a reference sound.
Seedance 2.0 handles dialogue-heavy scenes.
The joint audio-video pass produces lip-synced dialogue inside the same shot generation.
Seedance 2.0 covers text and image driven shots.
The same model handles text to video and image to video, so Scenema can call it for a range of shot types inside one scene.
Compare
How Seedance 2.0 compares to other video models
See where Seedance 2.0 slots in the lineup on the axes that matter most.
| Model | Best for | Shot length | Max resolution | Inputs | Quality tier |
|---|---|---|---|---|---|
| Veo 3.1 | Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video. | 4, 6, or 8 seconds | Up to 1080p | Text to video, Image to video | High |
| Kling 3.0 | Kling 3.0 is Kuaishou’s flagship generative video model, running as one of the shot generators inside a Scenema long-form video. | 4, 6, 8, 10, or 12 seconds | 1080p | Text to video, Image to video | High |
| Seedance 2.0 | Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video. | 4 to 15 seconds | Up to 720p | Text to video, Image to video, Audio to video | High |
| Seedance 1.5 | Seedance 1.5 is ByteDance’s image-to-video model, running as one of the shot generators inside a Scenema long-form video. | 4 to 12 seconds | Up to 720p | Image to video | Standard |
| Wan 3.0 | Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video. | 4 to 12 seconds | 720p | Image to video | High |
| MiniMax H3 | MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent. | 4 to 15 seconds | 768p | Image to video, Reference to video | High |
| Vidu Q3 Turbo | Vidu Q3 Turbo is a start-and-end-frame video model on Scenema, tuned for shots that must land on an exact ending composition. | 1 to 15 seconds | 720p | Image to video | Standard |
| Gemini Omni Flash | Gemini Omni Flash is Google’s multimodal video model, running as one of the shot generators inside a Scenema long-form video. | 4, 6, 8, or 10 seconds | Up to 4K | Image to video | High |
| LTX 2.3 | LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes. | 4 to 12 seconds | Up to 1080p | Text to video, Image to video, Audio to video | Fast and Pro tiers |
Where it fits
Where Scenema uses Seedance 2.0
Scenema runs a highly optimized workflow and uses the best model for the style and for each task under the hood. You choose the quality and the resolution. Here is where Seedance 2.0 fits and where a different model does.
- 1
Scenema uses Seedance 2.0 when a single sustained take has to run up to 15 seconds.
Seedance 2.0 generates up to 15 seconds in one pass, which is the longest single generation on Scenema. That is enough runway for a full spoken line, a longer action beat, or a slow camera move that would otherwise need a cut.
- 2
Scenema uses Seedance 2.0 when the video has to be driven by an audio track.
Seedance 2.0 accepts audio as an input alongside text and image. The picture is generated against the audio, so mouth movement, timing, and motion are shaped by the sound instead of dubbed on afterwards.
- 3
Scenema uses Seedance 2.0 when the shot needs its own audio in the same pass.
Seedance 2.0 uses a joint audio-video architecture, so ambience and effects are generated with the picture. That keeps sync tight on the longer takes it is designed for.
- 4
Scenema uses a different model when the shot has to be delivered at 1080p.
Seedance 2.0 tops out at 720p. If the same shot has to hit 1080p, use Kling 3.0 for cinematic takes up to 12 seconds or LTX 2.3 Pro for the widest aspect-ratio range.
- 5
Scenema uses a different model when the hero shot has to be the highest-quality frame in the video.
Seedance 2.0 trades some per-frame realism for length and audio-driven control. For the single shot the audience is meant to remember, use Veo 3.1 and let Scenema place the longer Seedance 2.0 shots around it.
FAQ
Frequently asked questions about Seedance 2.0
What is Seedance 2.0 on Scenema?+
Seedance 2.0 is ByteDance’s flagship video model, exposed inside Scenema as one of the shot generators available in the long-form pipeline. Each generation is a shot of 4 to 15 seconds at up to 720p, and Scenema packs those shots into full videos.
What resolution and aspect ratios does Seedance 2.0 support on Scenema?+
Seedance 2.0 on Scenema outputs up to 720p in 16:9, 9:16, and 1:1. Pick the aspect ratio to match the finished video, whether that is landscape, vertical, or square.
How long is a single Seedance 2.0 generation on Scenema?+
A single Seedance 2.0 generation runs 4, 5, 6, 8, 10, 12, or 15 seconds. That is the shot length. Scenema stitches many Seedance 2.0 generations together to build the finished video.
How does Scenema make long-form videos if Seedance 2.0 only generates up to 15 seconds at a time?+
Scenema treats each Seedance 2.0 generation as one shot inside a scene. The pipeline plans the shot list, generates each shot, and assembles them into scenes and full videos. Character consistency, voiceover, and music tracks run across every cut, so the finished piece plays as a single continuous video.
When does Scenema use Seedance 2.0 over another Scenema model?+
Scenema uses Seedance 2.0 when a shot needs to run longer than any other Scenema model allows in a single generation, when the input is audio, or when the scene needs joint audio-video generation in one pass.
Related
Other video models on Scenema
Veo 3.1
Google DeepMind- Google DeepMind’s flagship cinematic video model.
- 4, 6, or 8 second takes at up to 1080p, with audio generated in the same pass as the picture.
- Shot generator inside Scenema’s long-form pipeline.
Kling 3.0
Kuaishou- Kuaishou’s third-generation flagship video model.
- 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 1.5
ByteDance- ByteDance’s image-to-video model, driven from a reference frame.
- 4 to 12 second takes at up to 720p.
- Image-to-video shot generator inside Scenema’s long-form pipeline.
Wan 3.0
Alibaba- Alibaba Tongyi Wanxiang’s current-generation Wan model, replacing Wan 2.6 on Scenema.
- Image-to-video with last-frame conditioning; 4 to 12 second takes at 720p.
- Image-anchored shot generator inside Scenema’s long-form pipeline.
MiniMax H3
MiniMax- MiniMax’s next-generation video model.
- Reference-to-video; 4 to 15 second takes at 768p.
- Complex motion and multi-entity consistency inside Scenema’s long-form pipeline.
Vidu Q3 Turbo
Shengshu- Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
- 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
- Locks the shot to an exact ending composition for clean cut-to-cut choreography.
Gemini Omni Flash
Google- Google’s multimodal video model with native audio and 4K output.
- Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
- Highest-resolution shot generator inside Scenema’s long-form pipeline.
LTX 2.3
Lightricks- Lightricks’ open video foundation model with Fast and Pro tiers on Scenema.
- Text, image, and audio inputs; 4 to 12 second takes at up to 1080p.
- Widest aspect-ratio range of any Scenema video model.
Start generating with Seedance 2.0 in Scenema
Seedance 2.0 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.
200 credits on signup, then weekly refills through your first month. No credit card required.