Alibaba · Video models
Wan 3.0
Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video.
200 credits on signup, then weekly refills through your first month. No credit card required.
Wan 3.0 on Scenema
Wan 3.0 is Alibaba Tongyi Wanxiang’s third-generation video model, tuned for image-driven generation with clean motion and reliable identity carry across the take. On Scenema, Wan 3.0 runs as an image-to-video shot generator inside the long-form video pipeline, replacing the earlier Wan 2.6 in the lineup. Wan 3.0 delivers a 720p take from 4 to 12 seconds, with first-frame and last-frame conditioning so a shot can be locked to a defined beginning and a defined ending. Scenema packs those takes into scenes, and scenes into full-length videos with character consistency, voiceover, and music running across every cut.
Specifications
Wan 3.0 specifications on Scenema
- Supports
- Image to video
- Duration
- 4 to 12 seconds
- Resolution
- 720p
- Aspect ratios
- 16:9, 9:16, and 1:1
- Quality tier
- High
- Notes
- Supports last-frame conditioning: lock the shot to a defined ending frame.
Distinctive
What makes Wan 3.0 distinctive
Wan 3.0 supports last-frame conditioning.
Give Wan 3.0 both a first frame and a last frame, and the model generates the motion between them. That gives an editor exact control over where the shot ends up, which matters when the next cut has to land on a specific pose or composition.
Wan 3.0 is Alibaba’s current-generation Wan model on Scenema.
Wan 3.0 supersedes Wan 2.6 in the Scenema lineup. Faster generation, better identity carry, and last-frame conditioning are the practical differences.
Wan 3.0 handles the standard aspect ratios.
Shots can be generated in 16:9, 9:16, or 1:1, so the same reference frame can drive landscape, vertical, and square variants of a scene.
Scenema assembles Wan 3.0 shots into long-form video.
Each Wan 3.0 generation is one shot inside a Scenema scene. Scenema plans the shot list, feeds each shot its reference frames, and assembles the results into full videos with continuous character, voiceover, and music tracks.
Use cases
Use cases for Wan 3.0
Wan 3.0 fits image-anchored shots.
Turn an approved keyframe into a moving shot when the composition is already set and only the motion needs to be generated.
Wan 3.0 nails start-to-end shot control.
Use both first-frame and last-frame conditioning when the shot has to start on one composition and land on another.
Wan 3.0 handles character-consistent B-roll.
Reuse an entity manifest reference frame to keep the same character or product identifiable across many shots in a longer video.
Wan 3.0 works for social and standard cuts.
Landscape, vertical, and square outputs cover the common delivery formats without swapping models.
Compare
How Wan 3.0 compares to other video models
See where Wan 3.0 slots in the lineup on the axes that matter most.
| Model | Best for | Shot length | Max resolution | Inputs | Quality tier |
|---|---|---|---|---|---|
| Veo 3.1 | Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video. | 4, 6, or 8 seconds | Up to 1080p | Text to video, Image to video | High |
| Kling 3.0 | Kling 3.0 is Kuaishou’s flagship generative video model, running as one of the shot generators inside a Scenema long-form video. | 4, 6, 8, 10, or 12 seconds | 1080p | Text to video, Image to video | High |
| Seedance 2.0 | Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video. | 4 to 15 seconds | Up to 720p | Text to video, Image to video, Audio to video | High |
| Seedance 1.5 | Seedance 1.5 is ByteDance’s image-to-video model, running as one of the shot generators inside a Scenema long-form video. | 4 to 12 seconds | Up to 720p | Image to video | Standard |
| Wan 3.0 | Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video. | 4 to 12 seconds | 720p | Image to video | High |
| MiniMax H3 | MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent. | 4 to 15 seconds | 768p | Image to video, Reference to video | High |
| Vidu Q3 Turbo | Vidu Q3 Turbo is a start-and-end-frame video model on Scenema, tuned for shots that must land on an exact ending composition. | 1 to 15 seconds | 720p | Image to video | Standard |
| Gemini Omni Flash | Gemini Omni Flash is Google’s multimodal video model, running as one of the shot generators inside a Scenema long-form video. | 4, 6, 8, or 10 seconds | Up to 4K | Image to video | High |
| LTX 2.3 | LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes. | 4 to 12 seconds | Up to 1080p | Text to video, Image to video, Audio to video | Fast and Pro tiers |
FAQ
Frequently asked questions about Wan 3.0
What is Wan 3.0 on Scenema?+
Wan 3.0 is Alibaba Tongyi Wanxiang’s third-generation video model, exposed inside Scenema as an image-to-video shot generator. Each generation is a shot of 4 to 12 seconds at 720p, and Scenema packs those shots into full videos.
What happened to Wan 2.6 on Scenema?+
Wan 2.6 is retired from the active lineup. Wan 3.0 is the successor and is the default Wan-family choice on Scenema going forward. Projects that pinned Wan 2.6 still resolve to a valid config, but new work should pick Wan 3.0.
What is last-frame conditioning?+
You supply both a first frame and a last frame, and Wan 3.0 generates the motion between them. That lets an editor lock exactly where the shot ends up so the next cut lands on the intended composition.
How does Scenema make long-form videos if Wan 3.0 only generates up to 12 seconds at a time?+
Scenema treats each Wan 3.0 generation as one shot inside a scene. The pipeline plans the shot list, supplies the reference frames, and assembles the results into scenes and full videos. Character consistency, voiceover, and music tracks run across every cut, so the finished piece plays as a single continuous video.
When does Scenema use Wan 3.0 over another Scenema model?+
Scenema uses Wan 3.0 when the shot is image-anchored and the composition matters end to end. For text-driven shots it uses Seedance 2.0 or Kling 3.0, for 4K Gemini Omni Flash, and for complex motion across many entities MiniMax H3.
Related
Other video models on Scenema
Veo 3.1
Google DeepMind- Google DeepMind’s flagship cinematic video model.
- 4, 6, or 8 second takes at up to 1080p, with audio generated in the same pass as the picture.
- Shot generator inside Scenema’s long-form pipeline.
Kling 3.0
Kuaishou- Kuaishou’s third-generation flagship video model.
- 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 2.0
ByteDance- ByteDance’s multimodal video model with joint audio-video generation.
- Text, image, and audio inputs; up to 15 second takes, the longest on Scenema alongside MiniMax H3 and Vidu Q3 Turbo.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 1.5
ByteDance- ByteDance’s image-to-video model, driven from a reference frame.
- 4 to 12 second takes at up to 720p.
- Image-to-video shot generator inside Scenema’s long-form pipeline.
MiniMax H3
MiniMax- MiniMax’s next-generation video model.
- Reference-to-video; 4 to 15 second takes at 768p.
- Complex motion and multi-entity consistency inside Scenema’s long-form pipeline.
Vidu Q3 Turbo
Shengshu- Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
- 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
- Locks the shot to an exact ending composition for clean cut-to-cut choreography.
Gemini Omni Flash
Google- Google’s multimodal video model with native audio and 4K output.
- Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
- Highest-resolution shot generator inside Scenema’s long-form pipeline.
LTX 2.3
Lightricks- Lightricks’ open video foundation model with Fast and Pro tiers on Scenema.
- Text, image, and audio inputs; 4 to 12 second takes at up to 1080p.
- Widest aspect-ratio range of any Scenema video model.
Start generating with Wan 3.0 in Scenema
Wan 3.0 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.
200 credits on signup, then weekly refills through your first month. No credit card required.