Alibaba · Video models

Wan 3.0

Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video.

Try Wan 3.0 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

Wan 3.0 on Scenema

Wan 3.0 is Alibaba Tongyi Wanxiang’s third-generation video model, tuned for image-driven generation with clean motion and reliable identity carry across the take. On Scenema, Wan 3.0 runs as an image-to-video shot generator inside the long-form video pipeline, replacing the earlier Wan 2.6 in the lineup. Wan 3.0 delivers a 720p take from 4 to 12 seconds, with first-frame and last-frame conditioning so a shot can be locked to a defined beginning and a defined ending. Scenema packs those takes into scenes, and scenes into full-length videos with character consistency, voiceover, and music running across every cut.

Specifications

Wan 3.0 specifications on Scenema

Supports
Image to video
Duration
4 to 12 seconds
Resolution
720p
Aspect ratios
16:9, 9:16, and 1:1
Quality tier
High
Notes
Supports last-frame conditioning: lock the shot to a defined ending frame.

Distinctive

What makes Wan 3.0 distinctive

Wan 3.0 supports last-frame conditioning.

Give Wan 3.0 both a first frame and a last frame, and the model generates the motion between them. That gives an editor exact control over where the shot ends up, which matters when the next cut has to land on a specific pose or composition.

Wan 3.0 is Alibaba’s current-generation Wan model on Scenema.

Wan 3.0 supersedes Wan 2.6 in the Scenema lineup. Faster generation, better identity carry, and last-frame conditioning are the practical differences.

Wan 3.0 handles the standard aspect ratios.

Shots can be generated in 16:9, 9:16, or 1:1, so the same reference frame can drive landscape, vertical, and square variants of a scene.

Scenema assembles Wan 3.0 shots into long-form video.

Each Wan 3.0 generation is one shot inside a Scenema scene. Scenema plans the shot list, feeds each shot its reference frames, and assembles the results into full videos with continuous character, voiceover, and music tracks.

Use cases

Use cases for Wan 3.0

Wan 3.0 fits image-anchored shots.

Turn an approved keyframe into a moving shot when the composition is already set and only the motion needs to be generated.

Wan 3.0 nails start-to-end shot control.

Use both first-frame and last-frame conditioning when the shot has to start on one composition and land on another.

Wan 3.0 handles character-consistent B-roll.

Reuse an entity manifest reference frame to keep the same character or product identifiable across many shots in a longer video.

Wan 3.0 works for social and standard cuts.

Landscape, vertical, and square outputs cover the common delivery formats without swapping models.

Compare

How Wan 3.0 compares to other video models

See where Wan 3.0 slots in the lineup on the axes that matter most.

ModelBest forShot lengthMax resolutionInputsQuality tier
Veo 3.1Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video.4, 6, or 8 secondsUp to 1080pText to video, Image to videoHigh
Kling 3.0Kling 3.0 is Kuaishou’s flagship generative video model, running as one of the shot generators inside a Scenema long-form video.4, 6, 8, 10, or 12 seconds1080pText to video, Image to videoHigh
Seedance 2.0Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.4 to 15 secondsUp to 720pText to video, Image to video, Audio to videoHigh
Seedance 1.5Seedance 1.5 is ByteDance’s image-to-video model, running as one of the shot generators inside a Scenema long-form video.4 to 12 secondsUp to 720pImage to videoStandard
Wan 3.0Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video.4 to 12 seconds720pImage to videoHigh
MiniMax H3MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent.4 to 15 seconds768pImage to video, Reference to videoHigh
Vidu Q3 TurboVidu Q3 Turbo is a start-and-end-frame video model on Scenema, tuned for shots that must land on an exact ending composition.1 to 15 seconds720pImage to videoStandard
Gemini Omni FlashGemini Omni Flash is Google’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.4, 6, 8, or 10 secondsUp to 4KImage to videoHigh
LTX 2.3LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes.4 to 12 secondsUp to 1080pText to video, Image to video, Audio to videoFast and Pro tiers

FAQ

Frequently asked questions about Wan 3.0

What is Wan 3.0 on Scenema?+

Wan 3.0 is Alibaba Tongyi Wanxiang’s third-generation video model, exposed inside Scenema as an image-to-video shot generator. Each generation is a shot of 4 to 12 seconds at 720p, and Scenema packs those shots into full videos.

What happened to Wan 2.6 on Scenema?+

Wan 2.6 is retired from the active lineup. Wan 3.0 is the successor and is the default Wan-family choice on Scenema going forward. Projects that pinned Wan 2.6 still resolve to a valid config, but new work should pick Wan 3.0.

What is last-frame conditioning?+

You supply both a first frame and a last frame, and Wan 3.0 generates the motion between them. That lets an editor lock exactly where the shot ends up so the next cut lands on the intended composition.

How does Scenema make long-form videos if Wan 3.0 only generates up to 12 seconds at a time?+

Scenema treats each Wan 3.0 generation as one shot inside a scene. The pipeline plans the shot list, supplies the reference frames, and assembles the results into scenes and full videos. Character consistency, voiceover, and music tracks run across every cut, so the finished piece plays as a single continuous video.

When does Scenema use Wan 3.0 over another Scenema model?+

Scenema uses Wan 3.0 when the shot is image-anchored and the composition matters end to end. For text-driven shots it uses Seedance 2.0 or Kling 3.0, for 4K Gemini Omni Flash, and for complex motion across many entities MiniMax H3.

Related

Other video models on Scenema

Veo 3.1

Google DeepMind
  • Google DeepMind’s flagship cinematic video model.
  • 4, 6, or 8 second takes at up to 1080p, with audio generated in the same pass as the picture.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Kling 3.0

Kuaishou
  • Kuaishou’s third-generation flagship video model.
  • 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Seedance 2.0

ByteDance
  • ByteDance’s multimodal video model with joint audio-video generation.
  • Text, image, and audio inputs; up to 15 second takes, the longest on Scenema alongside MiniMax H3 and Vidu Q3 Turbo.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Seedance 1.5

ByteDance
  • ByteDance’s image-to-video model, driven from a reference frame.
  • 4 to 12 second takes at up to 720p.
  • Image-to-video shot generator inside Scenema’s long-form pipeline.
View model →

MiniMax H3

MiniMax
  • MiniMax’s next-generation video model.
  • Reference-to-video; 4 to 15 second takes at 768p.
  • Complex motion and multi-entity consistency inside Scenema’s long-form pipeline.
View model →

Vidu Q3 Turbo

Shengshu
  • Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
  • 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
  • Locks the shot to an exact ending composition for clean cut-to-cut choreography.
View model →

Gemini Omni Flash

Google
  • Google’s multimodal video model with native audio and 4K output.
  • Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
  • Highest-resolution shot generator inside Scenema’s long-form pipeline.
View model →

LTX 2.3

Lightricks
  • Lightricks’ open video foundation model with Fast and Pro tiers on Scenema.
  • Text, image, and audio inputs; 4 to 12 second takes at up to 1080p.
  • Widest aspect-ratio range of any Scenema video model.
View model →

Start generating with Wan 3.0 in Scenema

Wan 3.0 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.

Try Wan 3.0 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

Wan 3.0 on Scenema | Video models