Google DeepMind · Video models

Veo 3.1

Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video.

Try Veo 3.1 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

Veo 3.1 on Scenema

Veo 3.1 is Google DeepMind’s flagship generative video model, tuned for cinematic realism, coherent motion, and audio that is generated together with the picture. Scenema hosts Veo 3.1 as one of the shot generators inside its long-form video pipeline. Veo 3.1 handles the 8-second cinematic take, Scenema packs those takes into scenes, and scenes into full-length videos with character consistency, voiceover, and music running across every cut.

Specifications

Veo 3.1 specifications on Scenema

Supports
Text to video, Image to video
Duration
4, 6, or 8 seconds
Resolution
Up to 1080p
Aspect ratios
16:9 and 9:16
Quality tier
High
Notes
Audio generated in the same pass as the video. Supports first-frame and last-frame conditioning.

Distinctive

What makes Veo 3.1 distinctive

Native audio generated with the picture.

Veo 3.1 produces sound in the same pass as the video, so ambience, effects, and dialogue-adjacent audio land in sync with the motion. Scenema keeps the audio baked into the file rather than dubbing it on afterwards.

Cinematic-grade motion and lighting realism.

Veo 3.1 is trained for cinematographic style: parallax on camera moves, plausible physics on interactions, and lighting that holds shape across the shot. It is the model people reach for when a Scenema shot needs to feel cinematic.

Text-to-video and image-to-video from a single reference.

Point Veo 3.1 at a prompt for text-to-video, or at a starting frame for image-to-video. Scenema exposes both paths so you can either write the shot from scratch or anchor it to a keyframe you have already generated.

Scenema assembles Veo shots into long-form video.

A Veo 3.1 generation runs 4, 6, or 8 seconds. Scenema packs many Veo generations into a scene, and many scenes into a full-length video, holding character identity, voiceover, and music across every cut. Veo produces the cinematic take. Scenema produces the finished long-form video.

Use cases

Use cases for Veo 3.1

Cinematic hero shots inside a longer video.

Reach for Veo 3.1 when a specific shot in your project needs to be the visual peak. The rest of the scene can run on a faster model and Veo 3.1 handles the moment the audience is meant to remember.

Product and brand explainers.

Product explainers lean on realism when the product itself is on screen. Veo 3.1 gives Scenema the motion quality that a hero product shot needs inside a longer explainer.

Establishing shots and B-roll.

Wide establishing shots, aerial passes, and scene-setting B-roll are the shots where realism matters most. Veo 3.1 holds together on the big canvas where lower-tier models tend to reveal artefacts.

Chapter openers and transitions.

A chapter opener sets the tone for the minutes that follow. Veo 3.1 responds to camera direction in the prompt, so a Scenema shot written as a sweeping opener reads that way on screen.

Compare

How Veo 3.1 compares to other video models

See where Veo 3.1 slots in the lineup on the axes that matter most.

ModelBest forShot lengthMax resolutionInputsQuality tier
Veo 3.1Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video.4, 6, or 8 secondsUp to 1080pText to video, Image to videoHigh
Kling 3.0Kling 3.0 is Kuaishou’s flagship generative video model, running as one of the shot generators inside a Scenema long-form video.4, 6, 8, 10, or 12 seconds1080pText to video, Image to videoHigh
Seedance 2.0Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.4 to 15 secondsUp to 720pText to video, Image to video, Audio to videoHigh
Seedance 1.5Seedance 1.5 is ByteDance’s image-to-video model, running as one of the shot generators inside a Scenema long-form video.4 to 12 secondsUp to 720pImage to videoStandard
Wan 3.0Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video.4 to 12 seconds720pImage to videoHigh
MiniMax H3MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent.4 to 15 seconds768pImage to video, Reference to videoHigh
Vidu Q3 TurboVidu Q3 Turbo is a start-and-end-frame video model on Scenema, tuned for shots that must land on an exact ending composition.1 to 15 seconds720pImage to videoStandard
Gemini Omni FlashGemini Omni Flash is Google’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.4, 6, 8, or 10 secondsUp to 4KImage to videoHigh
LTX 2.3LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes.4 to 12 secondsUp to 1080pText to video, Image to video, Audio to videoFast and Pro tiers

Where it fits

Where Scenema uses Veo 3.1

Scenema runs a highly optimized workflow and uses the best model for the style and for each task under the hood. You choose the quality and the resolution. Here is where Veo 3.1 fits and where a different model does.

  1. 1

    Scenema uses Veo 3.1 when the shot has to carry the video.

    Veo 3.1 is the highest-quality video model on Scenema for cinematic realism. Save it for the shots the audience is meant to remember: the establishing shot, the money take in an ad, the closing beat of a narrative scene.

  2. 2

    Scenema uses Veo 3.1 when the shot has to carry its own audio.

    Veo 3.1 generates audio in the same pass as the picture. Ambience, effects, and impact hits land in sync with the motion instead of being dubbed on afterwards. That saves a separate sound-design pass.

  3. 3

    Scenema uses Veo 3.1 for hero product shots.

    A hero product shot is judged on motion realism and audio fidelity. Veo 3.1 delivers both in one generation, at up to 1080p.

  4. 4

    Scenema uses a faster model when you are still iterating on framing.

    Veo 3.1 is slower per generation than the standard-tier video models. For exploration passes on framing, blocking, or pacing, run the LTX 2.3 Fast tier and it returns a shot in a fraction of the time. Move back to Veo 3.1 once the shot is locked.

    See LTX 2.3 →

  5. 5

    Scenema uses another model when your shot needs to be longer than 8 seconds.

    A Veo 3.1 generation is 4, 6, or 8 seconds. If a single sustained take has to run 12 seconds, use Kling 3.0. If it has to run up to 15 seconds, use Seedance 2.0 or MiniMax H3. Scenema still assembles the longer shots into scenes and full videos the same way.

    See Seedance 2.0 →

FAQ

Frequently asked questions about Veo 3.1

What is Veo 3.1 on Scenema?+

Veo 3.1 is Google DeepMind’s flagship video generation model, hosted on Scenema for text-to-video and image-to-video generation. Every Veo 3.1 generation on Scenema produces a 4, 6, or 8 second shot at 720p or 1080p, with audio generated in the same pass as the picture.

What resolution and aspect ratios does Veo 3.1 support on Scenema?+

Veo 3.1 supports 720p and 1080p output on Scenema, in 16:9 and 9:16 aspect ratios. The default output resolution is 720p; step up to 1080p when the shot is the visual centrepiece of the scene.

How long is a single Veo 3.1 generation?+

A single Veo 3.1 generation on Scenema is a 4, 6, or 8 second shot. Long-form output comes from Scenema itself, which packs many shots into a scene and many scenes into a full video. The per-generation limit is what Veo returns, not the length of the video you can make on Scenema.

How does Scenema make long-form videos if Veo only generates up to 8 seconds at a time?+

Scenema is the layer above the model. It packs multiple Veo generations into a scene, and multiple scenes into a full-length video, holding character identity, voiceover, and music across every cut. Veo produces the individual cinematic shots; Scenema produces the finished long-form video.

When does Scenema use Veo 3.1 over a faster Scenema model?+

Scenema uses Veo 3.1 when the shot needs to be the visual peak of your video, when realism is critical, or when the shot should carry its own audio. For fast iterations, or for shots outside the Veo range, Scenema uses a faster model like Wan 3.0 or Seedance 2.0 and reserves Veo 3.1 for the shots that matter most.

Related

Other video models on Scenema

Kling 3.0

Kuaishou
  • Kuaishou’s third-generation flagship video model.
  • 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Seedance 2.0

ByteDance
  • ByteDance’s multimodal video model with joint audio-video generation.
  • Text, image, and audio inputs; up to 15 second takes, the longest on Scenema alongside MiniMax H3 and Vidu Q3 Turbo.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Seedance 1.5

ByteDance
  • ByteDance’s image-to-video model, driven from a reference frame.
  • 4 to 12 second takes at up to 720p.
  • Image-to-video shot generator inside Scenema’s long-form pipeline.
View model →

Wan 3.0

Alibaba
  • Alibaba Tongyi Wanxiang’s current-generation Wan model, replacing Wan 2.6 on Scenema.
  • Image-to-video with last-frame conditioning; 4 to 12 second takes at 720p.
  • Image-anchored shot generator inside Scenema’s long-form pipeline.
View model →

MiniMax H3

MiniMax
  • MiniMax’s next-generation video model.
  • Reference-to-video; 4 to 15 second takes at 768p.
  • Complex motion and multi-entity consistency inside Scenema’s long-form pipeline.
View model →

Vidu Q3 Turbo

Shengshu
  • Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
  • 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
  • Locks the shot to an exact ending composition for clean cut-to-cut choreography.
View model →

Gemini Omni Flash

Google
  • Google’s multimodal video model with native audio and 4K output.
  • Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
  • Highest-resolution shot generator inside Scenema’s long-form pipeline.
View model →

LTX 2.3

Lightricks
  • Lightricks’ open video foundation model with Fast and Pro tiers on Scenema.
  • Text, image, and audio inputs; 4 to 12 second takes at up to 1080p.
  • Widest aspect-ratio range of any Scenema video model.
View model →

Start generating with Veo 3.1 in Scenema

Veo 3.1 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.

Try Veo 3.1 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

Veo 3.1 on Scenema | Video models