Google DeepMind · Video models
Veo 3.1
Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video.
200 credits on signup, then weekly refills through your first month. No credit card required.
Veo 3.1 on Scenema
Veo 3.1 is Google DeepMind’s flagship generative video model, tuned for cinematic realism, coherent motion, and audio that is generated together with the picture. Scenema hosts Veo 3.1 as one of the shot generators inside its long-form video pipeline. Veo 3.1 handles the 8-second cinematic take, Scenema packs those takes into scenes, and scenes into full-length videos with character consistency, voiceover, and music running across every cut.
Specifications
Veo 3.1 specifications on Scenema
- Supports
- Text to video, Image to video
- Duration
- 4, 6, or 8 seconds
- Resolution
- Up to 1080p
- Aspect ratios
- 16:9 and 9:16
- Quality tier
- High
- Notes
- Audio generated in the same pass as the video. Supports first-frame and last-frame conditioning.
Distinctive
What makes Veo 3.1 distinctive
Native audio generated with the picture.
Veo 3.1 produces sound in the same pass as the video, so ambience, effects, and dialogue-adjacent audio land in sync with the motion. Scenema keeps the audio baked into the file rather than dubbing it on afterwards.
Cinematic-grade motion and lighting realism.
Veo 3.1 is trained for cinematographic style: parallax on camera moves, plausible physics on interactions, and lighting that holds shape across the shot. It is the model people reach for when a Scenema shot needs to feel cinematic.
Text-to-video and image-to-video from a single reference.
Point Veo 3.1 at a prompt for text-to-video, or at a starting frame for image-to-video. Scenema exposes both paths so you can either write the shot from scratch or anchor it to a keyframe you have already generated.
Scenema assembles Veo shots into long-form video.
A Veo 3.1 generation runs 4, 6, or 8 seconds. Scenema packs many Veo generations into a scene, and many scenes into a full-length video, holding character identity, voiceover, and music across every cut. Veo produces the cinematic take. Scenema produces the finished long-form video.
Use cases
Use cases for Veo 3.1
Cinematic hero shots inside a longer video.
Reach for Veo 3.1 when a specific shot in your project needs to be the visual peak. The rest of the scene can run on a faster model and Veo 3.1 handles the moment the audience is meant to remember.
Product and brand explainers.
Product explainers lean on realism when the product itself is on screen. Veo 3.1 gives Scenema the motion quality that a hero product shot needs inside a longer explainer.
Establishing shots and B-roll.
Wide establishing shots, aerial passes, and scene-setting B-roll are the shots where realism matters most. Veo 3.1 holds together on the big canvas where lower-tier models tend to reveal artefacts.
Chapter openers and transitions.
A chapter opener sets the tone for the minutes that follow. Veo 3.1 responds to camera direction in the prompt, so a Scenema shot written as a sweeping opener reads that way on screen.
Compare
How Veo 3.1 compares to other video models
See where Veo 3.1 slots in the lineup on the axes that matter most.
| Model | Best for | Shot length | Max resolution | Inputs | Quality tier |
|---|---|---|---|---|---|
| Veo 3.1 | Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video. | 4, 6, or 8 seconds | Up to 1080p | Text to video, Image to video | High |
| Kling 3.0 | Kling 3.0 is Kuaishou’s flagship generative video model, running as one of the shot generators inside a Scenema long-form video. | 4, 6, 8, 10, or 12 seconds | 1080p | Text to video, Image to video | High |
| Seedance 2.0 | Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video. | 4 to 15 seconds | Up to 720p | Text to video, Image to video, Audio to video | High |
| Seedance 1.5 | Seedance 1.5 is ByteDance’s image-to-video model, running as one of the shot generators inside a Scenema long-form video. | 4 to 12 seconds | Up to 720p | Image to video | Standard |
| Wan 3.0 | Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video. | 4 to 12 seconds | 720p | Image to video | High |
| MiniMax H3 | MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent. | 4 to 15 seconds | 768p | Image to video, Reference to video | High |
| Vidu Q3 Turbo | Vidu Q3 Turbo is a start-and-end-frame video model on Scenema, tuned for shots that must land on an exact ending composition. | 1 to 15 seconds | 720p | Image to video | Standard |
| Gemini Omni Flash | Gemini Omni Flash is Google’s multimodal video model, running as one of the shot generators inside a Scenema long-form video. | 4, 6, 8, or 10 seconds | Up to 4K | Image to video | High |
| LTX 2.3 | LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes. | 4 to 12 seconds | Up to 1080p | Text to video, Image to video, Audio to video | Fast and Pro tiers |
Where it fits
Where Scenema uses Veo 3.1
Scenema runs a highly optimized workflow and uses the best model for the style and for each task under the hood. You choose the quality and the resolution. Here is where Veo 3.1 fits and where a different model does.
- 1
Scenema uses Veo 3.1 when the shot has to carry the video.
Veo 3.1 is the highest-quality video model on Scenema for cinematic realism. Save it for the shots the audience is meant to remember: the establishing shot, the money take in an ad, the closing beat of a narrative scene.
- 2
Scenema uses Veo 3.1 when the shot has to carry its own audio.
Veo 3.1 generates audio in the same pass as the picture. Ambience, effects, and impact hits land in sync with the motion instead of being dubbed on afterwards. That saves a separate sound-design pass.
- 3
Scenema uses Veo 3.1 for hero product shots.
A hero product shot is judged on motion realism and audio fidelity. Veo 3.1 delivers both in one generation, at up to 1080p.
- 4
Scenema uses a faster model when you are still iterating on framing.
Veo 3.1 is slower per generation than the standard-tier video models. For exploration passes on framing, blocking, or pacing, run the LTX 2.3 Fast tier and it returns a shot in a fraction of the time. Move back to Veo 3.1 once the shot is locked.
- 5
Scenema uses another model when your shot needs to be longer than 8 seconds.
A Veo 3.1 generation is 4, 6, or 8 seconds. If a single sustained take has to run 12 seconds, use Kling 3.0. If it has to run up to 15 seconds, use Seedance 2.0 or MiniMax H3. Scenema still assembles the longer shots into scenes and full videos the same way.
FAQ
Frequently asked questions about Veo 3.1
What is Veo 3.1 on Scenema?+
Veo 3.1 is Google DeepMind’s flagship video generation model, hosted on Scenema for text-to-video and image-to-video generation. Every Veo 3.1 generation on Scenema produces a 4, 6, or 8 second shot at 720p or 1080p, with audio generated in the same pass as the picture.
What resolution and aspect ratios does Veo 3.1 support on Scenema?+
Veo 3.1 supports 720p and 1080p output on Scenema, in 16:9 and 9:16 aspect ratios. The default output resolution is 720p; step up to 1080p when the shot is the visual centrepiece of the scene.
How long is a single Veo 3.1 generation?+
A single Veo 3.1 generation on Scenema is a 4, 6, or 8 second shot. Long-form output comes from Scenema itself, which packs many shots into a scene and many scenes into a full video. The per-generation limit is what Veo returns, not the length of the video you can make on Scenema.
How does Scenema make long-form videos if Veo only generates up to 8 seconds at a time?+
Scenema is the layer above the model. It packs multiple Veo generations into a scene, and multiple scenes into a full-length video, holding character identity, voiceover, and music across every cut. Veo produces the individual cinematic shots; Scenema produces the finished long-form video.
When does Scenema use Veo 3.1 over a faster Scenema model?+
Scenema uses Veo 3.1 when the shot needs to be the visual peak of your video, when realism is critical, or when the shot should carry its own audio. For fast iterations, or for shots outside the Veo range, Scenema uses a faster model like Wan 3.0 or Seedance 2.0 and reserves Veo 3.1 for the shots that matter most.
Related
Other video models on Scenema
Kling 3.0
Kuaishou- Kuaishou’s third-generation flagship video model.
- 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 2.0
ByteDance- ByteDance’s multimodal video model with joint audio-video generation.
- Text, image, and audio inputs; up to 15 second takes, the longest on Scenema alongside MiniMax H3 and Vidu Q3 Turbo.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 1.5
ByteDance- ByteDance’s image-to-video model, driven from a reference frame.
- 4 to 12 second takes at up to 720p.
- Image-to-video shot generator inside Scenema’s long-form pipeline.
Wan 3.0
Alibaba- Alibaba Tongyi Wanxiang’s current-generation Wan model, replacing Wan 2.6 on Scenema.
- Image-to-video with last-frame conditioning; 4 to 12 second takes at 720p.
- Image-anchored shot generator inside Scenema’s long-form pipeline.
MiniMax H3
MiniMax- MiniMax’s next-generation video model.
- Reference-to-video; 4 to 15 second takes at 768p.
- Complex motion and multi-entity consistency inside Scenema’s long-form pipeline.
Vidu Q3 Turbo
Shengshu- Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
- 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
- Locks the shot to an exact ending composition for clean cut-to-cut choreography.
Gemini Omni Flash
Google- Google’s multimodal video model with native audio and 4K output.
- Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
- Highest-resolution shot generator inside Scenema’s long-form pipeline.
LTX 2.3
Lightricks- Lightricks’ open video foundation model with Fast and Pro tiers on Scenema.
- Text, image, and audio inputs; 4 to 12 second takes at up to 1080p.
- Widest aspect-ratio range of any Scenema video model.
Start generating with Veo 3.1 in Scenema
Veo 3.1 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.
200 credits on signup, then weekly refills through your first month. No credit card required.