AI models
Every AI model on Scenema
Scenema runs a highly optimized workflow and uses the best model for the style and for each task under the hood. You choose the quality and the resolution. Here is every model it runs, grouped by what it generates.
200 credits on signup, then weekly refills through your first month. No credit card required.
Video models
Video models on Scenema
Every video generation model Scenema hosts. Text to video, image to video, and audio to video, from cinematic 8-second takes on Veo 3.1 to 15-second shots on Seedance 2.0. Scenema packs generations from these models into scenes, and scenes into full-length videos with character consistency, voiceover, and music running across every cut.
Veo 3.1
Google DeepMind- Google DeepMind’s flagship cinematic video model.
- 4, 6, or 8 second takes at up to 1080p, with audio generated in the same pass as the picture.
- Shot generator inside Scenema’s long-form pipeline.
Kling 3.0
Kuaishou- Kuaishou’s third-generation flagship video model.
- 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 2.0
ByteDance- ByteDance’s multimodal video model with joint audio-video generation.
- Text, image, and audio inputs; up to 15 second takes, the longest on Scenema alongside MiniMax H3 and Vidu Q3 Turbo.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 1.5
ByteDance- ByteDance’s image-to-video model, driven from a reference frame.
- 4 to 12 second takes at up to 720p.
- Image-to-video shot generator inside Scenema’s long-form pipeline.
Wan 3.0
Alibaba- Alibaba Tongyi Wanxiang’s current-generation Wan model, replacing Wan 2.6 on Scenema.
- Image-to-video with last-frame conditioning; 4 to 12 second takes at 720p.
- Image-anchored shot generator inside Scenema’s long-form pipeline.
MiniMax H3
MiniMax- MiniMax’s next-generation video model.
- Reference-to-video; 4 to 15 second takes at 768p.
- Complex motion and multi-entity consistency inside Scenema’s long-form pipeline.
Vidu Q3 Turbo
Shengshu- Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
- 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
- Locks the shot to an exact ending composition for clean cut-to-cut choreography.
Gemini Omni Flash
Google- Google’s multimodal video model with native audio and 4K output.
- Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
- Highest-resolution shot generator inside Scenema’s long-form pipeline.
LTX 2.3
Lightricks- Lightricks’ open video foundation model with Fast and Pro tiers on Scenema.
- Text, image, and audio inputs; 4 to 12 second takes at up to 1080p.
- Widest aspect-ratio range of any Scenema video model.
Image models
Image models on Scenema
The image generation and image-to-image models Scenema hosts, used for keyframes that seed a video shot and reference images that lock character, wardrobe, and product identity across the project.

Nano Banana 2
Google- Google’s Gemini 3.1 Flash Image model.
- Up to 5 consistent characters and 14 reference objects per generation.
- Default keyframe and reference image generator on Scenema.

Seedream 4
ByteDance- ByteDance’s image model with unified text-to-image and image editing.
- Batch mode returns 4 candidate images per prompt, with identity held across the batch.
- Keyframe and reference image generator on Scenema.

GPT Image 2
OpenAI- OpenAI’s image model with a reasoning pass before rendering.
- Roughly 99% in-frame text accuracy across Latin, Japanese, Korean, Hindi, Bengali, and Arabic.
- Keyframe and reference image generator on Scenema.

Flux 2 Klein 9B
Black Forest Labs- Black Forest Labs’ distilled 9-billion-parameter Flux 2 variant.
- Renders a full image in about 4 inference steps.
- High-throughput keyframe and reference image generator on Scenema.

Z-Image Turbo
Alibaba Tongyi- Alibaba Tongyi Lab’s 6-billion-parameter distilled text-to-image model.
- Sub-second inference, text-to-image only, roughly 8-step photorealism.
- Reference image generator on Scenema, not used for keyframes.
Speech models
Speech models on Scenema
Scenema’s text-to-speech engines for voiceover and dialogue, with multilingual coverage and voice cloning across the Base, Turbo, and Pro tiers.
Start generating with any model in Scenema
Every model on this page is available inside Scenema on the free tier. Pick one, generate a shot, and iterate from there.
200 credits on signup, then weekly refills through your first month. No credit card required.