MiniMax · Video models

MiniMax H3

MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent.

Try MiniMax H3 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

MiniMax H3 on Scenema

MiniMax H3 is MiniMax’s next-generation video model. Scenema uses it for shots that need complex motion and character consistency across many different types of entities: people, animals, products, vehicles, and places sharing one frame. Every generation takes the shot’s keyframe as a reference, and a shot can continue the previous one as a single unbroken take. Takes run from 4 to 15 seconds at 768p. Scenema packs those takes into scenes, and scenes into full-length explainers with a consistent style, characters, locations, and props in every shot.

Specifications

MiniMax H3 specifications on Scenema

Supports
Image to video, Reference to video
Duration
4 to 15 seconds
Resolution
768p
Aspect ratios
16:9, 9:16, and 1:1
Quality tier
High
Notes
Complex motion and multi-entity consistency on Scenema. Continues the previous shot as one take.

Distinctive

What makes MiniMax H3 distinctive

MiniMax H3 handles complex motion.

Several subjects moving at once, hands at work, objects changing state, and camera moves through a busy scene hold together in a single take.

MiniMax H3 keeps many kinds of entities consistent.

One shot can hold characters, props, vehicles, and a location together, and each stays true to its reference. That is what a long explainer with a recurring cast needs.

MiniMax H3 continues the previous shot.

Scenema can hand MiniMax H3 the last seconds of the previous shot, so the next shot picks up as one unbroken take.

Scenema assembles MiniMax H3 shots into long-form video.

Each MiniMax H3 generation is one shot inside a Scenema scene. Scenema plans the shot list, gives each shot its keyframe and references, and assembles the results into a finished explainer cut to the narration.

Use cases

Use cases for MiniMax H3

Busy explainer scenes.

Scenes where several characters, props, and a location are all in motion, such as a factory floor, a kitchen, or a city street.

Recurring casts across a long video.

The same presenter, mascot, and product return shot after shot and stay recognizable.

Whiteboard and hand-drawn styles.

A drawing hand works across the board, and each element comes alive once it is finished.

Continuous takes across shots.

One camera move or one drawing carried across several shots as a single unbroken take.

Compare

How MiniMax H3 compares to other video models

See where MiniMax H3 slots in the lineup on the axes that matter most.

ModelBest forShot lengthMax resolutionInputsQuality tier
Veo 3.1Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video.4, 6, or 8 secondsUp to 1080pText to video, Image to videoHigh
Kling 3.0Kling 3.0 is Kuaishou’s flagship generative video model, running as one of the shot generators inside a Scenema long-form video.4, 6, 8, 10, or 12 seconds1080pText to video, Image to videoHigh
Seedance 2.0Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.4 to 15 secondsUp to 720pText to video, Image to video, Audio to videoHigh
Seedance 1.5Seedance 1.5 is ByteDance’s image-to-video model, running as one of the shot generators inside a Scenema long-form video.4 to 12 secondsUp to 720pImage to videoStandard
Wan 3.0Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video.4 to 12 seconds720pImage to videoHigh
MiniMax H3MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent.4 to 15 seconds768pImage to video, Reference to videoHigh
Vidu Q3 TurboVidu Q3 Turbo is a start-and-end-frame video model on Scenema, tuned for shots that must land on an exact ending composition.1 to 15 seconds720pImage to videoStandard
Gemini Omni FlashGemini Omni Flash is Google’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.4, 6, 8, or 10 secondsUp to 4KImage to videoHigh
LTX 2.3LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes.4 to 12 secondsUp to 1080pText to video, Image to video, Audio to videoFast and Pro tiers

FAQ

Frequently asked questions about MiniMax H3

What is MiniMax H3 on Scenema?+

MiniMax H3 is MiniMax’s next-generation video model. Scenema uses it as a reference-to-video shot generator for complex motion and for shots that hold many different kinds of entities. Each generation is a shot of 4 to 15 seconds at 768p, and Scenema packs those shots into full videos.

Can MiniMax H3 do text to video on Scenema?+

No. On Scenema every MiniMax H3 generation starts from the shot’s keyframe as a reference. For text-driven shots Scenema uses Seedance 2.0 or Kling 3.0.

How does MiniMax H3 keep entities consistent?+

Scenema composes each shot’s keyframe with every character, prop, and location in place, then gives that keyframe to MiniMax H3 as a reference. The prompt ties each entity to its place in the keyframe, so the model keeps them distinct and on model while they move.

How does Scenema make long-form videos if MiniMax H3 only generates up to 15 seconds at a time?+

Scenema treats each MiniMax H3 generation as one shot inside a scene. The pipeline plans the shot list, supplies each shot’s keyframe and references, and assembles the results into scenes and full videos. The style, characters, locations, and props stay consistent across every cut, and a shot can continue the previous one as a single take.

When does Scenema use MiniMax H3 over another Scenema model?+

Scenema uses MiniMax H3 when a shot needs complex motion and character consistency across many different types of entities, or when it has to continue the previous shot as one unbroken take.

Related

Other video models on Scenema

Veo 3.1

Google DeepMind
  • Google DeepMind’s flagship cinematic video model.
  • 4, 6, or 8 second takes at up to 1080p, with audio generated in the same pass as the picture.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Kling 3.0

Kuaishou
  • Kuaishou’s third-generation flagship video model.
  • 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Seedance 2.0

ByteDance
  • ByteDance’s multimodal video model with joint audio-video generation.
  • Text, image, and audio inputs; up to 15 second takes, the longest on Scenema alongside MiniMax H3 and Vidu Q3 Turbo.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Seedance 1.5

ByteDance
  • ByteDance’s image-to-video model, driven from a reference frame.
  • 4 to 12 second takes at up to 720p.
  • Image-to-video shot generator inside Scenema’s long-form pipeline.
View model →

Wan 3.0

Alibaba
  • Alibaba Tongyi Wanxiang’s current-generation Wan model, replacing Wan 2.6 on Scenema.
  • Image-to-video with last-frame conditioning; 4 to 12 second takes at 720p.
  • Image-anchored shot generator inside Scenema’s long-form pipeline.
View model →

Vidu Q3 Turbo

Shengshu
  • Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
  • 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
  • Locks the shot to an exact ending composition for clean cut-to-cut choreography.
View model →

Gemini Omni Flash

Google
  • Google’s multimodal video model with native audio and 4K output.
  • Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
  • Highest-resolution shot generator inside Scenema’s long-form pipeline.
View model →

LTX 2.3

Lightricks
  • Lightricks’ open video foundation model with Fast and Pro tiers on Scenema.
  • Text, image, and audio inputs; 4 to 12 second takes at up to 1080p.
  • Widest aspect-ratio range of any Scenema video model.
View model →

Start generating with MiniMax H3 in Scenema

MiniMax H3 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.

Try MiniMax H3 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

MiniMax H3 on Scenema | Video models