Lightricks · Video models

LTX 2.3

LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes.

Try LTX 2.3 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

LTX 2.3 on Scenema

LTX 2.3 is Lightricks’ open video foundation model, built on a diffusion transformer trained jointly on video and audio for synchronised output in a single pass. On Scenema, LTX 2.3 runs as one of the shot generators inside the long-form video pipeline, in text-to-video, image-to-video, and audio-to-video modes. The model ships in two tiers. Fast returns takes in about 5 seconds at 720p and 10 seconds at 1080p, tuned for exploration and iteration passes. Pro takes longer per generation and produces the higher-fidelity take that survives the final edit. Both tiers share the same input surface, the same 4 to 12 second range, and the widest aspect-ratio range of any Scenema video model. Scenema packs those takes into scenes, and scenes into full-length videos with character consistency, voiceover, and music running across every cut.

Specifications

LTX 2.3 specifications on Scenema

Supports
Text to video, Image to video, Audio to video
Duration
4 to 12 seconds
Resolution
Up to 1080p
Aspect ratios
16:9, 9:16, 4:3, 3:4, 3:2, 2:3, and 1:1
Quality tier
Fast and Pro tiers
Notes
Fast tier: about 5 seconds at 720p, 10 seconds at 1080p per generation. Pro tier: higher-fidelity output at longer render times.

Distinctive

What makes LTX 2.3 distinctive

LTX 2.3 ships in Fast and Pro tiers on Scenema.

Fast is the speed tier tuned for iteration; Pro is the quality tier tuned for delivery. Both share the same endpoint, input contract, and duration range, so a shot iterated on Fast can be re-run on Pro without changing any prompt or reference.

LTX 2.3 covers text, image, and audio inputs.

A single call can be driven from a written prompt, a reference frame, or an audio clip. That range is unusual inside one model and makes LTX 2.3 the versatile all-rounder in the Scenema video lineup.

LTX 2.3 offers the widest aspect-ratio range on Scenema.

Shots can be generated in 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, and 1:1, more than any other Scenema video model. Useful for print-adjacent and non-standard formats.

Scenema assembles LTX 2.3 shots into long-form video.

Each LTX 2.3 generation is one shot inside a Scenema scene. Scenema plans the shot list, runs each shot in the right input mode and tier, and assembles the results into full videos with continuous character, voiceover, and music tracks.

Use cases

Use cases for LTX 2.3

Iterate on Fast, deliver on Pro.

Explore framing, blocking, and pacing on the Fast tier at about 5 seconds per generation, then re-render the winner on the Pro tier for the delivery-quality take.

Unusual aspect ratios.

When a scene needs 4:3, 3:2, or 2:3, LTX 2.3 is the model that supports those formats natively.

Audio-driven scenes.

Feed LTX 2.3 an audio input and it generates the picture aligned to that sound in the same pass.

Mixed-input narrative video.

Combine text, image, and audio driven shots inside the same Scenema scene, all rendered through the same model for a consistent look.

Compare

How LTX 2.3 compares to other video models

See where LTX 2.3 slots in the lineup on the axes that matter most.

ModelBest forShot lengthMax resolutionInputsQuality tier
Veo 3.1Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video.4, 6, or 8 secondsUp to 1080pText to video, Image to videoHigh
Kling 3.0Kling 3.0 is Kuaishou’s flagship generative video model, running as one of the shot generators inside a Scenema long-form video.4, 6, 8, 10, or 12 seconds1080pText to video, Image to videoHigh
Seedance 2.0Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.4 to 15 secondsUp to 720pText to video, Image to video, Audio to videoHigh
Seedance 1.5Seedance 1.5 is ByteDance’s image-to-video model, running as one of the shot generators inside a Scenema long-form video.4 to 12 secondsUp to 720pImage to videoStandard
Wan 3.0Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video.4 to 12 seconds720pImage to videoHigh
MiniMax H3MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent.4 to 15 seconds768pImage to video, Reference to videoHigh
Vidu Q3 TurboVidu Q3 Turbo is a start-and-end-frame video model on Scenema, tuned for shots that must land on an exact ending composition.1 to 15 seconds720pImage to videoStandard
Gemini Omni FlashGemini Omni Flash is Google’s multimodal video model, running as one of the shot generators inside a Scenema long-form video.4, 6, 8, or 10 secondsUp to 4KImage to videoHigh
LTX 2.3LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes.4 to 12 secondsUp to 1080pText to video, Image to video, Audio to videoFast and Pro tiers

FAQ

Frequently asked questions about LTX 2.3

What is LTX 2.3 on Scenema?+

LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema in two tiers. Fast returns takes in about 5 seconds at 720p, tuned for iteration. Pro takes longer and produces the delivery-quality take. Both tiers accept text, image, and audio inputs and generate 4 to 12 second shots.

What is the difference between the Fast and Pro tiers of LTX 2.3?+

Same model, same endpoint, different speed and quality tradeoff. Fast returns shots in about 5 seconds at 720p and 10 seconds at 1080p. Pro takes longer per generation and produces higher-fidelity output. Use Fast to iterate, Pro to deliver.

What resolution and aspect ratios does LTX 2.3 support on Scenema?+

Up to 1080p in 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, and 1:1. That is the widest aspect-ratio range of any Scenema video model, on both Fast and Pro tiers.

How long is a single LTX 2.3 generation?+

Between 4 and 12 seconds. Scenema packs many LTX 2.3 generations into a scene, and many scenes into a full-length video, so long-form output comes from the pipeline above the model rather than from a single generation.

When does Scenema use LTX 2.3 over another Scenema model?+

Scenema uses LTX 2.3 when a shot needs the widest aspect-ratio range, when one model has to accept text, image, or audio input, or while framing is still being explored before a shot moves to a higher-tier model.

Related

Other video models on Scenema

Veo 3.1

Google DeepMind
  • Google DeepMind’s flagship cinematic video model.
  • 4, 6, or 8 second takes at up to 1080p, with audio generated in the same pass as the picture.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Kling 3.0

Kuaishou
  • Kuaishou’s third-generation flagship video model.
  • 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Seedance 2.0

ByteDance
  • ByteDance’s multimodal video model with joint audio-video generation.
  • Text, image, and audio inputs; up to 15 second takes, the longest on Scenema alongside MiniMax H3 and Vidu Q3 Turbo.
  • Shot generator inside Scenema’s long-form pipeline.
View model →

Seedance 1.5

ByteDance
  • ByteDance’s image-to-video model, driven from a reference frame.
  • 4 to 12 second takes at up to 720p.
  • Image-to-video shot generator inside Scenema’s long-form pipeline.
View model →

Wan 3.0

Alibaba
  • Alibaba Tongyi Wanxiang’s current-generation Wan model, replacing Wan 2.6 on Scenema.
  • Image-to-video with last-frame conditioning; 4 to 12 second takes at 720p.
  • Image-anchored shot generator inside Scenema’s long-form pipeline.
View model →

MiniMax H3

MiniMax
  • MiniMax’s next-generation video model.
  • Reference-to-video; 4 to 15 second takes at 768p.
  • Complex motion and multi-entity consistency inside Scenema’s long-form pipeline.
View model →

Vidu Q3 Turbo

Shengshu
  • Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
  • 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
  • Locks the shot to an exact ending composition for clean cut-to-cut choreography.
View model →

Gemini Omni Flash

Google
  • Google’s multimodal video model with native audio and 4K output.
  • Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
  • Highest-resolution shot generator inside Scenema’s long-form pipeline.
View model →

Start generating with LTX 2.3 in Scenema

LTX 2.3 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.

Try LTX 2.3 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

LTX 2.3 on Scenema | Video models