Lightricks · Video models
LTX 2.3
LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes.
200 credits on signup, then weekly refills through your first month. No credit card required.
LTX 2.3 on Scenema
LTX 2.3 is Lightricks’ open video foundation model, built on a diffusion transformer trained jointly on video and audio for synchronised output in a single pass. On Scenema, LTX 2.3 runs as one of the shot generators inside the long-form video pipeline, in text-to-video, image-to-video, and audio-to-video modes. The model ships in two tiers. Fast returns takes in about 5 seconds at 720p and 10 seconds at 1080p, tuned for exploration and iteration passes. Pro takes longer per generation and produces the higher-fidelity take that survives the final edit. Both tiers share the same input surface, the same 4 to 12 second range, and the widest aspect-ratio range of any Scenema video model. Scenema packs those takes into scenes, and scenes into full-length videos with character consistency, voiceover, and music running across every cut.
Specifications
LTX 2.3 specifications on Scenema
- Supports
- Text to video, Image to video, Audio to video
- Duration
- 4 to 12 seconds
- Resolution
- Up to 1080p
- Aspect ratios
- 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, and 1:1
- Quality tier
- Fast and Pro tiers
- Notes
- Fast tier: about 5 seconds at 720p, 10 seconds at 1080p per generation. Pro tier: higher-fidelity output at longer render times.
Distinctive
What makes LTX 2.3 distinctive
LTX 2.3 ships in Fast and Pro tiers on Scenema.
Fast is the speed tier tuned for iteration; Pro is the quality tier tuned for delivery. Both share the same endpoint, input contract, and duration range, so a shot iterated on Fast can be re-run on Pro without changing any prompt or reference.
LTX 2.3 covers text, image, and audio inputs.
A single call can be driven from a written prompt, a reference frame, or an audio clip. That range is unusual inside one model and makes LTX 2.3 the versatile all-rounder in the Scenema video lineup.
LTX 2.3 offers the widest aspect-ratio range on Scenema.
Shots can be generated in 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, and 1:1, more than any other Scenema video model. Useful for print-adjacent and non-standard formats.
Scenema assembles LTX 2.3 shots into long-form video.
Each LTX 2.3 generation is one shot inside a Scenema scene. Scenema plans the shot list, runs each shot in the right input mode and tier, and assembles the results into full videos with continuous character, voiceover, and music tracks.
Use cases
Use cases for LTX 2.3
Iterate on Fast, deliver on Pro.
Explore framing, blocking, and pacing on the Fast tier at about 5 seconds per generation, then re-render the winner on the Pro tier for the delivery-quality take.
Unusual aspect ratios.
When a scene needs 4:3, 3:2, or 2:3, LTX 2.3 is the model that supports those formats natively.
Audio-driven scenes.
Feed LTX 2.3 an audio input and it generates the picture aligned to that sound in the same pass.
Mixed-input narrative video.
Combine text, image, and audio driven shots inside the same Scenema scene, all rendered through the same model for a consistent look.
Compare
How LTX 2.3 compares to other video models
See where LTX 2.3 slots in the lineup on the axes that matter most.
| Model | Best for | Shot length | Max resolution | Inputs | Quality tier |
|---|---|---|---|---|---|
| Veo 3.1 | Google DeepMind Veo 3.1 as the cinematic shot generator inside a Scenema long-form video. | 4, 6, or 8 seconds | Up to 1080p | Text to video, Image to video | High |
| Kling 3.0 | Kling 3.0 is Kuaishou’s flagship generative video model, running as one of the shot generators inside a Scenema long-form video. | 4, 6, 8, 10, or 12 seconds | 1080p | Text to video, Image to video | High |
| Seedance 2.0 | Seedance 2.0 is ByteDance’s multimodal video model, running as one of the shot generators inside a Scenema long-form video. | 4 to 15 seconds | Up to 720p | Text to video, Image to video, Audio to video | High |
| Seedance 1.5 | Seedance 1.5 is ByteDance’s image-to-video model, running as one of the shot generators inside a Scenema long-form video. | 4 to 12 seconds | Up to 720p | Image to video | Standard |
| Wan 3.0 | Wan 3.0 is Alibaba’s third-generation Tongyi Wanxiang video model, running as one of the image-to-video shot generators inside a Scenema long-form video. | 4 to 12 seconds | 720p | Image to video | High |
| MiniMax H3 | MiniMax H3 is the video model Scenema uses for shots with complex motion and many different characters, props, and places that all have to stay consistent. | 4 to 15 seconds | 768p | Image to video, Reference to video | High |
| Vidu Q3 Turbo | Vidu Q3 Turbo is a start-and-end-frame video model on Scenema, tuned for shots that must land on an exact ending composition. | 1 to 15 seconds | 720p | Image to video | Standard |
| Gemini Omni Flash | Gemini Omni Flash is Google’s multimodal video model, running as one of the shot generators inside a Scenema long-form video. | 4, 6, 8, or 10 seconds | Up to 4K | Image to video | High |
| LTX 2.3 | LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema with a Fast tier for iteration and a Pro tier for delivery-quality takes. | 4 to 12 seconds | Up to 1080p | Text to video, Image to video, Audio to video | Fast and Pro tiers |
FAQ
Frequently asked questions about LTX 2.3
What is LTX 2.3 on Scenema?+
LTX 2.3 is Lightricks’ open video foundation model, hosted on Scenema in two tiers. Fast returns takes in about 5 seconds at 720p, tuned for iteration. Pro takes longer and produces the delivery-quality take. Both tiers accept text, image, and audio inputs and generate 4 to 12 second shots.
What is the difference between the Fast and Pro tiers of LTX 2.3?+
Same model, same endpoint, different speed and quality tradeoff. Fast returns shots in about 5 seconds at 720p and 10 seconds at 1080p. Pro takes longer per generation and produces higher-fidelity output. Use Fast to iterate, Pro to deliver.
What resolution and aspect ratios does LTX 2.3 support on Scenema?+
Up to 1080p in 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, and 1:1. That is the widest aspect-ratio range of any Scenema video model, on both Fast and Pro tiers.
How long is a single LTX 2.3 generation?+
Between 4 and 12 seconds. Scenema packs many LTX 2.3 generations into a scene, and many scenes into a full-length video, so long-form output comes from the pipeline above the model rather than from a single generation.
When does Scenema use LTX 2.3 over another Scenema model?+
Scenema uses LTX 2.3 when a shot needs the widest aspect-ratio range, when one model has to accept text, image, or audio input, or while framing is still being explored before a shot moves to a higher-tier model.
Related
Other video models on Scenema
Veo 3.1
Google DeepMind- Google DeepMind’s flagship cinematic video model.
- 4, 6, or 8 second takes at up to 1080p, with audio generated in the same pass as the picture.
- Shot generator inside Scenema’s long-form pipeline.
Kling 3.0
Kuaishou- Kuaishou’s third-generation flagship video model.
- 4 to 12 second takes at 1080p, with native audio and stronger character consistency.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 2.0
ByteDance- ByteDance’s multimodal video model with joint audio-video generation.
- Text, image, and audio inputs; up to 15 second takes, the longest on Scenema alongside MiniMax H3 and Vidu Q3 Turbo.
- Shot generator inside Scenema’s long-form pipeline.
Seedance 1.5
ByteDance- ByteDance’s image-to-video model, driven from a reference frame.
- 4 to 12 second takes at up to 720p.
- Image-to-video shot generator inside Scenema’s long-form pipeline.
Wan 3.0
Alibaba- Alibaba Tongyi Wanxiang’s current-generation Wan model, replacing Wan 2.6 on Scenema.
- Image-to-video with last-frame conditioning; 4 to 12 second takes at 720p.
- Image-anchored shot generator inside Scenema’s long-form pipeline.
MiniMax H3
MiniMax- MiniMax’s next-generation video model.
- Reference-to-video; 4 to 15 second takes at 768p.
- Complex motion and multi-entity consistency inside Scenema’s long-form pipeline.
Vidu Q3 Turbo
Shengshu- Shengshu’s Vidu Q3 Turbo, a start-and-end-frame image-to-video model.
- 1 to 15 second takes at 720p, driven by both a first frame and a last frame.
- Locks the shot to an exact ending composition for clean cut-to-cut choreography.
Gemini Omni Flash
Google- Google’s multimodal video model with native audio and 4K output.
- Image to video only; 4, 6, 8, or 10 second takes at up to 4K.
- Highest-resolution shot generator inside Scenema’s long-form pipeline.
Start generating with LTX 2.3 in Scenema
LTX 2.3 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.
200 credits on signup, then weekly refills through your first month. No credit card required.