Alibaba Tongyi · Image models
Z-Image Turbo
Alibaba Tongyi Z-Image Turbo is the fast text-to-image model Scenema uses for reference image generation only.
200 credits on signup, then weekly refills through your first month. No credit card required.
Z-Image Turbo on Scenema
Z-Image Turbo is Alibaba Tongyi Lab’s distilled text-to-image model, a 6-billion-parameter Single-Stream Diffusion Transformer released in late 2025 and built for sub-second inference on high-end GPUs. Scenema hosts Z-Image Turbo for reference image generation only. The model is not used as a keyframe generator on Scenema. When an operator needs a fast character, wardrobe, or product plate to seed an entity manifest, Z-Image Turbo turns the prompt into a reference image that can then feed a keyframe rendered by one of Scenema’s keyframe-tier models. Scenema exposes Z-Image Turbo at up to 1080p in 16:9, 9:16, and 1:1, text-to-image only.
Specifications
Z-Image Turbo specifications on Scenema
- Supports
- Text to image
- Resolution
- Up to 1080p
- Aspect ratios
- 16:9, 9:16, and 1:1
- Quality tier
- Standard
- Notes
- Used for reference image generation on Scenema, not for keyframes.
Distinctive
What makes Z-Image Turbo distinctive
Z-Image Turbo is a reference image generator on Scenema, not a keyframe generator.
Scenema uses the model to produce reference plates that go into the entity manifest. The plates then feed other models when rendering the keyframe for a shot.
Z-Image Turbo runs text-to-image only on Scenema.
There is no image-to-image path exposed for this model. An operator uses it to turn a written brief into a reference plate from scratch.
Z-Image Turbo produces a photoreal plate in about eight steps.
Alibaba’s distilled turbo variant reaches usable photorealism at very low step counts, which fits the moment in a Scenema project where an operator needs a reference plate quickly.
Z-Image Turbo holds strong prompt adherence at 6B parameters.
Public testing places Z-Image Turbo above other efficient open models on instruction following, so reference plates match the brief on the first pass more often than not.
Use cases
Use cases for Z-Image Turbo
Z-Image Turbo produces a character reference plate for the entity manifest.
A written character brief becomes a photoreal reference image, then Scenema stores that plate for reuse across the project.
Z-Image Turbo produces a wardrobe reference plate.
A garment description becomes a clean plate that other models pull from when rendering shots that involve the character wearing the item.
Z-Image Turbo produces a product reference plate.
A product brief becomes a reference image suitable for storing in the entity manifest, then reused across shots.
Z-Image Turbo produces a fast first-pass reference before an operator commits.
The sub-second latency lets an operator sketch several plate options before promoting one to the entity manifest.
Compare
How Z-Image Turbo compares to other image models
See where Z-Image Turbo slots in the lineup on the axes that matter most.
| Model | Best for | Max resolution | Inputs | Quality tier |
|---|---|---|---|---|
| Nano Banana 2 | Google Nano Banana 2 is the default image model behind every Scenema keyframe and reference image. | Up to 1080p | Text to image, Image to image | High |
| Seedream 4 | ByteDance Seedream 4 is the image model Scenema reaches for when a shot needs several candidate keyframes in one pass. | Up to 1080p | Text to image, Image to image | High |
| GPT Image 2 | OpenAI GPT Image 2 is the image model Scenema uses when a shot needs strict prompt adherence and precise on-screen text. | Up to 1080p | Text to image, Image to image | High |
| Flux 2 Klein 9B | Black Forest Labs Flux 2 Klein 9B is the fast, distilled image model Scenema uses for high-volume keyframe and reference passes. | Up to 1080p | Text to image, Image to image | Standard |
| Z-Image Turbo | Alibaba Tongyi Z-Image Turbo is the fast text-to-image model Scenema uses for reference image generation only. | Up to 1080p | Text to image | Standard |
Where it fits
Where Scenema uses Z-Image Turbo
Scenema runs a highly optimized workflow and uses the best model for the style and for each task under the hood. You choose the quality and the resolution. Here is where Z-Image Turbo fits and where a different model does.
- 1
Scenema uses Z-Image Turbo when you need a reference plate right now.
Z-Image Turbo runs at sub-second inference, which makes it the fastest way to get a reference image on Scenema. Use it to lock a character look or a product silhouette in seconds, then feed that plate into a keyframe generation.
- 2
Scenema uses Z-Image Turbo when you are drafting many reference variants.
The speed makes it practical to generate a large set of first-pass references, pick the strongest, and discard the rest. That is a workflow the slower High-tier models cannot support at the same pace.
- 3
Scenema uses a different model when the output is a keyframe, not a reference.
Z-Image Turbo is used for reference image generation only on Scenema. It is not wired up for keyframes. For the frames that actually seed video shots, Nano Banana 2 is the default keyframe generator.
- 4
Scenema uses a different model when the reference has to start from an existing image.
Z-Image Turbo is text-to-image only and does not accept image inputs. For image-to-image reference work, Flux 2 Klein 9B accepts reference plates and edits at Standard tier speed.
- 5
Scenema uses a different model when the reference has to read text correctly.
Z-Image Turbo is optimised for speed on visual identity, not for text fidelity. If the reference plate needs legible signage or a readable label, GPT Image 2 handles in-frame text accurately across multiple scripts.
FAQ
Frequently asked questions about Z-Image Turbo
How does Scenema use Z-Image Turbo inside a project?+
Scenema uses Z-Image Turbo strictly for reference image generation. The model produces reference plates for the entity manifest, and those plates then feed a separate keyframe-tier model when Scenema renders the actual shot.
Can Z-Image Turbo generate keyframes on Scenema?+
No. Z-Image Turbo is not exposed as a keyframe generator on Scenema. For keyframes, Scenema uses Nano Banana 2, Seedream 4, GPT Image 2, or Flux 2 Klein 9B.
Does Z-Image Turbo support image-to-image on Scenema?+
No. Scenema exposes Z-Image Turbo as text-to-image only.
What resolutions and aspect ratios does Scenema expose for Z-Image Turbo?+
Up to 1080p, in 16:9, 9:16, and 1:1.
Who built Z-Image Turbo?+
Alibaba’s Tongyi Lab. Z-Image Turbo is the distilled variant of the Z-Image series, published on Hugging Face and adopted quickly by the open-source community after release.
Related
Other image models on Scenema

Nano Banana 2
Google- Google’s Gemini 3.1 Flash Image model.
- Up to 5 consistent characters and 14 reference objects per generation.
- Default keyframe and reference image generator on Scenema.

Seedream 4
ByteDance- ByteDance’s image model with unified text-to-image and image editing.
- Batch mode returns 4 candidate images per prompt, with identity held across the batch.
- Keyframe and reference image generator on Scenema.

GPT Image 2
OpenAI- OpenAI’s image model with a reasoning pass before rendering.
- Roughly 99% in-frame text accuracy across Latin, Japanese, Korean, Hindi, Bengali, and Arabic.
- Keyframe and reference image generator on Scenema.

Flux 2 Klein 9B
Black Forest Labs- Black Forest Labs’ distilled 9-billion-parameter Flux 2 variant.
- Renders a full image in about 4 inference steps.
- High-throughput keyframe and reference image generator on Scenema.
Start generating with Z-Image Turbo in Scenema
Z-Image Turbo is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.
200 credits on signup, then weekly refills through your first month. No credit card required.
