Google · Image models
Nano Banana 2
Google Nano Banana 2 is the default image model behind every Scenema keyframe and reference image.
200 credits on signup, then weekly refills through your first month. No credit card required.
Nano Banana 2 on Scenema
Nano Banana 2 is Google’s Gemini 3.1 Flash Image model, released in February 2026 and made default across the Gemini app, Google Search AI Mode, and Google Flow. Scenema hosts Nano Banana 2 as the default image model for every project, covering both keyframes that anchor a video shot and reference images that lock character, wardrobe, and product identity. The model supports up to five consistent characters and up to fourteen reference objects in a single generation, which maps directly onto Scenema entity manifests. Nano Banana 2 handles text rendering with roughly 95 percent higher accuracy than the previous version, so on-screen signs, posters, and title cards survive the trip into video. Scenema exposes Nano Banana 2 at up to 1080p in 16:9, 9:16, and 1:1.
Specifications
Nano Banana 2 specifications on Scenema
- Supports
- Text to image, Image to image
- Resolution
- Up to 1080p
- Aspect ratios
- 16:9, 9:16, and 1:1
- Quality tier
- High
- Notes
- Default image model on Scenema for keyframes and reference images.
Distinctive
What makes Nano Banana 2 distinctive
Nano Banana 2 is the default image path on Scenema.
Every new project uses Nano Banana 2 for keyframes and reference images unless an operator picks a different model. The default reflects the model’s balance of speed, character consistency, and text legibility across most shot types.
Nano Banana 2 holds character identity across a scene.
Google’s model supports up to five consistent characters and up to fourteen consistent objects in a single generation. Scenema uses that behaviour when a shot needs the same actor, wardrobe, and prop pack rendered from a new angle for the next keyframe.
Nano Banana 2 renders legible in-image text.
Signs, menus, posters, and title cards render with roughly 95 percent higher accuracy than the previous Nano Banana. Text that survives the still frame usually survives the downstream video generation as well.
Nano Banana 2 pulls world knowledge from Gemini.
The model shares Gemini’s real-world grounding, so specific landmarks, product references, and named subjects render with fewer hallucinated details. That reduces the number of reference passes an operator needs before locking a shot.
Use cases
Use cases for Nano Banana 2
Nano Banana 2 anchors the opening keyframe of a shot.
A prompt with camera framing, lighting, and blocking becomes the still image that seeds video generation. This is the most common use of Nano Banana 2 inside Scenema.
Nano Banana 2 produces character reference images for entity manifests.
A single prompt yields a clean reference plate that Scenema saves to the entity manifest, then reuses across every shot in the project.
Nano Banana 2 renders on-screen text for posters, signs, and lower-thirds.
The model handles readable typography inside a shot without needing a separate compositing pass.
Nano Banana 2 iterates on a shot from an existing image.
Image-to-image passes let an operator revise wardrobe, lighting, or background on a locked plate before sending the still into video generation.
Compare
How Nano Banana 2 compares to other image models
See where Nano Banana 2 slots in the lineup on the axes that matter most.
| Model | Best for | Max resolution | Inputs | Quality tier |
|---|---|---|---|---|
| Nano Banana 2 | Google Nano Banana 2 is the default image model behind every Scenema keyframe and reference image. | Up to 1080p | Text to image, Image to image | High |
| Seedream 4 | ByteDance Seedream 4 is the image model Scenema reaches for when a shot needs several candidate keyframes in one pass. | Up to 1080p | Text to image, Image to image | High |
| GPT Image 2 | OpenAI GPT Image 2 is the image model Scenema uses when a shot needs strict prompt adherence and precise on-screen text. | Up to 1080p | Text to image, Image to image | High |
| Flux 2 Klein 9B | Black Forest Labs Flux 2 Klein 9B is the fast, distilled image model Scenema uses for high-volume keyframe and reference passes. | Up to 1080p | Text to image, Image to image | Standard |
| Z-Image Turbo | Alibaba Tongyi Z-Image Turbo is the fast text-to-image model Scenema uses for reference image generation only. | Up to 1080p | Text to image | Standard |
Where it fits
Where Scenema uses Nano Banana 2
Scenema runs a highly optimized workflow and uses the best model for the style and for each task under the hood. You choose the quality and the resolution. Here is where Nano Banana 2 fits and where a different model does.
- 1
Scenema uses Nano Banana 2 when you need identity to hold across a long video.
Nano Banana 2 is the default keyframe and reference image generator on Scenema because it can carry up to 5 consistent characters and 14 reference objects through a single generation. That is the widest identity window of any image model here, which matters when a scene has an ensemble cast or a hero product that has to look identical from shot one to shot forty.
- 2
Scenema uses Nano Banana 2 when the shot contains readable text.
Text rendering on Nano Banana 2 is roughly 95% more accurate than the original Nano Banana. Signage, book covers, phone screens, and product labels come out legible instead of garbled, so you can use those keyframes directly rather than fixing them in post.
- 3
Scenema uses Nano Banana 2 when the prompt leans on real-world knowledge.
The model is grounded in Gemini’s world knowledge, so specific references to real locations, historical costume, brand-adjacent styling, or accurate physical setups render correctly on the first pass.
- 4
Scenema uses a different model when you want batch variations of the same shot.
Nano Banana 2 returns one image per generation. If you want four candidate compositions of the same beat with identity locked across all four, Seedream 4 is built for that.
- 5
Scenema uses a different model when the prompt is a long multi-clause description.
Nano Banana 2 renders quickly but does not reason over the prompt before drawing. For dense prompts with layered spatial or narrative instructions, GPT Image 2 runs a reasoning pass first and follows the brief more literally.
FAQ
Frequently asked questions about Nano Banana 2
How does Scenema use Nano Banana 2 inside a project?+
Nano Banana 2 is the default image generator for both jobs Scenema needs an image for. First, it renders the keyframe that anchors a video shot. Second, it renders the reference images stored in the entity manifest for characters, wardrobe, and products.
Why is Nano Banana 2 the default image model on Scenema?+
The model produces the highest single-pass character consistency and text accuracy across the image models Scenema hosts, at Flash-tier speed. That mix fits the broadest range of Scenema shot types with the fewest revisions.
What resolutions and aspect ratios does Scenema expose for Nano Banana 2?+
Up to 1080p, in 16:9, 9:16, and 1:1. The image feeds into a video shot at the same aspect ratio.
Does Nano Banana 2 support image-to-image on Scenema?+
Yes. An operator can pass a reference plate and a prompt to iterate on an existing keyframe, which is useful when a shot needs a wardrobe or lighting change without re-rolling the whole composition.
Who built Nano Banana 2?+
Google DeepMind. The model is publicly branded as Nano Banana 2 and technically named Gemini 3.1 Flash Image, released February 2026.
Related
Other image models on Scenema

Seedream 4
ByteDance- ByteDance’s image model with unified text-to-image and image editing.
- Batch mode returns 4 candidate images per prompt, with identity held across the batch.
- Keyframe and reference image generator on Scenema.

GPT Image 2
OpenAI- OpenAI’s image model with a reasoning pass before rendering.
- Roughly 99% in-frame text accuracy across Latin, Japanese, Korean, Hindi, Bengali, and Arabic.
- Keyframe and reference image generator on Scenema.

Flux 2 Klein 9B
Black Forest Labs- Black Forest Labs’ distilled 9-billion-parameter Flux 2 variant.
- Renders a full image in about 4 inference steps.
- High-throughput keyframe and reference image generator on Scenema.

Z-Image Turbo
Alibaba Tongyi- Alibaba Tongyi Lab’s 6-billion-parameter distilled text-to-image model.
- Sub-second inference, text-to-image only, roughly 8-step photorealism.
- Reference image generator on Scenema, not used for keyframes.
Start generating with Nano Banana 2 in Scenema
Nano Banana 2 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.
200 credits on signup, then weekly refills through your first month. No credit card required.
