Google · Image models

Nano Banana 2

Google Nano Banana 2 is the default image model behind every Scenema keyframe and reference image.

Try Nano Banana 2 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

Nano Banana 2 on Scenema

Nano Banana 2 is Google’s Gemini 3.1 Flash Image model, released in February 2026 and made default across the Gemini app, Google Search AI Mode, and Google Flow. Scenema hosts Nano Banana 2 as the default image model for every project, covering both keyframes that anchor a video shot and reference images that lock character, wardrobe, and product identity. The model supports up to five consistent characters and up to fourteen reference objects in a single generation, which maps directly onto Scenema entity manifests. Nano Banana 2 handles text rendering with roughly 95 percent higher accuracy than the previous version, so on-screen signs, posters, and title cards survive the trip into video. Scenema exposes Nano Banana 2 at up to 1080p in 16:9, 9:16, and 1:1.

Specifications

Nano Banana 2 specifications on Scenema

Supports
Text to image, Image to image
Resolution
Up to 1080p
Aspect ratios
16:9, 9:16, and 1:1
Quality tier
High
Notes
Default image model on Scenema for keyframes and reference images.

Distinctive

What makes Nano Banana 2 distinctive

Nano Banana 2 is the default image path on Scenema.

Every new project uses Nano Banana 2 for keyframes and reference images unless an operator picks a different model. The default reflects the model’s balance of speed, character consistency, and text legibility across most shot types.

Nano Banana 2 holds character identity across a scene.

Google’s model supports up to five consistent characters and up to fourteen consistent objects in a single generation. Scenema uses that behaviour when a shot needs the same actor, wardrobe, and prop pack rendered from a new angle for the next keyframe.

Nano Banana 2 renders legible in-image text.

Signs, menus, posters, and title cards render with roughly 95 percent higher accuracy than the previous Nano Banana. Text that survives the still frame usually survives the downstream video generation as well.

Nano Banana 2 pulls world knowledge from Gemini.

The model shares Gemini’s real-world grounding, so specific landmarks, product references, and named subjects render with fewer hallucinated details. That reduces the number of reference passes an operator needs before locking a shot.

Use cases

Use cases for Nano Banana 2

Nano Banana 2 anchors the opening keyframe of a shot.

A prompt with camera framing, lighting, and blocking becomes the still image that seeds video generation. This is the most common use of Nano Banana 2 inside Scenema.

Nano Banana 2 produces character reference images for entity manifests.

A single prompt yields a clean reference plate that Scenema saves to the entity manifest, then reuses across every shot in the project.

Nano Banana 2 renders on-screen text for posters, signs, and lower-thirds.

The model handles readable typography inside a shot without needing a separate compositing pass.

Nano Banana 2 iterates on a shot from an existing image.

Image-to-image passes let an operator revise wardrobe, lighting, or background on a locked plate before sending the still into video generation.

Compare

How Nano Banana 2 compares to other image models

See where Nano Banana 2 slots in the lineup on the axes that matter most.

ModelBest forMax resolutionInputsQuality tier
Nano Banana 2Google Nano Banana 2 is the default image model behind every Scenema keyframe and reference image.Up to 1080pText to image, Image to imageHigh
Seedream 4ByteDance Seedream 4 is the image model Scenema reaches for when a shot needs several candidate keyframes in one pass.Up to 1080pText to image, Image to imageHigh
GPT Image 2OpenAI GPT Image 2 is the image model Scenema uses when a shot needs strict prompt adherence and precise on-screen text.Up to 1080pText to image, Image to imageHigh
Flux 2 Klein 9BBlack Forest Labs Flux 2 Klein 9B is the fast, distilled image model Scenema uses for high-volume keyframe and reference passes.Up to 1080pText to image, Image to imageStandard
Z-Image TurboAlibaba Tongyi Z-Image Turbo is the fast text-to-image model Scenema uses for reference image generation only.Up to 1080pText to imageStandard

Where it fits

Where Scenema uses Nano Banana 2

Scenema runs a highly optimized workflow and uses the best model for the style and for each task under the hood. You choose the quality and the resolution. Here is where Nano Banana 2 fits and where a different model does.

  1. 1

    Scenema uses Nano Banana 2 when you need identity to hold across a long video.

    Nano Banana 2 is the default keyframe and reference image generator on Scenema because it can carry up to 5 consistent characters and 14 reference objects through a single generation. That is the widest identity window of any image model here, which matters when a scene has an ensemble cast or a hero product that has to look identical from shot one to shot forty.

  2. 2

    Scenema uses Nano Banana 2 when the shot contains readable text.

    Text rendering on Nano Banana 2 is roughly 95% more accurate than the original Nano Banana. Signage, book covers, phone screens, and product labels come out legible instead of garbled, so you can use those keyframes directly rather than fixing them in post.

  3. 3

    Scenema uses Nano Banana 2 when the prompt leans on real-world knowledge.

    The model is grounded in Gemini’s world knowledge, so specific references to real locations, historical costume, brand-adjacent styling, or accurate physical setups render correctly on the first pass.

  4. 4

    Scenema uses a different model when you want batch variations of the same shot.

    Nano Banana 2 returns one image per generation. If you want four candidate compositions of the same beat with identity locked across all four, Seedream 4 is built for that.

    See Seedream 4 →

  5. 5

    Scenema uses a different model when the prompt is a long multi-clause description.

    Nano Banana 2 renders quickly but does not reason over the prompt before drawing. For dense prompts with layered spatial or narrative instructions, GPT Image 2 runs a reasoning pass first and follows the brief more literally.

    See GPT Image 2 →

FAQ

Frequently asked questions about Nano Banana 2

How does Scenema use Nano Banana 2 inside a project?+

Nano Banana 2 is the default image generator for both jobs Scenema needs an image for. First, it renders the keyframe that anchors a video shot. Second, it renders the reference images stored in the entity manifest for characters, wardrobe, and products.

Why is Nano Banana 2 the default image model on Scenema?+

The model produces the highest single-pass character consistency and text accuracy across the image models Scenema hosts, at Flash-tier speed. That mix fits the broadest range of Scenema shot types with the fewest revisions.

What resolutions and aspect ratios does Scenema expose for Nano Banana 2?+

Up to 1080p, in 16:9, 9:16, and 1:1. The image feeds into a video shot at the same aspect ratio.

Does Nano Banana 2 support image-to-image on Scenema?+

Yes. An operator can pass a reference plate and a prompt to iterate on an existing keyframe, which is useful when a shot needs a wardrobe or lighting change without re-rolling the whole composition.

Who built Nano Banana 2?+

Google DeepMind. The model is publicly branded as Nano Banana 2 and technically named Gemini 3.1 Flash Image, released February 2026.

Related

Other image models on Scenema

Start generating with Nano Banana 2 in Scenema

Nano Banana 2 is available on the Scenema free tier. Sign up and start your first explainer. Scenema uses it wherever it fits the style and the shot.

Try Nano Banana 2 in Scenema

200 credits on signup, then weekly refills through your first month. No credit card required.

Nano Banana 2 on Scenema | Image models