Nano Banana 2 (officially gemini-3.1-flash-image) is Google's reasoning-driven image model — it thinks before it renders, supports up to 14 reference images, and can pull real-time information from Google Search into the image. Most "Nano Banana prompt" lists online are unsourced screenshots. This one is built on Google's actual API documentation, DeepMind's official prompting framework, and the patterns that hold up in production.
Below is the practical guide: what Nano Banana 2 actually wants in a prompt, the 5 patterns that work, the 5 editing patterns Google's docs use as worked examples, and the mistakes that silently break results.
How Nano Banana 2 prompting differs from older Gemini models
Nano Banana 2 is a thinking model. It generates up to 2 interim images during a reasoning step before producing the final image. This means:
- Describe the scene, don't list keywords. Google's official guidance. A narrative paragraph produces a more coherent image than a comma-separated list of terms.
- The model is bilingual about text and image. It was trained to handle image-and-text prompts together, so you can mix a reference image, a text instruction, and another reference image in the same call.
- Camera and lighting language matters more than for Stable Diffusion or older models. "85mm lens, golden hour, soft daylight" lands as actual photoreal. The model knows what these terms mean and renders accordingly.
- Named film stocks work as color-grading cues. "Shot on Kodak Portra 400" gives warm natural skin tones and slight grain. The model recognizes the stock.
The 5-component framework (DeepMind's official structure)
When you need a detailed prompt that holds up, DeepMind's official framework is Style, Subject, Setting, Action, Composition. Use them in that order, and you get a prompt that doesn't fall apart on iteration.
Style — photoreal, illustration, oil painting, 3D render, anime, film noir, watercolor Subject — the person, object, or scene (be specific: age, expression, clothing) Setting — where it happens (kitchen, beach, Tokyo street, 1980s living room) Action — what's happening (running, looking at something, holding an object) Composition — framing, camera angle, focal point (close-up, wide, low-angle)
You don't need all 5 every time. A 3-line prompt with subject + setting + composition works for most social media. Add style and action for portraits and product shots.
5 production-ready patterns
These work on both gemini-3.1-flash-image (Nano Banana 2) and gemini-3-pro-image (Nano Banana Pro). The Pro model handles longer prompts and is the better pick for 4K output.
1. The describe-the-scene portrait (Google's official worked example)
This is the exact prompt Google uses as the worked example in their API docs. Use it as a template — substitute your own subject, keep the camera/lighting language.
A photorealistic close-up portrait of an elderly Japanese ceramicist with deep wrinkles and warm smile inspecting a freshly glazed tea bowl in his rustic workshop, illuminated by golden hour light through a window, captured with an 85mm lens.
Why it works: every visual element (subject, setting, action, lighting, lens) is specific. The model can't generalize because you've constrained everything.
2. The studio product shot (Google's official commercial example)
For ecommerce, ads, packaging mockups, and any product where clean lighting is the difference between "stock photo" and "expensive brand":
High-resolution studio photograph of minimalist matte black ceramic coffee mug on polished concrete, three-point softbox lighting creating soft highlights, 45-degree elevated shot focusing on steam rising from coffee.
Why it works: the "three-point softbox" + "45-degree elevated" are stock commercial photography terms. The model knows them, the output reads as professional.
3. Character consistency (upload a face, keep it across generations)
The workflow that broke Twitter in 2025 and still works in 2026. Upload a reference image of the person or character, then prompt a new scene that preserves the identity.
[Upload reference photo of the person as inline_data]
Create a portrait of this person as a Silicon Valley executive, professional headshot, soft studio lighting, blurred office background, confident expression, business casual attire, shot on 85mm lens, 4K.
The model holds facial features across the workflow. In the Gemini app, you can sustain up to 5 unique characters + 10 objects in a single workflow. The fal.ai endpoint exposes this as "up to 5 people per call." For a 50-frame storyboard, this is the pattern — generate frame 1, then use it as the reference for frame 2, and so on.
4. Film stock / aesthetic styling
Named film stocks work as color-grading cues. The model reproduces the look:
A portrait shot on Kodak Portra 400 film, warm natural skin tones, slight grain, soft daylight, the subject is a young woman laughing in a garden, 35mm analog photography, 1990s aesthetic.
Try: "Kodak Portra 400" (warm, natural), "Fuji Pro 400H" (cool, pastel), "Ilford HP5" (black and white, high grain), "Polaroid SX-70" (warm, faded, square), "1980s color film, slightly grainy" (the official Google Cloud suggestion).
5. Search-grounded factual visuals (Nano Banana 2's unique advantage)
This is what Nano Banana 2 does that the others don't. Add tools: [{"google_search": {}}] and the model pulls real-time information from Google Search and renders it as part of the image.
Visualize the current weather forecast for the next 5 days in San Francisco as a clean, modern weather chart. Add a visual on what I should wear each day.
It works for: current weather, today's news, recent product releases, stock prices, sports scores, anything that changes day-to-day. The API returns groundingMetadata with searchEntryPoint (Google Search chip) and groundingSupports (which sources the model used).
There's also Image Search grounding — set searchTypes: {"imageSearch": {}} and the model finds real images via Google Image Search and uses them as visual reference. Useful for "paint a Timareta butterfly" or "draw this specific landmark." (Note: can't be used to search for people.)
5 editing patterns (all from Google's official docs)
These are the patterns Google's own API documentation uses as worked examples. Copy the structure, swap the specifics.
Edit by description
Upload an image, describe the change, the model matches the original style and lighting.
[Upload image of your cat]
Using the provided image of my cat, please add a small, knitted wizard hat on its head. Make it look like it's sitting comfortably and matches the soft lighting of the photo.
Conversational mask (change one thing, keep everything else)
Specify the only thing that should change and the model holds the rest constant. Way more reliable than SD-style masking.
[Upload image of living room]
Using the provided image of a living room, change only the blue sofa to be a vintage, brown leather chesterfield sofa. Keep the rest of the room, including the pillows on the sofa and the lighting, unchanged.
Style transfer
Upload a photo, ask for it in a different art style. Composition is preserved.
[Upload photograph of city street at night]
Transform the provided photograph of a modern city street at night into the artistic style of Vincent van Gogh's 'Starry Night'. Preserve the original composition of buildings and cars, but render all elements with swirling, impasto brushstrokes and a dramatic palette of deep blues and bright yellows.
Multi-image composite
Pass multiple reference images and a text prompt. The model composites them into one coherent scene.
[Upload blue floral dress image]
[Upload woman image]
Create a professional e-commerce fashion photo. Take the blue floral dress from the first image and let the woman from the second image wear it. Generate a realistic, full-body shot of the woman wearing the dress, with the lighting and shadows adjusted to match the outdoor environment.
The fal.ai endpoint supports up to 14 reference images in one call. Use this for product mockups with multiple angle references, character consistency across scenes, or "put this person in this setting."
Sketch-to-image refinement
Upload a rough sketch, get a polished final image that follows the lines.
[Upload rough pencil sketch of a futuristic car]
Turn this rough pencil sketch of a futuristic car into a polished photo of the finished concept car in a showroom. Keep the sleek lines and low profile from the sketch but add metallic blue paint and neon rim lighting.
Control knobs worth knowing
Resolution and aspect ratio
The API accepts both. Use uppercase K, lowercase is rejected.
response = client.models.generate_content(
model="gemini-3.1-flash-image",
contents=prompt,
config=types.GenerateContentConfig(
response_format={"image": {"aspectRatio": "16:9", "imageSize": "2K"}}
)
)
Available resolutions: 512 (0.5K — Flash only), 1K, 2K, 4K.
Available aspect ratios: 1:1, 1:4, 1:8, 2:3, 3:2, 3:4, 4:1, 4:3, 4:5, 5:4, 8:1, 9:16, 16:9, 21:9.
Thinking level
thinkingLevel: "minimal" (default) is fastest. "high" produces meaningfully better results for complex prompts (compositions with multiple elements, prompts with text rendering, prompts with specific factual constraints). It costs more — thinking tokens are billed regardless of whether you set includeThoughts: true.
For social and iteration: minimal. For final hero assets, dense text rendering, or complex compositions: high.
Multi-turn editing
In a multi-turn conversation, every image response includes a thought_signature field. You must pass it back in the next turn exactly as received or the next response may fail. The SDKs handle this automatically when you pass the response contents into the next call, but if you're rolling raw HTTP, watch for this.
Common mistakes that silently break results
- Using lowercase
1kor4kfor resolution — rejected. Must be1Kor4K. - Using Stable Diffusion-style negative prompt weights like
(text:1.3)— Gemini uses semantic negatives: "no text, no watermark, no extra people." - Keyword lists instead of scene descriptions — "red dress, beach, sunset, smiling" produces worse results than "a woman in a red dress smiling at the camera on a beach at sunset, warm light, shot on 35mm film."
- Skipping camera/lighting language for photoreal — the model uses these as concrete constraints. "Photorealistic" alone is a vibe. "Photorealistic, 85mm portrait lens, soft natural daylight" is a constraint.
- Forgetting to pass back
thought_signaturein multi-turn edits — the next call may fail. - Using
gemini-2.5-flash-imageinstead ofgemini-3.1-flash-image— the older model doesn't support 4K, doesn't have the same thinking capabilities, and isn't being updated. Always use the 3.1 model for new work. - Treating Nano Banana 2 as a single-shot tool — multi-turn iteration is the design pattern. First generation is a draft; the next 2-3 turns refine it.
Quick start: a 3-line prompt that works for most social posts
If you want one template to remember:
A [style: photoreal / illustration / etc.] of [specific subject with age/clothing/expression], in [specific setting], [action if any]. Captured with [lens + lighting if photoreal], [aspect ratio] aspect ratio.
Example: "A photorealistic close-up of a 30-year-old woman in a cream linen blazer laughing at the camera in a sunlit cafe. Captured with 50mm lens, soft natural daylight, 1:1 aspect ratio."
Three lines. Subject, setting, composition. You can add a 4th line for "for [context: Instagram post / product card / editorial)" if you want to bias the model.
The summary
Nano Banana 2's prompt engineering isn't about magic keywords. It's about:
- Describing scenes in paragraphs, not lists.
- Using the 5-component framework (Style, Subject, Setting, Action, Composition) for any prompt that needs to hold up.
- Naming cameras, lenses, and film stocks as concrete constraints.
- Using search grounding for anything time-sensitive.
- Using
thinkingLevel: "high"for complex compositions and text rendering. - Iterating in multi-turn with thought signatures passed back.
Once you internalize those, the prompt patterns in this guide are just shortcuts. The underlying rules do the work.
Want to see these in action against other models? Compare Nano Banana 2 with GPT Image 2 for typography and editing differences, or browse the full Nano Banana 2 prompt library for ready-to-use templates.



