Text-only prompts have a ceiling. Once you need a specific product on a model, a consistent character across scenes, or a specific architectural style, you need a reference image — a photo you pass to the model alongside your text prompt.
The four primary image models all support reference images, but they handle them differently. Below is what each model supports, the 4 production patterns that work across all of them, the model-specific syntax, the common failure modes, and the workflow for picking the right reference.
What "reference image" means in 2026
A reference image is an input photo (or generated image) you pass to the model. The model uses it as conditioning for the new image. The text prompt is the change; the reference image is the anchor.
There are four kinds of references, and the model behavior differs for each:
- Style reference — the look, color grading, and rendering style. The model applies the style to a new subject.
- Subject reference — a face, body, or product. The model preserves the identity of the subject in a new scene.
- Composition reference — the layout, framing, and pose. The model copies the structure with new content.
- Material reference — the texture, surface, or material pattern. The model applies the material to a new object.
Most production workflows use 2-3 references at once: subject + style + composition. The 4-pattern section below covers how to combine them.
The model support matrix
| Model | Reference images per call | Types supported | How to pass |
|---|---|---|---|
Nano Banana 2 (gemini-3.1-flash-image) | Up to 14 | Style, subject, composition, material | inline_data in the same call; multi-image prompts |
GPT Image 2 (gpt-image-2) | Up to 4 (varies by endpoint) | Style, subject, composition | Image input + text prompt in the same call |
Seedream 5.0 Lite (bytedance/seedream-5-lite) | Multiple, exact count varies by route | Style, subject, composition | image parameter array in the API call |
Grok Imagine (grok-imagine-image) | Multiple | Style, subject | image_url or base64 in the request body |
Key takeaways:
- Nano Banana 2 has the highest reference-image count (up to 14). The other models cap lower.
- All four accept multiple references, so multi-image conditioning is portable across the field.
- None of the four require a model-specific format for the reference itself — JPG/PNG, base64 or URL, works on all of them. The difference is the prompt language you pair with the reference.
The 4 production patterns
These are the workflows where reference images are non-negotiable. Each one is the same pattern across all 4 models; the only difference is the prompt language.
Pattern 1: Product mockup (product on a model, in a scene, on a background)
This is the #1 ecommerce use case. You have a product photo (the product cut out on white). You want to see it on a person, in a scene, on a lifestyle background.
Reference roles:
- Reference 1: the product image (subject + material)
- Reference 2 (optional): a person or scene for the composition
Prompt structure (works on all 4 models):
[Reference 1: product photo, e.g. blue floral dress on white background]
Create a professional e-commerce fashion photo. Take the [product] from
the first reference image and place it on a woman in an outdoor park
setting with soft afternoon light. The woman is 30, with shoulder-
length brown hair, walking on a path, natural pose, soft smile. The
dress from the first reference must be preserved exactly — same color,
same pattern, same cut, same fabric. Match the lighting and shadows
to the outdoor environment. No watermark, no extra logos, no text.
Why it works: the reference image is named explicitly, the preservation rules are specific (color, pattern, cut, fabric), the new context is fully described, and the model has no permission to re-design the product.
On Nano Banana 2 specifically: you can pass 14 references in one call. Reference 1 = product, Reference 2 = scene reference photo, Reference 3-14 = alternative angles of the product for the model to "understand" the 3D shape. This is the only model where you can do this level of multi-reference conditioning.
On GPT Image 2 specifically: the prompt structure is the same. GPT Image 2 is stronger on text-in-image (if your product has a label) and on photoreal lighting matching.
On Seedream 5.0 Lite specifically: the reasoning model is good at inferring the "preserve list" from the reference image. You can be less explicit in the prompt and the model still preserves.
On Grok Imagine specifically: strong at the lifestyle/scene generation, may be less strict on the exact product preservation. Test before relying on it for ecommerce.
Pattern 2: Character consistency (same person, new scene)
This is the brand character / influencer clone use case. You have a face you want to use across multiple generated images — same face, different scenes, different outfits, different lighting.
Reference roles:
- Reference 1: the face (subject reference, the only anchor you need)
Prompt structure:
[Reference 1: face photo of the person, well-lit, neutral expression]
Create a portrait of this person as a Silicon Valley executive,
professional headshot, soft studio lighting, blurred office background,
confident expression, business casual attire, shot on 85mm lens, 4K.
The face must be preserved exactly from the reference.
The iteration loop: generate a candidate. If the face drifts, add "The face must remain identical to the reference image — same bone structure, same eye shape, same nose, same lip shape, same skin tone, same hairline." Iterate 3-4 passes until the face holds.
On Nano Banana 2 specifically: up to 14 references, and the model can sustain 5 unique characters + 10 objects in a single workflow. This is the strongest model for multi-character consistency.
On GPT Image 2 specifically: strong identity preservation, especially when you pass the face reference and explicitly name "preserve the face."
On Seedream 5.0 Lite specifically: the visual reasoning helps the model "understand" identity. Good for stylized consistency (same character, different art style).
On Grok Imagine specifically: consistency is improving but not at the same level. For headshots and serious identity work, use NB2 or GPT Image 2.
Pattern 3: Style transfer (your subject, the reference's look)
You have a photo of a subject. You want it rendered in the style of another image — a painter, a film stock, an era, a graphic design language.
Reference roles:
- Reference 1: the subject (the content)
- Reference 2: the style (the look)
Prompt structure:
[Reference 1: photograph of a city street at night]
[Reference 2: Van Gogh's Starry Night painting]
Transform the provided photograph of a modern city street at night into
the artistic style of the second reference (Starry Night). Preserve the
original composition of buildings and cars from the first reference,
but render all elements with swirling, impasto brushstrokes and a
dramatic palette of deep blues and bright yellows.
On Nano Banana 2 specifically: the official "Style transfer" pattern. Use it as-is.
On GPT Image 2 specifically: strong at photoreal style transfer. Specify the era, the artist, or the film stock explicitly.
On Seedream 5.0 Lite specifically: the reasoning model is good at "applying" a style while preserving content. Specify what to keep (composition) and what to change (rendering language).
On Grok Imagine specifically: style transfer is a strength of Grok Imagine. The "Fun" and "Spicy" modes on the consumer product are style transfer features.
Pattern 4: Material / texture application (the look, your object)
You have an object. You want to see it in a specific material — brushed aluminum, frosted glass, woven linen. The reference image provides the material; the text prompt provides the object and the scene.
Reference roles:
- Reference 1: the material (texture, surface, finish)
Prompt structure:
[Reference 1: close-up photo of brushed aluminum surface]
Apply the brushed aluminum texture from the reference to a vintage
handheld radio. The radio should have the same form factor as a
1980s Sony Walkman, with a small screen, two large dials, and a
speaker grille. The aluminum should follow the form of the radio —
catching light on the curves, shadowed in the recesses, with visible
brushed grain on the flat surfaces. Studio product photography, white
seamless background, soft overhead lighting.
On Nano Banana 2 specifically: can take the material reference plus a shape reference plus a style reference in the same call. The model composites intelligently.
On GPT Image 2 specifically: strong at photoreal material rendering. Specify the lighting interaction ("catching light on the curves, shadowed in the recesses") to push the model beyond a flat texture map.
On Seedream 5.0 Lite specifically: reasoning model is good at inferring "how light interacts with this material" from the reference. Specify the lighting context, not just the material.
On Grok Imagine specifically: less reliable for technical material work. Use NB2, GPT Image 2, or Seedream for product material workflows.
The common failure modes
1. The reference is too complex. If the reference has 5 objects, the model will not know which one to preserve. Crop or mask the reference to the single element you want the model to use. A product cut out on white is better than a product on a busy background.
2. The prompt does not name what to preserve. The reference shows the model the look, but the prompt must say "preserve the [product] from the reference exactly." Without the preserve line, the model has permission to re-interpret the reference.
3. The reference and the prompt contradict. If the reference shows a man and the prompt says "create a woman," the model will favor the prompt. Pick a reference that matches the prompt's subject, or change the prompt.
4. The reference is low quality. Models inherit the quality floor of the reference. A 256×256 reference produces a 256×256 quality output regardless of the target size. Always use a high-resolution, well-lit reference.
5. Too many references. If you pass 10 references and the model has to figure out which to use, the result is a blend of all of them. Use 1-3 references per call, each with a clear role (subject / style / composition).
6. The reference is not annotated. The model cannot tell which part of the reference is the "subject" vs the "background." Crop the reference to the subject, or explicitly say "the [object] in the upper-left of the reference."
The decision framework: which reference, when
| Goal | Reference count | Reference role | Pick this model |
|---|---|---|---|
| Product on a model | 1-2 | Subject (product), optional scene | Nano Banana 2, GPT Image 2 |
| Product on lifestyle background | 1-2 | Subject (product), optional scene | GPT Image 2, Seedream 5.0 Lite |
| Character in new scene | 1 | Subject (face) | Nano Banana 2, GPT Image 2 |
| Multi-character scene | 5+ | Subject (5 faces) | Nano Banana 2 (the only one that handles 5+) |
| Style transfer (your subject, new style) | 2 | Subject + style | Nano Banana 2, Grok Imagine |
| Material / texture on new object | 1-2 | Material + optional shape | GPT Image 2, Seedream 5.0 Lite |
| Architecture / scene composition | 1-2 | Composition + style | Nano Banana 2, GPT Image 2 |
| Multi-image composite (complex scene) | 4+ | Multiple subjects | Nano Banana 2 (up to 14) |
The reference image checklist
Before you ship a reference-based workflow, verify:
- Reference is high resolution (1024×1024 minimum, 2K+ preferred for hero assets).
- Reference is cropped to the subject (no busy background, no extra objects).
- Reference is well-lit (the model inherits the lighting quality).
- Reference matches the prompt's subject (no contradictions).
- Prompt explicitly names what to preserve ("the product from the reference must be preserved exactly").
- Reference count is appropriate (1-3 per call, 4+ only on Nano Banana 2).
- Model supports the reference count and type (check the support matrix above).
- Reference is base64 or URL, in the format the API expects (most accept both).
Skip any of these and the reference is more likely to cause drift than to prevent it.
The summary
Reference images are how you push past the text-only ceiling. They let you preserve identity, copy style, transfer materials, and compose complex scenes from separate parts.
- All four primary models support reference images. The difference is the count (NB2 supports up to 14) and the prompt language you pair with them.
- The 4 production patterns: product mockup, character consistency, style transfer, material application.
- The preserve list is mandatory. "Preserve the [subject] from the reference exactly" is the single most important sentence in any reference-based prompt.
- Crop, light, and resolve the reference before passing it. The model inherits the reference's quality floor.
- Pick the model by reference count and type. Nano Banana 2 for multi-reference complexity, GPT Image 2 for identity and text-in-image, Seedream 5.0 Lite for reasoning-heavy style transfer, Grok Imagine for consumer-grade style work.
The prompt defines the change. The reference defines the anchor. Pair them with an explicit preserve list and the model has no room to drift.



