Single-reference workflows hit a ceiling at "a product on a model" or "a character in a scene." The next level up is multi-image composition — combining 3, 5, 10, or 14 references into a single coherent scene where every element comes from a different source.
This is the workflow for ecommerce product mockups at scale, character + outfit + scene composites, marketing campaign visuals, and "make this look like a coherent photo" requests where the inputs are scattered across folders.
The four primary models all support multi-image input, but the number of references they accept and how they handle priority between them differs. Below is the priority model, the 5 production patterns, the model-specific syntax, the per-pattern prompt structure, the common failure modes, and the workflow for shipping multi-image compositions at scale.
The priority model: how models handle multiple references
When you pass 3+ references, the model has to decide which one to favor when they conflict. The way you describe each reference in the prompt sets the priority.
Priority 1 (anchor) — the element that must not change. Usually the subject (face, product) or the style. The model treats this as the load-bearing element.
Priority 2 (context) — the element that sets the scene or composition. Usually the background, the pose, or the lighting reference. The model applies this as the environment.
Priority 3 (detail) — supporting elements that refine the output. Outfit, materials, props, color palette. The model integrates these last.
Priority 4 (style transfer) — the look, not the content. Usually a film still, painting, or color-grade reference. The model applies this as the rendering language.
The way you write the prompt names the priority. "Take the [product] from reference 1, place it on the [person] from reference 2, in the [setting] from reference 3, with the lighting from reference 4" sets explicit priority. The model knows reference 1 is the load-bearing element.
The model support for multi-image composition
| Model | Max references | Sweet spot | Best use case |
|---|---|---|---|
Nano Banana 2 (gemini-3.1-flash-image) | Up to 14 | 4-8 | Ecommerce at scale, multi-character scenes, complex composites |
GPT Image 2 (gpt-image-2) | Up to 4 | 2-3 | Virtual try-on, focused 2-3 element composites |
Seedream 5.0 Lite (bytedance/seedream-5-lite) | Multiple, varies | 2-4 | Reasoning-heavy composites, scene composition |
Grok Imagine (grok-imagine-image) | Multiple | 2-3 | Consumer-grade multi-image, lifestyle composites |
| FLUX.2 Pro (for context) | Up to 9 | 4-6 | Production-grade editing, max flexibility |
Key takeaways:
- Nano Banana 2 has the highest reference count (14) and the most mature multi-image composition workflow. For ecommerce at scale, this is the default.
- GPT Image 2 has the most refined 2-3 element composite (virtual try-on, focused compositing). For product-on-model with strong identity, this is the default.
- Seedream 5.0 Lite is best when the composite requires reasoning ("combine the architecture from ref 1 with the lighting from ref 2, but make the materials match ref 3"). The reasoning model handles priority conflicts better.
- Grok Imagine is the consumer-grade option. It can do multi-image composition but the result is less predictable than the other three.
The 5 production patterns
Pattern 1: Virtual try-on (clothing on a person, in a scene)
The highest-value ecommerce pattern. You have a clothing reference (a product photo on white) and a person reference (a model or customer photo). You want the person wearing the clothing in a new scene.
Reference roles:
- Reference 1 (anchor): the person / face
- Reference 2 (anchor): the clothing product photo
- Reference 3 (context): the scene (lifestyle background)
Prompt structure (works on all 4 models):
[Reference 1: person photo, head-and-shoulders, neutral lighting]
[Reference 2: clothing product photo, on white background]
[Reference 3: lifestyle scene reference, e.g. outdoor cafe]
Create a fashion editorial photo. Place the clothing from reference 2 on
the person from reference 1, in the setting from reference 3. Preserve
the person's face, body, and pose from reference 1. Preserve the
clothing's design, color, pattern, and fabric from reference 2. Match
the lighting and color palette of reference 3. No watermark, no extra
text, no other people. Full body visible from head to feet.
On Nano Banana 2 specifically: can take the clothing from multiple angles (front, back, side, detail) as 3-4 separate references. The model understands the 3D form of the garment and renders it from the new angle correctly.
On GPT Image 2 specifically: the official "Virtual Try-On (Edit Mode)" pattern. Use the person from image 1, the jacket from image 2, the boots from image 3 in a single prompt.
On Seedream 5.0 Lite specifically: specify "the person from reference 1, the jacket from reference 2, the boots from reference 3" explicitly. The reasoning model maps each reference to a specific role.
On Grok Imagine specifically: consumer-grade virtual try-on is a built-in feature. Works for "try this outfit on me" but less reliable for production ecommerce.
Pattern 2: Multi-character scene (5+ characters, each with a face)
Storyboards, marketing campaign visuals with multiple models, character sheets for game assets. The pattern where you need 5 unique faces, each preserved across a single composite.
Reference roles:
- References 1-5 (anchor): 5 face photos
- Reference 6 (context): scene / environment
Prompt structure (NB2 only — the only model that handles 5+ characters):
[Reference 1-5: 5 different face photos]
[Reference 6: scene reference, e.g. a boardroom]
Create a boardroom scene with the 5 people from references 1-5 sitting
around a conference table. Preserve each face from the references.
Reference 1 is the CEO at the head of the table. Reference 2 is on
the left in a dark blazer. Reference 3 is on the right in a gray
suit. Reference 4 is at the far end with glasses. Reference 5 is
standing at a whiteboard. Match the lighting and setting from
reference 6. Photorealistic, editorial, 4K.
On Nano Banana 2 specifically: this is the model's signature use case. The Gemini app can sustain 5 unique characters + 10 objects in a single workflow.
On the other models: they handle 2-3 characters but start to fail at 4+. For multi-character scenes, NB2 is the only real option.
Pattern 3: Style + content composite (your subject, the reference's look, the third reference's lighting)
A 3-way composite: subject (your photo) + style (a film still or painting) + lighting (a third reference for the lighting setup). Used for high-end editorial work.
Reference roles:
- Reference 1 (anchor): the subject
- Reference 2 (style): the look (color grading, rendering language)
- Reference 3 (context): the lighting setup
Prompt structure:
[Reference 1: portrait photo of a person]
[Reference 2: film still with desired color grading, e.g. "Lost in Translation" still]
[Reference 3: reference photo with desired lighting, e.g. window light from left]
Create a portrait of the person from reference 1. Apply the color
grading and film aesthetic from reference 2 (muted teals, soft
highlights, slight grain). Match the lighting setup from reference 3
(soft window light from the left, gentle shadow on the right side
of the face). Preserve the person's face and identity from reference 1.
On Nano Banana 2 specifically: the official "Style transfer" pattern handles this naturally.
On Seedream 5.0 Lite specifically: the reasoning model is good at separating "what to copy" (color grading, lighting) from "what to preserve" (subject identity).
On GPT Image 2 specifically: strong at color grading and lighting matching. Specify the era, the film stock, or the cinematographer explicitly to push the result.
Pattern 4: Product collage (multiple products in one scene)
For ecommerce hero pages, gift guides, "shop the look" layouts, holiday collections. Multiple products in a single coherent image.
Reference roles:
- References 1-N (anchor): product photos
- Reference N+1 (context): a styled scene / surface
Prompt structure:
[Reference 1: handbag product photo on white]
[Reference 2: sunglasses product photo on white]
[Reference 3: scarf product photo on white]
[Reference 4: styled scene reference, e.g. flat-lay on marble]
Create a "shop the look" flat-lay composition. Arrange the handbag
from reference 1, the sunglasses from reference 2, and the scarf from
reference 3 on the marble surface from reference 4. Each product
must be clearly visible and product-accurate (preserved from its
reference). Soft overhead natural light, subtle shadows, photorealistic,
1080x1080. No watermark, no text, no extra items.
On Nano Banana 2 specifically: the highest reference count is critical for product collages. Pass 5-10 products + 1 scene reference.
On the other models: 2-3 products per call. For larger collages, generate in passes and composite manually.
Pattern 5: Character + outfit + scene (the brand asset pattern)
The most common pattern for brand systems: a consistent character, in a consistent outfit, in multiple scenes. This is what makes a brand recognizable.
Reference roles:
- Reference 1 (anchor, persistent): the character's face
- Reference 2 (anchor, persistent): the character's outfit
- Reference 3 (variable, per-scene): the scene
Prompt structure:
[Reference 1: face photo of the brand character]
[Reference 2: outfit product photo, the brand's signature look]
[Reference 3: scene reference, e.g. coffee shop, beach, office, gym]
Create a lifestyle photo of the character from reference 1 wearing the
outfit from reference 2, in the setting from reference 3. Preserve
the face, hair, and identity from reference 1. Preserve the outfit's
design, color, and material from reference 2. Match the lighting
and color palette of reference 3. Photorealistic, editorial, 4K.
The iteration loop: generate one scene. If the face drifts, add "the face from reference 1 must remain identical — same bone structure, same eye shape, same lip shape, same skin tone." If the outfit drifts, add "the outfit from reference 2 must be preserved exactly — same color, same pattern, same cut, same fabric." Iterate 3-4 passes per scene until both anchors hold. Then move to the next scene — the same character + outfit will hold across the series.
On Nano Banana 2 specifically: the multi-character + object support makes this the strongest model for sustained brand assets.
On GPT Image 2 specifically: strong identity and outfit preservation when you name both references explicitly.
On Seedream 5.0 Lite specifically: the reasoning model helps when the scene requires "preserve the character, change the lighting" type edits.
The "named-reference" prompt pattern
The single most important pattern for multi-image prompts is naming each reference explicitly in the prompt. The model cannot infer "the third image is the lighting" — you have to say it.
The structure:
[Reference 1: description of image 1]
[Reference 2: description of image 2]
[Reference 3: description of image 3]
Create [subject/scene] using [element] from reference 1,
[element] from reference 2, and [element] from reference 3.
Preserve [list of things that must not change].
The "from reference 1" pattern is the load-bearing phrase. It tells the model which reference to use for which element. Without it, the model treats all references as a single style hint and blends them.
The reference pre-flight checklist
Before you ship a multi-image composition:
- Each reference is cropped to its subject. No busy backgrounds, no extra objects.
- Each reference is high resolution (1024×1024 minimum, 2K+ preferred).
- The references do not contradict each other. A face from a man and a "woman" in the prompt will conflict.
- The prompt names each reference explicitly ("from reference 1," "from reference 2").
- The priority is set in the prompt (anchor > context > detail > style).
- The reference count is appropriate (sweet spot is 2-3 for most models, 4-8 for NB2).
- The model supports the reference count (check the support matrix).
- The preserve list is explicit (name what must not change).
- The lighting interaction is described (e.g., "match the lighting of reference 3" — without this, the model composites a daylight subject onto a night scene).
- A single fallback is planned (if the model returns 6 fingers, what is your iteration prompt?).
The common failure modes
1. The references are not named in the prompt. The model treats them as a blended style hint, not as discrete elements. Always say "from reference 1," "from reference 2."
2. The priority is not set. "Create a scene with these references" gives the model no guidance on what to preserve. Name the anchors first.
3. The lighting does not match. The model composites a daylight subject onto a night scene. Specify "match the lighting of reference N" or describe the lighting interaction in the text prompt.
4. The model is asked for too many anchors. "Preserve the face, the outfit, the background, the product, the material" — the model cannot preserve 5 anchors at once. Pick 1-2 anchors per call.
5. The references are too similar. Two face photos of the same person from slightly different angles. The model cannot tell which to preserve. Use one anchor reference, period.
6. The references are too different. A face from a person in 1990s film grain and a request for a modern 4K portrait. The model will pick one or blend awkwardly. Use references that are visually compatible.
7. The prompt does not name the output use case. "Create an image" is underspecified. "Create a 1:1 ecommerce hero, full body visible, no watermark, 4K" gives the model a target.
8. The model has a reference count limit you exceeded. GPT Image 2 caps at 4, others vary. Check the support matrix before passing 8 references.
The summary
Multi-image composition is how you ship production-grade composites. The 5 patterns — virtual try-on, multi-character scene, style + content, product collage, brand character — cover most production use cases.
- Name each reference explicitly in the prompt. "From reference 1," "from reference 2" — without this, the model blends.
- Set the priority in the prompt. Anchor > context > detail > style. The model needs to know which to favor.
- Pick the model by reference count. Nano Banana 2 for 5+ (up to 14), GPT Image 2 for 2-3 element composites, Seedream 5.0 Lite for reasoning-heavy composites, Grok Imagine for consumer-grade.
- Crop, light, and resolve each reference. The model inherits the reference's quality floor.
- Iterate with the preserve list. When the first pass drifts, add to the preserve list, re-run.
The prompt defines the change. The references define the anchors. The priority defines which anchor wins. Get all three right and multi-image composition becomes a reliable production workflow.



