AdX

Grok Imagine Multi-Image Editing: How to Combine Up to Three References

Grok Imagine Multi-Image Editing: How to Combine Up to Three References

Grok Imagine's deliberate constraint — 2-3 references with clear roles — is the strength. The 3-reference role pattern (subject + setting + style), the named-reference prompt structure, the 4 production patterns (product on setting, character in setting, style transfer, material on object), the 10 b

Grok Imagine's multi-image editing is the model-specific workflow for combining 2-3 reference images into a single coherent scene. Unlike the broader multi-image workflows (Nano Banana 2 with 14 refs, GPT Image 2 with 4, FLUX.2 Pro with 9), Grok Imagine is deliberately small — 2-3 references, each with a clear role. The constraint is the strength: the model has fewer references to balance, so the named-reference pattern works reliably.

This is the production guide: the 3-reference role pattern (subject + setting + style), the named-reference prompt structure, the 4 production patterns, the per-pattern prompts, the brand-safety rules, the iteration loop, and the workflow for shipping coherent multi-image composites with Grok Imagine.

What Grok Imagine's multi-image editing is (and isn't)

Grok Imagine's multi-image editing launched in February 2026 as a structured workflow. The model accepts up to 3 reference images in a single call. Each reference is named in the prompt, and the model uses each reference for a specific role.

The 3-reference role pattern:

  • Reference 1: Subject — the load-bearing element (the product, the face, the object that must be preserved)
  • Reference 2: Setting — the environment (the scene, the background, the context)
  • Reference 3: Style — the look (the color grading, the lighting, the rendering language)

The model is told explicitly: "from reference 1, take the subject. From reference 2, take the setting. From reference 3, take the style." The model has clear priorities and fewer conflicts to resolve.

What it isn't:

  • Not a "merge two photos pixel-by-pixel" workflow (that's Photoshop)
  • Not a "swap the face from one image to another" workflow (that's face-swap)
  • Not a 4+ reference workflow (Grok caps at 3; for more, use Nano Banana 2 or FLUX.2 Pro)

The Grok multi-image workflow is a compositional workflow — the model composes 3 inputs into a single coherent output, with explicit priority rules.

The named-reference prompt pattern

The single most important pattern for multi-image workflows. The model cannot infer "the third image is the lighting" — you have to say it.

The format:

[Reference 1: subject — short description of what image 1 is]
[Reference 2: setting — short description of what image 2 is]
[Reference 3 (optional): style — short description of what
image 3 is]

Create [output description] using the [subject] from reference 1,
in the [setting] from reference 2, with the [style] from
reference 3 (if used).

Preserve [list of things that must not change]. The [subject]
must remain exactly as in reference 1.

Example:

[Reference 1: a matte black ceramic coffee mug on white
background, product photo]
[Reference 2: a warm wooden kitchen counter with morning light
from a window]
[Reference 3: a 1990s editorial photograph with warm color
grading and soft contrast]

Create a lifestyle product photograph of the coffee mug from
reference 1, placed on the kitchen counter from reference 2,
with the color grading and film aesthetic from reference 3.

Preserve the coffee mug exactly as it appears in reference 1:
same color, same material, same size, same proportions, same
finish. The mug must look identical to the product photo. The
warm wood grain and morning light of reference 2 must be
preserved. The 1990s editorial color grading of reference 3
must be applied consistently.

Why it works: the model has explicit roles for each reference. "Subject from 1, setting from 2, style from 3" is unambiguous. The preserve list names the load-bearing element (the mug).

The 4 production patterns

Pattern 1: Product on a setting (the "lifestyle mockup" pattern)

The single highest-ROI multi-image pattern for ecommerce. Take a product photo (white background) and a setting photo (a place), and composite the product into the setting.

The 3-reference role assignment:

  • Reference 1 (subject): the product photo on white
  • Reference 2 (setting): the lifestyle scene (cafe, kitchen, gym, etc.)
  • Reference 3 (style): the color grading / lighting / film look (optional)

The prompt structure:

[Reference 1: matte black ceramic coffee mug on white
background, product photo]
[Reference 2: a warm wooden kitchen counter with morning light
from a window camera-left, kitchen in soft focus in background]
[Reference 3: a 1990s editorial photograph with warm color
grading and soft contrast, slight grain]

Create a lifestyle product photograph of the coffee mug from
reference 1, placed on the kitchen counter from reference 2,
with the color grading from reference 3.

Preserve the coffee mug exactly: same color (matte black
#1A1A1A), same material (ceramic with matte finish), same
size (8oz), same proportions, same details. The mug must
look identical to the product photo.

Match the lighting of the setting: morning light from
camera-left, soft shadow on the right side of the mug, warm
color temperature (4500K). The mug should ground on the
counter with a soft contact shadow.

Apply the 1990s editorial color grading consistently: warm
highlights, slightly desaturated, soft contrast, slight
grain.

Constraints: photorealistic, 1080x1080, no watermark, no fake
brand names, no fake text, no fake testimonials, the mug
must be the focal point in sharp focus.

Why it works: the product is named as the anchor, the setting is named as the context, the style is named as the look. The preserve list makes the product non-negotiable.

On Grok Imagine specifically: strong at this. The model handles product + setting + style well, with clear priority to the product.

The platform tweaks:

  • Amazon / Shopify: white background, even lighting, no style reference
  • Instagram Feed: lifestyle + warm grading, 4:5
  • Pinterest: lifestyle + soft grading, 2:3
  • TikTok / Reels: lifestyle + dynamic grading, 9:16

Pattern 2: Character in a setting (the "brand asset" pattern)

Take a face (Reference 1) and a setting (Reference 2), and place the character in the new setting. Used for brand characters, influencers, storyboards.

The 3-reference role assignment:

  • Reference 1 (subject): the face / character reference
  • Reference 2 (setting): the new setting
  • Reference 3 (style): the lighting / color grading (optional)

The prompt structure:

[Reference 1: front-view portrait of a 30-year-old woman with
short dark hair, cream linen blazer, neutral expression, soft
daylight]
[Reference 2: a sunlit modern office with golden hour light
through a west-facing window, large desk, plants]
[Reference 3: a 1990s editorial photograph with warm color
grading]

Create a professional portrait of the character from reference
1, in the office setting from reference 2, with the color
grading from reference 3.

Preserve the character exactly: same face (bone structure, eye
shape, lip shape, skin tone, hair color, hair length, hair
part), same outfit (cream linen blazer, white t-shirt),
same expression family (engaged, confident).

Match the office lighting: golden hour light from the west
window camera-left, soft fill from the right, 5500K color
temperature, gentle shadow on the right side of the face.

Apply the 1990s editorial color grading consistently.

Constraints: photorealistic, 1080x1080, no watermark, no fake
brand names, no fake testimonials, the face must be the
focal point in sharp focus.

Why it works: the character is the anchor (face, hair, outfit), the setting is the context, the style is the look. The preserve list names the load-bearing elements.

On Grok Imagine specifically: good at this. The model handles character + setting well. For multi-character scenes (3+ characters), use Nano Banana 2 (which supports 14 references).

Pattern 3: Style transfer on a subject (the "look" pattern)

Take a subject (Reference 1) and a style (Reference 2), and apply the style to the subject. Used for stylized portraits, brand-consistent content, art-direction.

The 3-reference role assignment:

  • Reference 1 (subject): the subject (person, product, scene)
  • Reference 2 (style): the style (the rendering language, the look)
  • Reference 3 (optional, setting): the new setting (if changing the scene)

The prompt structure:

[Reference 1: a photograph of a woman in a sunlit cafe, mid-
laugh, holding a coffee cup, 85mm portrait lens, soft
daylight]
[Reference 2: a film still from a 1970s Italian neorealist
movie, desaturated, natural light, available-source, grain]

Create a portrait of the subject from reference 1, in the
visual style of reference 2.

Preserve the subject exactly: same face, same expression (mid-
laugh, eyes crinkled, head back slightly), same outfit, same
coffee cup, same scene composition.

Apply the 1970s Italian neorealist aesthetic: desaturated
colors, natural available-source light, fine grain, soft
contrast, no studio polish. The result should look like a
film still, not a contemporary photograph.

Constraints: photorealistic, 1080x1080, no watermark, no fake
brand names, no fake text, the subject must be the focal
point in sharp focus.

Why it works: the subject is the anchor, the style is the look. The preserve list names what must not change. The style application is specific (desaturated, available-source, grain, soft contrast).

On Grok Imagine specifically: strong at style transfer. The model handles "subject from 1, style from 2" cleanly.

Pattern 4: Material on a new object (the "texture" pattern)

Take a material (Reference 1) and an object (Reference 2), and apply the material to the object. Used for product variants, material visualization, design exploration.

The 3-reference role assignment:

  • Reference 1 (subject): the material reference (close-up of the texture)
  • Reference 2 (subject): the object (the form factor)
  • Reference 3 (style): the lighting / setting (optional)

The prompt structure:

[Reference 1: close-up photo of brushed aluminum surface,
visible grain direction, soft studio light]
[Reference 2: a vintage handheld radio, 1980s form factor,
boxy body with rounded corners, large tuning dial, speaker
grille, simple antenna]
[Reference 3: a soft studio product photo with three-point
lighting on white seamless background]

Create a product photograph of the vintage radio from
reference 2, with the brushed aluminum material from
reference 1, in the studio lighting from reference 3.

Preserve the radio's form factor exactly: boxy body, rounded
corners, large tuning dial, speaker grille, simple antenna,
1980s proportions. The shape must match reference 2.

Apply the brushed aluminum texture from reference 1: visible
grain direction (horizontal), matte finish, no mirror
reflections. The aluminum must follow the form of the radio —
grain on the flat surfaces, catching light on the curves,
shadowed in the recesses.

Use the studio lighting from reference 3: three-point
softbox, key from camera-left, fill from camera-right, soft
overhead light, white seamless background.

Constraints: photorealistic, 1080x1080, no watermark, no fake
brand names, no fake logos, the radio must be the focal
point in sharp focus, color-accurate brushed aluminum.

Why it works: the material is the anchor (the texture, the grain, the finish), the object is the form, the lighting is the context. The preserve list names what must not change (the radio's form, the aluminum's grain).

On Grok Imagine specifically: good at this. The model handles material + form + lighting well.

The 2-reference variant (when 3 is too many)

For simpler composites, use only 2 references. The role assignment is:

  • Reference 1 (subject): the anchor element
  • Reference 2 (setting): the context

Style is described in the prompt, not passed as a reference.

The 2-reference prompt structure:

[Reference 1: subject description]
[Reference 2: setting description]

Create [output description] using the [subject] from reference
1, in the [setting] from reference 2. Apply [style description
in words, no reference image].

Preserve [list of things that must not change].

When to use 2 references vs 3:

  • 2 references: when style is generic or well-described in words
  • 3 references: when style is specific (a film, an artist, a specific look)

The 10 brand-safety rules for multi-image editing

  1. "Preserve the subject exactly as in reference 1" — the subject is the load-bearing element
  2. "No fake URLs" — no invented website addresses
  3. "No fake testimonials" — no invented customer quotes
  4. "No fake statistics" — no invented numbers
  5. "No fake brand names" — no invented competitors or partners
  6. "No fake user counts" — no "10,000+ users" type claims
  7. "All text must be exactly as written" — prevents paraphrasing
  8. "No watermark" — clean output
  9. "Color-accurate to the original subject" — material and color preservation
  10. "Disclose AI generation per platform rules" — per the safety post

The list is the same as the other recent posts. The rule is the same: the brand-safety rules are the walls that keep fake content out.

The iteration loop for multi-image editing

Pass 1: Generate with the 3 references and the named-reference prompt. Run the prompt. Inspect the output.

Pass 2: Verify the subject preservation. Is the subject from reference 1 preserved exactly? Color, material, shape, details? If not, add to the preserve list and regenerate.

Pass 3: Verify the setting. Is the setting from reference 2 applied? Lighting, environment, scene? If not, name the setting elements more explicitly.

Pass 4: Verify the style. Is the style from reference 3 applied? Color grading, lighting, rendering language? If not, describe the style more explicitly.

Pass 5: Verify the brand-safety rules. No fake URLs, no fake testimonials, no fake text. If the model added any, regenerate with explicit negatives.

Pass 6: Apply the post-processing pipeline. Color correction, color grading, sharpening, grain.

Pass 7: Final verification at target size. Open at the actual size and verify. If anything is off, regenerate.

The pre-flight checklist for multi-image editing

Before you ship a multi-image composite:

  1. The subject is preserved exactly (color, material, shape, details)
  2. The setting is applied correctly (lighting, environment, scene)
  3. The style is applied consistently (color grading, lighting, rendering)
  4. The references are named in the prompt ("from reference 1," "from reference 2," "from reference 3")
  5. The preserve list is explicit (specific elements that must not change)
  6. The brand-safety rules are applied (10 universal rules)
  7. The text is verified character by character
  8. The platform's aspect ratio is correct (1:1, 4:5, 9:16, 2:3, 16:9)
  9. The platform's AI label is applied (per the safety post)
  10. The post-processing pipeline is applied (color correction, grading, sharpening, grain)

Skip any of these and the composite is either wrong, off-brand, or non-compliant.

The model pick by pattern (Grok Imagine vs the alternatives)

PatternGrok ImagineBest alternativeWhy
Product on setting✅ StrongNano Banana 2 (14 refs), GPT Image 2 (4 refs)Grok is good for 2-3 refs; NB2 for more complex
Character in setting✅ StrongNano Banana 2 (5 characters), Midjourney v7 --crefGrok is good for single character; NB2 for multi-character
Style transfer✅ StrongGPT Image 2, Seedream 5.0 LiteGrok is good for style transfer; GPT Image 2 for dense text + style
Material on object✅ GoodGPT Image 2, Seedream 5.0 LiteGrok is good; GPT Image 2 for stronger material rendering

The rule: Grok Imagine is the right choice for simple 2-3 reference composites where the priority is clear (one subject, one setting, optional style). For more references, more characters, or more complex multi-input workflows, use Nano Banana 2 (14 refs) or GPT Image 2 (4 refs).

The 20-composites-per-week production cadence

Week 1: Establish the source library. Pick 20 source products + 20 source settings + 5 style references. The library compounds.

Week 2: Generate 20 composites. One product × one setting per composite. Run the batch. ~30 minutes.

Week 3: QA and select the top 10. Watch all 20. Identify the top 10 by subject preservation and visual quality. Re-generate the bottom 10 with refined prompts.

Week 4: Generate 20 more from a different set. Different products, different settings. The library grows.

By week 12: 240+ composites, a tested library, and a continuously-improving production engine. The cost is roughly $0.02-0.05 per composite (Grok Imagine's text-to-image pricing). 20 composites/week is under $2/week.

The summary

Grok Imagine's multi-image editing is the model-specific workflow for 2-3 reference composites. The 4 production patterns — product on setting, character in setting, style transfer, material on object — cover ~90% of multi-image use cases.

  • Use the 3-reference role pattern (subject + setting + style) for clear priority.
  • Use the named-reference prompt pattern ("from reference 1, take the subject; from reference 2, take the setting; from reference 3, take the style").
  • Write an explicit preserve list (specific elements that must not change).
  • Pick the model for the pattern (Grok for 2-3 refs, NB2 for more, GPT Image 2 for text + style).
  • Apply the 10 brand-safety rules (no fake anything, all text in quotes).
  • Run the 20-composites-per-week production cadence (~$2/week, 30 minutes of generation).
  • Iterate with the 7-pass loop (generate → verify subject → verify setting → verify style → verify safety → post-process → final check).

The model is not the bottleneck. The named-reference discipline is. Pick the pattern, name the references, write the preserve list, apply the brand-safety rules, and multi-image composites become a reliable, scalable production workflow.

Share this article: