AdX

How to Build a Reusable Prompt Card for Any AI Image Model

How to Build a Reusable Prompt Card for Any AI Image Model

A prompt card is a structured record with 11 spec fields, generation metadata, and outcome metadata. The spec is portable across models; the metadata is what makes it reusable. The format, the per-model renderings, the JSON template, and the 4-step workflow for turning a winning prompt into a librar

Most prompt "libraries" are folders of text blobs. The best ones are prompt cards — structured fields that name every part of the prompt so you can swap models, reuse for similar projects, and never lose the constraint that made a prompt work.

A prompt card is not a clever format invented for this blog. It is the working pattern of every prompt library that scales beyond one user and one project. Below is the format, the field definitions, the model-specific adapters, the metadata to capture, and the workflow for turning a winning prompt into a reusable card.

What a "prompt card" actually is

A prompt card is a structured record with three layers:

  • Spec fields — the parts of the prompt that describe the image (subject, setting, style, action, composition, lighting, materials, framing, use case, constraints, negatives)
  • Generation metadata — the model, version, size, quality, seed, and parameters that produced the accepted output
  • Outcome metadata — what stayed stable across re-runs, what drifted, the fix that worked, and the use cases the prompt serves

The spec fields become the prompt. The generation metadata lets you reproduce the result. The outcome metadata is what makes the card actually reusable — you know why it works, not just what it is.

The 11-field spec format

The format below works on Nano Banana 2, GPT Image 2, Seedream 5.0 Lite, Grok Imagine, Midjourney v7, Flux, and Stable Diffusion XL. Different models prefer different fields, but the full set is portable.

#FieldWhat it specifiesExample
1SubjectThe person, object, character, or sceneA 30-year-old woman with short dark hair, cream linen blazer, soft smile
2SettingWhere it happens (location, time, era)A sunlit modern office, golden hour through a west-facing window
3ActionWhat is happening (pose, gesture, event)Laughing at the camera, holding a coffee cup
4StyleVisual rendering languagePhotorealistic, editorial, 1980s Kodachrome
5Camera / lensPhotography or film language85mm portrait lens, f/2.8, shallow depth of field
6LightingWhere the light is and how it behavesKey light upper-left, soft fill right, 5500K daylight, gentle shadow under chin
7CompositionFraming, crop, where the subject sitsCentered subject, 2 feet negative space above head, square crop
8MaterialsSurfaces and textures (if relevant)Matte cotton blouse, brushed metal desk, frosted glass coffee cup
9Use caseWhat the image is forLinkedIn headshot, ecommerce hero, editorial magazine cover
10ConstraintsHard rules the model must followNo watermark, no extra text, photoreal only, full body visible
11NegativesWhat must NOT appearNo extra fingers, no warped text, no logos, no text, no second person

You do not need all 11 every time. A social post can ship with 4-5 fields. A hero asset uses 8-9. A production prompt card for a brand system uses all 11.

The full example, applied

Here is the same 11-field card, filled in for a specific use case (LinkedIn professional headshot), and the rendered prompt that each model receives.

The card:

Subject:    30-year-old woman, short dark hair, cream linen blazer, soft smile
Setting:    sunlit modern office, golden hour through a west-facing window
Action:     laughing at the camera, holding a white coffee cup
Style:      photorealistic, editorial
Camera:     85mm portrait lens, f/2.8, shallow depth of field
Lighting:   key light upper-left, soft fill right, 5500K daylight, gentle shadow under chin
Composition: centered subject, 2 feet negative space above head, square crop
Materials:  matte cotton blouse, brushed metal desk edge, frosted glass coffee cup
Use case:   LinkedIn headshot, professional but warm
Constraints: photoreal, no watermark, no extra text, no second person
Negatives:  no extra fingers, no warped anatomy, no logos, no text, no busy background

The same card, rendered for each model:

Nano Banana 2 (gemini-3.1-flash-image) — paragraph style, DeepMind's official structure:

A photorealistic editorial portrait of a 30-year-old woman with short dark
hair, wearing a cream linen blazer, soft smile, laughing at the camera
while holding a white coffee cup, in a sunlit modern office with golden
hour light through a west-facing window. Captured with an 85mm portrait
lens at f/2.8, shallow depth of field, key light from the upper-left with
soft fill from the right, 5500K daylight, gentle shadow under the chin.
Centered subject, 2 feet of negative space above the head, square crop.
Matte cotton blouse, brushed metal desk edge visible, frosted glass
coffee cup. For a LinkedIn headshot, professional but warm. No watermark,
no extra text, no second person. Do not include extra fingers, warped
anatomy, logos, or busy background elements.

NB2 prefers paragraphs and narrative structure. The card's fields become sentences.

GPT Image 2 (gpt-image-2) — three-section structure with explicit constraints:

Scene: A sunlit modern office, golden hour light through a west-facing
window, key light from the upper-left, soft fill from the right, 5500K
daylight, gentle shadow under the chin.

Subject: A 30-year-old woman with short dark hair, wearing a cream linen
blazer, soft smile, laughing at the camera, holding a white frosted
glass coffee cup. Matte cotton texture, brushed metal desk edge visible
in the foreground.

Constraints: Photorealistic, 85mm portrait lens, f/2.8, shallow depth of
field, centered subject with 2 feet of negative space above the head,
square crop. Use case: LinkedIn headshot, professional but warm. No
watermark, no extra text, no second person, no extra fingers, no warped
anatomy, no logos, no busy background.

GPT Image 2 prefers labeled sections. The card's fields become section headers.

Seedream 5.0 Lite — structured, explicit, reasoning-style:

Create a professional LinkedIn headshot. Subject: 30-year-old woman,
short dark hair, cream linen blazer (matte cotton), soft smile, laughing
at the camera, holding a frosted glass white coffee cup. Setting: sunlit
modern office, golden hour, west-facing window. Lighting: key light
upper-left, soft fill right, 5500K daylight, gentle shadow under chin.
Camera: 85mm portrait lens, f/2.8, shallow DoF. Composition: centered,
2 feet negative space above head, square crop. Visible materials: matte
cotton, brushed metal desk edge, frosted glass. Hard rules: photoreal
only, no watermark, no extra text, no second person, no extra fingers,
no warped anatomy, no logos, no busy background.

Seedream reads structured language well. The card's fields become ordered clauses.

Grok Imagine / Midjourney v7 — keyword + parameter style:

editorial portrait of a 30-year-old woman with short dark hair in a
cream linen blazer laughing at the camera holding a white coffee cup,
sunlit modern office, golden hour, west-facing window, 85mm portrait
lens, f/2.8, shallow depth of field, key light upper-left, soft fill
right, 5500K daylight, matte cotton, brushed metal desk edge, frosted
glass, square crop --ar 1:1 --style raw

These models prefer compressed keyword language with explicit modifiers. The card's fields become a comma list with --ar and --style parameters at the end.

Flux / Stable Diffusion XL — tag + weight style:

professional LinkedIn headshot, editorial portrait, 30 year old woman,
short dark hair, cream linen blazer, soft smile, laughing, holding
white coffee cup, sunlit modern office, golden hour, west-facing
window, 85mm portrait lens, f/2.8, shallow depth of field, key light
upper-left, soft fill right, matte cotton, brushed metal, frosted
glass, square crop, photorealistic

Flux and SDXL prefer tag-style language. The card's fields become a comma-separated list of weighted tags.

The point: one card, six model renderings, no rewriting from scratch. You maintain the spec once and translate per model.

The generation metadata (reproducibility)

A prompt card without generation metadata is just a text blob. Capture these fields on every accepted output:

  • Model name and version (e.g., gemini-3.1-flash-image, gpt-image-2, bytedance/seedream-5-lite)
  • Size / resolution / aspect ratio (e.g., 1024×1024 1:1, 2K 4:5)
  • Quality parameter (e.g., low, medium, high, 2K, 4K)
  • Seed (if the API exposes it — most do)
  • Generation count (how many candidates it took to land the accepted output)
  • Edit passes (for image-to-image or inpaint workflows, how many passes)
  • Date generated (model behavior changes; you want to know which version of the model produced this)

This metadata is the difference between "this prompt worked once" and "this prompt is reproducible." Without it, the model updates in 3 months and your card stops producing the same output.

The outcome metadata (reusability)

This is the part most prompt libraries skip. Capture on every accepted output:

  • Stability anchors — the 2-3 phrases in the spec that held identity / look / composition across re-rolls. Example: "85mm portrait lens" was the anchor for facial consistency.
  • Drift triggers — what caused the model to deviate when it did. Example: "Mentioning 'soft' lighting caused the model to drop contrast by 30%."
  • Fixes that worked — what you changed in a follow-up prompt to correct drift. Example: "Adding '5500K daylight' explicitly fixed the warm-tone drift on pass 2."
  • Use cases served — what the prompt is good for. Example: "LinkedIn headshots, blog author bios, About page portraits."
  • Variants that exist — links to derived cards (different subject, same style; different setting, same composition).

Without this, a card tells you what the prompt was but not why it works. With it, you can adapt the card to similar projects, predict what will break, and avoid re-discovering the same fix every time.

The format, in JSON

If you want to version-control your library (recommended), the card in JSON looks like this:

{
  "id": "linkedin-headshot-warm-editorial",
  "title": "Warm editorial LinkedIn headshot",
  "version": 3,
  "spec": {
    "subject": "30-year-old woman, short dark hair, cream linen blazer, soft smile",
    "setting": "sunlit modern office, golden hour through a west-facing window",
    "action": "laughing at the camera, holding a white coffee cup",
    "style": "photorealistic, editorial",
    "camera": "85mm portrait lens, f/2.8, shallow depth of field",
    "lighting": "key light upper-left, soft fill right, 5500K daylight, gentle shadow under chin",
    "composition": "centered subject, 2 feet negative space above head, square crop",
    "materials": "matte cotton blouse, brushed metal desk edge, frosted glass coffee cup",
    "use_case": "LinkedIn headshot, professional but warm",
    "constraints": "photoreal, no watermark, no extra text, no second person",
    "negatives": "no extra fingers, no warped anatomy, no logos, no text, no busy background"
  },
  "generation": {
    "model": "gemini-3.1-flash-image",
    "model_version": "gemini-3.1-flash-image-preview",
    "size": "1024x1024",
    "aspect_ratio": "1:1",
    "quality": "2K",
    "seed": 4231098571,
    "candidates": 4,
    "edit_passes": 0,
    "generated_at": "2026-06-15"
  },
  "outcome": {
    "stability_anchors": ["85mm portrait lens", "5500K daylight", "key light upper-left"],
    "drift_triggers": ["mentioning 'soft' lighting drops contrast", "removing 'frosted glass' loses material cue"],
    "fixes_that_worked": ["specifying 5500K explicitly fixes warm-tone drift"],
    "use_cases_served": ["LinkedIn headshots", "blog author bios", "About page portraits"],
    "variants": ["linkedin-headshot-warm-editorial-male", "linkedin-headshot-warm-editorial-outdoor"]
  }
}

JSON in a version-controlled repo gives you: diffs when you tweak a field, search across cards, automatic validation of field completeness, and a library that survives model updates.

The workflow: prompt to card in 4 steps

Step 1: Generate a candidate. Write a freeform prompt. Generate 4-8 candidates. Pick the one closest to your intent.

Step 2: Reverse-engineer the spec. Look at the accepted output and fill in the 11 fields. Use the image to identify the implicit camera, lighting, materials, and style. The reverse-engineer step is where you discover what actually worked — the implicit field values you did not write in your original prompt.

Step 3: Re-generate from the spec. Take the filled-in 11-field card, render it for your model, and confirm the re-rendered output matches the accepted output. If it does not, the card is missing a field. Add it.

Step 4: Capture generation + outcome metadata. Record the model, version, size, seed, candidate count, stability anchors, drift triggers, fixes, and use cases. This is what makes the card reusable next time.

The 4-step loop is: generate → spec → re-generate → metadata. Done well, you produce one card per accepted output, and the library compounds.

What makes a card actually reusable

A card is reusable if all three are true:

  1. Spec is complete — every field that mattered for the accepted output is filled in. Reverse-engineering from the image catches implicit fields.
  2. Generation is reproducible — the metadata lets you re-run the same model and get the same output. Seed matters.
  3. Outcome is documented — you know which fields are stability anchors, which trigger drift, and what fixes work. The next user of the card does not have to re-discover this.

A card missing any of the three is a text blob with extra steps. A card with all three is a piece of organizational knowledge that compounds over time.

The summary

A prompt card is a structured record with 11 spec fields, 7 generation metadata fields, and 5 outcome metadata fields. The spec is portable across models. The metadata is what makes the spec reusable.

  • Spec fields: subject, setting, action, style, camera, lighting, composition, materials, use case, constraints, negatives
  • Generation metadata: model, version, size, aspect ratio, quality, seed, candidate count
  • Outcome metadata: stability anchors, drift triggers, fixes, use cases, variants
  • Render per model: paragraph (NB2), labeled sections (GPT Image 2), ordered clauses (Seedream), keyword + parameters (Grok / Midjourney), tag list (Flux / SDXL)
  • Workflow: generate → spec → re-generate → metadata. One card per accepted output.

Maintain the spec once. Translate per model. Capture the metadata. The library compounds.

Share this article: