AdX

From 1.5 to 2: What OpenAI Quietly Changed in Its Flagship Image Model

From 1.5 to 2: What OpenAI Quietly Changed in Its Flagship Image Model

Moving from OpenAI's previous flagship to the current one: 8 things that changed at the prompt level, 4 head-to-head tests, the migration checklist, and when the older model is still the right choice. Plus the 1 upgrade that is a real downgrade.

GPT Image 2 (gpt-image-2) is the default in ChatGPT since May 12, 2026. GPT Image 1.5 is the previous generation, available until its December 1, 2026 deprecation. If you're a prompt writer moving from 1.5 to 2, here's what actually changes at the prompt level — what's the same, what's new, and the migration checklist that catches the gotchas.

The 7 fundamentals from OpenAI's cookbook still apply to both models. The structural prompt order (background/scene → subject → key details → constraints) carries over. The quality parameter (low, medium, high) works the same. But 8 specific things changed, and 1 of them (the loss of transparent backgrounds) is a real downgrade for some workflows.

This post is the prompt-writer's upgrade guide. 4 head-to-head tests with concrete results, 8 specific things that changed, the migration checklist, and when 1.5 is still the right choice.

The 8 things that changed for prompt writers

1. Multilingual text now works

GPT Image 1.5 was Latin-only. GPT Image 2 handles Chinese, Japanese, Korean, Hindi, Bengali natively. If your prompts include non-Latin scripts (think product packaging for an Asian market, scientific posters in Japanese, app UI in Korean), 2 is the upgrade. In the fal.ai head-to-head Test 1, a Japanese/Korean/English scientific poster rendered cleaner Hangul on 2 than on 1.5.

What to do: If you have prompts with non-Latin scripts, just re-run them on 2. Don't add "render the Japanese accurately" — the model now does it by default.

2. The "warm color cast" is gone

1.5's outputs had a slight yellow/orange bias. 2 is neutral. If you've been writing prompts that compensate ("cool blue tones, no yellow cast," "no warm lighting"), drop those instructions on 2. They're now actively wrong — they'll push the result too cold.

3. Custom dimensions are now possible

1.5 was 1024x1024, 1024x1536, or 1536x1024 only. 2 accepts any size satisfying the constraints:

  • Edges must be multiples of 16
  • Max edge length < 3840px
  • Aspect ratio (long/short) ≤ 3:1
  • Total pixels: 655,360 to 8,294,400

Practical: replace "1:1 aspect ratio" with "1024x1024" or your specific dimension. Replace "widescreen" with "1920x1080" or "2560x1440" (the recommended upper reliability boundary). Above 2560x1440 is marked "experimental" in OpenAI's prompting guide.

4. Transparent backgrounds no longer work

This is a real downgrade. 1.5 had background: transparent parameter for transparent PNG output. 2 doesn't expose it. If you depended on transparent PNG (for compositing in Figma, or generating icons/stickers), you have two options:

  • Workaround A: Use a different model for transparent outputs (Nano Banana 2 with the appropriate endpoint, or the GPT Image 1.5 endpoint until December 2026).
  • Workaround B: Use GPT Image 2's edit endpoint with a transparent-mask request to "remove the background" of an opaque image.

For icon/sticker workflows, this is the one thing that gets worse on 2.

5. Editing is now always high-fidelity

1.5's edit endpoint had input_fidelity: low|high — a parameter that controlled how strictly source pixels were preserved. low was much cheaper (135 tokens per reference image vs 3,050 at high). 2 doesn't expose this parameter. The model is always high-fidelity.

Practical: edit-heavy pipelines on 2 cost more in input tokens than 1.5 with input_fidelity: low. If you're doing 100+ edits per day, this is a meaningful cost shift. Test before migrating the pipeline.

6. 4K is now possible (but experimental above 2K)

1.5 topped out at 1536x1024. 2 goes to 3840x2160. But anything above 2560x1440 (the 2K threshold) is "experimental" per OpenAI's prompting guide. The results are more variable. For most production work, stay at or below 2K. For hero assets where you specifically need 4K (large prints, video stills), use 2K as the default and 4K as the explicit choice for the final.

7. Thinking is on by default

GPT Image 2 is the first OpenAI image model with thinking capabilities. Web search and self-checking are baked in. 1.5 had no thinking. Practical: 2 is slower per generation than 1.5, but more accurate on multi-step prompts. The thinking is what makes the 8 upgrades (text rendering, multilingual, complex compositions) work consistently.

You don't toggle thinking on 2. It's always on. If you need fast iteration, use quality: low to keep generation speed up.

8. Token billing unit changed

1.5 is per-thousand tokens. 2 is per-million. For typical prompts (a few hundred tokens), the change doesn't matter. For high-volume pipelines, the rounded-up-to-the-cent billing on 2 can add up at the edges (literally rounding up to the next cent on every request).

If you're doing a few hundred images per day, the per-million billing on 2 is fine. If you're doing 10,000+ images per day, re-test your cost model.

The 4 head-to-head tests (fal.ai, May 2026)

Fal.ai ran 4 tests on identical prompts against both models. The pattern: 2 wins on realism and multilingual text, 1.5 wins on prompt adherence for exact text. Sometimes that matters more.

Test 1: Multilingual scientific poster (mixed scripts, small subscript)

A clinical research poster with English, Japanese, and Korean columns, a Kaplan-Meier curve, and an 8-point footnote.

  • GPT Image 2: cleaner Korean (fewer Hangul mistakes), neutral colors, darker clinical tone
  • GPT Image 1.5: yellow-leaning colors, some Korean mistakes
  • Winner: GPT Image 2 (narrow, on text rendering)

Test 2: Three-temperature interior lighting

A studio with a north-facing window at 6500K, a tungsten lamp at 2700K, and overhead fluorescent at 4000K. A mug should be cool grey on its left side, warm cream on its right.

  • GPT Image 2: higher-quality background detail
  • GPT Image 1.5: both produced above-average photorealism
  • Both failed the mug test — 2 interpreted it as a shadow, 1.5 turned it half blue
  • Winner: GPT Image 2 (narrow, on background)

Test 3: Pharmaceutical product with regulatory small print

An amber dropper bottle with the label "NOCTURNA Rx" and a 4-line ingredient panel, 3 certification marks, batch code, expiration date, and 1D barcode.

  • GPT Image 2: better barcode rendering
  • GPT Image 2 also wrote "NOCTURNA Px" instead of "NOCTURNA Rx" — critical error for pharmaceutical labeling
  • Winner: GPT Image 1.5 (on prompt adherence for exact text)

This is the test that matters most for production work. If you're building a brand asset library or generating regulatory content, the "NOCTURNA Px" mistake would fail a review. 1.5's stricter prompt adherence is the safer choice here.

Test 4: Architectural interior with fine technical detail

A university library with 6 study tables, 6 lamps, varied book spines, brass section plates, fog through glass.

  • GPT Image 2: rendered 6 tables, more realistic, better prompt adherence
  • GPT Image 1.5: rendered only 2 tables (ignored the "6" instruction), less realistic
  • Both failed the "6 lamps, not 8" test (both rendered 8)
  • Winner: GPT Image 2 (clear, on realism and adherence)

Net of the 4 tests: GPT Image 2 won 3, GPT Image 1.5 won 1. The 1.5 win is the one that matters most for production work where exact text matters. If you're building a brand asset library, do not assume 2 will get your brand names right — verify each one.

Pricing comparison (fal.ai)

GPT Image 1.5 — 3 sizes, per-image + per-1k tokens

Quality1024x10241024x15361536x1024
Low$0.009$0.013$0.013
Medium$0.034$0.051$0.050
High$0.133$0.200$0.199

Token charges (1.5):

  • $0.005 per 1,000 input text tokens
  • $0.008 per 1,000 input image tokens (edit endpoint, where 1024x1024 reference = 135 tokens at low fidelity, 3,050 at high)
  • $0.010 per 1,000 output text tokens

GPT Image 2 — 6 sizes + custom, per-image + per-1M tokens

Quality1024x7681024x10241024x15361920x10802560x14403840x2160
Low$0.005$0.006$0.005$0.005$0.007$0.012
Medium$0.037$0.053$0.042$0.040$0.056$0.101
High$0.145$0.211$0.165$0.158$0.222$0.401

Token charges (2):

  • $5.00 per 1M input text tokens
  • $1.25 per 1M cached text tokens
  • $10.00 per 1M output text tokens
  • $8.00 per 1M input image tokens
  • $2.00 per 1M cached image tokens
  • $30.00 per 1M output image tokens

Cost shift

  • Per-image: 2 is more expensive at 1K-2K (e.g. 1024x1024 high: $0.211 vs $0.133). 2 unlocks 4K but 4K is $0.401/image.
  • Tokens: 2's per-million billing is cheaper for typical prompts and more expensive at the edges.
  • Edit cost: 1.5's input_fidelity: low was 22x cheaper per reference than high. 2 is always high. Edit-heavy pipelines cost more on 2.

Migration checklist (for prompt writers)

If you're moving prompts from 1.5 to 2, work through this list:

  • Remove "no warm cast" or "cool tones" instructions. 2 is neutral by default; these now push the result too cold.
  • Add size as 4 digits if you want a specific resolution. "Render at 1920x1080" instead of "widescreen."
  • Replace "1:1 aspect ratio" with "1024x1024" for predictable output. "16:9 aspect ratio" with "1920x1080" or "2560x1440."
  • Test multilingual text rendering if your prompts include non-Latin scripts — should now work without explicit instruction.
  • If you depended on background: transparent, plan a workaround before Dec 1, 2026 (when 1.5 is deprecated). Options: use 1.5 until then, use a different model for transparent outputs, or use the edit endpoint with a mask.
  • Re-test the cost of edit-heavy pipelines. 2's always-high-fidelity input is more expensive than 1.5's low-fidelity option.
  • Lock the quality tier at low for iteration, high for finals. This still works on 2.
  • Test brand-name and exact-text rendering. Don't trust 2 to spell brand names right. In the fal.ai Test 3, 2 wrote "NOCTURNA Px" instead of "NOCTURNA Rx." Verify all critical text in your output.
  • For 4K output: stay at 2560x1440 or below for production. Above that, results are experimental.
  • Don't migrate transparent-bg or strict-prompt-adherence workflows yet if you don't have a 1.5 fallback. Wait until you've validated the migration in a staging environment.

When to use which (for prompt writers)

Use GPT Image 1.5 if:Use GPT Image 2 if:
You need transparent PNG outputYou need 4K output (up to 3840x2160)
You're doing transparent-bg compositingYou have multilingual text in prompts (CJK, Hindi, Bengali)
You have cost-sensitive edit pipelines using input_fidelity: lowYou want better photorealism and neutral color baseline
You have a stable 1.5 workflow you don't want to retune (until Dec 1, 2026 deprecation)You need custom dimensions beyond 1024x1024 / 1536x1024
You need strict prompt adherence for exact brand names / exact textYou're starting a new workflow

For most prompt writers, GPT Image 2 is the right choice. It's the default in ChatGPT since May 12, 2026. The OpenAI cookbook is written for it. 1.5 is scheduled for deprecation on December 1, 2026. The transparent-bg limitation is the only real loss.

The summary

  • 8 things changed for prompt writers. Multilingual works. Warm cast is gone. Custom dimensions are possible. Transparent backgrounds no longer work. Editing is always high-fidelity. 4K is now possible (experimental above 2K). Thinking is on by default. Token billing switched to per-million.
  • The 7 fundamentals still apply. Order, format, specificity, latency-vs-fidelity, composition, people-pose-action, constraints. Carry over from 1.5 to 2.
  • GPT Image 2 won 3 of 4 head-to-head tests on realism, multilingual, and complex compositions. 1.5 won 1 — the test for exact-text prompt adherence (the "NOCTURNA Px" mistake). For brand-critical text, verify on 2.
  • Use 1.5 until December 1, 2026 if you depend on transparent backgrounds. After that, plan a workaround or use a different model.
  • Test your edit-heavy pipelines. 2's always-high-fidelity input is more expensive than 1.5's low option.

Want to see how this stacks up against the other primary models? Compare GPT Image 2 vs Nano Banana 2 for cross-vendor differences, or browse the full prompt library for prompts optimized for all 4 primary models.

Share this article: