Chapter 12 · GPT Image 2.5
Part 2 · Stills
GPT Image 2.5 is the second stills camera. Nano Banana Pro carries a signed face and the photographic keyframe; this camera makes what the first does not: an element on a transparent background, a frame led by text or layout, a packaging or key-visual study, and a bounded edit on a frame with no signed face. It is also the quick way to look at an idea. This chapter teaches how to write for it, how to set it, what to check on the file that comes back, and what it costs.
In this chapter
Its two models, and which job goes to which camera
What it reads, and the twelve elements in its dialect: the labelled brief, quoted copy, transparent elements, edits
Templates, and four examples across beauty, food and drink, telecom and tech, automotive and luxury
Settings, checks, failures and fixes, limits, prices and rights
Before you start. Chapter 9 teaches the twelve elements, the two grades and the frozen instant; this chapter maps them to one tool and does not repeat them. Chapter 10 teaches references, identity and the one-change edit contract. Chapter 11 is the first stills camera.
12.1 What it is
GPT Image 2.5 is OpenAI's image model, current on 30 Sep 2026. It comes as two models, both with snapshots dated 8 Sep 2026. Flare is the small model, "optimized for speed, with image quality comparable to GPT Image 2". Sunburst is the base model, "optimized for quality, with higher image quality than GPT Image 2", and OpenAI points to it where editing precision matters most. Both generate, edit and make transparent backgrounds. The names are models, not lighting effects. OpenAI says new work should use one of the two; GPT Image 2 is the previous generation and is not used here.
The book reaches it through Higgsfield (Chapter 33), on the same account and credits as Nano Banana Pro. OpenAI also offers it in its own API and in ChatGPT; those routes have their own fields (size, output format, mask), so never paste one route's settings into the other. When a request names no image model, the host picks GPT Image 2.5 (checked 27 Sep 2026): name the camera you chose in every request.
The jobs it takes.
An element on a transparent background. It is the only stills route with a real alpha channel.
A frame led by text or layout: a key visual, a poster or billboard study, a diagram, a packaging flat.
A bounded edit on a frame with no signed face: remove a stray object, add one element, change a garment or a state.
Fast exploration: several cheap looks at an idea before you commit to one.
The jobs it does not take.
A signed identity carried by several references, and the photographic keyframe, stay on Nano Banana Pro (Chapter 11). OpenAI documents identity preservation in edits, yet warns that the model "may occasionally struggle to maintain visual consistency for recurring characters or brand elements across multiple generations". No maker compares the two cameras. When a keyframe fails twice on one camera, the same brief on the other is a tactic, not a claim that either is better.
Type that ships, Arabic wordmarks, and a face with a headline in one frame. Type is set in post, and a person and a headline are two layers.
A masked edit. The host offers no mask.

Figure 12.1 — Choosing the stills camera. Ask the questions in order, stop at the first yes, and write the reason beside the job. The routing rests on the makers' documentation and on what the finished work used; no maker ranks the two cameras on a photographic keyframe.
12.2 How it reads what you give it
Text. It reads plain paragraphs, and it follows facts it can check in the picture (whose hand is under the box, where the eyes rest, what has not happened yet) better than moods. OpenAI asks for the same: name body framing, gaze and interaction with objects, as in "hands naturally gripping the handlebars". Words in quotation marks are read as required copy, with the position and typography you give them.
Pictures. Every picture goes through one input slot, whether it is the frame to edit, a next-shot reference or a style reference. The host states no maximum, and it does not state whether "the first image" means the first file attached, so name each picture by what it shows as well as by its number: "Image 1, the approved frame of the living room, controls the room and the camera." A picture is a strong authority. It hands over faces, clothes, objects and light, and it also hands over its own moment unless the prompt says otherwise. Say what each picture is for and what it must not bring (Chapter 10 teaches the manifest).
Settings. Variant, quality, resolution, aspect and background are controls of the route. Written into the prompt they do nothing, or they fight the real setting. The background needs both: set it to transparent, and say "isolated on a transparent background" in the words. There is no seed, no field for exclusions and no mask on the host; exclusions live in the prose, at most five, then the closing line.
Where its maker says it is weak. It "can still struggle with precise text placement and clarity"; it "may have difficulty placing elements precisely in structured or layout-sensitive compositions"; and complex prompts "may take up to 2 minutes". Plan for all three: fence the copy, write the layout as places and surfaces, and check every placement on the returned file.
12.3 The prompt, in this tool's dialect
A photographic frame is written exactly as Chapter 9 teaches: the twelve elements in their fixed order, world first, photography second, technical close, run together as one paragraph. This section gives only what is particular to this camera.
Labelled sections for layouts
OpenAI suggests that a complex request be organised as scene, subject, details and constraints, with labels. A layout, a diagram, a key visual, a text-heavy brief or a multi-part edit is easier to write and to check in parts. The four labels carry the twelve elements in the same order, so nothing is lost.
Label | Carries |
Scene | the medium, purpose and format (1); the place and its materials (6); the air (10) |
Subject | who or what, and the frozen instant (2); the identity clause (3); wardrobe and props (4); the frame inventory and edges (5); the must-show items (7) |
Details | the camera (8); each light (9); the finish (11); every quoted word with its position and typography |
Constraints | what must stay exact; at most five exclusions; the closing line (12) |
A frame with a face stays one paragraph. Labels earn their place where each part needs its own check. A placement drawing can lead a layout too: OpenAI's pattern is to preserve the exact layout, proportions and perspective, choose realistic materials and light, and add no new elements or text.
The instant, the camera and the light
Write the frozen instant as positions and contacts, never as an open verb: "the nearer hand is open beneath the free underside of the box, a finger's width below it" fixes the phase, where "is beginning to reach" leaves it to the model. Write the camera as height, distance, lens and stop, and each light as a source, a Kelvin figure and a ratio, with the visible result beside every number (9.3). OpenAI's caution belongs beside the rule: "Treat camera specifications as cues for appearance, not a guarantee of exact physical simulation."
Ask for the medium in words. OpenAI says to request "photorealistic" or "real photograph" explicitly when that is the goal, and to describe framing and texture: pores, fabric wear, the grain of a surface. Words that imply studio polish work against a real photograph. For a designed frame, name the medium: "a flat typographic key visual".
Copy in the frame
Put each required word in quotation marks, and give its position, its typography and how many times it appears. Fence the frame: "nothing else is written anywhere in the frame". Use medium or high quality for small or dense text. A frame with copy in it is a study: every word that ships is set in post from real art.
OpenAI makes no claim for Arabic. When a study needs an Arabic line, quote the approved string, copy it exactly, ask for correct right-to-left order and connected letterforms, then read every letter, dot and join against the string. A final ى where ي was approved fails the frame. Brand names, product names and legal lines are never generated; they are set in post.
Transparent elements
Describe only the element: its shape, its material and its own light, with no ground. Ask for a clean edge and no cast shadow, and name what must not appear: a floor, a white rectangle, a drawn checkerboard, which OpenAI says "is not transparency". Give a translucent material enough density to hold its colour over both black and white. On every edit of an element, repeat that the transparent background is preserved.
Edits
Open with the change, in physical terms. Name the consequences it may cause (a shadow, a contact, an occlusion), then what stays, region by region. Change one thing per edit, repeat the preserve list each time, and start every edit from the last approved frame, never from a damaged one. OpenAI adds that repeated edits can still change what you meant to keep, and that a region which must stay pixel-identical is composited into the original. A change that names no boundary spreads.
Length
Use Chapter 9's bands: about 400 to 600 words for a complex hero still, 150 to 300 for an insert or a layout, 40 to 150 for an edit. OpenAI states no limit; none of the bands is measured on this camera. An exploration is written at test grade (9.6) and is never the model of a finished frame; a good one earns a finished-grade rewrite before anything is signed.
12.4 Templates
The photographic template is Chapter 9's, used unchanged. Four more cover what this camera is chosen for.
Written for this book · template; not run · finished grade · GPT Image 2.5, a layout-led or text-led frame, from words
Template — a labelled brief for a layout-led frame
Scene: [what the picture is, what it is for, its format; held still or not]. [The place and its materials; the surface kept bare for type; the air, only where it shows]. Subject: [each element and its place in the frame, with its share of the frame height; wardrobe or prop truth; what touches or leaves each edge; what must be visible]. Details: [camera: height, distance, lens and stop, each with its result]. [Each light: source, place, Kelvin, effect on a named surface; one ratio]. [The finish as capture]. [Every quoted word: position, typography, count.] Constraints: [what must stay exact]; [at most five exclusions]; nothing else is written anywhere in the frame; no overlaid text, subtitles, logos, or interface graphics.Written for this book · template; not run · finished grade · GPT Image 2.5, transparent background, an element from words
Template — a transparent element
A photographed [element] for compositing in [the film or the card], isolated on a transparent background: [its shape and the instant it is caught, as positions]. [Its material, and the density that holds its colour over black and over white]. It sits in the centre with clear space on all four sides, about [share] of the frame height, touching no edge. The camera is [height, distance, lens and stop, each with its result]. [Each light, with Kelvin and one ratio]. [The finish as capture]. No ground, no cast shadow, no reflection, no [objects it must not carry], no checkerboard pattern, and no overlaid text, subtitles, logos, or interface graphics.Written for this book · template; not run · finished grade · GPT Image 2.5, an edit of an approved frame with no signed face
Template — a bounded edit
Edit image 1. Change one thing: [the change, in physical terms]. [The consequences: what the rebuilt surface must match; the shadow or contact that goes or arrives with the change]. Keep [protected region], [protected region] and [the camera, crop, light and colour] exactly as they are. Do not [the substitute you fear]; do not tidy or restyle anything else; no lettering or logos.Written for this book · template; not run · test grade · GPT Image 2.5, Flare, a first look at an idea
Template — an exploration at test grade
[What the picture is and what it is for]. [The subject and its instant, as positions]. [The place in two or three materials]. [The camera in words]. [One named light source and its side]. [Any quoted copy, once, with its place]. No overlaid text, subtitles, logos, or interface graphics.12.5 Examples at the standard
Written for this book · beauty and personal care · GPT Image 2.5, Flare, high, 2k, 4:5, opaque, text only · finished grade; not run
Example 12.1 — a night-cream key visual with a quiet zone for type
Scene: A photographed premium product key visual for a night-cream campaign, made for a treatment page and a poster; the frame is held still. A honed pale travertine shelf crosses the lower third in front of a warm off-white plaster wall at blue hour. The wall's upper left third is smooth, unmarked and shadowed, a bare surface for post-owned type.
Subject: One unbranded round frosted-glass jar, about nine centimetres wide, with a matte white ribbed lid and a blank front label panel facing the camera, stands upright at the shelf's centre right, filling about 35 percent of the frame height. A slim brass spatula lies on the shelf to its left, its tip clean. The shelf's edge crosses the lower frame edge in the foreground; the jar and spatula are the midground; the wall is the background.
Details: The camera sits at 0.25 meters, level with the jar's shoulder, 0.7 meters from it, on an 85mm lens at f/5.6, so the jar is sharp from lid to base and the wall softens to warm grey. A window out of frame at camera left, blue-hour sky at approximately 7000K, gives a broad soft key that skims the travertine and rolls along the frosted glass; the plaster wall at camera right bounces a warm fill at approximately 4200K into the shadow side, and contrast sits near 3:1 so the glass edge reads on both sides. The finish is a real photograph from a medium-format digital capture: fine, restrained grain and a halation only where the window light meets the frosted glass.
Constraints: Keep the label panel blank and every surface unprinted. No cream visible outside the jar, no second product, no ornament on the wall, and nothing else written anywhere in the frame; no overlaid text, subtitles, logos, or interface graphics.Why it works
Copy the quiet zone. It is written as a surface with its contents ("smooth, unmarked and shadowed"), so the model has something to draw where the type will go.
Copy the label panel: it is asked for blank, because the real label is applied in post from the approved artwork.
Copy the numbers: every camera and light figure sits beside the look it should give.
Written for this book · food and drink · GPT Image 2.5, Flare, high, 2k, 1:1, background transparent, text only · finished grade; not run
Example 12.2 — a hibiscus-tea splash as a transparent element
A photographed liquid element for compositing over a dark end card in a chilled-tea commercial, isolated on a transparent background: one crown splash of hibiscus tea caught at the instant its rim has risen about eleven centimetres and is breaking into droplets. The crown is one unbroken ring at the base, ruby red at the core and lighter where the liquid thins, with six thin fingers rising from the rim, four of them ending in a round droplet that has just detached, and five loose droplets in flight, two on the left and three on the right, none overlapping the crown. The whole element sits in the centre of the frame with clear space on all four sides, about sixty percent of the frame height, and no part of it touches an edge. The liquid is dense enough to hold its ruby colour over white as well as over black, with glossy edges and a clear interior. The camera is level with the base of the crown, 1.2 meters away, on a 100mm lens at f/8, so every droplet is sharp at full size. One large soft source overhead and slightly behind, at approximately 5600K, draws a bright rim on each finger and droplet; a white card at camera left lifts the front of the ring so the ruby reads; contrast sits near 4:1, so the thin fingers keep detail without clipping. The finish is a clean high-speed capture: crisp droplet edges, no motion blur, fine restrained grain, no glow. No glass or ice, no ground, no cast shadow, no reflection, no checkerboard pattern, and no overlaid text, subtitles, logos, or interface graphics.Why it works
Copy the empty frame. It holds the element and nothing else, and every object a transparent frame tends to bring (a glass, a floor, a shadow, a drawn checkerboard) is named and refused.
Copy the density sentence: a thin translucent liquid shows the ground it was made against, so the prompt writes the result you check for, colour that holds over white as well as over black.
Copy the space on all four sides: an element that touches the edge is cut off in the alpha.
Written for this book · telecom and tech · GPT Image 2.5, Sunburst for precision, high, 2k, the source's aspect, one picture: the approved frame · finished grade edit; not run
Example 12.3 — one cable removed from a router frame
Edit image 1. Change one thing: remove the white power cable that runs across the oak shelf in front of the router, from the router's base to the right edge of the frame. Rebuild the shelf where the cable lay, with the same grain direction, the same soft window light from camera left and the faint shadow line under the router's base; the cable's own shadow goes with it, and the shelf line under the router's base stays continuous. Keep the router (its shape, its status lights exactly as lit, its blank front panel), the plant, the books, the wall, the camera angle, the crop and the colour exactly as they are. Do not replace the cable with another object, do not tidy or restyle anything else, and add no lettering or logos.Why it works
Copy the order. One change, then its consequences (the cable's shadow, the shelf line under the base), then the protected regions by name.
Copy the ban on a substitute: "do not replace the cable with another object" stops the model filling the gap with a plant or a wire of its own.
Then: flick the original against the result across the whole frame. Where the router's pixels must not move, composite the rebuilt shelf onto the original.
Written for this book · automotive and luxury · GPT Image 2.5, Flare, high, 2k, 16:9, opaque, text only · finished grade; not run; the English and Arabic lines are placeholder copy for a study, not client-approved
Example 12.4 — a launch billboard study, with copy
Scene: A photographed out-of-home study for a luxury SUV launch, made for a client deck to judge the headline in place and never delivered as artwork; the frame is held, nothing in it moves. A wide roadside billboard stands beside a coastal highway on the Red Sea at dusk, seen from a car on the shoulder at a slight angle.
Subject: The billboard face fills the centre of the frame, about 55 percent of its width and 30 percent of its height, on two grey steel posts that cut the lower frame edge. On the face, a dark graphite SUV in three-quarter view stands low on wet asphalt in the right two thirds of the artwork, with no badge and no name. In the left third, the English line "QUIET IS THE NEW POWER" appears once, in white condensed capitals on three short lines, left-aligned; beneath the SUV, at the right, the Arabic line "الهدوء هو القوة الجديدة" appears once, in white, right-aligned on one line, in correct right-to-left order with connected letterforms.
Details: The camera sits at 1.3 meters, 40 meters from the face, on a 35mm lens at f/8, so the face is sharp from corner to corner while the sea softens into one dark band. A row of downward floodlights on the billboard, at approximately 4000K, washes the face evenly; the dusk sky at approximately 9000K lights the posts and the shoulder in cool blue; contrast across the face sits near 2:1 so the white type stays crisp. The finish is a real photograph with fine, restrained grain.
Constraints: The two quoted lines are the only words in the frame; no other text, plate characters, road signs or logos anywhere; no second vehicle, no person, no brand mark, and no overlaid text, subtitles, logos, or interface graphics.Why it works
Quote, place and count every word. Everything else is fenced, so a stray word is a visible fault.
Copy the split: the badge and brand name are absent because they are set in post from the real artwork; the two lines study scale and hierarchy and are never the delivered type.
Check the Arabic: every letter, dot and join against the approved string. One wrong letter fails the frame.
Copy the layout in places: thirds of the face and a named side for each line, where adjectives would leave it to the model.
12.6 Settings
Write all four on every job: variant, quality, resolution, aspect. The host's defaults are the lowest quality tier and the smallest size, and low at 1k is the cheapest render the model makes.
Setting | Value | Why |
Variant | Flare for exploration, layouts and elements; Sunburst for an edit where precision matters | Flare is the default. An unwritten variant falls back to it, so two runs of "the same" job can differ. |
Quality | high; medium for layout drafts; low only for a first glance | The default is low. Use xhigh or max only against a named, unmet requirement: OpenAI says a higher setting does not guarantee a better result. |
Resolution | 2k; 4k for a hero still you will crop or print | The default is 1k. Nano Banana Pro's default is 2k. |
Aspect | always set: 16:9, 9:16, 4:5, 1:1, 21:9, or the source's own | Fifteen are listed, including 27:16, 16:27, 9:8, 8:9 and an automatic choice. There is no 1.85:1: make 16:9 and crop. |
Background | transparent for an element; opaque for everything else | Left out, the model decides. |
Pictures | one slot, named in the prompt | No maximum stated. |
Not available | seed, exclusions field, mask, output format | Exclusions live in the prose. |
Aspect and resolution are settings, never words in the prompt (Chapter 9). Write the settings line into the run record with the price: gpt_image_2_5 · variant flare · quality high · resolution 2k · aspect 16:9 · background opaque.
Returned size. At 2k and 16:9 the host returned 2688 × 1520 in September 2026, where Nano Banana Pro's 2k returns 2752 × 1536. When a sequence mixes the two cameras, conform in the edit and never stretch. 2688 × 1520 is not exactly 16:9: scaled to a 1920-wide sequence (71.43 %) it stands about six lines taller than 1080; let the sequence crop them.
Variants. Compare them by changing the variant alone (same files, prompt, quality, resolution, aspect) and judge the task you need: protected regions, a garment's construction, small text.
12.7 Checks before you sign
Look at 100 % first, then 200 %, then flick the frame against its source where there is one.
Hands and contacts first. Every grip, support and gaze target, in any frame that hinges on a touch.
Invented marks. Lettering, logos or texture on any object you called unbranded, plain or blank. A panel you asked blank must be blank.
Copy. Each quoted word present once, spelled as given, in its place, and no other text anywhere. For Arabic, every letter, dot and join against the approved string, and the reading order.
Quiet zone. The surface you kept bare for type is bare.
Alpha. Open the returned file and look at the alpha over black, over white and over a saturated colour: shoulders, glass, liquid, fine detail. A drawn checkerboard is a failure. Then judge the element composited over the real plate, and again after it has been through your editor and your export. Note which file type came back.
Edits. Original and result at the same size across the whole frame: faces, hands, lettering, edges, colour, crop. Anything the edit was not asked to touch that moved is a fail; composite the approved region onto the original.
Product truth. A generated product is a stand-in. The label and the geometry that ship come from the real thing.
Size. Record the returned size and aspect against what you set.
12.8 Failures and fixes
You see | Why | Smallest fix |
A small, soft frame | the defaults (low, 1k) were left | write all four settings and re-run |
Marks or letters on a plain object | the model adds plausible detail | "blank label panel, no lettering or marks"; remove in an edit, or paint out |
Misspelt or extra words | placement and clarity are documented weaknesses | quote once, fence the frame, raise quality; after two controlled attempts set the words in post |
Arabic unjoined, wrongly dotted or reversed | OpenAI claims nothing for Arabic | blank the panel and set the line from the real font |
The type block off the brief, or the quiet zone filled | layout placement is a documented weakness | write the zone as a surface, cut elements, or set the layout in post and keep the frame for mood |
A fringe on black or white | a translucent edge made against another ground | denser material in the words; choke the matte in post |
A floor, shadow or checkerboard baked in | transparency not asked for in both places | set the background and say it in the words; try 1:1; check the file |
The edit changed more than asked | the change named no boundary | name consequences and protected regions; one change; composite |
A face or hand moved in an edit | edited from a damaged frame, or a chain of edits | return to the approved parent; composite the region |
The reference's moment came back | the picture's state beat the written state | make them agree, or say what not to copy |
The event is wrong (who holds what) | the instant was implied | positions and contacts |
The retry ladder, smallest lever first.
Check the four settings were written.
Rewrite the failing sentence as positions and checkable facts.
Tighten the change and preserve lists, and return to the approved parent.
Raise quality one step, only for fine detail or small text.
Change the variant alone.
Change the layer: composite, paint or set the type in post, or, for a photographic keyframe, run the same brief on Nano Banana Pro.
Never a third render of the same words. Two misses with the same diagnosis mean change a layer.
12.9 Limits, prices and rights
Limits (checked 30 Sep 2026).
Inputs. Text and pictures in; one picture slot on the host with no stated maximum. OpenAI's own service also takes a mask, which it describes as prompt-based guidance that "may not follow its exact shape with complete precision"; the host offers none.
Time. Complex prompts may take up to 2 minutes (OpenAI).
Size. On the host: 1k, 2k or 4k, with fifteen aspects. On OpenAI's own service, custom sizes in multiples of 16, an aspect between 1:3 and 3:1, no edge over 3840 pixels, and sizes above 2560 × 1440 marked experimental.
Transparency. PNG or WebP on OpenAI's side; check which file the host returns.
Moderation. OpenAI filters prompts and images under its usage policy; a blocked job is not retried unchanged.
Prices on Higgsfield (Flare, credits per frame, read 27 Sep 2026 by a price check at no cost; a dash means not checked).
1k | 2k | 4k | |
low | 0.25 (the defaults) | — | — |
medium | — | 1 | — |
high | 1.5 | 2.75 | 4.25 |
xhigh | — | 4.5 | — |
max | — | 9 | — |
Sunburst at high and 2k quoted the same 2.75, as did a transparent 1:1 element; one attached picture added nothing; Nano Banana Pro at 2k quoted 2. At the planning rate of about five US cents a credit (Chapter 34), a high 2k frame is about 14 cents. The same shape quoted 5.5, 3 and 2.75 credits between 11 and 27 Sep, so ask for the exact shape's price on the day, write it beside the job, and read the charge afterwards. On OpenAI's own service, image tokens cost USD 8 per million in (2 cached) and 30 per million out (checked 30 Sep 2026).
Rights. OpenAI's terms of use (effective 1 Jan 2026, read 30 Sep 2026) say the user owns the output, and forbid representing output as human-generated when it was not. On Higgsfield the host's own terms govern the account (Chapters 33 and 34). A real face, a brand mark and a real product label are rights decisions made before generation, never after. What may ship: a signed frame or element, with the type set in post, through the rights and acceptance gates. The host's own background remover, relight and outpaint tools are not used, because none names the engine that runs it.
12.10 Version notes
Current. GPT Image 2.5, in Flare and Sunburst, snapshots dated 8 Sep 2026 (checked 30 Sep 2026). OpenAI's guide names them for all new work.
What changed from GPT Image 2. Both models add the xhigh and max quality settings and, in OpenAI's words, "offer improvements in precise editing and subject preservation". Token rates are unchanged. The input-fidelity setting is documented for earlier models only, and OpenAI's cookbook is still written for GPT Image 2: use it for patterns, not for 2.5 settings.
Re-check when the next version ships. The variant names and which one is the host's default; the quality tiers and their prices; whether the host returns transparent files as PNG or WebP; whether a mask arrives; the number of pictures a job takes; the image sizes returned.
Open questions
Which camera is better at a photographic keyframe from text, and at a signed face carried through references. No maker compares them; the routing follows the documentation and the finished work.
How the model weighs a camera number against the words beside it, and the order of the twelve elements. Neither is published.
What it does with Arabic, and which file a transparent job returns on the host. Read both on the returned file.
Whether the host's Sunburst is exactly OpenAI's 8 Sep snapshot, the maximum number of pictures, and whether their order binds.
Sunburst against Flare on the edits a commercial needs. Before a client job depends on Sunburst, run one edit on both with the same file and prompt and only the variant changed; about 5.5 credits at 27 Sep prices.




Comments