top of page

Chapter 11 · Nano Banana Pro

Writer: Yasser Ashour
Yasser Ashour
2 hours ago
44 min read

Part 2 · Stills

Nano Banana Pro is Google's image model and the main stills camera of this book. It makes the hero frames of a commercial, the first picture of a character, the keyframes that the video models animate, and the repairs that save a frame that is nearly right. It reads a prompt once, keeps no memory between jobs, and takes every word and every picture literally. This chapter teaches it completely, as checked on 30 Sep 2026: what it reads, the twelve elements written in its dialect, templates, seven frames at the finished standard, settings, checks, failures, and its limits, prices and rights.

In this chapter

  • What Nano Banana Pro is today, and when another tool is the right one

  • What it reads: words, up to fourteen pictures, and how attachment order binds them

  • The twelve elements in its dialect, the four-part order for a frame made from pictures, and four templates

  • Seven examples: beauty, food and drink, telecom in a home, an automotive exterior, Mastorna's establishing shot, a frame built from an actor's photograph, and a one-change edit

  • Settings, checks, failures and fixes, limits, prices and rights, and version notes

Before you start. Chapter 9 teaches the finished frame this tool writes to, and Chapter 10 teaches references, identity and the one-change edit in general. This chapter maps both onto Nano Banana Pro and does not teach them again. Chapter 33 covers the host platform, and Chapter 34 the money and rights of a whole job.

11.1 What it is

Nano Banana Pro is Google's Gemini 3 Pro Image model (model id gemini-3-pro-image). Google released it in November 2025 and made it generally available on 28 May 2026. Google's model list, its model page and its changelog show no newer Pro image model (checked 30 Sep 2026). The two newer image models are Nano Banana 2 and Nano Banana 2 Lite, which are Flash-class models built for speed and volume, not the Pro model. This chapter therefore teaches Nano Banana Pro as the current Pro image model. Google describes it as "designed for professional asset production and complex instructions", with a default thinking process that refines the composition before it renders, output up to 4K, and up to 14 reference pictures.

Names cause mistakes here. Community posts and some prompt libraries call Nano Banana Pro "Nano Banana 2". In Google's own naming, Nano Banana 2 is the Flash model (id gemini-3.1-flash-image). When a tutorial's advice names a model, check the id before you copy the advice.

What it is for. Nano Banana Pro is the stills camera of the production. It makes the hero frames of a commercial, the master (the first neutral picture of a character, 11.2) and the anchors built from it, the keyframes that the video models animate (Chapter 9), the plates of rooms and products that recur, the boards of a treatment, and the repairs that save a frame that is 80 % right. It does best with faces, rooms, products in rooms, the blending of several reference pictures, and keeping an instant through an edit. Google adds text rendering and world knowledge. It is weak where the host gives it no controls: no mask, no seed, no field for things to exclude, no transparent background (11.6). An edit redraws the whole frame, so things you did not mention can move. Small text and fine detail may not come out right, as Google says itself.

When to reach for something else.

You need

Use

Why

A typographic plate, a layout-led graphic, a cut-out on a transparent background

GPT Image 2.5 (Chapter 12)

it is the only route here with a transparent background, and its maker documents reliable text rendering; never ask one model for a person and a headline

A fast look at an idea, or a bounded edit of a frame with no signed face

GPT Image 2.5 (Chapter 12)

its maker positions it for precise editing and speed

The same photographic frame, when two attempts here miss the same way

the same brief on GPT Image 2.5 (Chapter 12)

a cheap tactic, not a claim that it is better

A brand mark, a wordmark, a legal line

the real artwork, placed in post (Chapter 30)

generated brand text cannot be trusted

A real product's exact label

the photograph of the product as a reference (Chapter 10), and the artwork in post

the real thing is the authority

Anything that moves

the video models (Chapters 16 to 19)

this model makes stills only

More pixels than 4k

an upscaler (Chapter 32), and only if the delivery needs it

this model stops at 4k; an upscaler changes size, not composition

Nano Banana Pro is the default for a photographic keyframe and for a signed identity carried through several references.

How it is reached. There are two ways. Through Higgsfield (Chapter 33), which hosts the model under one account and one price list in credits; the book works this way. Or directly at Google: the Gemini app, Google AI Studio, the Gemini API and Vertex AI, billed in dollars. Google's chat and API keep the conversation across edits. A Higgsfield job starts from nothing, so a fact that matters goes in the prompt or in a picture attached to that job. On Higgsfield, always choose the model by name: a request that names no model runs GPT Image 2.5 (checked 27 Sep 2026).

11.2 How it reads what you give it

Treat Nano Banana Pro as a sharp stills photographer who reads your brief once, has no memory, and takes everything you hand it literally.

What it accepts.

Input

Accepted

What it owns

Text

yes; no required syntax and no word limit in Google's guidance; the input cap is 65,536 tokens

the change, the frame, and every fact the pictures do not carry

Pictures

yes; up to 14 in all. Google's developer table divides them into 6 objects, 5 characters and 3 style references, and its other pages state the split differently, so 5 is the safe limit for people (checked 30 Sep 2026)

everything they show, unless the words fence them

Video, audio

no; Google lists audio as unsupported, and video input only on its Flash image models

nothing

Search

the model supports grounding in Google Search; Higgsfield does not expose it

nothing on Higgsfield: a prompt cannot ask the model to look something up

Each input picture costs 560 tokens against the cap, so the cap is never the limit in practice.

A picture owns everything it shows. Attach a photograph of an actor and you receive the face, and also the room, the light, the pose and the clothes, unless the prompt says otherwise. The fence is written in the prompt, in words (11.3). This is the most common way a reference goes wrong, and the first fix is always a sentence, not a new picture.

Three words are used with one meaning each in this chapter and in Chapter 10.

  • Master: the first, neutral picture of a character, a place or a product: a photograph of the actor, or a generated face on a plain grey ground in even light. It is the authority on identity, and it is first-generation: made directly for its job, never derived from another output.

  • Anchor: any approved picture a later job is built on and checked against: a master, an in-look anchor, a place plate.

  • In-look anchor: the character approved in the film's own light, made from the master with the master attached.

How pictures are bound. On Higgsfield there is one slot for pictures, and the attachment order gives "Image 1", "Image 2" and so on. The role of each picture lives in the prompt, not in the slot. Google's own formula puts the reference pictures first, then the instruction that relates them, then the new scenario, and advises a distinct name for each character or object. So the prompt opens with a manifest: one sentence for each picture, in attachment order, naming what to take and what to take nothing of (11.3). Attach an earlier generation from the platform's own library rather than downloading and re-uploading it: an upload is stored as a resized JPEG, and an earlier job passes as a PNG (checked 27 Sep 2026).

Words and pictures must agree. When a prompt says one thing and a picture shows another, expect a result you did not choose, and do not assume which side wins. Make them agree. If the picture shows a grey T-shirt, do not ask for a collared shirt over it without saying that the shirt changes.

It thinks before it draws. Thinking is on by default and cannot be switched off. Google says the model may draw up to two interim images to test composition and logic, and the last of them is the final render. You cannot see this or control it. The consequence for a director is that a complete prompt is welcome, because it settles decisions the model would otherwise make itself.

It paints what the words say is there. A state mentioned once, and as a negative, may not hold: a ceiling light described as "off" can come back lit. Name what is there ("the only light in the room is the low warm bulb"). Google gives the same advice in general: describe the scene you want, and use a positive description in place of a negative one.

Text and language. Put the exact words to be rendered in quotation marks, as Google advises, and describe the lettering. Google lists ar-EG (Egyptian Arabic) among the languages in which the model performs best. Prompts in this book are written in English; an Arabic string inside a prompt is copied letter for letter from the approved text. Small text may not come out right, and Arabic in a frame is checked glyph by glyph (11.7).

11.3 The prompt, in this tool's dialect

Chapter 9 teaches the twelve elements, their order and their reasons. Nano Banana Pro takes them in that order, written as one paragraph of plain prose with no labels, bullets or headings. This table maps each element to what it does here.

#

Element

How Nano Banana Pro reads it, and how to write it here

1

Operation, medium, purpose

Opens the prompt. It colours every noun after it. "A digital-cinema frame from a premium athletic-wear commercial." The tool fixes the operation, so no "create" is needed

2

Subject and frozen instant

Give the exact number of people, the share of frame height, and the phase of every moving system as positions and participles with a direction readable in them

3

Identity clause

Six to ten fixed facts, asymmetries first. With pictures attached it follows the manifest and is pasted, never retyped

4

Wardrobe and props

Material, wear and state of everything you will check on the returned frame

5

Frame inventory and edges

Foreground, midground, background, and what touches or leaves each edge. It leaves the model nothing to invent

6

Environment and materials

Surfaces and systems, not mood. Text that exists in the world is production design: its language, script, wear and distance

7

Must-show anchors

The continuity items this frame needs. If the story needs it visible, name it

8

Camera placement

Height, distance, angle, lens, stop and framing, each with its result; a cue for the look, not a simulation

9

Light

Each source, where it sits, hard or soft, its Kelvin, what it does to a named surface; one contrast ratio

10

Atmosphere

Only where the air would show

11

Finish

The capture: stock or sensor character, grain by zone, base tone, halation, edge softness. Never damage. The colour goes on named surfaces

12

Exclusions, closing line

At most five, aimed at drifts you expect, then "no overlaid text, subtitles, logos, or interface graphics"

Syntax. The prompt is plain English. Kelvin is written "approximately 5600 Kelvin" or "5600K", a ratio "3:1", a stop "T2.8", a lens "75mm". Exact text goes in double quotation marks. An anchored prompt has short preamble paragraphs in front of the frame; the frame itself stays one paragraph. Keyword lists and JSON show no advantage on this model.

Aspect and resolution never go in the words. They are settings (11.6). Written in the prompt, they are at best ignored and at worst fight the real setting, and the frame comes back in the wrong shape.

What it responds to.

  • Specific matter: "damp-stained concrete, a hairline crack running from the floor joint", not "a gritty tunnel".

  • Positions and participles: "the index finger lifted a finger's width and the middle finger just touching down".

  • Asymmetries, skin facts, and light with a source and a named surface it falls on.

  • A result beside every number.

What it ignores or misreads.

  • Quality stacks ("8K, masterpiece") describe nothing.

  • "Cinematic" alone tends to bring a teal-and-orange grade. Say the format, the lens or the grade you mean.

  • Emotion words ("distracted", "tense") give the model a state to guess. Write the muscle, the light and the matter (Chapter 9).

  • Camera-body and lens-brand names, and the name of any director, photographer or painter used as a look, are style stickers, not decisions (Chapter 5). They stay out of Nano Banana Pro prompts entirely, although Google's own style-transfer example names a painter. The house amendment that lets a director's name stand as a technique token in Seedance 2.5 motion prompts (Chapter 17) does not reach stills: here it stays banned, even as the name of an editing style.

  • A stack of negatives hands the model the ideas you wanted out. Describe the world positively, then add up to five exclusions.

Check your camera numbers before you run

A number is a cue for the look, and it should not contradict the frame you wrote beside it. Use one formula: the height of the scene the lens sees at the subject, in millimetres, is about the distance in millimetres times the sensor height, divided by the focal length. On a full-frame equivalent the sensor height is 24 mm for 3:2, 20.25 mm for 16:9 and 15.4 mm for 21:9.

  1. Pick the lens, the distance and the aspect.

  2. Work out the height of the scene at the subject. 75mm at 15 metres on 16:9 sees 15,000 × 20.25 ÷ 75 = about 4,050 mm.

  3. Divide the subject's height by it. A sprinter 1.7 metres tall fills about 42 % of the frame, so "roughly 40 percent of the frame height" is honest.

  4. If the share you wanted is wrong, move the distance, not the words.

The frozen instant. Chapter 9 teaches it. On this tool the rule is practical: the model draws one instant, so an open verb ("runs", "pours") gives a pose with no direction. Freeze the action 20 to 40 per cent in, and write the evidence of where it is heading: grit hanging behind a spike that has just left the ground, a cube caught above the surface of a pour, a fingertip with a wet trail behind it.

Length. Neither maker sets a limit. The working bands are these, and they are declared for each shot in the shot record.

Kind of frame

Words

A complex hero still that settles all twelve elements

about 400 to 600

An anchored hero: the manifest and the clause add about 70 to 130 words before the frame

about 500 to 730

A single-subject insert, or an environment

about 150 to 300

A one-change edit of an approved frame

about 40 to 150

A test frame

about 60 to 200

Below the band, look for a missing element. Above it, look for repetition and cut it.

The anchored order. A frame made from pictures is written in four parts, in this order, and Chapter 10 teaches the reasons.

  1. The manifest: "Image N is [name]: take [what]; take nothing of its [what]".

  2. The locked identity clause, pasted from the character's record: six to ten facts, asymmetries first, the name, the wardrobe now, one small unique mark.

  3. The change contract, only where the frame differs from a picture: "Now: …. What still holds: …".

  4. The frame: the twelve elements, opening with the medium.

The master wins any disagreement between the clause and the picture. Write the clause from the master, at 200 %, and change it only when the master changes.

Exact text. Text that exists in the world (a shop sign, a street name) is generated inside quotation marks, with its language, script and wear, and checked letter by letter. Brand names, product names and legal lines are never generated: the frame carries a clean panel, and the real artwork is composited in post (Chapter 30). Arabic in a frame must match the approved string exactly; a final ى where ي was asked fails. The closing line bans the editorial layer only, so a frame that carries a shop sign never gets a "no text" exclusion.

Keyframes. A frame made to be animated also states the mouth (lips visible, corners clear and lit, neutral or barely parted) if the person will speak, and the room left in the frame for the move (Chapter 9).

11.4 Templates

Four fill-in blocks. Replace each bracket with words a camera could see; delete the bracket labels, which are there to be read, not sent. The finished-grade template is the standard. The test template is for proving a route or a look cheaply and is never the model of a finished frame.

Written for this book · template; not run · finished grade, a frame written from words alone

Template — the finished frame

[1 Medium and purpose: a digital-cinema frame / a photograph, from a (kind of film or campaign), captured how.] [2 Subject and frozen instant: how many, who or what, the phase of every moving system as positions and participles 20 to 40 percent into the action with the evidence of its direction; the subject's share of the frame height and where it sits.] [3 Identity clause, if a person: six to ten fixed facts, asymmetries first, skin facts last.] [4 Wardrobe and props: material, wear and state of everything you will check.] [5 Frame inventory: foreground, midground, background, and what touches or leaves each edge.] [6 Environment: architecture, surfaces, fittings, deliberate wear; any world text with its language, script and wear.] [7 Must-show anchors: the items this frame must show.] [8 Camera: height, distance, angle, lens and stop, each followed by its result.] [9 Light: each source, where it sits, hard or soft, its Kelvin, what it does to a named surface; one contrast ratio.] [10 Atmosphere, only where the air shows.] [11 Finish: the capture, grain by zone, base tone, halation, edge softness, what stays sharp; the color on named surfaces.] [12 Up to five exclusions, then:] no overlaid text, subtitles, logos, or interface graphics.

Written for this book · template; not run · finished grade, a frame made from pictures, four parts in order

Template — the anchored frame

Image 1 is [name]'s master, [what it is]: take [what]; take nothing of its [ground, light, clothes, pose]. Image 2 is [name] in [state], [what it is]: take [what] only; take nothing of its [ground, light, layout], and no face, because [name]'s face is Image 1's alone.

[Name]: [six to ten fixed facts pasted from the record, asymmetries first, one small unique mark]. Wardrobe now: [state].

Now: [what differs from the pictures]. What still holds: [what is protected].

[The frame: elements 1 to 12 as one paragraph, without the face, which the pictures and the clause carry.]

Written for this book · template; not run · a one-change edit of an approved frame

Template — the one-change edit

Image 1 is the approved frame: keep everything in it exactly — [the instant], [the framing and camera position], [the person's face, expression and hair], [both hands and where they rest], [the light and where it falls], [the room and surfaces], the grain — and change only this one thing: [the change]. [Its physical consequence: how light falls on it; what is revealed where something left.] Image 2 is [name]'s master: [face, hair, skin] stay exactly as in Image 2; take nothing of its framing, light or clothes. Nothing else moves or changes. No overlaid text, subtitles, logos, or interface graphics.

Written for this book · template; not run · test grade, a route or a look proved cheaply

Template — the test frame

[Medium and purpose]: [subject, one action, one place], [one light source and its side], [lens feel]. [Two things that must be visible.] No overlaid text, subtitles, logos, or interface graphics.

A test frame proves a direction in one job. It earns a rewrite at the finished grade before anything is signed from it.

11.5 Examples at the standard

Seven frames follow. The first five are finished-grade stills, each written to the twelve elements in order. The sixth is one whole anchored frame, built from an actor's photograph. The seventh is a one-change edit of an approved frame. The products are generic and unbranded on purpose: with a client's real product, attach photographs of it, write the frame in the anchored order of Example 11.6, and set the brand marks in post. None of these prompts has been run.

Written for this book · beauty and personal care · a serum close-up in a Cairo bathroom · Nano Banana Pro, text only, 4:5, 2k, one variant · not run · 606 words

Example 11.1 — a serum, a third of the way through the stroke

A digital-cinema close-up from a premium skincare commercial, captured on a large-format sensor, showing one Egyptian woman in her early thirties at the instant her fingertips are a third of the way through smoothing a serum across her right cheekbone: the drop has landed beside the nose, the right index and middle fingertips are gliding toward the ear and leaving a glossy wet trail about two centimeters long behind them, the wrist bent and the elbow outside the frame, her eyes lowered to a point just below the lens, the brow smooth and the mouth closed and relaxed. Her face is turned about twenty degrees toward camera-right, presenting the right cheek to the lens and the window, and from chin to hairline it fills about seventy percent of the frame height. She has the left eyebrow a few millimeters higher than the right, the left upper lid fractionally heavier over dark brown eyes, a small dark mole on the right cheekbone, full lips with the lower lip slightly fuller on the left, a straight nose with a soft rounded tip, a narrow oval face with a slightly wide forehead, long straight dark brown hair parted on the left and tucked behind the right ear, and warm light-olive skin with visible pores across the nose and cheeks and a faint natural sheen on the forehead. She wears a plain ivory cotton robe with a soft shawl collar and no jewelry. In her left hand, held at chin height at the right edge of the frame, is a small frosted-glass dropper bottle with a matte white cap and a blank pale label panel facing the camera. Nothing crosses the lens in the foreground; in the midground are her face, the two fingertips and the bottle, her hair cut by the top edge and her right wrist leaving through the lower left edge; behind her a hand-plastered ivory wall and the edge of a white marble vanity fall soft. The serum trail must read as a wet line catching the window, the mole must be visible, and the bottle's frosted body and cap must be readable. The camera sits level with her cheekbones, 0.9 meters from her face, aimed straight at it, on an 85mm spherical lens at T2.8, which holds her near eye and the serum trail sharp while her far ear, the hair behind it and the room fall soft. Soft morning daylight from a frosted window at camera-left, about one meter from her face, at approximately 5600 Kelvin wraps the right cheekbone and gives the wet trail a bright line of highlight; a white plaster wall at camera-right returns a weak soft fill at the same 5600 Kelvin so the shadow side stays open; a small warm wall lamp at approximately 3200 Kelvin, out of focus at the upper right, adds an amber disc; overall contrast sits near 2.5:1. A trace of warm humidity from a recent shower softens the edge of a mirror at the upper right. The finish is that of a large-format digital cinema capture: a neutral base tone, soft highlight roll-off on the cheekbone that never clips, fine photographic grain, finer in the highlights and coarser in the shadows, gentle halation around the lamp and the window's edge, and slight softness toward the frame edges while her eyes and the serum trail stay sharp. The color sits on warm olive skin, ivory plaster and the pale frost of the glass, with one amber accent from the lamp. No airbrushed skin, no glitter or shimmer particles in the serum, no extra fingers, no second person, no overlaid text, subtitles, logos, or interface graphics.

Why it works

  • Copy the frozen instant. The fingertips sit a third of the way along the stroke and the wet trail behind them shows which way they are going. The brow and mouth are written as muscle, not as a mood.

  • Copy the clause. Eight facts, asymmetries first (brow, lid, mole, lip), then the skin facts that stop the plastic look. Written once, it is pasted into every frame of her.

  • Copy the numbers with their results. 85mm at 0.9 meters on a 4:5 frame sees a field about 32 centimeters high, which is why the face fills 70 percent. T2.8 is followed by what it keeps sharp and what falls soft.

  • A product with no marks. The label panel is blank because brand marks are set in post. With the client's real bottle, attach its photographs and use Example 11.6's order.

Written for this book · food and drink · an iced hibiscus tea, poured in a Downtown Cairo café · Nano Banana Pro, text only, 16:9, 2k, one variant · not run · 599 words

Example 11.2 — the pour, a third of the way done

A digital-cinema hero frame from a premium beverage commercial, captured on a large-format sensor with a natural, restrained grade, showing one tall ribbed glass of deep ruby hibiscus tea over ice at the instant the pour is a third finished: a brushed-steel jug enters from the top left with its spout tipped down, and one unbroken ribbon of ruby liquid falls about twelve centimeters from the spout into the glass, which is a third full; three ice cubes turn in the liquid and a fourth is caught just above the surface with a low crown of red liquid rising about two centimeters where the stream lands, a trail of small bubbles running down from the impact. The pouring hand is outside the frame above; no person appears. The glass stands right of center and occupies roughly 55 percent of the frame height. The glass is clear and ribbed with a thick rim and fine beads of condensation on its outer wall. On a small round brass tray beside it sit a lemon wedge with a pale pith edge, a sprig of fresh mint with three leaves, and a small brass bowl of dried hibiscus calyces, dark red and crinkled, with a folded white linen napkin at the left. In the foreground the near edge of the marble table is sharp at the bottom of the frame with a faint ring of moisture; in the midground are the glass, the stream and the tray, the tray cut by the right edge and the jug's handle leaving through the top edge; in the background an old Downtown Cairo café falls soft: a tall arched window with green shutters half open burns at the upper left, dark bentwood chairs stand beyond, and a mirror with an aged frame catches a blur of light. The tabletop is cream marble with grey veining and old ring stains, the floor is worn wood, and no lettering is legible anywhere. The unbroken stream from spout to liquid, the airborne cube with its crown, and the ice seen through the glass wall must all be visible. The camera sits 30 centimeters above the marble, 1.6 meters from the glass, tilted down about four degrees, on a 100mm spherical lens at T2.8, so the liquid's surface reads as a thin ellipse and the glass's front wall, the ice and the stream stay sharp while the café dissolves into soft shapes. Low late-afternoon sun at approximately 4300 Kelvin enters from camera-left behind the glass through the half-open shutter, backlighting the liquid so the tea glows ruby and drawing a bright rim line down the glass's left edge; open-shade bounce from a cream wall at camera-right at approximately 6500 Kelvin lifts the marble's near edge and the front wall of the glass; overall contrast sits near 4:1 with the ice highlights bright and the marble's shadow side falling to a cool dark. A thin breath of cold vapor rises from the ice at the rim, visible only where the backlight crosses it. The finish is that of a large-format digital cinema capture: a neutral base tone, clean highlights on the ice and the rim that roll off without clipping, fine photographic grain, finer in the highlights and coarser in the shadows, gentle halation around the bright window, and slight softness toward the frame edges while the stream and the crown stay sharp. The color sits on ruby-magenta liquid, cream marble and warm brass, with a green accent from the mint. No straw, no exaggerated splash shapes, no visible hand, no legible signage, no overlaid text, subtitles, logos, or interface graphics.

Why it works

  • Copy the pour. An unbroken ribbon, a cube caught above the surface and a crown two centimeters high are three positions, and together they show the direction of the pour. "Pouring" alone would give a pose.

  • Say where the hand is. The hand stays outside the frame and the prompt says so; a jug in the air with no arm invites the model to add one.

  • Copy the light. Backlight makes the liquid glow, bounce lifts the near edge of the marble, and the ratio of 4:1 keeps the ice bright. Each source has a place, a Kelvin figure and a surface.

  • Check the numbers. 100mm at 1.6 meters on 16:9 sees a field about 32 centimeters high, so an 18-centimeter glass fills about 55 percent.

Written for this book · telecom and tech · a home-internet router in a Riyadh living room · Nano Banana Pro, text only, 16:9, 2k, one variant · not run · 613 words

Example 11.3 — the glance at the green ring

A digital-cinema frame from a home-internet commercial, captured on a large-format sensor, showing Tariq, a Saudi man in his early forties, seated on a cream sofa in a Riyadh apartment at the instant he glances up from a tablet toward the router on the console behind him: his head is turned about fifteen degrees toward frame-right, his eyes have lifted from the screen and are a third of the way to the router, the brow relaxed and the right corner of his mouth beginning to lift by about two millimeters, the tablet still tilted in his right hand with its back to the camera and his left forearm resting along the sofa's arm. He sits a third of the way in from the left edge and occupies roughly 55 percent of the frame height. He has the right eyebrow thicker and slightly lower than the left, a prominent straight nose with a slight bump at the bridge, a short black beard trimmed close with grey concentrated at the chin and in a patch on the left cheek, deep-set dark brown eyes with faint crow's feet, a broad face with a high forehead, a receding hairline at the temples with short black hair, and olive-brown skin with visible pores and faint smile lines. He wears a pressed white thobe with a plain collar and two covered buttons. The tablet has a dark grey case. In the foreground the sofa's arm cuts the lower left corner as a soft cream shape; in the midground he sits among cushions with a folded throw beside him; in the background at frame-right a walnut console against a warm off-white plaster wall carries a brass tray and, at its center, the router, while a floor-to-ceiling window with a sheer curtain stands at the left edge. The living room has a cream linen sofa with a hand-woven throw in red, black and cream stripes, a woven wool rug in sand and charcoal, and beyond the sheer a pale haze of city light. The router is a matte white cylinder about 25 centimeters tall, a fine ring of perforations near its base and a single thin ring light near the top, steady green. His glance, the router with its green ring lit, and the back of the tablet must all be visible. The camera sits at 1.1 meters, level, 3.5 meters from the man, on a 35mm spherical lens at T4, so that the man and the router 1.1 meters behind him both read clearly while the window and the far wall soften. Afternoon daylight through the sheer curtain from camera-left at approximately 5600 Kelvin falls across the right side of his face and the sofa; a floor lamp behind the sofa at approximately 3200 Kelvin lays a low amber rim on his shoulder and the throw; the router's ring light gives off a small green glow that pools about fifteen centimeters wide on the console top; overall contrast sits near 3:1. Fine dust hangs in the curtain's light beam near the window and is visible only there. The finish is that of a large-format digital capture: a neutral base tone, highlights on the thobe that roll off without clipping, fine photographic grain, finer in the highlights and coarser in the shadows, gentle halation around the sheer and the ring light, and slight softness toward the frame edges while his face and the router stay sharp. The color sits on cream, warm walnut and olive skin, with the router's green as the single accent. No visible cables, no glow from the tablet on his face, no second person, no wall art with lettering, no overlaid text, subtitles, logos, or interface graphics.

Why it works

  • Copy the smile as millimeters. The eyes a third of the way to the router and the mouth corner lifting by about two millimeters give the glance a direction and a size, where "pleased" would give the model a mood to act.

  • Write the device as an object. Shape, height, ring light and its color are all things to check on the returned frame. The ring's color is a must-show anchor, and Example 11.7 changes exactly that.

  • Choose the stop for the depth you need. The router stands 1.1 meters behind him, so T4 keeps both readable; T2.8 would soften it.

  • Check the numbers. 35mm at 3.5 meters on 16:9 sees a scene about 2.0 meters high, and a seated man of 1.1 meters fills 55 percent. A name for the character gives later frames and the edit something to bind.

Written for this book · automotive and luxury · a luxury saloon on a canyon road at dusk, a sandstone landscape in northwest Saudi Arabia · Nano Banana Pro, text only, 21:9, 2k, one variant · not run · 588 words

Example 11.4 — the saloon, a third of the way through the bend

A wide digital-cinema frame from a luxury saloon commercial, captured on a large-format sensor with a restrained professional grade, showing one long four-door luxury saloon in deep emerald metallic paint as it rounds a left-hand bend on a sandstone canyon road at dusk, caught a third of the way through the bend: the nose is pointed toward the left of the frame, the front wheels are steered about eight degrees toward the camera, the body has rolled slightly onto its outer suspension, and a low ribbon of pale dust lifted about twenty centimeters behind the left rear wheel hangs in the air. The car is seen from the front three-quarter, sits right of center with the open road ahead of its nose filling the left third, and occupies roughly 60 percent of the frame width. The car is alone in the frame. It has a long bonnet, a low tapering roofline, a tall upright grille of fine vertical slats in dark chrome with no badge, slim LED headlamps with a thin horizontal daytime strip, a strong shoulder line running from the headlamp to the tail, flush door handles, 21-inch multi-spoke alloy wheels in dark graphite with a bronze brake caliper visible behind the spokes, and a blank white number plate with no characters. In the foreground dark asphalt with a scatter of sand grains falls soft; in the midground is the car; in the background sandstone walls rise on both sides, the right wall cutting the right edge in deep shade, the road curves away out of the frame at the left, and a strip of deep blue-violet sky lies along the top. The canyon has ochre and rust rock carved into layered strata by wind, smooth dark asphalt with a faint worn centerline, fine sand drifting at the road's edges and a few dry desert shrubs. The grille, the headlamps with their daytime strips, the shoulder line and a wheel with its caliper must all be visible. The camera sits 0.5 meters above the road, 9.5 meters from the car's front wheel, angled up about two degrees, on a 50mm spherical lens at T2.8, which keeps the whole car sharp from the nose to the rear door while the canyon walls soften, and the low angle gives the car a planted, heavy stance. The low sun at approximately 3600 Kelvin sits near the horizon behind camera-left and rakes along the car's near flank, so the shoulder line and the doors carry one long warm highlight; the open dusk sky at approximately 7500 Kelvin fills the bonnet and roof with cool reflected blue; the two thin daytime strips at approximately 6000 Kelvin draw fine white lines in the headlamps; overall contrast sits near 4:1 with the rock face on the right falling into deep shade. Fine dust hangs in the low sunlight above the road between the car and the canyon wall, a warm haze visible only where the sun crosses it. The finish is that of a large-format digital cinema capture: a neutral base tone, deep shadows and clean warm highlights that roll off without clipping, fine photographic grain, finer in the highlights and coarser in the shadows, gentle halation on the sun-struck shoulder line, and slight softness toward the frame edges while the car stays sharp. The color sits on emerald paint, ochre rock and blue-violet sky, with a bronze accent at the calipers. No motion blur, no lens flare across the bodywork, no other vehicles, no people, no overlaid text, subtitles, logos, or interface graphics.

Why it works

  • A moving car is a still. The wheel angle, the body roll and a ribbon of dust at a stated height fix the direction without motion blur, which is why "no motion blur" is one of the exclusions.

  • A text-only car is the model's car. For a client's car, attach photographs of the real one (front three-quarter, side, a wheel, the rear), each with one job, and set the badge and the plate in post.

  • Copy the warm and cool split. Sun at 3600 Kelvin on the flank, sky at 7500 Kelvin on the bonnet and roof, and a ratio of 4:1, each source on a named surface.

  • Check the numbers. 50mm at 9.5 meters on 21:9 sees a scene 6.8 meters wide, so a car about 4.2 meters wide in three-quarter view fills about 60 percent.

From Mastorna, corrected · film trailer · shot 1.1, the establishing frame, a 1960s cabin in black and white · Nano Banana Pro, text only, made at 16:9 and cropped to 1.85:1, 2k, one variant · corrected for this book; not run · 644 words

Example 11.5 — the establishing frame of a 1966 trailer, in corrected form

A black-and-white photograph from a 1966 Italian film, capturing a lean Italian man in his early forties modeled on Marcello Mastroianni circa 1965 seated in a 1960s commercial airliner window seat, his body slightly reclined, the brow smooth and the mouth slack, his eyes caught mid-drift toward the oval aircraft window on his right, the head turned only a few degrees that way, where pale clouds are visible outside. He has dark brown almost-black thick slightly wavy hair swept back from a high forehead with a receding hairline at the temples, parted roughly on the left with a few strands loosened across the forehead, prominent angular cheekbones, a straight nose with a barely perceptible leftward deviation at the bridge, deep-set dark brown slightly hooded eyes with the left upper lid fractionally heavier than the right, visible under-eye darkness, a network of fine expression lines around the eyes and mouth, clean-shaven jaw with a slight asymmetry where the left jaw is marginally stronger, Mediterranean olive skin rendered in warm mid-grey tones showing visible pores on the nose and cheeks and a faint natural shine on the forehead. He wears a dark wool single-breasted suit jacket over a white dress shirt with a loosely knotted dark tie, the collar slightly wilted from travel. His left hand lies on the armrest beside a large dark hard-shell cello case propped upright in the adjacent seat, the index finger lifted a finger's width and the middle finger just touching down, caught mid-tap, the case showing scuffed corners and tarnished metal latches from years of use, its curved body following the shape of the instrument within. A dark heavy wool overcoat is folded on the seat beside the cello case. Behind him in the dim cabin, two or three other passengers are visible: one sleeps with a cloth sleeping mask over their eyes, the chest caught a little raised at the top of an in-breath, another reads a folded newspaper. The cabin has fabric-upholstered seats in neutral tones, open overhead luggage racks without doors, built-in armrest ashtrays, and curtain dividers between sections. At the far end of the cabin a small projection screen casts a faint bluish-white flicker onto nearby faces. In the foreground the blurred edge of an aisle seat back cuts the lower-left corner and the top of the cello case's shoulder enters beside his hand, while the curved frame of the oval window cuts the right edge. The camera is positioned in the aisle at seated eye level approximately one hundred and ten centimeters from the cabin floor, one and a half meters from the subject, held level, using a 40mm lens at T2.8 that holds his face sharp while the passengers and the screen behind him fall soft, with the man framed slightly right of center and cut at mid-chest by the lower edge, his head and neck occupying roughly forty percent of the frame height. Overhead cabin lights at approximately 3200 Kelvin warm tungsten provide the key light from above-left, while cooler daylight at approximately 5600 Kelvin enters through the oval window from camera-right creating a soft cross-light on his face, with a contrast ratio of approximately 3:1. Fine atmospheric haze of pressurized cabin air is faintly visible in the light from the window. The image has the tonal quality of a 1960s black-and-white release print, with fine organic silver-gelatin grain that is slightly finer in the brighter highlight areas and marginally coarser in the shadow areas, a faint warm ivory base tone inherent to the film stock, rich mid-tones dominating the cabin interior, natural halation blooming softly around the bright window, and slight optical softness at the edges of the frame characteristic of period spherical lenses while maintaining sharp focus at center. No color, no modern aircraft interior details, no digital sharpness, no plastic overhead bins, no overlaid text, subtitles, logos, or interface graphics.

Why it works

  • The order is the twelve-element order, unchanged. The medium and the frozen instant open the paragraph, the identity clause follows at once, and the wardrobe, props, cabin, camera, light, haze, finish and exclusions follow.

  • Every system the record freezes is now a position. The fingers are mid-tap, the sleeper's chest is at the top of a breath, and the eyes are mid-drift toward the window; "distracted" is replaced by "the brow smooth and the mouth slack".

  • Copy the clause method. The face is written from asymmetries (the heavier left upper lid, the leftward deviation of the nose, the stronger left jaw), so it can be pasted into every later frame of him. The real name is a rights decision for the whole project (Chapter 34); the clause carries the face without it (11.9).

  • Two elements are completed, and the length is earned. The frame has its edges (the seat back, the oval window, the cello case, the cut at mid-chest), and its camera has an angle and the result of the stop. The finish names a stock that existed, and the prompt ends on four exclusions and the closing line. At about 640 words it sits above the usual band because the cabin holds three frozen systems and a full clause.

Written for this book · sport and apparel, a colour commercial built from an actor's photograph · Nano Banana Pro, two references in order (her approved photograph, then the race-kit sheet), 16:9, 4k, one variant · not run · 638 words

Example 11.6 — one whole anchored prompt: manifest, clause, change, frame

The four paragraphs are the four parts, in order. Copy the block whole.

Image 1 is Salma's master, the approved neutral photograph: take her face, hair and skin; take nothing of its ground, light, clothes or pose. Image 2 is Salma in the race kit, the approved state sheet: take the kit only, its cut, colors, bib and spikes; take nothing of its ground, light or panels, and no face, because her face is Image 1's alone.

Salma: the right brow fractionally higher than the left; a small healed scar above the left eyebrow; a straight nose, subtly flattened at the bridge; full lips with a faint asymmetry favoring the right corner; a narrow face with high cheekbones; dark brown eyes set slightly deep; black hair in a tight low bun with fine flyaway strands at the hairline; true skin texture. Wardrobe now: the race kit.

Now: she is mid-stride out of a stadium tunnel in morning sun, not at rest on a plain ground as in Image 1, and she wears the kit of Image 2, not Image 1's clothes. What still holds: her face, hair and skin, as written above.

A digital-cinema frame from a premium athletic-wear commercial, large-format sensor look with shallow-focus falloff and restrained professional grade, showing Salma, an Egyptian sprinter in her mid-twenties, captured at the exact instant her second stride leaves a stadium service tunnel into open morning light, the drive already roughly a third into its motion so her rear foot has just broken contact with a visible scatter of grit hanging mid-air behind the spikes, her torso leaning well past her base of support with weight fully committed onto the bent front leg, both arms mid-swing in opposition; she occupies roughly 40 percent of the frame height, framed left of center and carrying the image's visual weight against the tunnel's dark mass. Her face shows the facial musculature of full exertion: jaw set, nostrils flared, eyes fixed on a point beyond the frame. She wears a deep-teal performance top and charcoal shorts with matte technical fabric that creases and pulls realistically at the hip and shoulder, a fabric race bib slightly curled at one corner, and worn track spikes with fine abrasion at the toe box. In the foreground the tunnel's concrete edge cuts the lower left corner as a soft dark shape; in the midground she breaks the threshold between shadow and light; in the background the empty stadium curve sits defocused, its rows reading as texture with no legible signage. The tunnel interior shows damp-stained concrete, a hairline crack running from the floor joint, and a single caged work lamp; the track surface outside is worn synthetic red with faded lane paint and fine surface grit. The camera sits at 0.9 meters, 15 meters from the subject, angled slightly upward, on a 75mm spherical lens at T2.0, which holds her sharp inside a depth of about five meters while the stadium dissolves into soft shapes. Morning sun from camera right at approximately 5600K rims her arm and cheek; the tunnel's caged lamp behind her burns warm at approximately 3200K as a motivated practical; overall contrast sits near 4:1 with hard-edged shadows on the tunnel floor and soft fill from the sky. Fine grit and dust hang in the sunlit air at the threshold, visible only inside the beam. The finish is that of a large-format digital cinema capture: a neutral base tone, deep shadows and clean warm highlights that roll off without clipping, fine photographic grain, finer in the highlights and coarser in the shadows, gentle halation on the sun rim, and slight softness toward the frame edges while the center stays sharp. The color sits on charcoal, warm skin and track red, with a teal accent. No motion blur, no lens flare across the face, no other athletes, no modern LED signage, no overlaid text, subtitles, logos, or interface graphics.

Why it works

  • Copy the order and the fences. Each picture has one job and one fence, the clause has eight facts with the asymmetries first and one small unique mark, and the contract says what differs from Image 1 and what still holds.

  • Read the clause off the photograph. View it at 200 %; write the asymmetries first, each as the subject's own side; then the hairline, the set of the eyes and the skin; write only what the photograph shows. Store the clause with a version and paste it into every frame of her. "Salma" is the character's name, not the actress's, whose name stays out of the prompt (11.9).

  • Keep the frame complete. The frame paragraph carries its own camera, light and finish; her face is out of it, because Image 1 and the clause carry it. The numbers hold: 75mm at 15 meters on 16:9 sees a scene 4.05 meters high, and she fills about 42 percent of it.

Written for this book · telecom and tech · a one-change edit of the approved frame of Example 11.3 · Nano Banana Pro, the approved frame first, Tariq's master second, 2k, 16:9, one variant · not run · 152 words

Example 11.7 — the ring light turned from green to white

Image 1 is the approved frame: keep everything in it exactly — the instant, the framing and the camera position, Tariq's face, expression and hair, the white thobe, the tablet held in his right hand with its back to the camera, the sofa, the throw and the rug, the sheer curtain and its light, the walnut console with the brass tray, the router's shape and position, the grain — and change only this one thing: the router's ring light is steady white, not green. The small glow it casts on the console top and on the wall behind is white too, in the same small pool, and nothing else in the room takes its color. Image 2 is Tariq's master: his face, hair and skin stay exactly as in Image 2; take nothing of its framing, light or clothes. Nothing else moves or changes. No overlaid text, subtitles, logos, or interface graphics.

Why it works

  • Order and job. The approved frame is first and owns everything except the change; the master is second and owns the face only. There is no mask on this tool, so the words fence what stays.

  • Name the change and its consequence. The ring light's color, and the color of the pool it casts, are one change; a change with no stated consequence leaves the model to decide what else follows.

  • Repeat the preserve list on every attempt. The list is the definition of "everything", and the drift a redraw brings is found by comparing against it.

  • If the result drifts, go back to the approved frame, never edit the edit. The router is small and the camera has not moved, so the changed area can be composited onto the approved frame in post.

11.6 Settings

Everything in this section is chosen in the interface, never written in the prompt.

Setting

Values on Higgsfield (checked 27 Sep 2026)

Use

Why

Model

nano_banana_pro

name it every time

a request with no model runs GPT Image 2.5 (checked 27 Sep 2026)

Resolution

1k, 2k, 4k; the default is 2k

2k as the baseline

1k costs the same as 2k, so there is no cheaper draft tier on this model (checked 27 Sep 2026)

Aspect

ten ratios, Figure 11.1; no automatic choice

set it on every job, from the delivery aspect

nothing is chosen for you

Variants

1 to 4

1 for a finished prompt

each variant is a job and a charge

Pictures

one list; the order is Image 1, Image 2

two to five, each with one job

the role lives in the prompt

Resolution. A 2k 16:9 frame returns 2752 × 1536 pixels, wider than a 1920 timeline. Choose 4k when the use needs the pixels: a planned crop or push-in, a large screen or print, a dense sheet, a small detail of a face or a product. A bigger file does not make a better face, so inspect the returned detail and do not assume it.

fig11-1

Figure 11.1 — The ten aspect ratios at 2k, drawn to scale (sizes from Google's table, checked 30 Sep 2026). 1k is half the width and height of each, and 4k is twice. There is no automatic choice, no 1.85:1 and no 2.39:1; those are cropped from 16:9 and 21:9.

Aspect. Choose the delivery aspect first, and make each format as its own frame. The two cinema shapes are crops: 1.85:1 from a 2k 16:9 frame keeps 2752 × 1488 pixels, and 2.39:1 from a 2k 21:9 frame keeps 3168 × 1326. Never crop a vertical from a wide frame. A crop loses the second person, the receiving hand or the path of an object. Recompose the vertical as its own frame, and name the planes again for the tall shape (Chapter 9). To widen a frame, do not stretch it: make the wider frame as its own job with the parent attached as a picture, and write the new inventory of planes.

Variants. Run one variant of a finished prompt. To compare two looks, change one thing between jobs so that you can tell what made the difference. Several variants of a test frame are fine for choosing a look.

Pictures. Attach them in this order: identity first, then the state or the body, then the frame or the plate, then the one object the task depends on, then style. When slots run short, the master face and the current state are the last to be cut. Two pictures with the same job get averaged, and every extra picture is another room, light and pose to fence off. A sheet is made from the first-generation master alone, at 2k and 16:9, as one variant (Chapter 10). The working range is two to five. Google's ceiling is 14 in all, of which 5 may be characters. Higgsfield states no maximum. The number of pictures does not change the price: two references cost the same as one (checked 27 Sep 2026).

What the host does not give you.

You would want

Do this instead

A seed, for the same result twice

Keep the approved frame. It is your seed: edit it, never re-roll it

A field for things to exclude

Describe the world positively; write at most five exclusions in the prompt

A mask

The whole frame is redrawn, so fence what stays in words, and composite the changed area onto the parent in post. (The platform lists a mask slot on Nano Banana 2, Google's Flash model, which defaults to 1k; it has not been tried.)

A transparent background

GPT Image 2.5 (Chapter 12)

A thinking level, or search

Not exposed; thinking is always on

Defaults that bite.

Default

Consequence

No model named

GPT Image 2.5 runs, not Nano Banana Pro

Every job starts from nothing

Anything that must carry over is in the prompt or in a picture attached to this job

Uploads are stored as resized JPEG; earlier generations pass as PNG

Attach earlier frames from the platform's library

The job history records nano_banana_2 for a job submitted as Nano Banana Pro, while the bill reads "Nano Banana Pro" (checked 27 Sep 2026)

Log both labels in the run record (Chapter 33)

11.7 Checks before you sign

Chapter 9 teaches how to read a returned frame, at the size it will be seen and then at full size. These are the checks that belong to this tool, in the order to make them.

  1. The file. Open it at 100 %. The size and aspect must be the ones you set: a 2k 16:9 frame is 2752 × 1536. Every output carries Google's invisible SynthID watermark; there is nothing to see and nothing to remove.

  2. The picture at 100 % and 200 %. Eyes, ears, teeth, hands (count the fingers and trace each wrist to a person), closures, jewellery, labels.

  3. The frame against its prompt, element by element. Every frozen system is visible in the phase written. Every light is in the state asked for, and the shadows fall as the ratio says. The depth of field looks like the stop you wrote. The edges you named are cut as named.

  4. A frame made from pictures. Put the face at 200 % beside the master, never beside another frame made from it, because faces drift in small steps and a copy of a copy hides the drift. Look for what leaked from a reference: its ground, its light, its clothes, its pose, a second face.

  5. Stray marks. Faint printing or a logo on an object described as plain, or on a screen. In-frame text is read letter by letter, and an Arabic string is compared with the approved string: a final ى where ي was asked fails, and so does a broken join.

  6. Products. Compare with the real product photographed: proportions, closure, label position.

  7. An edit. Flick the parent and the result at full size and list what moved. Decide whether the cut can live with it.

  8. A composite. Look at the seam at 100 % and match the grain.

  9. A keyframe. The mouth state and the room left for the move (Chapter 9).

  10. A real person. A face that has drifted from the consented photograph is a hard failure, however good the rest is.

Before you sign a Nano Banana Pro frame

  • The size and aspect are the ones set; the job label is logged beside the model asked for.

  • The person reads as themselves at viewing size; at 200 % the face matches the master.

  • Every frozen system, light and named edge is as the prompt wrote it.

  • Nothing leaked from a reference: no second face, no borrowed room, light or clothes.

  • No stray printing, logo or text; every in-frame string is exact.

  • Any product matches the real one; brand and legal text is set in post.

  • After an edit, what moved has been listed and accepted or repaired.

11.8 Failures and fixes

You see

Likely cause

Smallest fix

A light is on that should be off

the state was mentioned once, and as a negative

an edit that says it firmly and names the only light: "the only light in the room is the low warm bulb"

An object you never asked for

a surface was left unnamed

name every surface and what sits on it; describe emptiness as a surface: "the tabletop between the glass and the tray is bare"

Plastic skin

no skin facts

add one to three: pores on the nose and cheeks, a sheen on the forehead, the lines from nose to mouth

The face is acting

an emotion word

write the muscles, or "her expression does not change"

The frame goes teal and orange

"cinematic" or a grade word

delete it; name three real surfaces and their colors

A younger, prettier face than you wrote

the face was written from a type

put the fixed facts early, asymmetries first

The reference's room, light or clothes came along; a second face; a grey ground

the picture owns everything it shows and no fence was written

write the fence ("take nothing of its ground, light, clothes or pose; no face"), or crop the picture to its one job

The face drifts from the photograph over several frames

the clause was retyped, the master was not attached, or an edit was made of an edit

re-attach the master, paste the clause, and start again from the last approved frame

The depth of field or the scale does not match the numbers

numbers are cues, not physics

write the result beside the number; recheck the distance with the formula (11.3); judge the frame

A clip started from this frame begins with a photograph waking up

the instant carried no evidence of direction

write the phase of each moving system as positions

The frame comes back in the wrong shape

an aspect written in the prompt fights the setting

set it in the interface and keep it out of the words

Logos, faint printing or extra text on plain objects

the model adds print to surfaces

describe the surface as plain and unmarked; clean it in post

In-frame text is misspelt, or Arabic letters are wrong

small text and fine detail may not come out right

quotation marks and the exact string; two tries at most; then a blank panel and the real text set in post (Chapter 30)

A product's shape or label is wrong

words cannot hold a product

attach front and side photographs; then composite the photographed product

An edit moved something you did not mention

no mask: the whole frame is redrawn

repeat the preserve list and name the thing that moved, or composite only the changed area onto the parent

A relit frame looks pasted in

Google notes that complex edits, lighting changes among them, can produce unnatural artifacts

a one-change edit that names the new light's source, direction and Kelvin and what it does to each named surface; inspect shadows and reflections

The finish looks like damage: dirt, scratches, heavy noise

the finish block asked for age or damage

describe capture only; add delivery grain and halation in post (Chapter 30)

Captions or watermark-like text appear

the closing line is missing

end on "no overlaid text, subtitles, logos, or interface graphics"

The retry ladder, smallest lever first.

  1. Fix the sentence that failed and run once.

  2. If it fails the same way twice, stop rewriting. Make a one-change edit of the best attempt.

  3. Change the evidence, not the adjectives: a new picture, a photograph of the object, a blocking photo of a hand or a grip.

  4. Composite in post: keep the parent and take only the changed area.

  5. Run the same brief on GPT Image 2.5 (Chapter 12).

  6. Change the method: photograph it, or build it as a plate.

Always return to the last approved frame; never edit an edit that is drifting. To remove a logo or a mark, go straight to post. In-frame Arabic that is wrong twice gets a blank panel and the text set in post.

11.9 Limits, prices and rights

Limits (checked 30 Sep 2026 unless dated otherwise).

Item

Value

Output sizes

1K, 2K and 4K in ten aspects; Figure 11.1 shows 2k. 16:9 is 1376 × 768, 2752 × 1536 and 5504 × 3072; 21:9 is 1584 × 672, 3168 × 1344 and 6336 × 2688

Input and output

input cap 65,536 tokens, output cap 32,768 tokens; each input picture costs 560 tokens

Pictures

14 in all: up to 6 objects, 5 characters and 3 style references (Google); Higgsfield states no maximum

Variants

1 to 4 per request

Prompt length

no word limit; the working bands are in 11.3

Thinking

always on, cannot be disabled

Search grounding

supported by the model, not exposed on Higgsfield

Watermark

SynthID on every output

Language

Google lists English and ar-EG among the languages of best performance

Prices. These are planning figures. Ask for the quote on the day.

  • On Higgsfield (checked 27 Sep 2026): 1k and 2k cost 2 credits, 4k costs 4, at 16:9 with text only; two references at 2k also cost 2. Each variant is charged. A credit is worth about $0.05, so a 2k frame costs about $0.10 and a 4k frame about $0.20. Credits have no cash value and unused credits do not roll over (Chapter 33).

  • At Google (Gemini API paid tier, pricing page updated 24 Sep 2026): $0.134 per 1K or 2K image and $0.24 per 4K image; the batch price is $0.067 and $0.12; an input picture costs $0.0011. Nano Banana Pro has no free tier, and Google does not use paid-tier inputs to improve its products.

Rights and terms.

  • Higgsfield's terms (last updated 26 Jul 2026, read 30 Sep 2026). The company does not claim ownership of your inputs or outputs and does not restrict commercial use of outputs. Your rights in exported outputs survive cancellation, and you may transfer or sublicense them to clients. You confirm that you hold the rights to what you upload, including releases from anyone whose likeness appears. Inputs and outputs may be used to train Higgsfield's models unless you are under an enterprise agreement, which is a decision to take before client material is uploaded (Chapter 34). The terms ask that sensitive personal information, full names included, stays out of text prompts.

  • The provider's policy applies too. When a model comes from a third party, you agree to that provider's acceptable-use policy in addition, and the stricter of the two governs. Google's Generative AI Prohibited Use Policy (last modified 17 Dec 2024) names, among other things, using personal data or biometrics without the consent the law requires.

  • Google's terms (Gemini API Additional Terms, effective 23 Mar 2026). Google does not claim ownership of generated content, and it may generate the same or similar content for others. Higgsfield says likewise that outputs may not be unique across users. A hero frame is therefore not exclusive, and a client who needs an exclusive look is told so.

  • Real people. A living person's face comes from a consented photograph as the master, the consent is in writing, and the name stays out of the prompt. A face that keeps drifting from that photograph is a hard failure. A deceased person's likeness is a rights decision made once for the whole project (Chapter 34).

  • What may ship to a client. A frame that has been signed after 11.7, with brand marks, legal lines and Arabic text either exact in the frame or set in post. A test frame is never shipped. The rights gate of Chapter 8 comes before the frame leaves the building.

11.10 Version notes

What is current (checked 30 Sep 2026). Nano Banana Pro is Gemini 3 Pro Image, model id gemini-3-pro-image: released in November 2025, generally available since 28 May 2026, with November 2025 as the latest update Google lists for it. Google's Vertex AI page lists its retirement date as 28 May 2027 or later, and Google announces no successor on the pages read. Nano Banana 2 (gemini-3.1-flash-image) and Nano Banana 2 Lite (gemini-3.1-flash-lite-image) are newer Flash models with a different profile: they expose a thinking level, accept video input, and are not the Pro model. Google recommends Nano Banana 2 as its general-purpose image model and describes the Pro as the one for professional asset production and complex instructions, which is why this book keeps the Pro for finished frames.

What to re-check when a newer Pro ships.

  • The model id, and whether the platform's Nano Banana Pro is the new model.

  • The settings: whether a mask, a seed or a negative field appears.

  • The reference limits, the language list and the raster sizes in Figure 11.1.

  • The price in credits at 1k, 2k and 4k, and Google's and Higgsfield's terms.

  • The four category frames of 11.5: run each once and hold the result to the checks of 11.7 before the new model replaces this one.

Open questions that change what you do

  • What runs under the platform's "Nano Banana Pro". The job history records the label nano_banana_2 while the bill reads "Nano Banana Pro" (27 Sep 2026). Whether the two are the same model as Google's Pro is not established from outside. Log both labels on every job and ask Higgsfield (Chapter 33).

  • The platform's reference ceiling. Higgsfield publishes none. Plan on five pictures or fewer, and treat six or more as an experiment.

  • How the model reads a camera height and distance in real units. No maker publishes it. Write the visible result beside every number and judge the returned frame.

  • Which sizes the platform returns outside 2k 16:9. That size has been measured (2752 × 1536); the others in Figure 11.1 are Google's table. Read the size of the first file you make at any new setting.

  • Whether Higgsfield's mask on Nano Banana 2 is worth a route. It has not been tried. Until it is, the fence in words and the composite in post are the routes.

What to remember

  1. Nano Banana Pro (Gemini 3 Pro Image) is still Google's current Pro image model as of 30 Sep 2026, and it is the stills camera of the book. Name it on every request.

  2. It is a literal photographer with no memory: everything it needs is in this prompt or in a picture attached to this job.

  3. Write the twelve elements as one paragraph in the order of Chapter 9; put aspect and resolution in the settings and never in the words.

  4. A picture owns everything it shows. A frame made from pictures opens with the manifest, the locked clause and the change contract, and the frame follows.

  5. Numbers are cues for the look: write the result beside each, and check lens, distance and share of frame against each other.

  6. There is no mask and no seed. An approved frame is the seed: change one thing at a time, from the parent, and composite in post where pixels must not move.

  7. Brand marks, legal lines and exact Arabic are composited in post or verified letter by letter; the real product is the authority.

  8. Prices, limits and terms change: quote on the day, and log the job label beside the model you asked for.

Comments


bottom of page