Chapter 10 · Identity, references and the world in the frame
Part 2 · Stills
A finished still is seldom made from words alone. It rests on a face that must stay the same face, a room that must stay the same room, a product that must stay the client's product, a sign that must say exactly what was approved, and a signed frame that must change in one respect and no other. This chapter teaches how those things are held, whichever tool makes the picture.
In this chapter
What a reference does, and the sentence that stops it doing more
Master, anchor and in-look anchor: one meaning each
The identity stack, the identity clause and the anchored prompt
Sheets, states and props; real people; crowds
Places and plates; products; text in the frame, with Arabic always exact
The one-change edit; relight, extend, new angles, and a look taken from a picture
Before you start. Chapter 8 keeps the records this chapter feeds: the reference manifest, the continuity plan and the order of work. Chapter 9 teaches the frame that every prompt here ends in, and the repair ladder. Chapters 11 (Nano Banana Pro) and 12 (GPT Image 2.5) say how each tool takes pictures and in what syntax; this chapter teaches what holds for both.
10.1 What a reference does
On a set, a reference is something you show the crew, and each person looks at it for the thing you point at. A model looks at all of it: the face you wanted, and also the room behind it, the colour of the light, the pose and the framing. On the routes this book uses, every picture goes into one list, and the role of each exists only in your words.
A reference carries everything in it, and the model takes all of it unless you say otherwise. Give every reference one job, and say what it must not give.
Both makers ask for this. Google's guidance gives each image a role ("Use Image A for the character's pose, Image B for the art style, and Image C for the background environment") and says to "assign a distinct name to each character or object in your prompt". OpenAI's is to "assign roles to references", identifying each input "by number and purpose: subject, style, clothing, or background", and to explain "how the inputs should combine and which elements should move where" (both checked 30 Sep 2026). An unassigned reference bleeds into the composition and the style. Two references given the same job are averaged, and the average reads as "it ignored my references".
Weak: "Use Image 1 for the character." The picture's kitchen light comes along into the new room.
Strong: "Image 1 is Karim: take his face, hair, stubble, build and grey T-shirt; take nothing of its kitchen, laptop, lamp or his seated pose."

Figure 10.1 — Everything in the picture arrives with it. One sentence divides the picture into what it gives and what it is fenced from giving.
Name each thing to take, as a list you can check on the returned frame: "his face, hair, stubble, build and grey T-shirt", not "the character". Write the fence as "take nothing of its …", not as a list of banned objects: a banned thing named in a prompt puts its word into the prompt, and "take nothing of its background" fences the whole background without naming an object in it. Name every reference, and give no two the same job. Give each person one primary identity source. A reference that shows several views of one face may be read as several people, so add a count line ("Exactly one person is present: …") wherever twins are a risk. An unwanted second person is a conditioning problem, not an adjective problem: it comes from a picture that carries a face, or from two images that claim the same person, and only removing or fencing the image cures it.
Decide which job a picture will do: identity (who is this?), state (which version exists now?), shot start frame (what is true when the clip begins?) or edit source (which pixels are authoritative?). The commonest preparation error is a picture used for a job it cannot prove: a sheet (10.6) never proves a reverse angle's geography, which hand supports an object in a handoff, or where the camera starts.
Kind of reference | It gives | It must not give | The fence |
Master face or portrait | face structure, hairline, marks, skin | its ground, light, pose, framing | "take nothing of its kitchen, lamp or seated pose" |
State or body master | garment cut, colour, wear; build | its light; a face, unless it is the master | "take the kit only; take no face from it" |
Character sheet | the views and state it shows | its grey, panels, flat light | "take nothing of its grey backdrop, panels or flat light" |
Place plate (10.10) | layout, surfaces, wear | its camera position, people, light state | "take the room; take nothing of its framing" |
Product source (10.11) | geometry, marks, material, printed copy | its ground, angle, light | "take nothing of its ground, light or framing" |
Layout map | positions and distances | its drawn look | "objects appear exactly per the scheme; use the scheme as a reference, not a keyframe" |
Style picture (10.14) | one dominant quality, such as the grain | people, palette, geography unless named | name what to borrow and what to suppress |
Parent frame, the edit source (10.13) | every pixel outside the change | "keep everything in it exactly, and change only …" |
Crop a reference to its job when a fence is not obeyed. If a sheet's grey still leaks after "take nothing of its grey", crop the face into its own file, so the grey is not there to take. Crop in an image editor, not by asking a model to "isolate" a panel, and keep the original of every crop.
For each question a frame must answer, one file is the authority, and every new view is approved against it, never against the view made just before: drift accumulates in small steps, and a copy of a copy hides it. An invented person's authority is the master (10.2). A cast actor's is the approved photographs. A sold product's is the supplied photography, dimensions and artwork, which outrank any generated beauty image. A room's is its geography master (10.10).
Make the words and the pictures agree before you generate. Either describe what the reference shows, or leave that item to the reference and say nothing about it. No maker says which side wins when they differ, and it varies by route. A difference you want is written as a change contract (10.5).
10.2 Master, anchor and in-look anchor
Three words carry the weight of this chapter, and each has one meaning.
Word | In this book |
Master | the first, neutral picture of a character, a place or a product: a photograph of the actor, or a generated face on a plain grey ground in even light. It is the authority on identity, and it is first-generation: made directly for its job, never derived from another output |
Anchor | any approved picture a later job is built on and checked against: a master, an in-look anchor, a place plate |
In-look anchor | the character approved in the film's own light, made from the master with the master attached |
Plate | a still of a room or a product, with no action, made to be reused as a reference. In motion the word also names a clip made for a later pass: a talking plate for lip sync, a wordless plate (Chapters 13 and 22) |
A commercial makes the master first, then the in-look anchor. The master carries the face and no look; the in-look anchor carries the look, and the shots that follow are built on it. The master stays the authority on identity: later faces are checked against it, never against the in-look anchor made from it.
A film with one look may take a shortcut. Let the in-look anchor stand in for the master, and fence it ("take nothing of its light or grade") whenever it is reused. The Mastorna trailer did this, its first shot made in the film's black-and-white look and used for the rest. A neutral master carries no look to leak, which matters on a commercial with changing light, so there the neutral master first is the safer course.
Anchor every new generation to the first-generation masters, never to derived outputs. A chain of copies compounds drift the way tape dubs compound noise. When a keyframe's face has moved away from the master, the frame is re-anchored, not edited: attach the master again (as the second image of an edit, 10.13), restate the identity clause (10.4), and if the frame keeps degrading, remake it from the masters. Every few scenes, make a clean frame from the masters and set it beside what the chain now produces; when the two no longer match, the fix is upstream.
A derived frame carries scene state, light and composition, and never becomes the authority for a face or a product. A derivation is right when it is deliberate: a held face made from the frame before it, with the record saying "same face, same framing, held". The last frame of an accepted clip can start the next shot on the same terms: give it one job, scene state (who holds what, where the objects are, which way the light falls), fence its motion blur and compression, and attach the master second. Record it: "derived from [clip] last frame; identity from [master]".
Two things are called a manifest. The reference manifest of Chapter 8 is the production record, with a row for every recurring picture. The manifest of 10.5 is the opening sentences of one prompt, written from that record.
10.3 The identity stack
The master runs out when a shot reveals something it does not show. A head turn reveals a profile, a walk away the back, a wide shot the shoes and the length of a coat; rain changes the hair. Each needs a file of its own, made in a fixed order, each with one job.
S1, the master face. The entire head, nothing cropped, frontal or a few degrees off it, on a plain neutral grey sweep (seamless studio paper) in even light. Grey keeps the success rate high because nothing competes with the face. It is made with no reference, and it is where the identity clause is minted for the whole film.
S2, the angle set. The three-quarter view and the profile, made once from S1, on the same sweep in the same light.
S3, the body master. Full body from the front and from behind. Behind shows no face, so the picture still holds exactly one face.
S4, the state variants. A separate master for each look-state the story needs: wet, dusty, injured, race kit, formal (10.7).
S5, the world plates. The character inside the film's own light and grade, framed three-quarter with the environment. Identity is proved inside the look, because S1 to S4 deliberately carry none.
S6, the stress set. The identity under the film's hardest conditions: heavy smoke or steam, hard low-angle light, frames where the person fills under 10 per cent of the height. An anchor that survives only portraiture is not an anchor. Make it before the shot list locks.
S7, the chain addition. After the first hero keyframe is approved and tested in motion, it joins what you attach, beside the masters and never in their place.

Figure 10.2 — The identity stack. Start from the shot and build only what it reveals; S1 is always first.
Decide from the shot. Ask what the shot will reveal that the masters do not show. A profile or a head turn needs S2. The back, the body or the shoes need S3. A new state needs S4. A very large face, a lip-sync close-up or a poster is what S1 was made for, its head filling about 70 per cent of the frame. Then check both ways: every element the shot list needs has a master, and every master built is used by a shot. An orphan on either side is an error. Chapter 8 holds the order of work and the record.
10.4 The identity clause
A picture is one signal in a prompt that also has a scene to build. The words tell the model which features of the picture are the identity. OpenAI's guidance for keeping a character consistent is to repeat the appearance constraints and to say "Do not redesign the character" (checked 30 Sep 2026). The book adds that the repetition must be exact, which is why the words are written once and pasted.
What goes in it. Five parts: a distinct name, used the same way every time; the forensic description; the wardrobe state now, in the words of the state record; one detail anchor, a small unique fact ("a small healed scar above the left eyebrow") that gives the model a handle it cannot miss (a period film may choose to have none); and the fixed marks, the age bracket and the build. The forensic description covers the hairline and part, the brow, the eye set, the nose bridge, the cheekbones, the jaw, the skin and the grooming, and it writes the asymmetries first. A face written from its asymmetries ("the left upper lid fractionally heavier than the right") stays the same face; one written from a type ("handsome Italian man") drifts.
A shot carries an extract of six to ten load-bearing facts, about 30 to 70 words (a working band). The master's own likeness paragraph runs longer, because it is the source.
Research exemplar · the July 2026 research guide · written in the research, never generated · Nano Banana Pro, no reference · [excerpt]: the likeness paragraph of a 411-word master-face prompt, 4:5, 4k · 118 words
Example 10.1 — a master face that mints the clause
Her likeness is specific and fixed: a narrow athletic face with high cheekbones, dark brown eyes set slightly deep beneath straight brows, the right brow fractionally higher than the left, a small healed scar above the left eyebrow, a straight nose with a subtle flattening at the bridge, full lips at rest with a faint asymmetry favoring the right corner, black hair pulled into a tight low bun with fine flyaway strands at the hairline and temple, small unadorned ear studs, skin with true texture — visible pores across the nose and cheeks, faint natural shine at the forehead, a light scatter of darker pigment beneath the eyes, and fine expression lines beginning at the outer eye corners.Why it works
Copy the shape: the asymmetries named with their sides ("the right brow fractionally higher than the left"), one detail anchor (the scar), the hair as a fixed structure, the skin as evidence (pores, pigment, expression lines).
Watch for: beauty smoothing that erases the pores, and symmetry correction that erases the minted asymmetries. Either is a hard fail.
Where the clause lives. It is written once, after the master is approved, from what the approved picture shows and not from what was asked for. It is stored as one string with an identifier, a version and a validity window (the locked wording of the reference manifest), and pasted, never retyped: "emerald eyes" must never become "green eyes" in shot 14. A change is a new version, dated to the beat where it takes effect.
When the clause and the master disagree, the master is the authority. If the clause is wrong against the master, correct the clause. If the master is wrong against the character's intent, remake the master; do not paper over it in words. If they differ on purpose (a new state), say so as a change contract (10.5). Never ship a contradiction.
10.5 The manifest and the anchored prompt
A frame made from references is not a frame made from words with a picture attached. It is written in four parts, in this order.

Figure 10.3 — The anatomy of an anchored prompt. The first two parts add roughly 70 to 130 words in front of the frame (a working figure); the third is present only when it is needed.
A house rule of 18 July 2026 decides when these words are owed.
References and verbal anchors are workflow-scoped. Full verbal description is mandatory when the call establishes or changes pixels: text-to-image, sheet and anchor generation, text-to-video, and image or video editing. In ordinary image-to-video, the approved source frame owns composition, appearance, and identity; the motion prompt names preservation requirements and adds only motion, timing, physics, endpoint, and surgical negatives. It does not fully re-describe the frame and fight the source. Exact model and provider reference-role syntax still governs.
This chapter is the first half: a call that makes or changes a still is written in full. Chapter 14 takes the second half, where the anchor enters a clip. The manifest and the clause address the pictures and the face, not the frame, so the medium still opens the frame proper; the preamble may sit on its own lines, and the frame stays one paragraph.
Part 1, the manifest: one sentence for each image, in the order you attach them. Say what it is, by number and by content ("Image 1 is Karim's master face"), so that a change of order cannot silently swap two roles. Say what to take, as a list you can check, and what to take nothing of. Give one job to each image: when a state picture also shows the face, say that the face comes from the master alone. State each conflict ("the mannequin in Image 2 is not the actor"). A manifest says what each picture is for, not what the face is like; that is the clause's work.
Part 2, the clause, pasted from the record (10.4).
Part 3, the change contract. Most frames need none. Where this frame differs from a reference in a way the reference carries, name the difference as a decision: "Now: a wet overcoat, replacing the dry one in Image 2; the collar turned up." The grammar is three verbs. Preserve the face and body proportions from the identity images. Take the wardrobe from the state image (or describe a new state in words when no state image exists yet, and let the returned frame become the state master). Set this shot's light and mood as free description. What is not named as preserved is negotiable, so the list is explicit every time.
Part 4, the frame, in full. A master carries likeness, not look, so this shot's light, lens and moment are written every time. An in-look anchor, a frame or a plate carries its own light and grade and hands them over unless it is fenced. With the manifest and the clause, a complex identity-bearing hero still lands around 400 to 600 words in all.
Template: the anchored still prompt (written for this book; not run)
Image 1 is [name]'s master face: take [face, hair, skin]; take nothing of its [ground, light, pose]. Image 2 is [name] in [the state]: take [the garment: cut, colour, wear, build]; take no face from it. Image 3 is [the place]: take [layout, surfaces]; take nothing of its [camera position, light state, people]. [The identity clause, pasted from the record: the asymmetries first, the hairline, the skin, the detail anchor, the wardrobe now.] [Only if needed: Now: what differs from a reference, replacing what it shows; what still holds.]
[The frame: the medium, the frozen instant and share of frame, the inventory, the place, the camera and each light with their visible results, the finish, at most five exclusions, and: no overlaid text, subtitles, logos, or interface graphics.]Order and how many. For stills, attach identity first, then the state or body, then the frame or plate, then the one task-critical object, then style; when slots run short, the master face and the current state are the last to be cut. Binding is the tool's syntax (Chapters 11, 12, 17 and 18 give each), and whatever the syntax, name the image by its content too. The working range is two to five images, each with a distinct job. Makers' ceilings are higher and differ by tool and surface, so they belong to the tool sections; every extra image is another room, light and pose to fence off, so remove rather than add. A shot with more than five people is split: make the scene with the principal cast anchored, then add or correct the rest in a controlled edit (10.13).
10.6 The master's grammar, sheets and single views
A master or a sheet is a still prompt whose world is a studio void. Chapter 9's twelve elements bend as follows.
The medium names the artifact honestly: "a studio identity plate photographed for a film's character reference system".
The frame is STATIC: a neutral standing pose, weight even, arms at rest, gaze to the lens or the named angle. A master encodes identity, not an instant. Give the head's or the body's share of the frame as a number ("the head occupies roughly 70 percent of frame height").
The forensic paragraph and the wardrobe are written at full depth, and the sweep is the environment: the paper's tone, the floor curve, the falloff toward the edges, the absence of props stated as design.
The camera is standardised: a portrait lens (85 to 105 mm) at eye height for faces; a normal lens (40 to 50 mm) at chest height for full-body views, with even headroom and floor space.
The light is neutral by rule: an even wrap at about 5500K, contrast near 2:1, soft open shadows, no practicals. The film's look is applied shot by shot and proved in the world plates.
The finish is neutral and honest (true skin, fine grain, no grade), and the exclusions add one: no baked panel labels or captions. View names live in file names, because text painted into a reference leaks into outputs and crops.
Sheets
A character sheet puts several views of one person in one image. It is a way of making the body master, and the form you attach when a later frame must show a character from an angle or in a state the master does not show, on a route that reads it.
The rule on sheets. One sheet per character. The 3-panel is the default: a headless body from the front, the body from behind, and a chest-up face close-up, on flat mid-grey under soft, even light. The 5- or 6-panel is the variant for a character seen in profile or turning.
And a house check on every sheet: count the panels when it comes back, and crop what you need.
The 3-panel holds one face. If one reference shows several views of the same face, a model may treat them as several people, and a grid of faces hands it many chances to grab the wrong one, at a fraction of the resolution for each. The 5- and 6-panel variants hold three faces (front, three-quarter, profile), so they carry the risk. Attach a sheet only where the route reads it; otherwise crop each view into its own file, and inspect the crop at 200 per cent for softness and for contamination from the grid. A grid is a crop factory: useful upstream, left out of what you attach. Build only the views the planned shots need.
Count the panels, and look at the torso. A request for six panels can come back as seven, with two near-identical profiles, and a headless torso can carry a faint rectangular mark on the chest. Both are checks for every sheet, not limits of the method; a mark is cleaned in post.
Single views
A profile or a full-body view can be made as its own file from the approved portrait. Compare each with the portrait at full size: nose to chin, ear position, hairline, neck length, garment end points. On an assembled sheet, a panel too small to read an ear, an eye, a fastener or a logo fails, and the single view goes to the shot instead. A three-quarter where a profile was asked is a one-change edit; a new facial structure goes back to the portrait. Leave the step with each approved view as its own file and a handoff note: "front portrait controls identity; front and back body views control the dry charcoal-shirt wardrobe; none of the sheet's grey background, separators or labels belongs in the scene."
10.7 States, props and expressions
Stories change people: rain, a jacket taken off, a seal broken. Each change is a state, made as a variant of the clean master and never of the last shot, because a state built from a shot inherits that shot's drift.
Build each state in three steps: make the face, make the outfit separately, then fuse them in one master. Separating them gives more control than prompting both at once, and it is the method when a costume will not hold. Make the state a one-change edit of the clean master, and say where the change lives: "damp across the shoulders and upper sleeves" locates it; "wet" lets it spread. Name the state in the record ("state: wet after rain"), give it the beat where it takes effect, rename the file when the state changes, keep the old versions, and give the file only to the shots that need it.
Write wardrobe states in material words, built on the locked face: "washed pale-blue chambray shirt, open, over a faded dark-grey T-shirt, both sleeves rolled twice", never "the same jacket". Paste each wardrobe and prop line into every prompt, unchanged, so that the words and the file agree. When a change alters how the face reads, it is a new master: glasses on, a beard grown.
Keep a prop that comes off on screen out of the master, in its own file: glasses, a hat or a bag then go on or come off by reference. If the character smiles on screen, sign a smile with teeth; a closed-mouth master leaves the teeth to invention.
Make expression states as separate frames from S1, written as muscles and breath, never the emotion. "Lower lids tighten; nostrils widen slightly; the jaw shifts forward" is an observable state; "angry" is not. One frame for each state the script needs, in the same neutral light. If an expression changes the perceived age, the jaw or the eye spacing, reduce its amplitude and regenerate against S1 (Chapter 15 teaches the language of performance).
Check a state by comparing the face with the master at 200 per cent and finding the state where you placed it and nowhere else. A state that spreads is located again; a face that drifts goes back to the master with the clause restated once.
10.8 Real people, and an actor's name as a rights decision
Two different questions arrive under this heading. A house rule of 18 July 2026 answers the first.
Actor-anchor scoping. Real-name likeness anchoring ("modeled on …") is a project-level rights decision: natural for historical reconstruction, default OFF for branded and commercial work where likeness rights govern. The forensic micro-anatomy layer is the universal likeness engine either way; the Real-Person Protocol governs living persons above this guide.
A name in the prompt is a project decision, made once, in writing, with a note of who decided and why. It suits a historical reconstruction; on commercial work it is off. When it is on, the name sits inside the locked clause, used the same way every time, and never alone: the forensic facts stay with it, so that the name is the option and the facts are the engine. If the name stays in every prompt while the clause thins to a loose description, the order has inverted.
A real person as themselves (a testimonial, an adaptation, a colleague) is the second case. It is allowed only with the person's written consent, or the estate's release, covering both AI likeness and the training of the platform's models. The consented photographs are the master and the only identity source. No name goes in the prompt. The clause is written from the photographs and approved by the subject. A generated view of a real person, such as a side view nobody photographed, is only a proposal: its likeness is approved before any motion is made from it. Build the prompt as an anchored prompt (10.5).
For a real person: a written release first, the consented photographs as the only identity source, no name in the prompt, and any drift is a failure. No retries on a real face: stop, and go back to the photographs. A near-likeness of a real person is worse than a stranger. Where there is no release, invent a character.
What the makers say (checked 30 Sep 2026). Google's Generative AI Prohibited Use Policy (last modified 17 Dec 2024) bars impersonating an individual, living or dead, without explicit disclosure and in order to deceive, and using personal data or biometrics without the consent the law requires. OpenAI's usage policies (effective 29 Oct 2025) address using someone's likeness, image or voice, without consent in ways that could confuse authenticity. Read the whole policy of the tool you use before a rights decision; video models differ on whether they accept a real face as a reference at all (Chapter 17), and Chapter 34 gives the platform's terms. Whether a deceased public figure's name may be used in a private reconstruction is a legal question: ask counsel.
Before a real person's frame leaves the room
The release covers AI likeness and platform training.
No name in the prompt.
The forensic clause is written from the photographs and approved by the subject.
The frame is checked against the photographs at 200 per cent.
The subject or the family has signed it off.
10.9 Crowds and extras
A street, a market or a café asks the model to invent many faces at once, and it fails in a particular way: asked for many individuals, it clones faces, and a crowd written as a number invites a grid of the same person. Direct a crowd as a first assistant directs extras: in groups, each with its business.
Write crowds as clusters with a task, never as a head count, and keep background faces small and soft by distance.
Clusters, not counts. Models miscount. An ahwa table of three backgammon players, a queue at a microbus stop, a souk stall and its customers: each cluster gets a task and a direction, staged in depth. If a number matters to the idea, test it before a client shot, or set the count in post.
Individuate a few, and describe the rest by type. A finished station scene names three figures by age, dress, gait and face, and everyone else by kind (men in dark overcoats, women in headscarves, porters with trolleys, a priest, an elderly couple), with one statistic ("at least sixty percent of the men wearing hats") and one line that every visible face is distinct in age, bone structure and expression. Write every mover's phase as a position, and say who stands still: a crowd that "flows" leaves the instant open (Chapter 9).
Give extras independent activity, and keep a crowd's geography across cuts: the two banks of spectators at a parade stay on their sides. Let distance and focus do the rest. Faces too small to read leave the model no individual face to repeat. Still check that no background face resembles your lead, and clean any that does.
Template: a background written as clusters (written for this book; not run)
Behind them, soft with distance, [clusters as positions: two men facing each other over a crate, one hand raised; a woman with shopping bags half-turned into a doorway; a boy on a bicycle side-on at the far end], faces too small to read.The same thinking applies to objects. Build density in zones: a working centre, a support layer and a human trace, grouped by function and owner ("the tea corner", "the paperwork spill"), not by count. Answer "make it real" with fewer, more legible layers, not more stuff (Chapter 7).
10.10 Places and plates
A plate is the most inherited frame in a production. A video model lifts its texture and its light from the frame it is given, so a plastic-looking plate spoils every clip built on it. Make plates at finished grade (Chapter 9), with every number beside its result.
A location is built like a character.
Plate | Its job |
L1 · The geography master | the establishing wide that fixes the landmarks; every later shot's geography sentence is written against it |
L2 · Coverage plates | the key set-ups at a three-quarter angle, never flat head-on: the oblique gives the model depth to hold when the camera moves |
L3 · Light-state plates | the same geography under each time of day and weather the script needs, each locked |
L4 · Material close-ups | the surfaces that must stay true at insert distance: patched cobbles, chipped tile, the worn counter |
Write a geography sentence for every shot in the place, against the geography master: what stands where in the world, then where the camera stands among it. A piazza does it in a clause: the camera looks "down and across the square toward an enormous Gothic cathedral at the far end", with "the flanking buildings creating walls on either side". Keep the identity masters neutral and let the plate carry the look: a grey master asked for "a warm Cairo kitchen at night" invents a new kitchen for each shot. A place whose identity is its light, such as a hotel lobby lit by candles alone, names that light the only source, at about 1800 Kelvin with contrast near 5:1.
A general plate gives the room (its space, surfaces, wear and practical lights) and never sets the framing of a new shot: take the room, its surfaces and its wear, and take nothing of its camera position, people or light state. A station plate is made for one camera position and may set the framing on purpose. Describe the room from the camera, with world positions fixed: screen left and right change with the camera, the fridge and the stove do not.
Empty an approved frame to make a general plate. The room, its light and its wear are already signed; a one-change edit (10.13) that removes the people keeps them. Name everything that goes, and what each vacated surface becomes, in positive words ("the table top where she sat is bare wood"). Edit a plate to fit the action; do not regenerate it, and when an edit softens a crisp plate, composite the original texture back into the untouched regions.
Making a station plate. Draw a floor plan first: where the camera stands, which wall is behind the subject, where the landmarks are in the world. Write the landmarks as the new station sees them, and say that the view is not a mirror of the master (Figure 10.5). Generate the empty plate with the room master as image 1 and any accepted view of the destination wall as image 2, and check it against the floor plan before it enters the location pack. Keep the light's source in its world position: the window does not move because the camera did. Never let a prompt design a wall: a wall nobody has seen is a production-design decision, made and approved as its own plate.
Use a straight-on diagnostic view when the job is geometry, and a cinematic view when the job is look: a dramatic frame cannot also serve as a precise layout map. A diagnostic prompt gives the room as a compass and measured positions ("platform windows along north wall; entrance door west; … two-person table at coordinate C3, 1.2 metres from north window") and forbids signage beyond blank approved sign zones.
When prose keeps misplacing the actor, the product and the space for copy, draw a map. Say that the drawing controls positions and relative scale only, not its lines, labels, flat colours or style, and ask for every sketch mark to be removed. Identity and product keep their own authorities. If sketch lines survive, use a photographed block-in.
You see | Why | Smallest fix |
A figure or object left in the "empty" room | the removal list was incomplete | name everything that goes, and what the surface becomes |
Surfaces smooth and plastic after an edit | the edit softened a crisp plate | composite the original texture back into the untouched regions |
The plate brings its framing into a new shot | it was fed without a fence | "take the room; take nothing of its camera position" |
The room comes back mirrored | landmarks were described as screen sides | state them in world terms first, then as seen from this station |
10.11 Products: the real thing is the authority
A sold product is the one element of a commercial whose identity is not negotiable. A model left to itself makes a plausible bottle, not the client's bottle.
A product the model half-knows is a product it will redesign. Give it the real object from every side, tell it to keep it, and check the geometry before anything else.
Build the product as you build a place. Prepare the real front, three-quarter and side photographs, the packaging artwork (kept for post), the dimensions and the approved brand mark as an image. Where the product is complex, add a multi-angle sheet at matched scale, an exploded view where assembly matters, a material macro, and for a hero prop a wear map. A prop's condition follows the object ledger, so that a case damaged in scene four stays damaged in scene nine.
Carry a fidelity clause in every shot the product appears in, verbatim: "use the product exactly as shown; preserve shape, design and proportions; do not redesign". Fence each product picture by its job: the bottle plate "controls only bottle silhouette, cap and blank label proportions; do not inherit its grey studio background or turntable reflection". A product plate is diagnostic: neutral, straight-on, free of decoration, so its geometry is readable.
Name the geometry that must hold, and bind a second view "only to resolve that geometry". A second view with its own job resolves thickness without bringing a competing look. Say what each glass edge does under the one light: a long soft highlight down one edge and a darker opposite edge show the thickness, while an over-wide white stripe makes rectangular glass read as a metal cylinder.
Lock the product's scale in words, against a human anchor, and bias it to the forgivable error. Generated objects drift in size between shots: "if the box reads longer than her forearm, the frame fails; render the box smaller, never larger."
A real brand or model name in the prompt is not product authority. A car's or a phone's name draws the model's idea of it, not your pack shot.
The product route
Generate the product in context, with the real photographs attached and the fidelity clause in the prompt.
Inspect geometry first (silhouette, shoulder radius, cap-to-body scale, contact with the ground, liquid level, both glass edges), then the label letter by letter (10.12), then light and material. Count the products.
If the label is wrong, composite the supplied artwork onto a clean surface with correct perspective, texture and occlusion, and keep the generated clean plate.
Write the list of parts that may not change in animation. The director approves it.
When a product is placed from a cut-out, the cut-out's job is separation and orientation and the original photograph's job is detail: attach the original again as the authority in any finishing pass, with an explicit keep-list (geometry, proportions, label layout, logo, cap, colour, finish, count, orientation; Chapter 12). A composite on a still frame still needs tracking and occlusion if the product moves later. Never hand the video an attractive approximation of a sold product.
You see | Why | Smallest fix |
The product grows or shrinks between shots | no scale lock | the forearm test, and "smaller, never larger" |
Cap or shoulders redesigned | the model's idea of the product | attach the side view again "only to resolve that geometry"; after two controlled attempts, a photographed or composited product |
A second bottle, badge or seal nobody asked for | the model's habit | name it in the exclusions |
10.12 Text in the frame: two layers
Every word in a commercial belongs to one of two layers, made in different places. A house rule of 18 July 2026 divides them.
In-world text may be generated. Editorial text is always composited in post. Signs, packaging, books and screens belong to the world: they are production design, specified like any material (language, script, era, style, wear), given honest distance-based legibility and checked character by character. Titles, subtitles, supers, end cards, watermarks, legal lines and logo lockups used as graphics sit on top of the image and come from real assets in the edit.
The closing line bans the overlay layer only. Every still ends on "no overlaid text, subtitles, logos, or interface graphics". Never write "No text" on a frame that must carry a sign: a blanket ban fights the sign you have specified. Faint printing appears on plain packaging, and a third-party logo on a laptop lid, whatever the closing line says: look for both. Typography a video model renders is pixels and cannot be retyped; if the client may change copy, a number or a claim, build the graphics as editable layers (Chapter 30).
In-world text and Arabic signs
The model may paint a shop sign, a label, a headline or a screen; you must read it. Never use this method for text that ships without a glyph check, or for a wordmark you cannot check; the fallback is always a blank surface, with the words set in post on its plane.
Approve the exact string outside the model: Arabic in Arabic script, as a canonical Unicode string with its language and locale, approved by a qualified Arabic reader for wording, spelling, direction and meaning. Never give the model two spellings.
Specify the sign like a material, then give the string once, in quotation marks, with its position and typography, and say that nothing else is written. Google says to "enclose your desired words in quotes" and to describe the typography; OpenAI says to "put required wording in quotes and describe its position and typography" and to "ask for no extra text, then check spelling and legibility in the output". Google adds that "rendering small text, fine details, and producing accurate spellings may not work perfectly" (checked 30 Sep 2026).
Choose legibility by distance, and light the letters to read. A sign the story needs is legible, in a raking light across the paint, not glare. Asked for exact words at a distance, the model invents near-words, so let distant signs "recede into honest illegibility with distance and haze", and fold newspapers "so no headline resolves".
Read every glyph on the returned original at delivery scale, never on a scaled preview: dots, final forms, joins, direction, word order, numerals. Make two independent character-by-character transcriptions, compare them with the canonical string, and write the read down. Arabic text generated inside a frame is always exact: it must match the approved string letter for letter, and a final ى where ي was asked fails, and is fixed in post or run again. A sign that came out right proves nothing about the next string. Remove painted quotation marks and guillemets in post.
Two controlled attempts, then climb a level. A failure that survives them is not a wording problem: blank the panel, set the words in post on its plane, and keep the clean plate for a moving shot.

Figure 10.4 — The in-world text method. The read in step 5 is the only proof; a sign that looks right can still have one letter changed.
For a sign the production designer needs to see, make it straight-on first; a cinematic view comes after, for look.
Research exemplar · the July 2026 research, an instructional example for a weathered café sign · written in the research, never generated · an exact-text condition, 4:3, 2K · its own status: not reliability evidence; the strings are fixtures, not approved copy · 56 words
Example 10.2 — a straight-on world-text reference
A straight-on production-design reference of a weathered enamel café sign mounted on plaster. Exact Arabic line: مقهى المحطة. Exact English line beneath: STATION CAFÉ. Preserve right-to-left Arabic letter joining, line order, spacing, and the accented É; off-white enamel, chipped navy border, rust around four screws, even neutral light, no perspective distortion, no extra words or symbols.Why it works
Copy the shape: each line's language stated; what must be preserved (right-to-left joining, line order, spacing, the accented É); the material and its wear; even light; a ban on extra words or symbols. For a finished frame, add the closing line and record the blank-panel fallback.
Brand marks and packaging copy
A brand mark is either exactly right or unusable. It is generated from the real asset, attached as a reference, because that gives the best chance of a true mark in true geometry. The composite is the legal-grade fallback: whatever ships as a legal mark or as regulated copy is exact by construction, so exact logos, claims and regulated copy go in post whatever the sample looks like. Where the label will be set in post, ask for a plain unprinted panel on the face that will carry it, at the scale and angle of the real label.
Label copy has ground truth outside the image: a canonical transcription of every string, the Arabic, the English, the numerals and their style, approved by the client and a qualified reader. Name the word from the start. An unnamed brand word is filled with a frequent one, and editing a frame that already shows the wrong word keeps it. When the model puts pseudo-text on a label that should be blank, the repair is an edit that removes it and leaves a clean surface. A brand mark replaced on an approved garment is a controlled edit whose scale is checked by a landmark: "smaller than the width between the inner edges of the black shoulder panels".
Before you sign a frame that carries words
The string was approved outside the model, by a qualified reader if it is Arabic.
It is quoted once, with position and typography, and "nothing else is written".
The letters are lit to read, at a size honest for the distance; distant text is texture.
The closing line is the fixed one, not "No text".
Every glyph was read on the returned original; no painted quotation marks; packaging and screens checked for stray printing.
10.13 The one-change edit
An edit changes pixels of a frame that is already approved. The approved frame is the parent. The preserve list names what stays, region by region; the causal scope is what the change may move as a consequence (a new light moves shadows, a new garment moves folds).
An edit changes one thing, and the parent is the source of truth. Name the change, name what it may move, list what stays, and say all of it again on every edit.
The edit contract, in order
Attach the parent first, and say that it is the source of truth for the composition, the camera, the pose, the light and the finish. With a face in frame, attach the master second and restate the identity clause (10.4): "Do not beautify, average, age, de-age, or re-cast the person."
Name the one change, in physical terms: what, and where.
Name what the change may move: the occlusion, contact and shadows it causes.
List what stays, region by region, and repeat the list on every edit in a chain. What is not named as preserved is negotiable.
Inspect the result against the parent at full size: faces, hands, edges, lettering, then the framing.
When drift builds, return to the last approved parent and make a smaller edit. Where pixels must not move, composite the approved change back into the parent. If the frame is structurally wrong, make it again.
An almost-right keyframe is edited, not rolled again: a re-roll renegotiates identity, an edit preserves it. Exact preservation is never guaranteed, so an edit is judged like any new frame. The words describe the change in full, as the house rule quoted in 10.5 requires: the picture owns the pixels, the words own the change.
Two grades of edit. A test-grade edit is the change and the preserve list, and it is right while you prove a route or make a small bounded edit on a frame with no signed face. Any frame that enters the cut, carries a face, a product or a sign, or will be the parent of more edits is edited at the finished grade: the parent named as the source of truth, the identity clause restated where a face is in frame, the one change, its consequences, the preserve list by object, the camera and light restated with their numbers, and the closing line. An edit is as long as its preserve list needs: 40 to 150 words is the band for a bounded edit, and a finished edit with a signed face runs to several hundred.
When only a small area must change and every other pixel must stay, use the method that never regenerates the other pixels. Clean it up in post first (a mark, faint printing, a logo, with the grain restored). Second, edit the whole frame and composite the changed area back onto the parent. Third, a tool with a mask, where a tool section teaches one: its pixels come from another engine and can look as if another camera took them.
Removal, addition, garment and correction edits name the change, its consequences and what stays. A blanket "keep everything unchanged" leaves the model to guess what everything includes when the change itself must move something, and a change that names no boundary spreads. A removal names the object and its place, what the rebuilt surface must match (texture, perspective, light, focus) and "do not replace it with another object". Which tool: a frame with a signed face is changed on Nano Banana Pro (Chapter 11), because the face travels through the parent and the master; a frame with no signed identity can also be changed on GPT Image 2.5 (Chapter 12).
Template: an edit (written for this book; not run)
Image 1 is the approved [frame] and the source of truth for its [composition, camera, pose, light, focus and finish]. [Image 2 is (name)'s master: restate the face, hair and skin in the locked wording.] Change only [the one change, in physical terms: what, and where]. It may move only [its own shadows, contact and occlusion; the folds it makes]. Preserve, unchanged: [each protected region by name: faces, hands, product and label, set, camera, crop, light, grain]. The result reads as the same photographed instant after one change, not a new rendering. No overlaid text, subtitles, logos, or interface graphics.10.14 Relight, extend, new angles, and a look from a picture
These are edits with a special reason, and each has one trap.
Relight. The light changes inside a scene: a laptop becomes the key, the pendant goes off, evening comes. The composition is accepted and the light state is new. It is a one-change edit whose change is the light. If only exposure and colour change, it is not a relight but a job for the grade. Host relight and outpaint tools are not used, because their engines are not named (tool text last read 28 Sep 2026).
A relight is the riskiest edit, so judge it hardest. Google says: "Advanced editing tasks like blending or lighting changes can sometimes produce unnatural artifacts" (checked 30 Sep 2026). Describe the new light by its source and what it does, and give the old source a new role: "the warm household bulb behind him falls back to a soft warm rim on his hair, ear and one shoulder" tells the model where the old light went, so it neither vanishes nor doubles. Admit the dependent changes, the shadows, reflections and exposure, and fix everything that is not light. Judge the lamp's state, the light's direction on skin, the highlights on glass and the background's brightness together, then compare the framing with the parent, which a relight can return slightly wider. Later shots take the relit frame as their parent for scene state; their faces still come from the masters.
Extend. Extending is for static pictures only. For a speaking or travelling shot, make the other shape as its own frame: a new composition serves a shot better than a stretched one. The routes, in order: originate the other shape from the approved frame, with the target aspect set and the new area described; extend by hand in post from the frame's own material; a named outpaint model third, where a tool section teaches one. The arithmetic decides: taking a 1920 × 1080 frame to 4:5 needs 1920 × 2400, so 1,320 of 2,400 rows, 55 per cent, are invented. Say what fills the new area, as a physical surface ("more of the white wall tiles beyond the dark stove"), and judge the extended edge like a new frame: look for strangers, signs, doubled objects and seams.
New angles and reverses. Coverage of a signed room or person needs new camera positions. Each is a new frame built from the same authorities, the room and the person, with the camera described from its new station. Attach the room's master, an accepted view of the wall the new camera will see, and the actor's master. If that wall was never established, make and approve it first (10.10). Anchor the new angle to the masters, not to a neighbouring shot, which may lend scene state and nothing else. Describe the camera relative to a frame you already have: "at his eye level a hand's width camera-right of the first position, slightly tighter". A distance from a known position can be checked; a camera described from scratch cannot. Keep the line: record who owns screen-left, where each gaze points and which side of the 180-degree line between two people the production photographs; later coverage crosses it only as an explicit transformation.
A reverse moves the camera, not the room. Place every landmark in world terms first, then say where it falls as the new camera sees it, and say that this is a new viewpoint, not a mirrored copy. Screen left in the master is not a place in the room.

Figure 10.5 — A shot and its reverse, both from the same side of the line. Every landmark keeps its place in the room; the reverse sees the other wall, which must be designed before it can be shot. From camera 2 the doorway falls at the left of frame though it stands at the right of the plan; a mirrored copy of camera 1's frame would put the window behind B.
Keep practicals in their place in the room and name what they now light: a lamp that moves with the camera breaks the geography. Never flip an image with text or an asymmetric prop in it to make a reverse. The prompt has the shape of 10.5: the room master with "take the room, its surfaces and its light; take nothing of its framing", the face and clothes "only", then the camera's station, what is now left, right and behind, and the sentence "this is a new viewpoint in the same room, not a mirrored copy of image 1".
A look from a picture. Put the look into words first, as a finish that describes behaviour (Chapter 9). A style picture brings its content along with its colour. If the words miss twice, add one style picture and give it the look only, fenced: "take only its tonal range, contrast and grain; take nothing of its people, room, objects or framing." A style picture never goes on an identity master or a sheet, because a master goes to every frame of its character. For a material idea, write the material's construction and the state it is photographed in ("caught mid-collapse"), not "in the style of".
Open questions
Neutral master or in-look anchor. Which holds a face better over a long run, and whether a locked clause improves identity on top of the pictures, is not yet known. On a commercial with changing light, start from the neutral master.
Sheet layout for a profile. Which layout holds a character best in profile is not yet settled; until it is, the 3-panel is the default.
A real person's likeness. The release route has not been shown to hold a likeness. Before a client job, make one frame from a consenting colleague's photographs and check it at 200 per cent.
Edits that keep a face. Whether a finished-grade edit holds a signed face is unproved; judge each edit against the parent.
GPT Image 2.5 for a signed identity is not yet qualified (Chapter 12); signed faces stay on Nano Banana Pro.
What to remember
A reference carries everything in it. Give each one job, name what to take, and fence the rest with "take nothing of its …". One file is the authority for each question; approve against it, never against the last view.
Master, anchor and in-look anchor have one meaning each. A drifting frame is re-anchored to the first-generation masters, never edited in place.
The identity clause is six to ten fixed facts, asymmetries first, written once from the approved master and pasted unchanged. The master wins any disagreement.
An anchored prompt runs manifest, clause, change contract, frame. A master carries likeness, not look.
Sheets and states are built from the clean master. The 3-panel holds one face; count the panels; attach a sheet only where the route reads it.
A real actor's name is a written project decision, off for commercial work. A real person as themselves needs a release, consented photographs and no name; any drift fails.
Crowds are clusters with a task, faces small by distance. Places are built from a geography master and plates; a reverse is never a mirror. A sold product is the authority on itself.
In-world text is specified and read glyph by glyph; editorial text is composited; Arabic is always exact.
An edit changes one thing, names what it may move, and repeats the preserve list. A look goes into words first.




Comments