Chapter 2 · Directing a model in words
Part 1 · Before the camera
A prompt is a brief for a reader who has never met you, reads it once, takes every word literally and forgets it the moment the job is done. This chapter teaches the principles that hold for every model: which work belongs to the words and which to the inputs you attach, the two grades of prompt, what a finished prompt settles, how to write numbers so they help, what to translate before you borrow other people's words, how to say "no" without handing the model the thing you forbade, how long to write, and when to stop writing and change what you give the model. Each tool chapter then teaches its own model's dialect.
In this chapter
Every model is a different reader, and what the makers document about how theirs reads
The inputs own what they carry; the words carry the rest
Two grades: the finished grade is the standard, the test grade only proves a route
What every prompt settles, with two short finished examples: a still, and the clip made from it
Choosing the writing format by the job
Numbers: both words and numbers, each number beside the result it should give
Borrowed words and other people's prompts: what to translate, what to leave behind
Exclusions: positives first, at most five, then the closing line
How long a prompt should be, where to stop, and briefing an AI assistant
Before you start. Chapter 1. Chapters 9 and 14 teach the anatomy of a still and of a clip in full, and the tool chapters (11, 12, 16 to 18) teach each model's dialect; this chapter holds what they share.
2.1 Every model is a different reader
On a set you brief people who share your language and your intentions. A model shares neither. What it has in front of it is your words and whatever you attached, and each job starts from nothing, with no memory of the last. Each model also reads in its own way. Keep the picture in mind while you write.
Model | The reader | What that means for the words | Chapter |
Nano Banana Pro | a sharp stills photographer who reads your brief once, has no memory, and takes everything literally | every sentence after the first names something a camera could see | 11 |
GPT Image 2.5 | a photographer reading a brief, who follows facts it can check in the picture better than moods | write the instant as physical facts: who holds what, what has not happened yet | 12 |
Kling 3.0 | the frame owns everything that exists; the prompt owns everything that changes | write only the change: motion, camera, interaction, the end state | 16 |
Seedance 2.5 | a crew reading a shot description, top to bottom, once; the frame you give it is the set, dressed and lit | write who moves, in what order, where the camera is, what the room sounds like and when to stop | 17 |
Wan 3.0 | a director's shot list: an overall line, then the subject, scene, motion and look of each shot | say "one continuous shot" when you want no cuts; state dialogue and music in words | 18 |
ElevenLabs | the voice decides the dialect and the range; the text is the score | direction lives in the line, its punctuation and spelling, and the tags the model accepts | 23 |
Suno | a session band reading a brief once: the form, the players, what to avoid | one directed paragraph; the structure in section tags | 26 |
Sync Lipsync 3 | no reader at all: there is no prompt | all the direction goes into the plate it redraws | 24 |
The reader has only this job. Whatever the model needs must be in this prompt or in a picture, clip or recording attached to this job. Your intentions, your last prompt and the frame you liked yesterday do not carry over.
What the makers document
Two makers publish prompting advice for the stills cameras, and it is worth knowing what they say and what they do not, because much of what circulates as rules is neither. Both were re-read on 30 Sep 2026.
Maker and page | What it says | Where it matters |
Google, on Nano Banana Pro (prompting tips, 20 Nov 2025; the Google Cloud guide, 6 Mar 2026; its best-practice page, updated 28 Sep 2026) | "Provide context and intent: Explain the purpose of the image to help the model understand the context." | 2.3 |
"Use positive framing: Describe what you want, not what you don't want (e.g. “empty street” instead of “no cars”)." | 2.7 | |
"Direct the shot like a cinematographer", with its own example "a shallow depth of field (f/1.8)". | 2.5 | |
"A simple list of keywords won't cut it; you need to describe the scene narratively." | 2.4 | |
"When using uploaded images, clearly define the role of each." | 2.2 | |
OpenAI, on GPT Image 2.5 (image prompting guide) | "Name the subject and intended use, such as a product photograph, advertisement, or diagram." | 2.3 |
"Choose the format that makes the requirements easiest to read and update rather than relying on special syntax." For complex requests, "labeled sections". | 2.4 | |
"Treat camera specifications as cues for appearance, not a guarantee of exact physical simulation." Its own portrait example says "using a 50mm lens". | 2.5 | |
For wide, low-light or neon scenes, "specify scale, atmosphere, and color instead of relying on mood words alone." "State exclusions such as unwanted text, logos, or watermarks." | 2.5, 2.7 | |
"Assign roles to references. Identify each input by number and purpose: subject, style, clothing, or background." | 2.2 |
Neither maker requires a syntax, a template, a JSON schema, a word count or an order of elements. Google lists elements (subject, composition, action, location, style); OpenAI lists scene, subject, details and constraints. They describe decisions, not formulas. Google's API page (updated 23 Sep 2026) adds that its Gemini 3 image models reason before they draw and that this cannot be disabled in the API, which is one more reason to write a single brief that does not contradict itself. The order this book teaches in Chapter 9 is a craft practice, and its reason is craft, not a maker's rule. The video makers' formulas are in the tool chapters.
2.2 The inputs own what they carry; the words carry the rest
Every job gives the model two kinds of evidence: the words, and the inputs (the pictures, clips or recordings you attach). The first principle of writing a prompt is to know which carries what, and to write only what the inputs do not already carry.

Figure 2.1 — The inputs own what they carry; the words carry the rest. The further down a production you are, the more the inputs carry and the less the words need to say.
Read the figure from top to bottom and a production's shape appears. The first frame is made from words alone. After that every signed frame, take and clip becomes an input, and each later prompt shrinks to what the inputs cannot say.
A call that makes or changes pixels needs the full description. A still from words, a character sheet, a clip from text, an edit of an image or a clip: nothing else carries the look, so the words must. A clip from a signed start frame does not. The frame owns composition, appearance and identity. The prompt names what must be preserved and adds only motion, timing, physics, the endpoint and surgical exclusions; it does not re-describe the frame and fight the source. The exact reference-role syntax of each model and provider still governs (house rule; Reference E).
How an approved still enters a clip. It enters as the start frame: you attach it in the model's start-frame slot, and the prompt never mentions it. The prompt names the mover, says what moves, and protects by name whatever the motion could break. When references stand in for a start frame, each is bound by number and name, placed, and protected. Chapter 14 teaches the pattern and each video chapter shows its model's slots.
When pictures are attached to a still, the words carry each picture's job. Google's formula is reference images, then the relationship instruction, then the new scene; OpenAI says to assign roles to references. The weak version says "Use Image 1 for the character." The strong version names the picture, gives it one job and fences what it must not bring: "Image 1 is the presenter, the approved master: take her face, hair and skin; take nothing of its room or its light." Chapter 10 teaches the manifest.
When the words and the inputs disagree, make them agree. They do not simply add up, and no maker says which wins. Never let them disagree, because you cannot predict the winner. In a clip, describe what must hold correctly, or not at all.
2.3 What every prompt settles
A prompt works when it settles the decisions the shot needs. Every decision left open is a decision the model makes for you, from its defaults. How many decisions you settle depends on what the frame or clip is for, and this book teaches two grades.
Test grade | Finished grade | |
What it is for | proving that a route works, that a look is reachable, or exploring cheaply | any frame or clip that will be delivered, or that other work is built on: the standard the book teaches |
A still holds | the medium and purpose, the subject and its frozen instant, the place, the camera and one named light source in words, and the closing line: roughly 60 to 200 words (a working figure) | all twelve elements below; a hero still lands near 400 to 600 words |
A clip prompt holds | a declared camera or one move, one action, an end and a hold, one protection line, and the sound setting | the seven elements below |
Numbers | words alone are enough | both words and numbers, each number beside the result it should give (2.5) |
Never left out | the frozen instant, the light source, the closing line; for a clip, the endpoint and the sound setting | any element the shot needs |
What it proves | a route, not a finished look | nothing: it is the frame or clip that ships |
Use the test grade to find out cheaply whether a model can do something before you spend on a finished attempt. Use the finished grade for anything that will be delivered and anything other work is built on. Never present a test-grade prompt as the model of a finished one, and never promote one by adding adjectives: rewrite it through the anatomy.
A finished still settles twelve things, in a fixed order: world first, photography second, technical close. The order is a writing and checking order, invisible in the finished prose. (1) What the image physically is and what it is for; (2) the subject and its frozen instant; (3) the identity clause; (4) wardrobe and props; (5) the frame inventory and its edges; (6) place and materials; (7) the must-show items; (8) the camera's placement; (9) the light; (10) the atmosphere; (11) the finish; (12) the exclusions and the closing line. Chapter 9 teaches each. No maker requires this order. The reason is craft: the medium and capture decide how every later noun renders, identity binds early, and camera, light and finish then describe a world that already exists.
A finished clip settles seven things, in a functional order: (1) the declaration, the camera move with its rig, or a locked camera with all the motion given to the world; (2) one primary action; (3) the world's life; (4) what is protected, by name; (5) the endpoint, a cuttable composition and its hold; (6) the pace, woven into 1 and 2; (7) earned exclusions and the sound line. Chapter 14 teaches each, and the tool chapters give each model's order.
Four points hold for every prompt.
Open with what the image physically is and what it is for. The first sentence names the kind of image, its era or production context and its capture, and its purpose in physical terms: "a hero frame for a premium athletic-wear commercial", "a photograph from a 1966 Italian film". Google says to explain the purpose of the image and OpenAI to name the intended use. That the earliest words weigh most is a production synthesis, not a measured law; the reason to open on the medium is craft. On an image endpoint the tool fixes the operation, so the medium opens the prompt; in a chat surface one operation phrase ("Create an image of…") leads in, because Google warns that a chat model may otherwise answer in text.
Keep the intention in the record and its physical result in the prompt. "Emotional beat: quiet awe" is not something a camera can see, and the model will draw a stock version of it. "Eyes down on the tea, mouth closed and relaxed" is. The story meaning, the emotional beat and the goal of the shot stay in the shot record (Chapter 8); the prompt stays physical, and whatever the frame must show has to be written into the prompt as well, because the model cannot read the record (house rule).
Keep the settings out of the prose. Aspect ratio, resolution, the number of results, the reference list and the sound switch are chosen in the interface. "8k" or "16:9" written into a still prompt sets nothing and can fight the real setting (house rule).
Keep one authoritative version of the brief. Never add a summary that restates the brief in other words: two versions give the model two places to contradict you, and it will not say which it followed.
Two more cautions, stated once. A keyframe for a clip is frozen early in the action, with each mover's direction written as positions and never as an open verb (Chapter 9). A keyframe of a face that will speak also states the mouth (Chapters 9 and 22).
Here is the pair in miniature: a finished insert still, and the clip made from it. The still carries everything a camera could see. The clip prompt says only what changes. The model chapters print more, in each tool's own dialect (Chapters 11, 12 and 16 to 18).
Written for this book · Food and drink · Nano Banana Pro, text only, 2k, 16:9, one variant · finished grade; not run
Example 2.1 — a finished insert still: the pour
A digital-cinema frame from a premium loose-leaf tea commercial, a hero insert with a large-format sensor look, shallow-focus falloff and a restrained professional grade, showing black tea leaving the spout of a dented brass pot into a small tulip-shaped glass on a marble table in an old Cairo café, caught about a third of the way into the pour: the pot is tilted about forty degrees, the stream is a smooth amber rope some twelve centimetres long, the glass is a quarter full with its surface just breaking into first ripples, and the first thread of steam has begun to lift from the rim. A man's right hand, seen from the wrist, holds the pot by its black wooden handle: dark hair on the back of the fingers, a plain steel watch, a short clean nail, the knuckles slightly whitened by the weight. The glass has a thin gold band round its rim; the pot is hand-hammered brass, dented near the spout, its tarnish worn bright along the handle bracket. In the foreground the saucer's edge crosses the lower right corner as a soft bright shape; in the midground the stream, glass and hand sit left of centre, the glass about a third of the frame height, with the pot's body cut by the upper left edge; in the background a green-shuttered window and the back of a bentwood chair dissolve into soft vertical shapes, and nothing else enters the frame. The tabletop is grey-veined marble with pale ring stains and one chipped corner. The frame must show the whole stream from spout to glass, the gold band on the rim, the dent near the spout and the watch on the wrist. The camera sits 25 centimetres above the table and 75 centimetres from the glass, angled 12 degrees down, so the spout falls just inside the top edge and the foot of the glass just inside the bottom, on a 100mm macro lens at T/2.8, holding the glass and the stream sharp while the shutters dissolve. Late sun through the window at camera left, about 4500K, lights the stream from behind and rims the steam and the gold band; a bare tungsten bulb out of frame at camera right adds a dim 2700K glow on the brass; fill comes from the pale marble, and overall contrast holds near 3:1, deep enough that the amber keeps its depth and soft enough that the shutters do not clip. Steam is visible only where the backlight catches it. Captured on digital cinema: fine, restrained grain, a neutral base tone, gentle halation on the backlit steam, softness only at the extreme corners. Warm amber and brass fill most of the frame, the grey marble and green-grey shutters the second field, and the black handle and deep shadow the smallest accent. No spilled drops on the table, no second glass, no lens flare across the stream, no overlaid text, subtitles, logos, or interface graphics.Why it works
Every number sits beside the result it should give: the 100mm lens at T/2.8 holds the glass sharp while the shutters dissolve, and the 3:1 ratio is chosen so the amber keeps depth without clipping the shutters.
The instant is written as positions and quantities (tilt, stream length, glass level, first ripples, first steam), so a clip has a starting state and a direction to run.
The hand is evidence, not feeling: hair, watch, whitened knuckles. The purpose opens the prompt in physical terms and the story meaning stays in the shot record.
Every plane and every edge is named, and the palette is written as proportions on named surfaces.
Written for this book · Food and drink · Kling 3.0, standard, 5 s, sound off, start frame = Example 2.1, 16:9 · finished grade; not run
Example 2.2 — a finished clip prompt: the pour continues
The camera pushes in on a low slider, advancing approximately seven centimetres over the five seconds, enough to swell the glass by a tenth in the frame, with a slight mechanical lag at the start and a firm, damped arrival at the end, never digital glide. The hand rights the pot from about forty degrees toward upright across the five seconds, thinning the stream to a thread that cuts clean at about three seconds, the amber level climbing from a quarter to about three-quarters of the glass by that moment while the ripples widen, then settle. Steam thickens and drifts up and to the right on the warm air from the window, catching the backlight in curls; the bulb's glow on the brass stays steady. Keep the glass's shape and gold rim, the pot's dents, the watch and the soft shutters exactly as in the start frame, and the tea's amber colour unchanged. The shot lands on the glass three-quarters full at left of centre, the surface almost still, the last steam curling from the rim and the pot righted, and holds for one second. No overflow, no second stream, no new hands entering the frame.Why it works
It says only what changes. Nothing the start frame carries (the room, the colours, the hand) is described again, and what the motion could break is protected by name.
One measured magnitude per moving system, each with its duration and its visible result: the camera's seven centimetres becomes a tenth more glass in the frame, because the glass is 75 centimetres away and the camera closes to 68.
It lands: a cuttable composition and a one-second hold. Sound is a setting (off), not a sentence.
2.4 Choose the writing format by the job
Prompts circulate in many forms: a line of comma-separated fragments, a paragraph, a numbered list, labelled sections, blocks of JSON copied from forums. The makers require none of them. OpenAI's advice is to choose the format that makes the requirements easiest to read and update, and Google's to describe the scene narratively. Choose the form by what the job has to carry.
Format | Good use | What carries the result | Trap |
One short edit sentence | add, remove or change one thing in an image that already shows everything | the input image | surroundings you did not state get redesigned |
One descriptive paragraph | one hero subject, or one coherent shot | ordered nouns, physical relations, the camera | chains of adjectives with no visible decision |
Reference sentences, then a paragraph | an actor or a product scene built from several sources | each input's job, stated before the style | vague or conflicting jobs |
A numbered panel list | a sheet or a grid whose cells genuinely differ | distinct assignments plus shared constants | near-identical cells; contradictory states of the world |
Labelled sections (scene, subject, details, constraints) | a layout, a packaging flat, a text-heavy brief, a multi-part edit | each obligation stated once, where it can be ticked off | a structured block plus a prose summary of it |
Structured fields (JSON-like) | software assembling many variants | fields that are easy to maintain | pretend parameters, a duplicated summary, irrelevant numbers |
The default for a photographic still is one prose paragraph, not a keyword string. A line of fragments is not wrong: it is the right form for open exploration, when nothing is fixed and the model's own defaults are welcome. For a grid, give each cell one distinct visible mechanism and state the shared constants once; shorthand leaves the model to choose. Each tool chapter's N.3 gives the format its model wants, and N.4 the template.
How to write a labelled brief
List the brief's obligations once, including the exact copy, and sort them under scene, subject, details and constraints.
Check that nothing appears twice or contradicts itself. Section names are not interface parameters, and a structured block plus a prose summary gives two places to contradict yourself.
Generate, then check each section's obligations one by one on the frame, reading every glyph at 100 %. The director approves.
2.5 Numbers, each beside its result
Planning numbers (focal length, stop, height and distance, colour temperature, contrast ratio, share of frame, palette shares, a move's magnitude) are a director's working language. At finished grade they go into the prompt, and each one goes next to the visible result it should give.
The reasons are in the makers' own pages. Google's example uses "f/1.8" and OpenAI's "using a 50mm lens", and no maker forbids numbers. But OpenAI also says to treat camera specifications as cues for appearance, not a guarantee of exact physical simulation. A number alone leaves the model to guess the look; a word alone leaves out the decision. Both together give it a cue and the result to aim at, and give you something to check on the returned frame. In the author's finished trailer (Chapter 35) lens and stop were written in 30 of 32 stills, colour temperature in 31 and a contrast ratio in 28.
The number | Written beside its result |
lens and stop | "on a 100mm macro lens at T/2.8, holding the glass and the stream sharp while the shutters dissolve" (Example 2.1) |
camera position | "25 centimetres above the table and 75 centimetres from the glass, angled 12 degrees down, so the spout falls just inside the top edge and the foot of the glass just inside the bottom" (height and distance in real units or in body-and-furniture units: either fixes a camera) |
a source's colour temperature | "Late sun through the window at camera left, about 4500K, lights the stream from behind and rims the steam" (Example 2.1) |
a contrast ratio | "near 3:1, deep enough that the amber keeps its depth and soft enough that the shutters do not clip" (Example 2.1) |
a share of the frame | "the glass about a third of the frame height" |
a move's magnitude | "advancing approximately seven centimetres over the five seconds, enough to swell the glass by a tenth in the frame" (Example 2.2) |
Treat every number as a cue for the look, never as physics, and never let one carry a decision alone. Keep the plan and its reasons in the record. Name a source by its role in the world (a bulb, a window, headlights), not by studio equipment: a softbox named in a prompt can turn up in the picture, a risk to check on the returned frame. At test grade, words alone are enough: "a low warm bulb, a portrait lens feel". For music, a tempo written "about 72 BPM" is allowed, because the maker treats tempo as approximate.
2.6 Borrowed words, and other people's prompts
You will meet prompts full of vocabulary that sounds precise and sets nothing. Translate each term into what it actually asks the camera to show before you use it.
Borrowed term | What it is in practice |
"Identity anchor" | the image you supply, not a guarantee that the face will hold |
"Chiaroscuro" | organised light and dark, not orange light: write the source and the shadow |
"Rim light" | separation at the edge, from a source that could plausibly be there |
"Frozen motion" | one chosen visible state, early in the action, with its phase written as positions |
"8K", "masterpiece", "ultra-realistic" | nothing: quality tags describe no look, and resolution is a setting; a house rule bans quality-tag spam |
"Cinematic", "Fellini-esque" | a genre or a person, not a look: say the format, the lens or the grade you mean; a house rule bans both in authored description |
A painter or photographer as style ("Caravaggio", "Brassai") | the same fault: write the light and the finish you mean |
A stock name ("Vision3 250D") | useful only if the name implies a visible behaviour (a daylight balance, reduced grain, a red halation); write the behaviour |
A camera body ("Arri Alexa") | less load-bearing than the lens, the stop and the light, and it can carry a look you did not choose; describe the capture |
--iw and JSON field names | text, unless the interface documents them as controls |
A named director's technique is a separate case. A house amendment permits a technique token (for example "Guy Ritchie speed-ramping") in Seedance 2.5 motion prompts only, strictly as the name of an editing or motion style. It stays banned as a visual-style adjective everywhere, and banned in Nano Banana Pro stills and on Kling (Chapters 11, 16 and 17).
Published prompts. Keep the original text unchanged beside your version, then translate. Settings written as flags go into the settings; @Name or [Character A] tokens become a real attached image or plain words; a negative field becomes earned exclusions in prose (2.7 sets the count); a lens or distance number stays only if a result sits beside it. A published prompt is evidence of how someone else worked, not a template. Typical traps in a copied structured prompt: --iw 2 is not a Nano Banana Pro control, "8k" sets no resolution, a JSON block plus a second summary gives two places to contradict, and field names are not interface parameters.
2.7 Exclusions and the closing line
Every model is told what not to do in two ways: a very few exclusions, each earned by a drift you expect, and one fixed closing line that keeps overlaid text, subtitles and stray music out of every result.

Figure 2.2 — Where each kind of "no" belongs. The prompt says what is there first, spends at most five exclusions, and ends on the closing line; silence and music refusals live in the settings.
Say what is there before you say what is not. A positive description gives the model something to draw; a negation hands it the banned word. Google's own example is "empty street" instead of "no cars". So "the woman, her glass of tea and the spoon are gone, and the table top where she sat is bare wood" beats "no people on the table, no tea".
Spend at most five, each aimed at a drift you expect; on a clip, work to three or fewer. Every generated world drifts towards its genre: in the author's Mastorna trailer nearly every prompt closes on locks against the period going modern ("no plastic overhead bins", "no modern signage"). That is what an exclusion is for. A long list reads as the generated house style and dilutes the one exclusion that matters, and the cap holds even where a maker's own examples carry long blocks. Where the maker documents a likely defect, name it as one targeted negative; it counts inside your five. Five is the ceiling for any prompt, still or clip (a house rule of July 2026). On a clip the working number is three or fewer, because a prose exclusion is the weakest control there (below) and each one adds a fault to check on the returned file: that is advice, not a limit.
Keep exclusions physical and clean. Word each as a checkable fact about the frame ("no second face", "no hand near the mouth"). Never name an emotion in a negative ("no friendly customer-service expression" hands the model the banned idea), never fight identity ("no asymmetric face" erases the marks that make a face this person), and never add quality tags ("hyper-realistic, ultra detailed, 8K HDR" does negative work on a reasoning model; "no AI look" is not a direction anyone can photograph). A stack of a dozen negatives copied from a forum does nothing on a reasoning model and fights the asymmetries that hold a likeness: negatives correct observed, persistent artefacts; they are not a ritual.
Never let a negative stand in for the camera. "No camera movement" alone leaves the camera undecided; declare it first, in positive words ("The camera is locked."). When the model would otherwise invent coverage, lock the structure in words, in the model's own form (Chapter 18 gives Wan's).
Expect prose exclusions to be weak on a clip. A "does not tip or turn" may not stop a box rotating, and naming an off-frame object can put it in the frame. Treat these as risks to check on the returned clip, and describe what does happen instead.
The closing line
Stills end on a fixed closing line, and it sits outside the five: "no overlaid text, subtitles, logos, or interface graphics" (a house rule). It bans the overlay layer only; whether the world carries legible text is a per-shot production-design decision. Titles, subtitles, supers, end cards and logo lockups are composited in post from real assets (Chapter 30); a world that carries a sign or a shop front writes that text as production design and never writes "No text" (Chapter 10). The book counts five drift locks and then the line.
A clip ends on its model's own sound and subtitle line. Native sound is set explicitly on every job, because a model that is not told decides for itself, and the wording is the maker's where the maker gives one. Make silence a setting, never a sentence. Each tool chapter gives the line for its model (N.3) and where its exclusions live (N.6): the image and video models the book uses expose no negative field, so exclusions live in the prose; on Suno, refusals go in the Exclude Styles field; Sync Lipsync 3 has no prompt at all.
2.8 How long
Complete, not padded. There is no universal floor and no maker's word limit. A prompt is complete when it satisfies the grammar of its workflow, keeps every required anchor and state, and passes the element audit without padding or contradiction; completeness outranks length (house rule). The 32 stills of the author's finished trailer measured 390 to 598 words with a median of 500, a useful target for complex, identity-sensitive cinematic stills and not a universal law. Every other figure below is a working figure.
Job | Working band | Standing |
A complex, identity-bearing hero still, all twelve elements | about 400 to 600 words | measured in finished work: 390 to 598 |
A single-subject insert or an environment | about 150 to 300 | working figure; an insert that settles all twelve elements, like Example 2.1 (about 490 words), belongs in the row above |
An edit of an approved frame | about 40 to 150; the preserve list decides | working figure; Google's own edit examples are one line |
A test-grade still | about 60 to 200 | working figure |
A Kling clip that is a camera move through a populated set | about 120 to 200 | measured in finished work: 137 to 177 words, median 152 |
An action-led Kling clip on a good frame | fewer | the frame carries the rest |
A Seedance clip | compact, about 55 to 130, unless it carries several beats | a drafting range |
Words buy control until they start repeating what an input or a setting already carries; after that they buy nothing and add places to contradict yourself. Below the band, audit for a missing element; above it, audit for repetition. Declare a band per shot in the record. For a scene too complex for one pass, Google's staged build ("First, create a background…") is the maker's alternative, at the cost of edit drift (Chapter 10). Seedance's maker warns that too little plot lets the model improvise, and that too much packed into a time range may drop parts of the plot; Wan's page says a single sentence is enough to make a video, but the more complete and precise the description, the closer the result. The later in a production, the shorter the prompt, because the inputs carry more (Figure 2.1).
2.9 Where to stop prompting
Stop prompting when the input, not the wording, is missing. More words help only an ambiguous instruction. They cannot create a missing reference, and they cannot supply a control that the route does not have. Reword once, changing one thing; if the same defect returns, ask which of the three things is missing and supply that.

Figure 2.3 — Where to stop prompting. When a reworded attempt returns the same defect, more adjectives will not help: find which of the three things is missing and supply that.
The table gives the commonest cases, for any model.
Defect | First action | Stop prompting when | Hand on to |
The wrong identity in a good scene | give the identity picture its authority again; remove competing faces | the face still differs in a view that matters | a new accepted frame, or another route |
A wrong profile the model has never seen | make and approve that view first | a real actor's likeness cannot be reached | an approved side view, before any motion |
A wrong grip or contact | fix what supports what; add a blocking photograph | a reworded attempt returns the same grip | a new keyframe, or different coverage |
A product that keeps its wrong shape | supply real photographs of its front and side | the model keeps redesigning it | a photographed or composited product |
A scene drifting after several edits | return to the last approved frame | you are editing the damaged copy | one new edit from the accepted source |
Chapter 9 gives the stills version of this table, with the cheapest repair of all, a composite in post. Chapter 29 teaches the whole ladder of repairs across a production, and when to stop repairing a shot altogether.
2.10 When the reader is an assistant
Directors increasingly brief a language model (an AI assistant such as Claude or ChatGPT) to draft treatments, shot lists and the prompts themselves, with the director approving each step. A language model is a different reader from an image model. The practices below come from software work with language models and from published reports, not from film prompts; read them as other people's experience.
Structure the brief in labelled blocks: role and standing rules first, then the task with its limits and the form of the answer, then context or examples, then the request.
Say what happens next at every stage, and check each against criteria set in advance. Many-stage prompts are reported to stop early or skip a stage when the next step is left implicit: end each stage with an instruction to continue or, when the director is the gate, to stop and hand the result back, and ask for a pass or fail with what is missing named.
Separate exploring from deciding. Asked for one answer, a model converges too early on something safe: ask first for several widely different ideas with no constraints, then choose one and develop it against every requirement.
Keep working notes apart from the paste-ready text, ask for concrete physical detail and state the form of the answer. The numbers the assistant proposes belong in your shot record; how they enter the prompt is 2.5.
Imperatives, conditions and checklists help a language model follow rules. An explicit "Do not …" works on a language model, whereas an image model handed a "no" is handed the forbidden word (2.7).
One rule protects the director's authority: a conflict the assistant resolved by itself is returned to the director as a decision, never settled quietly.
Template: one stage of a brief to an assistant (written for this book; not run)
Role and standing rules: [who the assistant is on this job; the rules that hold at every stage, such as "never change a signed frame" and "return any conflict to me as a decision"].
This stage: [the one task of this stage, and nothing else].
Inputs you may use: [the treatment, the shot list, the signed frames, the locked takes].
Form of the answer: [the table and its columns; each paste-ready prompt in its own block; your notes kept separately and labelled as notes].
Before you hand it back: [the criteria set now]. Report each as pass or fail, and name what is missing.
Then stop and wait for my approval before the next stage.Open questions
How a model reads camera height and distance in real units, how it weighs the order of the elements, and whether 20 to 40 per cent is the best freeze for a keyframe. Until this is known, write each number beside its visible result and check the returned frame against it.
Whether labelled sections beat one paragraph for a photographic frame on GPT Image 2.5. Until this is known, use one paragraph for a photographic frame and labelled sections for layouts and text-heavy briefs.
What to remember
The model has only this job: your words and what you attached. Nothing carries over.
Each model is a different reader; write for the one in front of you, in the dialect its tool chapter teaches.
The inputs own what they carry; the words carry the rest. A call that makes or changes pixels needs the full description; a clip from a signed frame protects it and adds the change. Never let words and inputs disagree.
Two grades: the finished grade is the standard, and the test grade only proves a route. Never present a test-grade prompt as the model of a finished one.
A finished still settles twelve elements in order, world first, photography second, technical close; a finished clip settles seven and always lands.
Both words and numbers, each number beside the result it should give; numbers are cues for the look, never physics and never alone.
Keep the intention in the record and the settings in the interface; write each instruction once.
Translate borrowed words before you use them; keep a published prompt's original beside your version.
Say what is there first; at most five exclusions (three or fewer on a clip), each aimed at a drift you expect; the closing line last; silence in the settings.
Complete, not padded. When a reworded attempt returns the same defect, stop writing and change what you give the model.
Next: Chapter 3 · The concept




Comments