Chapter 22 · The speaking shot: two routes
Part 4 · Voice and speaking faces
A visible face saying a line is made in two halves, a voice and a picture, and the join decides which audio the client hears. This chapter teaches both ways of making the join, with no default between them: the exact take, where your locked recording ships and a lip-sync pass fits the mouth to it, and the new performance, where a video model performs the line. You will learn to design shots so that few mouths need help, to make the talking plate a pass needs, to write the prompt for a performance, and to check the result before anyone cuts it.
In this chapter
The two routes, the questions that choose between them, and what each costs
The contract line, and designing lip-sync exposure down before the shots lock
The talking plate, and how to time it to a take
Route A, the exact take, and Route B, a new performance
The checks after the pass or the performance
Before you start. Chapter 21 makes and locks the take; Chapter 8 (8.2) decides the route for the project; Chapter 9 (9.10) teaches the speaking keyframe; Chapters 14 and 15 the moving shot and the acting. The tools are Kling (16), Seedance (17), Wan (18), Veo and MiniMax (19), ElevenLabs (23) and Sync Lipsync 3 (24).
22.1 What this stage decides, and who decides it
A speaking shot is one in which a visible face says a line. It is made in two halves. The voice comes first: a native voice, directed, repaired and locked by ear (Chapter 21). The picture comes second, from a keyframe the director has signed (Chapter 9). This chapter is the join, and the join decides the one thing a client will hear: whose audio ships.
The director decides it on the first day of the project, before any picture is paid for, and writes it on the project card (8.2). There is no default. Two routes are equally legitimate.
Route A, the exact take. The locked take ships. You make a talking plate for it, a clip of one face talking quietly with no words, and Sync Lipsync 3 fits the mouth to the take (Chapter 24).
Route B, a new performance. A video model performs the line from the signed frame and the voice, and its audio ships once you have judged it on the clip. On an English project Seedance 2.5 can speak the line itself in the chosen voice. Egyptian and Saudi dialogue always comes from a voice actor or ElevenLabs: the model performs a voice you cast and never invents a dialect from text.
Choose A when the exact reading is approved, the line recurs, captions and repairs must follow a waveform you own, or dialect or legal accountability decide; choose B when a fresh rendition is acceptable (8.2 sets the reasons side by side). You may decide by shot class, hero lines on A and background lines on B, but write the classes down.
Word | In this book |
Take | the locked recording of the line, prepared for the pass |
Talking plate | a clip of one face talking quietly with no words, made for a lip-sync pass to redraw; "plate" alone in this Part means this |
Pass | one Sync Lipsync 3 job: a video and a take go in; the same video comes back with a new mouth |
Speech window, talk window | where the speech starts and ends in the take; where the actor's own lips move in the plate. Both are measured, and the speech should start where the talk starts |
Hold | the frames after the last word, where the face stays still; the edit trims it |
Exposure | how much a shot asks of lip-sync: the mouth's size in frame, its angle to the lens, the length of the line |

Figure 22.1 — Two questions, in this order, before any picture is made. Costs are per 5 s at 720p: the plate and the performance as quoted on 27 Sep 2026, the pass as charged on 12 Sep; check them on the day.
What each route costs (per 5 s at 720p; checked 27 Sep 2026; check on the day):
Route | Credits | What ships |
A: Kling 3.0 standard plate, sound on, plus the pass | 8.75 + 18.15 = 26.9 | your take |
A: Kling 3.0 standard plate, sound off, plus the pass (not yet run with a talking plate) | 6.25 + 18.15 = 24.4 | your take |
A: Seedance 2.5 720p silent plate, plus the pass | 35 + 18.15 = 53.15 | your take |
B: Seedance 2.5, 720p, its own sound on | 35 (6 s: 42) | the model's audio |
B: Wan 3.0, 720p, its own sound on | 8.75 | its own track |
Route A costs two jobs for each visible mouth; Route B costs one job and a second judgement of every word on the clip. Chapter 34 turns credits into money.
Write the route for every speaking shot before any picture is made, and change it only as a written decision with a reason. A shot that was judged as an exact take never quietly becomes a new performance, because that changes what the client hears.
22.2 Write the contract before any picture
The contract is one line in the shot's record. It also says how the prompt uses the voice: on Route A the take ships and the plate prompt carries no words; on Route B the take, or the chosen voice, is a reference and the picture prompt performs the line.
Template: the contract line (written for this book; not run)
Shot: <id> · Contract: exact take / new performance · Take or voice: <lock file and measured length / voice reference and language> · What ships: <the take, re-encoded / the model's audio> · Judged on: <the take's verdict carried, and the pass / the clip> · Captions timed to: <the take / the clip's own audio> · Changes: <date, from → to, reason, approved by>Whose voice ships. The voice in a client final is a person who signed: a consenting native actor's recording first; a clone only of that person, under a written release; never a clone of a library voice or of a public clip. A library voice is for casting, screening and animatics, and enters a final only after the checks of Chapters 21 and 34. Higgsfield's own voice-change and dubbing tools are not used: they name no engine.
The whole route for one spoken line runs in eight steps. Steps 1 to 3 belong to Chapter 21; this chapter teaches steps 4 to 8.
# | Step | What it produces |
1 | The identity and words lines of the dialogue-ledger row | an approved line; nothing is voiced without one |
2 | Casting, audition and the rights record | the voice, and its source class |
3 | The take prompt, repairs, lock and measurement | the locked take and its speech window (on Route B with no take, the chosen voice reference) |
4 | The contract line and the exposure design (22.2, 22.3) | exact take or new performance; the strategy and its fallback |
5 | The plate (22.4) or the performance prompt (22.6) | one clip |
6 | The pass (Chapter 24) or the performance job (Chapter 17) | the synced or performed clip |
7 | The checks (22.7) | a verdict from eye and ear |
8 | The captions, timed to the audio that shipped | timed captions |
The test grade may use one voice and two takes, the director's ear labelled "not certified", and checks limited to playback and frame-stepping. It may never drop the approved line, the contract line, or the rule that a library voice, a public clip or a clone of a stranger never reaches a client.
22.3 Design the sync out first
The cheapest lip-sync is the one you never run. A pass can move a mouth, but it cannot add a head, eyes, breath or thought, and no source establishes a model that guarantees an imperceptible edit on every face and angle. So decide, shot by shot and before the shots lock, where a visible mouth is worth its risk.
Tag every shot that might speak with its line's status, its exposure (none, low, medium or high), a preferred strategy and a fallback if the sync fails. Exposure rises with the mouth's size in the frame, its angle to the lens and the length of the line, and falls with a trembling, whispered or thrown-away delivery. Figure 22.2 sets out the six strategies; the last is kept for the moments the story needs the face.

Figure 22.2 — Six ways to keep a mouth out of trouble, from none to most exposed. Every shot with possible speech takes one, and a fallback if the sync fails. The grades are those of the Mastorna shot table.
Set a written ceiling, and put a break shot between visible moments. A ceiling is a project decision: "visible sync is limited to two short moments in the locked cut" is the shape of one. Between the two most exposed moments plan a break, a shot with voices or atmosphere and no visible mouth to match: an audience, a reaction, a wide.
Write the tolerance. The craft transfers from any era; the tolerance does not. A period picture such as the Mastorna trailer can be dubbed in post (its brief asks for speech "historically dry, close, and post-dubbed rather than modern-naturalistic"). A modern sync-sound commercial has no such licence, so the same shot is low exposure in one and high in the other.
Let the maker's claims prompt a test, not skip the design. sync.so documents that profile shots and over-the-shoulder angles "are now handled with confidence" and that Sync Lipsync 3 "detects and generates around obstructions without manual intervention" (checked 30 Sep 2026), while its media tips still say "Avoid profile views and obstructions covering the face". Test the actual pose on a spare plate; do not keep a difficult mouth in the frame because of the claims.
From Mastorna, corrected · trailer, shot 4.3, the hotel desk · Kling 3.0 standard, start frame, sound off, 5 s, 16:9 · low-medium exposure, voice post-dubbed in the edit, so no plate and no pass · corrected for this book; not run
Example 22.1 — four clipped words, in partial profile
The camera pushes forward on a dolly by about ten centimetres over four seconds, from the over-the-shoulder framing toward the porter's face, with a soft start and a soft stop, and settles for the last second. The man in the foreground leans a few degrees toward the desk and speaks four clipped words in his left profile, his jaw opening and closing four times in about a second and a half, his mouth turned away from the lens. Behind the desk the porter extends his right hand with a brass room key over about two seconds, and his professional smile widens slightly over a second until a flash of gold teeth catches the warm light of the desk candle. The candle flame between them flickers, and the shadows on both faces shift softly. In the deep background a small figure crosses the dark corridor from right to left, far back and out of focus. Keep the man's left profile and jaw, and the porter's smile and gold teeth, clear and true throughout the push, and the key a single object. The clip lands on the two figures across the desk, the key held out between them and the corridor dark behind, and holds there for the last second. No captions, no subtitles, no on-screen text.Why it works: the mouth is a small part of what the shot shows. The jaw is written as a count and a duration, the profile keeps it turned from the lens, and the porter's smile and the candle carry the rest.
What to copy: the push stated as a rig with one magnitude and its duration; protection by name for the faces that will be watched, placed after the action; a landing composition and a hold; a closing line that matches the sound switch, which is off.
The route: fallback a full rear view that removes the mouth. In a period picture the voice is laid in the edit, so there is no pass; a modern commercial with the same shot would make a talking plate.
The dialogue ledger holds the picture line. The ledger of Chapter 21 has one row per spoken moment, in four lines: identity, the words, the voice and the picture. This chapter writes the picture line, with one extra field, the performance brief. Above the rows sit the avoidance strategy, the visible-sync ceiling and one line stating the tolerance.
Template: the picture line of the dialogue-ledger row, with the performance brief (written for this book; not run)
Line ID <id> · Shot <id> · Speaker <name> · Listener <name, or off-screen> · Exposure <none / low / medium / high> · Strategy <off-screen / profile / rear / occlusion / wide / visible sync> · Framing <size, angle, distance> · Fallback if sync fails <strategy> · Contract <exact take / new performance> · Route <plate model + Sync Lipsync 3 / Seedance / Wan> · Captions timed to <the take / the clip's own audio> · Brief <the state; what the body does; the address; the pace>Write a performance brief for each visible beat: one physical line with the state, what the body does, the address and the pace. It becomes the bracket for the voice (Chapter 21) and the prompt for the plate's face (22.4). The register: "delivery fast, offhand, almost thrown away"; "whisper: intimate, close-mic'd, stammering".
Written for this book · telecom and tech · the contract line and the picture line for the line of Example 24.1 · Route A · not run
Example 22.2 — one line, recorded before the picture
Shot: 12 · Contract: exact take · Take or voice: T-04_lock_v02.wav, 4.6 s padded (1.2 s head, 3.0 s speech, 0.4 s tail) · What ships: the take, re-encoded · Judged on: the take's verdict carried, and the pass · Captions timed to: the take · Changes: none
Line ID T-04 · Shot 12 · Speaker Adel · Listener Karim, on the phone screen · Exposure medium · Strategy visible sync · Framing medium close-up, near-frontal, 1.4 m · Fallback if sync fails: the line as a voice over Karim's face on the screen · Contract exact take · Route Kling 3.0 standard plate, sound on, + Sync Lipsync 3, cut_off · Captions timed to the take · Brief: tired and relieved; leans back into the sofa and holds the phone at chest height; eyes on the screen; brisk, the last word thrown away
Canonical line: «شايفاني كويس دلوقتي؟ الصورة نضيفة أهي!»Why it works: the contract names what ships and whose verdict it carries, so the edit knows the take is the timing truth; the picture line records the exposure and a fallback before a credit is spent.
What to copy: the take's three lengths written as a sum; the fallback stated as a strategy; a brief a plate prompt can compile (a lean, a held phone, the eyes on the screen).
Check: every shot that might speak is tagged and has a fallback; the ceiling and the tolerance are written; no two visible moments meet without a break.
22.4 The keyframe and the talking plate
The keyframe. Chapter 9 (9.10) teaches the speaking frame: the mouth state written into the prompt, a face large enough to judge, an unobscured mouth under even light, one speaker, nothing crossing the lips, room for a landing. On Route A that frame is frame 0 of the plate; on Route B it is the mouth the model starts from.
Build the plate for the pass. One face, fully visible, frontal or near-frontal (sync.so documents that Sync Lipsync 3 handles profiles and over-the-shoulder angles, but that front-facing angles still produce the best results); mouth, chin and cheeks unobscured and evenly lit; lips neutral or barely parted at the start; a calm or locked camera, because body motion competes with facial synchronisation; nothing crossing the mouth. One line or breath group to a clip. After the line, a physical landing: lip-sync with no acting beat after it looks like an interface demo.
Prompt the plate to talk quietly, with no words, for about as long as your line: "talks quietly toward the screen for about three seconds". A plate whose mouth stays closed can still be synced, but the maker says Sync Lipsync 3 "can open silent lips to match audio, though results are generic rather than speaker-style matched". A plate that already moves like speech gives the new mouth a face that is speaking, so never tell it to keep its mouth rigidly closed. Never give the plate the words: the plate model would perform its own line, and Kling turns dialogue in unsupported languages into English.
Time the plate to the take before the pass, because the pass will not. It starts plate and take together at the first frame and redraws the mouth for as long as the audio lasts, with no offset between the files, so the speech lands where the take's own head puts it. If the plate leans, taps and only then talks, and the speech starts 200 ms in, the redrawn mouth speaks over the business. Line them up on purpose.
Measure the plate. Note when its lips first move, when they stop and when the mouth is closed again: the talk window. A time written into a Kling prompt is an ask, not a control.
Put the speech where the talk starts. Pad the take's head: add silence (or the low room tone under the take) until the first sound falls on the talk onset, and keep the 400 ms tail; the mouth is then asked to match a silent stretch while the business runs, and the pauses check of 22.7 looks at whether it stays still. Or trim the plate to open on its talk onset and keep the take's 200 ms head; the business you cut belongs to the edit, taken from the plate's own frames. A trimmed clip must be uploaded, so it can no longer go in as the finished job.
Check the sum. Head, speech and tail together must be 0.3 to 0.5 s shorter than the plate (Figure 22.3). A five-second Kling plate returns as 5.04 s, so a 4 s line with its 0.4 s tail leaves at most a third of a second for a head: a long line on a five-second plate has no business before it. Shorten the line, keep the business small and under the talk, or take the beat from the edit.
Write both windows in the shot record, measured, and which way you lined them up.

Figure 22.3 — The clock of Example 24.1. The take's head is padded to 1.2 s, the measured talk onset of the plate, so the first sound falls where the lips first move. Seconds run left to right; the scale is the same for the three rows.
The landing has the same limit: after the last word the pass shows only the take's tail. Pad the tail inside the same sum, or take the rest from the plate's own frames in the edit.
Make the plate 0.3 to 0.5 s longer than the padded take, at your delivery size. The result keeps the plate's frame size: Kling 3.0 standard returns 1284 × 716 and pro 1912 × 1080 (checked 27 Sep 2026); Seedance 2.5 returns 1280 × 720.
Which model makes the plate. Kling 3.0 standard, five seconds, your keyframe as the start frame, sound on: 8.75 credits, the cheapest plate (sound off is 6.25 and is not yet run with a talking plate). The pass discards the plate's own sound and lays the take, so plate sound follows what has passed: Kling 3.0 on, Seedance 2.5 off, Wan 3.0 off (24.3). Seedance 2.5 at 720p is the second choice, when Kling's acting or face fails twice or the plate must match Seedance neighbours (35 credits); 480p is for screening only. Wan 3.0, Veo 3.1 and MiniMax H3 have not made a plate. Seedance's maker delivers 1080p as 10-bit HEVC, so such a plate may need conforming first (Chapter 30). Until Kling 4.0 ships (announced 28 Sep 2026, due in October), plates are made on 3.0 (16.10).
Two grades. The test grade proves a route cheaply: the shot and camera, who and where, the business, the quiet talking mouth, the landing, the face protection and the hygiene line; it is right for a line that will not ship. The finished grade is written on the seven elements of Chapter 14, with one approximate magnitude for each moving system. Templates are in 24.4 and three complete plate prompts, timed to stated takes, in 24.5.
Before the plate goes to the pass
"Talks quietly … for about [n] seconds", with no words; business before, a landing after (Chapter 15).
One face, fully visible and lit; nothing crosses the mouth.
The plate 0.3 to 0.5 s longer than the whole padded take, head included.
The talk window measured, and the take's head padded (or the plate trimmed) so the speech starts where the talk starts.
The plate watched, and approved as a performance, before the pass.
After two flat plates on one model, go to the other; after two on both, follow the stop rules of 22.5.
22.5 Route A: the exact take
Lock and prepare the take. Chapter 21 makes the take and locks it by ear. For the pass it must be clean speech from a single speaker with no music, prepared the same way every time, to the figures of 24.6 (with more silence at the head when the business comes first, 22.4).
Make the plate, then run the pass (Chapter 24): the plate, the take, and sync_mode set to cut_off, with no prompt. Never leave the default, which reverses the actor's motion when the take outlasts the plate. The result returns at the plate's frame size, cut to the take's length, with your take as the audio, re-encoded. Watch it at real speed, then run the checks of 22.7.
Fix each fault where it lives. A pronunciation fault is in the take: re-lock it with the spelling repaired (Chapter 21) and run the pass again. A dead face is the plate's: re-plate. A missing first sound is the audio window: restore the head silence. Never spend video credits on a pronunciation fault, and never regenerate an accepted take because the picture is weak. Treat a changed line as a new take.
Symptom | Cause | Smallest fix |
A word sounds wrong | the take | re-lock the take |
Sync right, face dead | the plate | re-plate (22.4) |
The first sound is missing | the audio window or its trim | restore the head silence and the onset |
Head and jaw fight the new line | the approved clip spoke the old line | a talking plate of the same shot |
The mouth is blurred or looks pasted on | the face is small, turned or occluded | a tighter, nearer-frontal plate with nothing at the mouth; else redesign the shot (22.3) |
Stop rules. If two passes on one plate fail, re-plate. If two plates fail, redesign the shot (22.3), record to picture (below), or, only where a new performance is acceptable, change to Route B as a written change of contract.
Captions follow the take. They are timed to its speech window, measured in the delivered clip, and their words come from the canonical line, never the engine spelling. Premiere's speech-to-text does not list Arabic, so Arabic captions are typed or pasted from the script (Chapter 30).
A native actor's recording as the take. Where dialect, acting or legal accountability decide, the take is a consenting native actor's recording. Either the actor records first and the plate is made to the recording's length, or the actor watches the approved clip and records to picture, matching its rhythm, and the pass fits the mouth to that recording. Recording to picture changes the voice route, so it is a recorded decision with the actor's consent, and its fallback is a new plate and a new pass with the actor's take. The new take becomes the caption authority. This has not yet been run: try it on a scratch clip first.
A line changed after picture lock. A new locked take can replace a line on an approved picture without regenerating the shot. Supply a new complete take, never a pasted word, because the stress of the whole sentence changes with the wording, and compare its length with the old speech window. The approved clip already speaks the old line, so its head and jaw rhythm may fight the new mouth; where they do, or the new line runs longer, make one longer replacement plate in the same continuity state. Do not force a length match by retiming the video. Re-time the captions and any music ducking, and name the new take's version in every dependent file. This route has not yet been run.
22.6 Route B: a new performance
Use this route when a new performance is acceptable: previz, a shot where the client accepts a fresh rendition, or an English project where the model can speak the line itself. Seedance 2.5 is the model (Chapter 17); picture and sound come from one job.
The audience hears the model's performance of the line, not your take. The voice and the words are meant to survive; the timing is the model's own; the returned sound is new audio, not your waveform. That sound becomes the timing truth for the edit and the captions.
Two ways to give the model a voice. The maker's notation (checked 30 Sep 2026) puts dialogue in curly braces, sound effects in angle brackets, music in parentheses and subtitles in corner brackets, and advises naming the language before any dialogue that is not Chinese. Its English example: <Lina> speaks in natural American English, using a young, soft, curious child's voice: {It's showing us the way.}
Describe the voice, and write the line. On an English project the model speaks the line from text in the voice you describe: age, weight, pace, register. It suits a voice-over-like or background voice, where the voice need not be one particular person.
Attach a voice reference and bind it. The maker defines an audio reference as one that "References audio information such as music, dialogue, voice, tone, or timbre", and binds it in the prompt: "Image 1 depicts the protagonist John and uses the voice timbre from Audio 1." Clips run 2 to 30 s, WAV or MP3, up to 15 MB; five to ten seconds of one voice works better. The reference is the chosen voice: an ElevenLabs take, or a consenting actor's recording.
Egyptian and Saudi dialogue takes the second way. The voice is a take from an actor or ElevenLabs, attached as the audio reference, and the model performs it. Never ask a video model for Arabic from text: Kling translates dialogue in any language outside Chinese, English, Japanese, Korean and Spanish into English (checked 30 Sep 2026); Seedance 2.5 lists Arabic among eleven languages, but a language list is not proof of a dialect; Wan 3.0 gives no list of spoken languages. Judge the clip word by word, because a repaired word can be lost in a new performance. If one is, regenerate once, then move the shot to Route A in writing, which keeps your take.
What the prompt carries.
Protection, then binding, then the line. The frame owns face, room and light. Name who speaks, bind the voice to the speaker, name the language (for Arabic, the dialect, in words) and write the line once.
Business first when you want a beat before the line, in order: "she sets the glass down, and only then speaks". With nothing before the words, the line starts on frame 0. Hold the timing by order, not by quoting the line's words.
Write the mouth: moving with the words, still in the silences and after the last word. Give two to four visible cues and place the eyes (Chapter 15). Keep ElevenLabs bracket tags out: they direct another tool.
Close the sound: two or three real sources of the room, only the dialogue and those sounds, no music. The maker's forms are "No BGM; generate only environmental sounds and action sounds." and "No subtitles." Keep the model's sound on: a clip made silent with a take laid over it has a mouth timed to nothing, and that is a plate for Route A's pass.
Design the keyframe as the state before the business, and buy duration for the line, its business and the hold (Seedance clips run 4 to 30 s). A gesture the frame already shows finished will not happen.
The subtitle warning. The maker warns that business tied to a line's words can trigger burned-in subtitles: "Avoid repeating dialogue words after a line or attaching separate tone, facial-expression, or action instructions to specific words." Its form for different deliveries is "Character's line (emotion): content". If subtitles appear, move the business outside the quoted line and use clean references.
Template: the speaking lines of a Seedance motion prompt, finished grade (written for this book; not run)
Image 1 is [Name], the first frame exactly; [one clause of what the frame carries, no more]. [The camera: locked, or one named move with its size and duration.] [Business before the line, as ordered beats.] [Name] uses the voice timbre from Audio 1 and speaks [language, and dialect if Arabic] in [voice: age, weight, pace]: {[the line, once]}. [Name]'s lips and jaw move with the words and are still when [he/she] is not speaking; [the eyes go to ...]. [The listener's task, mouth closed.] After the line [the mouth closes, the breath releases, the gaze holds]. The clip lands on [the composition] and holds for the last [n] seconds. Sound: [two or three real sources of the room]; only the dialogue and those sounds, no music. No subtitles, captions or on-screen text.Where a voice must be one particular person, or Arabic is spoken, drop "in [voice]" and attach the reference. Chapter 17 (17.5) has complete Seedance examples.
Other models. Kling 3.0 speaks its five languages from text with its sound on and takes no audio input on the connector, so no cast voice can be given to it there (the maker's own site can bind a voice tone to a character element, a different route): it is unfit through the connector when a particular voice must ship, and never for Arabic. Kling 4.0 (announced 28 Sep 2026, due in October) lists dialogue in Chinese, English, Japanese, Korean, Spanish, Portuguese, German and French "and more", and voice references among its inputs (checked 30 Sep 2026); Arabic is on no list, so treat it as unproven until auditioned (16.10). Veo 3.1 is evaluated by its maker in English only, with its sound always on and no audio input (19.3). MiniMax H3 is not a speech route. Wan 3.0 is the economy version: whether your take or a new rendition comes back must be settled by comparing its track with the upload (18.3).
What goes to the edit: the probed clip, and a record with the voice's version and spelling, where the line landed, the hold, and the listening verdict. The edit cuts to the clip's own sound, and the captions are timed to it.
22.7 Checks after the pass or the performance
A completed job is not a pass. The pass changed the mouth, and possibly the face around it, for the length of the take: check that region and that stretch of time closely, and the rest against the plate. The order is normal playback first, then the onsets, the closed-lip consonants, the pauses, the teeth and the first and last frames, compared before and after. A stronger mouth match can still be a worse shot.
The checks after the pass or the performance, in order
Normal playback: watch at real speed before anything else.
File: the frame size is the plate's; the length about the take's; 24 fps.
Audio: on Route A, decoded and compared with the upload for content and timing (expect a re-encode that matches closely, not a new reading); on Route B, listened to as new audio: every word present and right, ع ح ق ج, stress and vowel length, the same voice, no extra sounds.
Onsets: the first visible articulation against the first audible sound.
Closures: the lips meet on every ب and م.
Pauses and the hold: no mouth movement in the silences; the mouth closes after the last word.
Anatomy: no tearing at the corners, no blurred teeth, no pasted-on mouth.
Identity: the face still matches the master; on Route A the eyes, head, breath and hands are the plate's.
First and last frames, and the handle: where the plate's own tail can carry the landing, noted.
Side by side with the plate (Route A), then the native ear on the audio.
Again in the final export, in the player the client will use.
The hard fails on Route B are a wrong or missing word, a changed voice, a drift towards classical Arabic, extra speech, and burned-in subtitles. Do not let a measurement pass a shot: the eye decides. A failed check is traced to the plate, the take or the mode, and never fixed by running the same job again. Write the verdict into the shot record.
A rough cut that carries a synced shot is labelled "DRAFT — NOT FOR APPROVAL" while the checks run, and "ROUGH CUT vN — FOR EDIT DECISION: [the question]" when a decision is wanted, for example "ROUGH CUT v3 — FOR EDIT DECISION: does the mouth in shot 12 hold at 200 %?" Every export carries its version.
22.8 Which tool for which job
Job | Tool | Note | Section |
The take: a cast voice, a consenting clone, a converted read | ElevenLabs, v4 and v3 | audition both on each new Arabic voice | |
A talking plate | Kling 3.0 standard | cheapest; sound on; never the words | 16.3 |
A talking plate when Kling fails twice, or must match Seedance | Seedance 2.5 720p, sound off | 480p for screening only | 17.3 |
The pass | Sync Lipsync 3 | three fields, no prompt; cut_off | 24.3 |
A new performance, English speech in a described voice | Seedance 2.5 | the maker's dialogue notation | 17.3 |
A new performance from a take (Egyptian, Saudi) | Seedance 2.5 | the take is a voice reference; judged word by word | 17.3 |
An economy new performance | Wan 3.0 | audio semantics unsettled; not a route for Arabic | 18.3 |
Native speech in five languages, no cast voice | Kling 3.0 | never Arabic; Kling 4.0 announced | 16.3, 16.10 |
A sound-led beat with English speech | Veo 3.1 | sound always on; no audio input | 19.3 |
22.9 The records this stage writes, and the gates
In the shot record: the contract line, the exposure and strategy with its fallback, the plate's talk window and the take's speech window with which way they were lined up, and the caption line. In the run record: the mode, the plate and take versions, the frame count and raster, the charge, and the verdict of 22.7. Three gates apply: LOCK, because the take is locked before the plate is made; SPEND, because a plate and a pass are two charges and the route is written first; and ACCEPT, the director's eye and a native ear, then again in the running cut (Chapter 29).
Open questions
A head padded to a second or more, room tone in the pad, a trimmed plate joined to its own business frames, and a sound-off Kling plate: not yet run. Neither has a plate over 5 s, a take over about 4.7 s, or a 1080p plate, and upscaling after the pass is untried.
A talking plate against a closed-mouth plate on Sync Lipsync 3: the maker's wording favours the talking plate; one scratch pair on your own shot would settle it.
Two speaking faces in one frame, on either route: prove it on a scratch clip before a client shot depends on it.
Route B in Arabic: whether a repaired word survives a new performance, and how Seedance 2.5 and Kling 4.0 handle Egyptian and Saudi Arabic, are not yet known; only a native ear settles them.
What to remember
Decide the route at the start of the project and write it as a contract line before any picture: the exact take ships, or a new performance is acceptable. There is no default, and a change is a written decision.
Design the sync out first, with a fallback per shot, a written ceiling for visible sync and a break shot between visible moments. State the tolerance.
Route A is a locked take, a plate that talks quietly with no words, and a Sync Lipsync 3 pass with cut_off. The plate's acting is the ceiling.
Time the plate to the take: measure the talk window, pad the take's head (or trim the plate) until the speech starts where the talk does, and keep head, speech and tail 0.3 to 0.5 s shorter than the plate.
Route B is a video model performing the line: on English projects Seedance 2.5 can speak it in the chosen voice; Egyptian and Saudi dialogue is a cast take given as the voice reference, judged word by word, and never Arabic from text.
The voice in a client final is a person who signed. Fix pronunciation in the take and a dead face in the plate; the eye and the native ear pass a shot.
Captions follow the audio that shipped: the take on Route A, the clip's own sound on Route B.




Comments