top of page

Chapter 21 · The voice department

Writer: Yasser Ashour
Yasser Ashour
2 hours ago
26 min read

Part 4 · Voice and speaking faces

In a commercial the edit owns every sound, and the voice comes first among them. This chapter settles who owns sound and how a shot's sound contract is written, then teaches where a voice comes from (recorded, generated or converted) and who may speak in a client final. It teaches the craft of one spoken line: written the way the character says it, cast before any tag is written, directed with one bracket, repaired only on the word that fails, and locked and measured before any picture is made. Egyptian and Saudi dialogue have their own workflows, and the chapter ends with the ledgers that keep them honest.

In this chapter

  • The edit owns sound, and the four sound contracts a shot can have

  • The three roads to a voice, and who may speak in a final

  • Writing a line as two texts, and directing it with one bracket

  • Casting a voice by audition

  • The Egyptian line workflow: repair by rungs, measure, lock; a Saudi character in a named variety

  • Numbers, dates, brand names and English terms

  • The dialogue ledger, the take ledger and the other records of the voice

Before you start. Chapter 8 (8.2) decides each project's speaking-shot contract. This chapter makes and locks the take; Chapter 22 makes the picture for it and Chapter 24 fits the mouth to it. The tool that makes the take is ElevenLabs (Chapter 23). Chapter 25 places music and effects around the voice, and Chapter 34 covers rights.

21.1 What this stage decides, and who decides it

The voice department makes every spoken line and decides who owns each sound in the film. The director decides which lines exist and are approved, whose voice speaks them, each shot's sound contract, and the lock of every take. The operator writes and runs the jobs. A native speaker of the variety approves the wording before any voice reads it, and hears every take. In a commercial the client-approved copy is the authority for the words.

The edit owns sound. Treat every generated soundtrack as scratch until you have judged it, and build the final sound in the edit. Video models can return speech, room tone, effects and music nobody asked for. That serves previz, rhythm and scratch ambience. The final track is chosen, repaired and mixed outside the generator, because a generator returns one finished mix that cannot be taken apart, re-timed or replaced line by line.

Judge a video model's own audio in five passes before you keep any of it: its levels and silences; contamination (music, a second voice, humming under the line); whether each contact sound lands on its frame; whether the ambience matches across the cuts; and a written decision, kept or discarded. Beds differ from clip to clip, so a scene that keeps each clip's room jumps at every cut. Lay one continuous room bed under the scene instead (Chapter 25).

Word

In this book

Canonical line

the script's words for a line, in the spoken dialect, with its punctuation; it never changes for the engine's sake

Engine text

the same words as typed into the voice tool, with a bracket and any repairs; it never reaches captions

Bracket (audio tag)

a direction in square brackets inside the text, which the voice plays and does not speak

Take

one generated or recorded reading of a line; lock is the accepted take, normalised, named and measured

Source class

how a voice came to be: actor recording, consenting clone, library, designed or converted

Tashkeel

Arabic diacritics: short-vowel marks, sukūn (no vowel) and shadda (a doubled consonant)

Rung

one step of the repair ladder (21.7); a line climbs only as far as the failing word requires

Najdi, Hejazi

the Arabic of Riyadh and of Jeddah, two Saudi varieties, cast separately

21.2 The sound contract, decided first

Write down each shot's sound contract before its prompt: a silent picture; the model's own sound; the exact take; or a new performance. Each is judged differently. For a speaking shot (the last two) there is no default: the speaking-shot contract is decided at the start, by shot class, in writing (8.2), and a change is a written route change with a reason.

fig21-1

Figure 21.1 — The four contracts. The third and fourth are the two speaking-shot routes of Chapter 22; the first two cover every other shot.

Split every line of dialogue across four owners, and never ask one prompt to solve all four. The writing owns meaning, dialect, register and cultural fit. The voice owns pronunciation, prosody, acting, speaker identity, audio quality and rights (this chapter and Chapter 23). The picture owns identity, composition, body performance, camera and environment (Parts 2 and 3). Sync and post own the timing of sound to the mouth, whether the audio is kept, ambience, captions and mix (Chapters 22, 24, 29 and 30). When a speaking shot sounds wrong, the four owners tell you where to look.

Lock the line before you spend on picture. A repair to a line costs characters on a voice account. Re-rolling a speaking shot costs video credits (Chapter 22 quotes them). If every picture route sounds wrong in the same way, the fault is in the take: stop video spend and repair the take.

Know which promise an audio input makes. A video model's audio reference is a guide to a voice, not a copy of your take, so a clip performed from it carries new timing and needs new judgement (Chapter 22). A lip-sync pass takes your audio and changes the mouth to it (Chapter 24). Which of the two you want is the contract.

21.3 Three roads to a voice

Every voice in a film reaches it by one of three roads. It is generated: text to speech from a library voice, a designed voice or a clone. It is converted: a native performance recorded by a person, then carried into another voice's timbre. Or it is recorded: an actor's own read, kept as it is.

fig21-2

Figure 21.2 — The roads to a voice. A change from one box to another is a new decision with its own consent, never a silent swap.

For a client final the order is fixed: a person who signed. Record a consenting native actor first. That recording is the master, or the source of that person's own clone. A library voice serves casting, screening and animatics, and a designed or converted voice is an audition until a native ear passes it (21.6).

  • Generated lines come from ElevenLabs Text to Speech (Chapter 23). The line is repaired at no picture cost and can be regenerated, but a take cannot be reproduced exactly. Text to Dialogue is a rhythm pass for a scene, and every line that ships is still locked on its own.

  • Converted lines start as a native performance, recorded by a person under a written consent and carried into another voice by Voice Changer. It keeps the performer's dialect and cadence, and so their slips. Use it when the acting matters more than the words can carry, or when a voice fails the acting but the words pass. The maker lists Arabic for Saudi Arabia and the UAE only.

  • Recorded lines are the actor's own read, normalised and locked. Pronunciation is the surest, and there is nothing to repair.

  • A designed voice is made from a written description. The maker's own Arabic example is Emirati-influenced, and replacing "UAE" with "Saudi" does not make a Saudi voice. It is an audition until a native ear passes it.

No road is a fallback that can be swapped in silently. A recorded read in place of the AI voice you asked for, or a conversion in place of an approved take, is a new decision with its own consent, written down before the new take is used.

Converting a native performance

  1. Get the performer's written consent, naming conversion into another voice, AI generation and the platforms (21.11).

  2. Direct and record the performer in a quiet room, the gain set so that the take neither hides nor clips: WAV, 48 kHz, the canonical line, three takes. The acting is the point of the road.

  3. Choose the target voice by the audition of 21.6. It supplies timbre only: the dialect, cadence and breath are the performer's.

  4. Convert each take on the model named in 23.6, with noise removal on only for a room you could not treat.

  5. Score it like any take. The performer's slips survive, and a mark cannot be applied to a recording, so a wrong word is re-recorded, not repaired.

  6. Lock and measure as in 21.7, and keep the performer's source recording with the conversion record.

The speech models Higgsfield hosts run a named maker's engine (ElevenLabs, version not stated) and serve English scratch lines only (33.9). Higgsfield's own voice-change, dubbing and voice-creation tools name no engine and are not used in this book.

Written for this book · template; not run · the route record, one line for each voice

Template — the route record

Voice <name or character> · Road: recorded / generated (library, designed, clone) / converted · Source class · Why this road · Fallback (another road, only as a new decision) · Consent record: <ref> · Model and voice ID (if generated) · Decided by, date

21.4 Writing the line: two texts

Every spoken line starts as words. Write it as the character says it, in the variety the character speaks: Egyptian colloquial, or Najdi or Hejazi (21.8). A line that only works in Modern Standard Arabic is a failed line, and so is a translated one. In a commercial the client-approved copy is the authority, and no line reaches a voice, a plate or a sync pass until it is approved.

A native speaker of the variety approves the spoken wording before any voice reads it. Slang, the brand name, and how numbers and dates are said are that speaker's call, not a translation's.

Write a line note before you write the bracket. The note holds the beat's objective, the emotion at its start and its end, the energy, the distance to the listener (whisper-close, conversational, stage-forward, announcer), the pace (clipped, measured, breathy, rushed, halting), the words that carry stress, the pauses and breaths, and any pronunciation risk. The note is yours. The bracket compresses it into the few words a voice can play. In the finished package the note is one physical line per beat: the state, what the body does, the address and the pace. Chapter 22 writes the same line for the face (22.4).

Keep two texts for every line: the canonical line and the engine text. The canonical line is what the script says. It never changes for the engine's sake, and it is what captions, decks and the client see. The engine text is what you paste into the voice tool: the canonical words plus a bracket and, only where a word fails, a repair. It never reaches captions or an actor's script. A pronunciation workaround never silently becomes the written line.

fig21-3

Figure 21.3 — One line, two texts. The repairs that make a voice say the line right never leave the voice department.

Plain dialect spelling first. Write the words the way the variety is ordinarily written, with the punctuation the delivery needs. Do not add marks in advance and do not swap a word for an MSA form to make the engine comfortable. The dialogue ledger begins here: the identity and words lines of the row are written now, before any voice is cast (21.11). A line with no approval never reaches a voice.

Written for this book · food and drink · an Egyptian baker's line at dawn, from line note to engine text · ElevenLabs (Chapter 23), settings at their defaults · written for this book; not run

Example 21.1 — one line note, compressed to one bracket

CANONICAL    العيش لسه سخن… ومقرمش من بره.
LINE NOTE    Objective: to make a customer wait one more minute. Starts tired, ends proud. Energy low. Distance: across a counter, conversational. Pace: unhurried. Stress: «سخن». One pause, after «سخن». Risk: the ق in «مقرمش».
BRACKET      [warm, husky with the early hour, quietly proud, across the counter]
ENGINE TEXT  [warm, husky with the early hour, quietly proud, across the counter] العيش لسه سخن… ومقرمش من بره.
  • The note is long; the bracket is short. It keeps a state (husky with the early hour), an attitude (quietly proud) and an address (across the counter). The pace and the pause are already in the punctuation, so the bracket does not repeat them.

  • The risk stays in the note. The ق is not respelled in advance. It is heard in the first takes and repaired only if it fails (21.7).

  • The canonical line is untouched. Everything the voice needs is added around it.

21.5 Directing a line

Every spoken line goes through here: an Egyptian dialogue take, a Saudi audition line, an English voice-over, a scratch read for timing. One voice, one line, one breath group per generation. The inputs are the canonical line, approved in its spoken form; the cast voice, with its name, ID and source class on the take ledger (21.6); and what the character is trying to do with the line, from the shot record. That last is what the bracket translates. The tool's syntax is in 23.3.

The finished grade is the standard: an approved line, a cast voice, a line note, one bracket, three takes a round, every take on the ledger, repairs by the ladder, and a lock. The test grade proves a voice, a length or a rhythm cheaply: one voice, the plain line with no bracket or one documented tag, two or three takes, the director's ear labelled "not certified", the files named TEMP. It never reaches a client, and nothing is built on it. It may never drop the approved line, the consent rule of 21.6, or the ledger row.

Directing a line

  1. Baseline, for a voice or a kind of line you have not directed before: the engine text with its punctuation and no bracket. At least three takes. Note what the voice does unasked.

  2. Directed read: one bracket at the head of the line. Three takes. Between rounds change one variable only: the voice, the punctuation, the bracket, stability, the model, or the spelling of one word.

  3. Listen twice to each take, alone: first for every word, vowel, consonant and dialect form; then for intention, timing, breath, emphasis and whether the change of thought is audible. One take-ledger row per take, with the failing word and where it falls.

  4. Decide: keep it, repair the failing word (21.7), or re-direct. A beautiful voice never cancels a wrong word.

  • One bracket at the head: voice quality and physical state, then attitude, then the address if the line has one. Describe the voice, never a sound. A second bracket only where the line turns. No style adjective, accent claim or ambience.

  • The voice sets what a bracket can do. A voice that has never shouted will strain if asked to; recast rather than pile on tags.

  • Hear a new voice first without your bracket. A plain baseline tells you what the bracket adds.

  • A single emotion word is valid but weak for a dramatic line, because it leaves the intention to the voice. Stack sparingly: one to three cues that fit the voice's range.

  • Punctuation is the score. An ellipsis for a trailing or weighted beat, a dash for a break, capitals only for a truly stressed word, a pause tag where the scene breathes. A pause tag is a direction, not a duration; if a silence must be exact, make the two halves as two takes and place the gap in the edit.

  • Split a hard line into smaller units and keep each take to one breath group.

  • English tags inside an Arabic line are fine; everything else stays out of the box. Stage directions written as prose get spoken, and effect words invite noises into a dialogue take.

  • Never paste a bracket into a video prompt. Brackets are the voice tool's syntax. In a video model the performance goes into a prose sentence (Chapter 22).

  • Write numbers, dates and symbols as words, the way the character says them (21.9). An English voice-over uses the same grammar; the voice does the accent.

  • Three takes, one variable. Never change all five things and call the result a diagnosis.

A voice-over is a run of takes. Copy longer than one breath group is written as groups, each its own take with its own ledger row. Measure the windows on the cut first, the stretches where the voice may speak, and write the groups to fit them. Join the takes in the edit, with the neighbouring text passed to the tool so that level and pace carry across the joins (23.6). Keep one bracket across the groups and change it only where the thought turns: a voice-over that changes mood at every line sounds like an advertisement. Example 23.3 is a voice-over in three groups.

21.6 Casting a voice, and who may speak in a final

A voice library holds ready-made voices, filtered by language and then by accent. It is a casting filter, not a setting you can pass to another tool. You cast from each voice's samples; a label is not a listening verdict.

Cast the voice before you write a single tag. The voice decides dialect, age, weight, register, pace and range. A bright commercial voice does not become a tired parent through twelve brackets. The reliable dialect route is a native voice, a script in ordinary dialect spelling, and a native ear on every take.

Who may speak in a client final is settled by consent, not by taste, and the order matters.

  1. A person who signed. Record a consenting native actor first. That recording is the master, or the source of that person's own clone. The release names the voice, the project, the uses (including AI generation and every platform it is uploaded to) and the period.

  2. A clone only of that person, under the maker's consent confirmation, or a Professional clone made on that person's own verified account. Never a clone of a library voice, and never of a public clip: a public clip is no authorisation (23.9).

  3. A library voice is for casting, screening and animatics. It enters a final only if its notice-period status is read and filed, every generated file is archived at delivery, the client is told, and the client's legal side accepts that the licence texts say nothing about advertising use. A video clip that re-performs a library voice carries that voice's limits.

  4. A designed or a converted voice is an audition for dialect work until a native ear passes it.

  5. The source class goes on every take-ledger row and on the voice line of the dialogue ledger: actor recording, consenting clone, library, designed, converted.

The actor plan

  1. Cast a native actor of the variety by the audition below, and get the signed release before anything is recorded.

  2. Record the locked canonical lines with the actor, directed as in 21.5, where the bracket becomes your note to her: WAV, 48 kHz, 24-bit, a quiet room.

  3. Those recordings are the final lines. Normalise and lock them (21.7, steps 9 to 12). Her recorded performance is the surest pronunciation, and it needs no repair.

  4. Fallback, in the same session: record one to two minutes of clean speech in one consistent style for an Instant clone, or have her build a Professional clone in her own account and share it. Clone only for lines she cannot come back for.

  5. Any new line from the clone is scored like any take, and its ledger row says the source class is a consenting clone.

  6. File the release, the source files and the voice ID with the project.

Casting a voice

  1. Filter: the language, then the accent, then gender and age for the character.

  2. Shortlist three voices from their samples. Names, voice IDs and source classes go on the ledger.

  3. Run the five-line audition on all three, in plain spelling with no bracket, at identical settings, two takes each.

  4. Score the words and the dialect first, then the acting. For a final, two native listeners score blind with names hidden and their disagreements are kept. For a test, the director's ear alone is enough, written down as "director's ear, not certified".

  5. Audition the model on the shortlist: the hardest-feature line, the same voice and text on v4 and on v3, with the model written on each take (23.1).

  6. Choose a winner and a runner-up. The casting record keeps the search, the shortlist with IDs, both listeners' scores, the models heard, the winner, the runner-up and the reason. The winner's ID goes on every ledger row of that character.

The script is identical for every candidate, so that only the voice changes. Its five lines come from the character's own script. Judge word accuracy, pronunciation, dialect and register, breath, pace, emotional credibility, noise and editability. Resemblance to a reference is not the test: a voice can match a source and perform badly, and performance is the deliverable. A clearly wrong dialect fails whatever the average, and a perfect match cannot overrule a native ear's "wrong register".

Written for this book · template; not run · the five-line audition

Template — the five-line audition

1. Neutral:             <a plain declarative line from the script>
2. Hardest feature:     <the line with the character's hardest sound, proper noun or dialect feature: a ق word, the brand, a number>
3. Trailing:            <an interrupted or trailing line, ending on an ellipsis>
4. Low intensity:       <a quiet emotional line>
5. Peak, still clear:   <the character's loudest line, with its punctuation>
Every line in plain spelling, no bracket, on every candidate voice. A second round adds one bracket per line.

Written for this book · template; not run · the blind score sheet

Template — the blind score sheet

Code | Words (pass/fail) | Dialect 1–5 | Pronunciation 1–5 | Breath and pace 1–5 | Emotional truth 1–5 | Noise | Cuttable | Listener | Notes (failing word and where)

Judge a voice line by line, not once for the film. A voice that passes one line can fail the next, and a repair that helps one voice can hurt another. Keep a cast sheet for each voice: its source class, the languages it suits, its range, its known strengths and failures, its stability setting and the role it fits. If a voice you saved disappears, the maker may have removed it after its notice period: keep the voice ID, the source class and your locked files, and re-audition. When no library voice passes, the answer is a person who signed, or a designed voice as an audition.

21.7 The Egyptian line workflow: repair, measure, lock

This is the heart of the chapter: the finished-grade package for a line in Egyptian Arabic. Keep the parts apart and never put them all in the speech box.

Part

What it holds

Where it lives

1. Line record

the canonical line, its literal meaning, speaker, listener, intention, register, and the approval

the script, the dialogue ledger, the caption source

2. Casting and audition

variety, age, weight, pace and range; the five lines; the native listener's choice; source class and rights

casting and native review

3. The take prompt

the cast voice; the model and settings; one bracket; the canonical text with punctuation as score; one breath group

the voice tool's box, and your notes

4. Repair record

the failing word, the rung used, the reviewer's approval, the change in length

the take ledger

5. Lock and measure

the untouched export, its checksum, the speech window, head and tail silence

the lock and the media record

6. Contract line

exact take or new performance; what ships; which audio the captions follow

the shot record (22.2)

For an Arabic line the packet adds pronunciation notes for names, loanwords, numbers, abbreviations and code switches, and a dialect note written in linguistic and performance terms and not as a city stereotype. The reviewer judges lexical accuracy, vowels and emphatic consonants, dialect markers, natural stress, the code-switch transition, emotional truth and breath, and whether the line sounds like a person from the intended social and age context rather than generic broadcast MSA.

The line workflow

  1. Write the canonical line in عامية as the character says it, with its punctuation, and open its line record. The director and a native speaker approve it; in a commercial the client's approved copy is the authority. An unapproved line goes no further.

  2. Write the first engine text: the bracket (21.5) plus the canonical spelling, plain, with no marks, for the model and voice the audition chose.

  3. Take it: at least three takes, the model and every setting written down.

  4. Listen and name the word. For each take, write which word is wrong, where it falls and what you hear ("ق said as a hard q", "ساعة said as MSA sāʿa"). If you cannot name it, play it again word by word. The ear checks ع ح ق ج, stress and vowel length.

  5. Score it: wording, dialect (Cairene, no MSA, no Gulf or Levantine), pronunciation, performance, each 1 to 5. All four at 4 or above to lock. Fail it at once for a wrong or missing word, a changed timbre, a slide into MSA, or any sound outside the line.

  6. Make the three free repairs first: correct the written line or a normalisation error, adjust punctuation and phrase length, change stability one step.

  7. Then climb the rungs (Figure 21.4), one at a time, on the named word. New takes after each. Compare the whole repaired sentence with the pass before, at the same level: a corrected word can bring a new accent, pause or stress elsewhere. Keep both durations.

  8. Stop when a rung passes, or when all have failed on this voice: the runner-up voice, then a person who signed (21.3). Record every variant and its verdict.

  9. Lock. Pick the accepted take and normalise it: WAV, 48 kHz, mono, at the one working level written for the project (about −20 LUFS with peaks at −3 dBTP), with 200 ms of silence at the head and 400 ms at the tail. Name it <id>_lock_v01.wav. Keep the source file, since normalising a lossy file restores nothing, and record its checksum.

  10. Measure the raw length, the lock file's length, and where the speech starts and ends. Re-time the stillomatic or the edit to the lock.

  11. Mark what went stale: every earlier take of the line, every cut built on it, every clip made from it.

  12. Hand to picture with the contract (22.2), and ask the ear again on the clip (22.7): if a video model re-performed the take, a repair can be lost.

fig21-4

Figure 21.4 — The repair ladder. A line climbs only as far as the failing word requires, and the reviewer approves every mark before it runs.

The rules of the spelling. The voice carries the dialect, the engine text carries the sounds the voice gets wrong, and the bracket carries the intention.

  • Plain dialect spelling, punctuation and phrase length come first. Repair only the word you can name, one rung at a time.

  • Never substitute a word to make the engine happy. A different word, or an MSA form, is a failed line, however well it is read.

  • The dialect reviewer approves every mark before it runs. A mark can change the word, not only its sound. One missing shadda can turn «بيصلّح» ("repairs") into «بيصلح» ("suits"), which is why Example 23.5 keeps that shadda in the canonical line.

  • Write ق as a hamza only in a word where this voice still says the ق, as «اتقطعش» becomes «اتأطعش» in the engine text. The Cairene ق is a glottal stop, but a voice may or may not need the respelling, and the same respelling that mends one voice can worsen another. Decide word by word, by ear, with the reviewer.

  • Plain spelling is not wrong by default. Repair what you hear, not what you fear.

  • A repair that helps one word can hurt the performance. Change only the mark, respelling, alias or punctuation around the target word.

  • Measure every line after a repair. A repair can lengthen or shorten a line by a large margin, and you cannot predict which. Never stretch or squeeze the file to fit. Trim the head and tail silence, then give the shot the time the line needs; if the shot cannot grow, ask the picture for a slower read and check the pronunciation again.

  • The working level is for handling. Delivery loudness is measured on the finished mix against the named specification: EBU R 128 asks −23 LUFS and −1 dBTP, ATSC A/85 asks −24 LKFS and −2 dBTP (30.6).

What the lock hands on. To picture (Chapter 22): the file, its checksum, the speech window, the head and the tail. To captions (Chapter 30): the canonical line, timed to the audio that shipped. To the mix (Chapter 25): the take as a dialogue stem at its own level, with its ledger row.

21.8 Saudi characters: Najdi and Hejazi

No maker evidences a Najdi or Hejazi voice on either ElevenLabs model. The makers list only "Arabic (Saudi Arabia, UAE)" for the older multilingual model and for Voice Changer, and no page speaks to a variety. The method is the Egyptian one with two changes: the variety is named, and the reviewer comes from it. Every Arabic line is candidate script until a speaker of that variety approves it.

"Saudi" alone is not a casting specification. Cast in a named variety, Najdi (Riyadh) or Hejazi (Jeddah), from samples in that variety, and never from a "Gulf" label. The varieties are separate profiles, not interchangeable presets, and a decision is made for each: a Najdi decision and a Hejazi decision, never one "Saudi winner", and never an Egyptian verdict inferred from either.

The voice carries the region, and so do the words. «وش» ("what") is Najdi where Hejazi says «إيش», and the ق is another marker a reviewer listens for; in Najdi speech it is commonly a /g/. Never borrow one variety's word into the other's line to make it "more Saudi", and never swap a colloquial word for an MSA one to make the engine comfortable. A voice that reads a Najdi line with a Cairene glottal stop or an MSA q is the wrong voice: recast. Use one documented tag at most for the baseline, then one bracket if the baseline needs direction. Whisper only when the scene is private, because consonants are the first thing a whisper loses.

Casting a Saudi voice

  1. A reviewer from the variety approves the audition lines: one restrained exchange, one with traps (a brand, a number, a hard sound), one emotional turn, and the other lines of the five-line audition.

  2. Shortlist two or three voices from samples in that variety. Use the accent filter where it is offered, but cast from the samples.

  3. Run the same approved lines on each voice, at the same settings, two takes each, on v4 and on v3.

  4. The reviewer scores blind, marking regional vocabulary, vowels, stress and the restrained intention separately.

  5. Decide per variety, then lock as in 21.7, steps 9 to 12.

Written for this book · template; not run · the Saudi line packet

Template — the Saudi line packet

Variety: <Najdi (Riyadh) / Hejazi (Jeddah)> · Character: <age, class, role> · Reviewer: <name, variety>
Line 1, restrained exchange: <approved line>
Line 2, traps (a brand, a number, a hard sound): <approved line, with its spoken form>
Line 3, emotional turn: <approved line>
Voices: <name + ID> ×2–3 · Models: <v4 / v3> · Settings: <held constant> · Scores: vocabulary / vowels / stress / intention, per listener

The output is one voice decision per variety: the voice's name, ID and source class, the settings, the reviewer's scores and the rights status, then the lock record as in 21.7. When no voice passes in the variety, the answer is a native actor of that variety, recorded (21.3). Example 23.2 writes a finished Najdi line.

21.9 Numbers, dates, brand names and English terms

Every commercial has a price, a date or a product name, and each is where a voice fails first.

Keep the written and the spoken form of a number in separate fields. The screen and the captions take digits. The voice takes the words, as the character says them, and in Egyptian the spoken form differs from the written one. The maker advises writing numbers and symbols as words, and its normaliser (automatic in the app, auto in the API) is not documented for Egyptian prices, so do not rely on it. The native speaker approves the spoken form.

A dictionary before a respelling. For a brand that recurs across a campaign, an alias in a pronunciation dictionary (23.3) is worth more than marks repeated in every line. Never speed up a take because the layout was timed first. Measure the read, then lengthen the card.

A character who moves between Arabic and English is harder still. Splitting the line into segments, or giving each language its own voice, is a route to try, not a proven one, and the rule still holds: the whole sentence is judged as one. The native speaker decides whether an English term takes English or Arabic stress.

A line with a number or a brand

  1. Write the canonical line with digits, for the screen and the captions.

  2. Write the spoken expansion, the way the character says it. The script owner and the native speaker approve it.

  3. If the brand might fail, audition it alone as a diagnostic. If it fails, use an alias for a recurring brand, or marks on that word only (21.7).

  4. Go back to the full line: two complete takes. An isolated pass does not certify the sentence.

  5. Measure the speech before you fix the end card's length.

Written for this book · template; not run · two fields for one line

Template — two fields for one line

Canonical (screen and captions):  <the line with digits>
Spoken (engine):                  <the same line, numbers and dates written as the character says them>
Brand probe (only if wrong):      <the spoken line, marks on the brand only>

Written for this book · telecom and tech · a data pack's price and end date, in Egyptian, as two fields · ElevenLabs (Chapter 23), the spoken field is what the voice reads · written for this book; not run

Example 21.2 — a price and a date, said the way people say them

CANONICAL (screen and captions)  باقة نيلو ١٤٩ جنيه، ولحد ١٥ نوفمبر.
SPOKEN (engine)                  باقة نيلو مية وتسعة وأربعين جنيه، ولحد خمستاشر نوفمبر.
  • The screen keeps the digits; the voice reads the words, with every component of the price present.

  • The spoken form is colloquial all the way through: «خمستاشر» for fifteen, «ولحد» for "until". A numeral or a connective read in MSA fails the line.

  • Captions come from the canonical field. The measured length of the spoken read sets the end card, not the reverse.

21.10 Which tool for which job

The job

Use

Where

A spoken line: dialogue, voice-over, scratch

ElevenLabs Text to Speech, Eleven v4 and v3

Hearing a scene's rhythm before the lines are locked

Text to Dialogue

23.3

A native performance in another voice

Voice Changer

A voice for audition, from a description

Voice Design

23.3

The voice of a consenting person

an Instant or Professional clone

23.9

A wordless plate for a locked take

Kling, Seedance or Wan

Chapters 16 to 18, 22

The mouth on a locked take

Sync Lipsync 3

A new performance on an English project

Seedance 2.5

A single sound event or a short bed

ElevenLabs sound effects

Music, and the sound around the voice

Suno, Eleven Music, a composer

Captions

timed to the audio that shipped

English scratch on the host platform

the ElevenLabs engine on Higgsfield, preset voices

33.9

21.11 Checks, the gate and the records

This stage writes to two gates. LOCK: no picture is paid for on an unlocked line. RIGHTS: no final without a person who signed, or a library voice that meets 21.6. ACCEPT is a native ear on the take, and again on the clip (22.7).

Before you lock

  • The line record is complete and approved, with the canonical line and the engine text in separate columns.

  • The model and source class are on the take-ledger row, and the failing word was named, with where it falls, before any repair.

  • Repairs climbed rung by rung; every mark approved by the dialect reviewer; ق to hamza only in that word and only for this voice.

  • The whole sentence was compared with the pass before after every repair.

  • Four scores at 4 or above and no hard fail.

  • The lock is normalised, named, measured and checksummed, and the old takes, cuts and clips are marked stale.

  • The consent record is filed, and a library voice's notice-period status is filed.

  • The contract is written, exact take or new performance.

The records of the voice department are these. The dialogue ledger holds one row for each spoken moment, in four lines, and it is the one ledger of lines. Chapter 22 writes its picture line. The take ledger holds one row for every generation, and feeds the ledger its file, its measured window and its verdict. When a take is locked, its model, voice ID, settings, engine text, file, checksum, speech window and verdict are copied into the voice line of the line's ledger row. Rejected takes stay on the take ledger, each with its failing word.

Written for this book · template; not run · the dialogue ledger row

Template — the dialogue ledger row

IDENTITY     line ID · shot · speaker · listener · scene intention
THE WORDS    approved line in the spoken dialect · literal meaning · authority (the client-approved copy) · approval status (blocked if unapproved)
THE VOICE    source class and rights record · model, voice ID, settings · engine text with marks, aliases, IPA · take file and checksum · measured speech window · native reviewer's verdict [accepted / rejected: what failed / pending]
THE PICTURE  written in Chapter 22: exposure · strategy · framing · fallback if sync fails · contract · route · which audio the captions follow
ABOVE ROWS   the master avoidance strategy · the visible-sync ceiling · the blocked-line policy · one line stating the tolerance

Written for this book · template; not run · the take-ledger row, one for every generation

Template — the take-ledger row

Line ID | Speaker | Canonical (never changes) | Model (v4 / v3) and language | Voice (name, ID, source class) | Engine text (bracket, marks, variant) | Settings (stability, similarity, seed) | Failing word and where | Rung and reviewer | Take file and checksum | Measured (raw / lock file / speech start–end) | Verdict [accepted / rejected: what failed / pending] | Listener, date

Written for this book · template; not run · the consent, casting, conversion and dialogue-pass records

Template — the four voice records

CONSENT     Voice: <name> · Source: <session date, files> · Consent: <signed release, date> · Allowed: <project, media, territory, period> · AI generation: yes / no · Platforms: <ElevenLabs / other> · Clone type: none / Instant / Professional (her own account) · Voice ID
CASTING     Character · Variety · Search (filters) · Shortlist with IDs and source classes · Models heard · Blind scores, per listener · Winner, runner-up, reason · Decided by, date
CONVERSION  Source: <speaker, consent ref, file, length> → Model: eleven_multilingual_sts_v2 · Target voice: <name, ID> · Noise removal: off / on · Seed: <n / none> · Output: <file, format> · Score: wording / dialect / pronunciation / performance · Source kept: yes
DIALOGUE    Turn (line ID) | Speaker | Approved line | Voice: source class, model, voice ID, settings | Engine text and pronunciation notes | Take number and file | Intended overlap | Verdict

The dialogue-pass record is scratch, but it carries what the pass taught into the single-line locks: written before the first take, one row for each turn, and judged against. It holds no lock.

Open questions

  • Which model reads Egyptian, Najdi and Hejazi best. No maker says. The audition and a native ear decide, voice by voice (23.1).

  • Whether a repaired word survives a video model's new performance. Only a native ear on the clip can say (22.7).

  • Whether a library voice may be used in an advertisement. The licence texts are silent; the client's legal side decides in writing (23.9).

What to remember

  1. The edit owns sound. Write each shot's sound contract before its prompt; for a speaking shot there is no default, and a change is a written route change.

  2. Every voice has a road (recorded, generated, converted) and a source class. In a client final the voice is a person who signed, or a consenting clone of that person. A library voice is for casting and screening.

  3. Keep two texts for every line: the canonical line, which never changes, and the engine text, which carries the bracket and any repair.

  4. Cast the voice before you write a tag. Audition with five lines, score the words and dialect first, and let a native ear choose.

  5. Direct with one bracket of voice quality, attitude and address; punctuation is the score; three takes a round; change one variable at a time.

  6. Repair only the word that fails, by the ladder: three free repairs, a mark, a respelling for this voice, phonetic help, a new voice, and whole-line vocalisation last. The reviewer approves every mark; a word is never swapped.

  7. Lock the take before any picture is paid for: WAV 48 kHz mono, about −20 LUFS, peaks −3 dBTP, 200 ms head, 400 ms tail; measure it, keep the source and its checksum, and mark what went stale.

  8. Saudi means Najdi or Hejazi, cast from samples in the variety with a reviewer from it. Never "Gulf".

  9. Numbers, dates and brands have a canonical field for the screen and a spoken field for the voice.

  10. Keep the ledgers: the dialogue ledger for the line, the take ledger for every generation.

Comments


bottom of page