Chapter 23 · ElevenLabs: v4 and v3
Part 4 · Voice and speaking faces
ElevenLabs is where this book's spoken lines are made, repaired and locked: text to speech for dialogue and voice-over, Text to Dialogue for hearing a scene, Voice Changer for carrying a native performance into another voice, Voice Design and cloning for voices, and a sound-effects generator for single events. Eleven v4 was released on 28 and 29 September 2026 and Eleven v3 is now the maker's previous generation. Both stay in this chapter, because nothing yet shows how v4 handles Egyptian or Saudi Arabic. It teaches the tool in its own syntax, then the settings, checks, limits, prices and rights, and ends with five finished lines.
In this chapter
What ElevenLabs offers today, v4 beside v3, and when to reach for another route
What each input owns: the voice, the model, the text, the tags, the recording
A line in ElevenLabs' own syntax: one bracket, punctuation as score, pronunciation levers
Five lines at the standard: an Egyptian telecom line, a Saudi automotive line, an English beauty voice-over, a food voice-over with its one sound, and a two-voice scene
Settings, checks, failures, and limits, prices and rights, with sound effects in each
Before you start. Chapter 21 casts the voice, directs the line, repairs it and locks it, and says who may speak in a client final. Chapter 22 makes the picture for a locked take, and Chapter 24 fits the mouth to it. Chapter 25 decides which sounds are recorded, licensed or generated. Chapter 34 covers rights across the whole job.
23.1 What it is
ElevenLabs makes speech and other sound from text or from a recording. Its current speech model is Eleven v4 (eleven_v4), which the maker calls "Our most emotionally rich, expressive speech synthesis model" (models page, checked 30 Sep 2026). It speaks 90+ languages and takes up to 10,000 characters in a request, about ten minutes. Eleven v3 (eleven_v3) is now the maker's "previous generation" model: 70+ languages and 5,000 characters, about five minutes. Eleven v4 Turbo (eleven_v4_turbo) is built for real-time agents, at a median inference latency of about 100 ms. A finished line does not need it, and this book does not use it.
Both models stay. The maker calls v4 "a net upgrade over Eleven v3" in almost every case and allows that a few edge cases may call for a different fit. It lists Arabic with no dialect named, and no maker page says how either model reads Egyptian, Najdi or Hejazi. So neither is preferred in advance. For every new Arabic voice, run the same line, voice and seed on v4 and on v3, let a native ear choose, and write the model on every take (Chapter 21).
Job | ElevenLabs tool | Craft taught in |
A spoken line: dialogue, voice-over, scratch | Text to Speech, on v4 or v3 | |
Hearing a scene's rhythm | Text to Dialogue, on v4 or v3 | 21.11 |
A native performance carried into another voice | Voice Changer | 21.3 |
A voice from a description, for audition | Voice Design | 21.3 |
The voice of a consenting person | Instant or Professional voice clone | 21.6 |
One sound event or a short bed | Sound effects |
Reach for something else when:
A person who signed can record it. The voice of a client final is an actor's recording, or the consenting clone of that person (21.6). ElevenLabs then makes the screening reads, the repairs and the clone.
The shot may carry a new performance on an English project. A video model can speak the line (Chapter 22, Chapter 17). Egyptian and Saudi dialogue never comes from a video model.
The line is sung. Use a singer (Chapter 26).
You need music or a mouth. Music is Suno (Chapter 26) or Eleven Music (Chapter 27); the mouth is Sync Lipsync 3 (Chapter 24).
How it is reached. On the ElevenLabs site, in the ElevenCreative app, on your own account, and through the ElevenLabs API for an operator who scripts it. Every product draws on one monthly credit pool (23.9). It is not reached through Higgsfield. Higgsfield's speech tools run an ElevenLabs engine of unstated version with preset voices only, none Arabic, and serve as English scratch (Chapter 33).
Sound effects
ElevenLabs also turns a written description into a sound effect, on the model eleven_text_to_sound_v2. Reach for it for one isolated event or a short bed that cannot be recorded or found in a licensed library, in that order (25.8). It makes no music for a cue and does not place anything: the edit places every sound (Chapter 25). The app makes four variations for each Generate; the API returns one per request. The maker's overview describes multi-part prompts ("Footsteps on gravel, then a metallic door opens"), while its best-practices page says to break complex effects into smaller, sequential elements and combine the results. For a commercial, make each event alone, so that the edit can put it on its frame.
23.2 How it reads what you give it
ElevenLabs reads a request as separate fields, and each field owns one thing. That tells you where a fault is fixed: a wrong accent is fixed in the voice, never in the text.
Input | What it owns | How it binds (checked 30 Sep 2026) |
Voice | dialect, accent, timbre, age, natural pace, and the deliveries it can play | chosen first, by ear (21.6) |
Model | how the text is read | named on every request; the API default is eleven_multilingual_v2, which is neither v4 nor v3 |
Text | the words and, in square brackets, their delivery | tags are natural-language instructions, not a fixed list |
Language | which language's rules read the text | on the website it is detected from the text; in the API language_code, ISO 639-1, such as ar |
Stability, Similarity | how far delivery varies from take to take; how closely the voice is held | v4 has both; v3's controls are in 23.6 |
Seed | the random path | best effort: "Determinism is not guaranteed" |
Neighbouring text or requests | continuity across separate generations | previous_text, next_text, and up to three request IDs each side |
Pronunciation dictionary | replacements for named words | up to three locators per request (23.3) |
Dialogue turns | who says what, in order | each turn has its own text and voice_id |
A recording (Voice Changer) | the performance: cadence, accent, breath, whisper | the chosen voice supplies timbre only |
Description and preview text (Voice Design) | a new voice | three previews for each generation |
Prompt, duration, loop, prompt influence (sound effects) | one sound | the prompt describes it; the other three are settings |
The voice decides the dialect, and the language setting names only a language. The maker: "The accent used when generating audio comes from the voice that you use." A voice not native to the language "might retain its native accent, or drift between different accents". Arabic is a language on both models' lists, so ar reads text as Arabic and says nothing of Cairo or Riyadh. Cast a voice from samples in the variety. On the website avoid mixing languages in one prompt, because it detects the language from the text.
v4 changed one thing about accent. When the language you generate matches the language of the reference voice, the accent is preserved. When it differs, v4 "generates fluent, natural-sounding speech in the target language" rather than carrying the reference voice's accent over. What that gives a non-Arabic voice reading an Egyptian line is not stated, and the rule stands: cast a native voice.
Tags are read as language, and sometimes as sound. The maker: "Audio tags are natural-language instructions, not an enum parameter." Because v4 is trained to generate both voice and sound effects, "a tag can occasionally be interpreted as a request for a sound effect rather than a delivery instruction". Write the voice you want, [low, gravelly voice], and not a noise.
Digits are read by a normaliser. The app converts numbers and symbols to words, and the API's apply_text_normalization defaults to auto. The maker's advice is to write numbers as words. How it reads an Egyptian price is not documented for Arabic, so write the spoken form (21.9).
A dialogue is one request. Text to Dialogue takes a list of turns; each has its own text and voice, and the tags for a turn go inside that turn's text. The speaker's name is not part of the text. The output is one file with the overlaps and gaps the model chose (23.3).
Voice Changer keeps the input's accent and cadence. The maker: it excels at preserving "accents and natural speech cadences", and "The input sample is crucial, as it determines the output characteristics." The performer's dialect therefore survives, and so do the performer's mistakes.
23.3 The prompt, in this tool's dialect
Chapter 21 teaches what a line is (21.4, 21.5). This section writes it in ElevenLabs' syntax, in the order the tool wants.

Figure 23.1 — Where each decision about a line lives. The voice carries what no words can add; the box carries the intention, the beats and the sounds the voice gets wrong.
A line, in order.
Outside the box: the voice and the model. The voice is chosen before any tag is written (21.6). The model is named: eleven_v4 or eleven_v3.
One bracket at the head: voice quality and physical state, then attitude, then the address if the line has one. [warm, a little out of breath, amazed, into the phone]. Describe the voice and never a sound. A second bracket appears only where the line turns.
The canonical words, in the spoken dialect, in plain spelling, with their punctuation (21.4).
Beat marks: an ellipsis for a trailing or weighted beat, a full stop for a stop, an em dash for a break, [pause] where the scene needs to breathe, capitals only on a word that is truly stressed.
A repair on the failing word only, by the ladder of 21.7: a mark, a respelling for this voice, IPA, a dictionary alias.
One breath group per take. A short sentence or two that a speaker would say on one breath. If you have to choose where the breath falls, the copy is two groups and two takes (21.5).
Words for the bracket. A tag works when a voice can perform it. Pick words a director could say to an actor.
Part | Words a voice can play | Words it cannot |
Voice and body | low and close, hoarse, dry, tired, a little out of breath, half-whispered, clipped | cinematic, powerful, perfect, authentic |
Attitude | amazed, doubtful, amused, certain, half sceptical, unhurried, warm | very emotional, intense |
Address | into the phone, to his son, to the viewer, half to himself, close to the microphone | perfect Egyptian accent, Saudi, Gulf |
Pace and breath | unhurried, halting, rushed, one small breath before the name | fast-paced energy |
What it answers to.
Tags. Examples the maker gives: [whispers], [sighs], [exhales], [sarcastic], [curious], [laughs]. They come in three families: voice-related, sound effects such as [applause] and [gunshot], and experimental ones such as [strong X accent] and [sings]. This book uses the first only. A sound belongs to the edit (Chapter 25), and a dialect belongs to the voice.
Voice fit. "The voice still matters." A delivery already in the voice's training data is easier to reproduce; v4 follows tags more reliably than earlier models, "including when the voice wasn't trained on that delivery", less reliably than for a voice that was.
Punctuation. Ellipses "add pauses and weight", capitals increase emphasis, and standard punctuation gives the rhythm.
Turn tags in a dialogue. The maker's forms are [starting to speak], [jumping in], [interrupting] and [overlapping], with a dash where a thought is cut.
What it ignores, or does differently.
"Eleven v4 and Eleven v3 do not support SSML break tags." Silence comes from ellipses, dashes, short phrases and [pause]. An exact silence is set in the edit between two takes.
v4 has no Style or Speed slider. Fix pace in the writing and in the edit.
Stage directions written as prose, and words such as "rain" or "door slam", are spoken or turn into noise. They go in the picture prompt or the sound spotting.
The Enhance button in the app has an LLM add tags, and its own instructions allow it to add capitals, question or exclamation marks and ellipses. Do not run a canonical line through it unseen. Compare the result with the approved line, and keep only the tags.
Pronunciation has three levers, in this order.
A mark or a respelling on the word (21.7). Nothing on the maker's pages covers Arabic diacritics.
IPA, on v4: the transcription in forward slashes, inside double quotes, with stress marks: "/ˌbaɪoʊˈkemɪstri/". The maker: results "can still vary by voice and phrase". Whether IPA works for Arabic words is not stated; audition it on the failing word.
A pronunciation dictionary, a TXT or .PLS file of rules, each an alias (a word read as another) or a phoneme. The first matching rule wins, searches are case sensitive, and up to three dictionaries apply to a request. The maker documents dictionaries in ElevenCreative Studio and in the API. SSML phoneme tags are documented for eleven_flash_v2 only. Use an alias for a brand that recurs.
Eleven v4 and Eleven v3 side by side (checked 30 Sep 2026).
Eleven v4 | Eleven v3 | |
Maker's line | "Our most emotionally rich, expressive speech synthesis model" | a "previous generation" model |
Languages | 90+ | 70+ |
Characters a request | 10,000 | 5,000 |
Voice controls | Stability, Similarity | Stability; the app page lists Similarity, Speed and Speaker Boost as not available |
Style and Speed | not available | Style is set to 0; Speed pages disagree (23.6) |
SSML | not supported | not supported |
IPA in slashes | "improved native support" | not documented |
Audio tags | followed "more reliably" | supported |
Professional clones | fully supported | "not fully optimized" |
Text to Dialogue | supported; name it | supported; the API default |
Text to Dialogue. One turn per speaker change: voice_id, then the text with its own bracket. Keep the total of all turns' text at or below 2,000 characters, and use no more than ten distinct voices. Name the model, because the API default is eleven_v3. It is a rhythm pass: the overlaps and timing it chooses are for listening, and every line that ships is remade as its own take and locked (21.7).
Voice Design: the description. The maker's order is language and dialect first, then gender and age range, quality, persona and emotion, then a sentence on timbre, pacing and distance from the microphone. Its own Arabic example opens "Native Arabic, soft Gulf (UAE) accent influence." That is Emirati influence, and replacing "UAE" with "Saudi" does not make a Saudi voice. The maker warns against "accent" when you mean intonation, and against effect words such as "reverb" or "phone". Write the preview text in the target language, 100 to 1,000 characters, in the register of the voice. A designed voice is an audition (21.3).
Sound effects: the prompt. Follow 25.7: object, material, action, force, surface, perspective, space, attack, decay and refusals, in a sentence or two, within 450 characters. The maker's words help once they have a scale: "impact", "whoosh", "ambience", "one-shot", "loop", "stem", "braam", "glitch", "drone". "Impact" alone is any collision, so write what struck what. Duration, looping and prompt influence are settings, not words (23.6).
23.4 Templates
Written for this book · template; not run · finished grade, one spoken line
Template — the finished line
LINE [id] · [speaker] · [listener, or "to camera"] · [language and variety]
CANONICAL [the approved line in the spoken dialect, with its punctuation; it never changes for the engine]
MODEL [eleven_v4 / eleven_v3] · language_code [ar / en] · audition [the other model] on the same text, voice and seed for each new Arabic voice
VOICE [name, voice ID, source class]; cast by the five-line audition (21.6)
SETTINGS stability [0.5] · similarity [0.75, v4 only] · seed [n] · normalisation [auto] · [three] takes per round, one change between rounds
ENGINE TEXT
[voice quality and physical state, attitude, address] [the canonical words; a mark, respelling or IPA only on a word that failed; one breath group]
CONTINUITY previous_text [the words before, if this is one take of a run] · next_text [the words after]Written for this book · template; not run · test grade, a scratch read to find a voice or a length
Template — the test line
[the canonical line, no bracket] · [eleven_v4] · default settings · three takes · MP3 · file named TEMP_[id]_g1 to g3A test line proves a voice, a length or a rhythm. It is labelled TEMP, it never reaches a client as a take, and nothing is built on it (21.5).
Written for this book · template; not run · Text to Dialogue, a rhythm pass
Template — the dialogue pass
MODEL [eleven_v4 / eleven_v3, named] · language_code [ar] · stability [0.5] · similarity [0.75, v4 only] · seed [n]
VOICES [speaker A: name, voice ID] · [speaker B: name, voice ID] (at most ten voices)
INPUTS one turn each, speaker names outside the text, total at or below 2,000 characters
1 A [state, attitude] [line]
2 B [state, attitude; cuts in, if it does] [line]
3 A [after a beat; state] [line]Written for this book · template; not run · Voice Design and Voice Changer
Template — a designed voice, and a conversion
DESIGN Native [language], [dialect and city]. [Gender], [age range]. [Quality]. Persona: [two to five words]. Emotion: [two or three adjectives]. [One sentence: timbre, pitch, pace, distance from the microphone.]
PREVIEW [100 to 1,000 characters in the target language, in the register of the voice]
CONTROLS guidance scale [5, the default; higher only for accent accuracy] · loudness [0, about −24 LUFS] · seed [n] · three previews; save one only after a native ear has heard it
CONVERT model eleven_multilingual_sts_v2 named · input [the performer's recording, under 5 minutes and 50 MB, consent on file] · target voice [name, ID] · remove background noise [on / off] · seed [n]Written for this book · template; not run · one sound effect
Template — the sound event
PROMPT [Object] of [material] [action] on/against [surface], [force], [close / medium / distant], [space]; [attack], [decay]. No music, no voices. (450 characters at most)
SETTINGS model eleven_text_to_sound_v2 · duration [seconds, or auto] · prompt influence [0.3 default; higher for a literal single event] · loop [off; on for a bed]
ROW [sound row id] · placed on [frame or time] · source: generated, ElevenLabs, [date] · kept: [variation]Written for this book · template; not run · a pronunciation dictionary for a name that recurs, one alias rule
Template — the alias rule
<?xml version="1.0" encoding="UTF-8"?>
<lexicon version="1.0"
xmlns="http://www.w3.org/2005/01/pronunciation-lexicon"
alphabet="ipa" xml:lang="en-GB">
<lexeme>
<grapheme>[the word as the script writes it]</grapheme>
<alias>[the word respelled as the voice should read it]</alias>
</lexeme>
</lexicon>23.5 Examples at the standard
Each example is a complete request in the anatomy of 23.3: the line, its model, its voice and its settings, with the engine text last. The canonical line is the client-approved copy, and a native reviewer of the variety has approved its wording before any voice reads it (21.4). Brand names are fictional. Seeds are start points.
Written for this book · telecom and tech · a father's line for a 20-second metro spot, Cairo · Eleven v4, with Eleven v3 as the audition twin, language ar, stability 0.5, similarity 0.75 · written for this book; not run
Example 23.1 — the line that holds underground
LINE N03 · father, about 45 · to his daughter, by phone · Egyptian (Cairene) colloquial
CANONICAL ألو؟ سامعني كويس؟ أنا في المترو… والخط ما اتقطعش.
MODEL eleven_v4, then eleven_v3 on the same text, voice and seed · language_code ar
VOICE a native Cairene man of about 45, warm, close-miked; cast by the five-line audition (21.6); source class and voice ID on the take ledger
SETTINGS stability 0.5 · similarity 0.75 (v4 only) · seed 40311 · normalisation auto · three takes per model
ENGINE TEXT
[warm, a little out of breath, amazed, into the phone] ألو؟ سامعني كويس؟ أنا في المترو… والخط ما اتقطعش.One bracket, three parts. A physical state (a little out of breath), an attitude (amazed) and an address (into the phone). Each can be heard. Nothing in it is a style word or an accent claim.
Plain spelling first. The line carries no mark. The ق in «اتقطعش» is the hardest feature and sits in the last word, so the audition and the first takes expose it. If it fails, the repair goes on that word alone, in the engine text, by the ladder of 21.7.
The ellipsis is the only beat mark: the moment he hears that the signal holds. No [pause] and no break tag.
Both models, one seed. Three takes each, the ear chooses, and the model goes on the take-ledger row (21.11).
Written for this book · automotive and luxury · a father's line for a 30-second premium SUV film, Riyadh at dawn · Eleven v4, with Eleven v3 as the audition twin, language ar, stability 0.5, similarity 0.75 · written for this book; not run
Example 23.2 — a dry line in Najdi
LINE H02 · father, about 50 · to his son at the wheel · Najdi (Riyadh) colloquial
CANONICAL تبي تسوق؟ سوق… بس لا تخلي الطريق يحس إنك مستعجل.
MODEL eleven_v4, then eleven_v3 on the same text, voice and seed · language_code ar
VOICE a Najdi man of about 50, dry and unhurried; cast from samples in the variety and passed by a Najdi reviewer (21.8); source class and voice ID on the take ledger
SETTINGS stability 0.5 · similarity 0.75 (v4 only) · seed 22718 · normalisation auto · three takes per model
ENGINE TEXT
[dry, unhurried, half a smile in the voice, to his son] تبي تسوق؟ سوق… بس لا تخلي الطريق يحس إنك مستعجل.The variety lives in the voice and the words. «تبي» and the ق of «سوق» and «الطريق» are what a Najdi reviewer listens for. No tag says "Saudi" or "Gulf".
A dry line is a casting test. In a voice that sells, it comes back as an announcer. Recast the voice; do not pile tags on it.
The bracket colours and does not dress. Dry, unhurried, and an address. The half smile is a hint for the last clause, not a mood for the whole line.
The brand is not in the line. It arrives on the end card in post, so there is no name to pronounce.
Written for this book · beauty and personal care · the voice-over of the 30-second hair-repair film of Example 26.1, Cairo studio, window 3 to 24 s · Eleven v4, English voice, language en, stability 0.5, similarity 0.75, three takes each · written for this book; not run
Example 23.3 — a close voice-over in three breath groups
LINES V01 (window 3.0 to 10.0 s) · V02 (10.5 to 17.5 s) · V03 (18.5 to 24.0 s) · one woman's voice-over
MODEL eleven_v4 · language_code en
VOICE an English voice, female, about 40, low and close, no announcer's lift; cast by the five-line audition (21.6); source class and voice ID on the take ledger
SETTINGS stability 0.5 · similarity 0.75 · seed 5190 · normalisation auto · three takes per group
V01 [low and close, unhurried, quietly certain] Damage doesn't arrive in a day. It builds… a little heat, a little hurry, a lot of years.
V02 [low and close, unhurried, a little warmer] Repair doesn't arrive in a day either… but it can begin tonight.
V03 [low and close, a small smile, no lift on the last word] Nahla Repair Mask. Give it three washes.
CONTINUITY V02: previous_text = the words of V01, next_text = the words of V03 · V03: previous_text = the words of V02
BRAND "Nahla" read plainly first; only if it fails, "/ˈnɑːlɑː/" on v4, or an alias in a pronunciation dictionary if the name recurs across the campaignOne breath group per take. Three groups are three takes, joined in the edit; previous_text and next_text keep the voice level across the joins.
The same bracket each time, changed only where the thought turns: warmer on the promise, a smile on the name. A voice-over that changes mood at every line sounds like an advertisement.
Beats are punctuation. Ellipses carry the weight; no break tag, which v4 and v3 do not support. Any exact silence is placed in the edit, against the windows written beside each group.
The brand is read plainly first, with IPA and then an alias as the ladder's later rungs, and the measured length of each take sets the cut (21.7).
Written for this book · food and drink · a 15-second film for a date-and-tahini bar, Cairo kitchen: the voice-over in two groups and the one sound between them · Eleven v4 (stability 0.5, similarity 0.75) and eleven_text_to_sound_v2 (1.5 s, prompt influence 0.6) · written for this book; not run
Example 23.4 — a voice-over and the snap it waits for
VOICE-OVER
LINES R01 (window 1.0 to 6.0 s) · R02 (9.0 to 13.0 s) · man, about 40 · to the viewer · English with a light Cairene accent
MODEL eleven_v4 · language_code en · stability 0.5 · similarity 0.75 · seed 8802 · three takes per group
VOICE an English voice with a Cairene accent, warm, close-miked; cast by the five-line audition (21.6); source class and voice ID on the take ledger
R01 [warm, close to the microphone, a smile you can hear] Dates from Siwa. Tahini from Alexandria. That is the whole list.
R02 [warm, close to the microphone, a smile you can hear] Rutab. Break one in two.
CONTINUITY R02: previous_text = the words of R01
SOUND EFFECT (one event, placed in the edit on the frame the hands part, inside the three seconds between R01 and R02)
MODEL eleven_text_to_sound_v2 · duration 1.5 s · prompt influence 0.6 · loop off · four variations in the app, one per request in the API
PROMPT A firm date-and-sesame bar snaps in two in one hand: a short dry crack, then a soft tacky pull of sticky fruit as the halves part. Close dry recording, quiet kitchen scale, sharp attack, brief natural tail. No music, no voices.Short sentences carry a warm voice. The accent comes from the voice; the bracket adds a smile and a distance from the microphone. "One" and "two" are words, not digits (21.9).
The silence is exact, so it is made in the edit. The two groups leave three seconds, and the snap sits in them. No [pause] or break tag is asked to hold a time; the measured length of each take sets the cut.
The sound is a physical event heard from a place: object, material, action, force, perspective, space, attack and decay, with the refusals last, in 227 of the 450 characters allowed (25.7).
A high prompt influence for one literal event. The four variations are four candidates, all kept until one is chosen, and the crack is placed on the frame in the edit. The voice takes and the sound row are separate records.
Written for this book · beauty and personal care · a 15-second salon scene, Zamalek, two Egyptian voices: hearing the rhythm before any line is locked · Text to Dialogue, eleven_v4 named, language ar, stability 0.5, similarity 0.75 · written for this book; not run
Example 23.5 — a client and a stylist, four turns
MODEL eleven_v4 (named: the API default is eleven_v3) · language_code ar · stability 0.5 · similarity 0.75 · seed 6117
VOICES S: a native Cairene woman, about 30, the client · N: a native Cairene woman, about 50, the stylist · two voice IDs, on the take ledger
INPUTS one turn each; speaker names stay out of the text; 228 characters in all, against the 2,000 the maker advises
1 S [curious, close, half sceptical] هو ده اللي بيصلّح الشعر من أول غسلة؟
2 N [warm, amused, unhurried] مش من أول غسلة يا مدام… بس من التالتة هتشوفي الفرق.
3 S [after a beat, doubtful] تلاتة بس؟
4 N [a quiet smile, certain] تلاتة بس. جربي وقوليلي.Each turn has its own voice and its own bracket. The maker: "Add audio tags inside the text for the turn they should affect." Turn 3 is two words, so the beat goes into the bracket.
The claim is honest in the dialogue, three washes, and the same three washes close the voice-over of Example 23.3.
The shadda in «بيصلّح» belongs to the approved script. Without it the verb can be read as "suits" rather than "repairs". It is a meaning mark, not an engine repair, and it is in the canonical line.
It is a rhythm pass. Choose the run whose gaps and interruptions work, then remake each turn as a single-line take, repair it by the ladder (the ق of «وقوليلي» is the likely failure), lock it, and lay it on its own track (21.7).
23.6 Settings
Which controls exist depends on the model and the surface, the app or the API. Read the label you see and write it on the take. Change one setting at a time, after the voice and the plain line are right.
Speech and dialogue (checked 30 Sep 2026).
Setting | Use | Why, and the default that bites |
Model | eleven_v4; for each new Arabic voice, v4 and v3 | The API default is eleven_multilingual_v2, an older model. Text to Dialogue defaults to eleven_v3. Name the model every time. |
Language | ar or en in the API; the website detects it | A language, never a dialect. Not supported on eleven_multilingual_v2. |
Stability | 0.5 (the app: "around 50") | Lower widens the emotional range; too low gives odd, rushed performances; higher flattens. The maker: the sliders work as a range and do not guarantee a result. Move one step at a time. Older guides describe three named modes on v3 (Creative, Natural, Robust); the maker's pages show a slider. |
Similarity | 0.75 on v4 | Higher holds the source voice closer and can reproduce its flaws. The app page lists it as not available on v3. |
Style | 0 | Not on v4. Above 0 the maker warns of less stability and more latency. |
Speed | 1.0 | Range 0.7 to 1.2, extremes degrade. Not on v4; the pages disagree for v3. Fix pace in the writing and the cut. |
Speaker Boost | leave on | Not on v3 (app page). The maker calls its differences "rather subtle". |
Seed | one per round, written down | 0 to 4294967295. Best effort only. |
Normalisation | auto, and numbers written as words | on and off exist. |
Continuity | previous_text, next_text for a run of groups | Up to three request IDs each side; best when the same model is used throughout. |
Output | test: mp3_44100_128 (the default); lock source: WAV or PCM | The app downloads MP3 or WAV from History. 44.1 kHz PCM and WAV need Pro or above. The lock is made from the source file (21.7); converting an MP3 restores nothing. |
Free regenerations | use both | Two, in the web app only, for identical content and settings. Not in the API. |
Voices and conversion (checked 30 Sep 2026).
Setting | Use | Why |
Voice Changer model | name eleven_multilingual_sts_v2 | The API's default is not stated. The list has 29 languages, Arabic as "Arabic (Saudi Arabia, UAE)"; Egyptian is not listed. eleven_english_sts_v2 is English only. |
Voice Changer input | under 5 minutes and 50 MB | Longer material is split. The performer's gain matters: a quiet take hinders recognition, a loud one clips. |
Remove background noise | on for a room you could not treat | Off for a clean take, so nothing is processed twice. |
Voice Design | guidance 5; loudness 0; a seed; one preview text | Higher guidance can sound "artificial or robotic". Loudness 0 is about −24 LUFS, and the API default is 0.5. should_enhance is off by default. |
use_pvc_as_ivc | off | It uses the Instant version of a Professional voice and "may improve expressiveness"; audition it before use. |
Sound effects (checked 30 Sep 2026).
Setting | Use | Why |
Duration | the event plus its tail; auto for a bed | API 0.5 to 30 s; the overview says 0.1 to 30. Specifying it may change the cost (23.9). |
Prompt influence | 0.3 default; 0.5 to 0.7 for one literal event | 0 to 1 in the API, 30 % in the app. Higher follows the prompt and varies less. |
Loop | on for a bed only | eleven_text_to_sound_v2 only. Check the seam at 100 %. |
Output | WAV 48 kHz for events | The overview: MP3 for all effects, WAV at 48 kHz for non-looping ones. The app offers MP3 44.1 kHz or WAV 48 kHz. |
23.7 Checks before you sign
Listen alone, on headphones, to the whole take, in this order. Judge a take by ear against the canonical line, never against the engine text.
Listening checks
Every word. The words are the canonical words: none added, none dropped, none swapped for an easier one. A different word is a failed line, however well it is read.
Dialect and register. A native reviewer of the variety hears the vowels, the ق, the stress. A clearly wrong dialect fails whatever the average (21.6).
The performance. The bracket was obeyed: the state, the attitude, the address. Breath and pace fit the picture the line will meet.
Nothing else in the file. No click, hiss or second voice, and no sound a tag was taken for.
The whole sentence after a repair. Compare it with the pass before, at the same level: a corrected word can bring a new stress or pause elsewhere.
The record. Model, voice ID, source class, settings, seed and the engine text are on the take-ledger row (21.11), for the rejected takes too.
A dialogue. Each turn is in the right voice, no name is read aloud, and the overlaps are ones you would keep. Then each line is remade alone.
A conversion. The words are the performer's, unchanged; the timbre is the target's; the room is not audible.
A sound effect. Alone, at 100 %: the object, material and force are right for the picture, the tail is as long as asked, no music or voice leaked in, and a loop's seam has no click. All four variations are kept until one is chosen.
The file. Probe its format and sample rate; the download may be MP3 where you expected WAV.
23.8 Failures and fixes
Symptom | Cause | Fix |
A word changed or added | low stability, or a long line | one step toward higher stability; one sentence per take |
The bracket is ignored, or the read is flat | the voice cannot play it, or stability is high | recast the voice; one simpler bracket; one step lower stability |
A tag or direction is spoken aloud, or turns into a noise | the tag read as a sound cue; direction written as prose | describe the voice quality; fewer tags; direction to the picture prompt |
Foreign or wrong dialect | the voice, not the text | recast from samples in the variety; audition the other model |
One word fails: ق, ع, a vowel | this voice on this word | the ladder of 21.7 |
IPA is ignored or wrong | IPA "can still vary by voice and phrase"; Arabic is not documented | generate again; another voice; an alias |
A number or date misread | the normaliser | the spoken form, in words (21.9) |
The accent drifts, the level shifts | a long line, or a voice not native to the language | one breath group per take |
A clone sounds different on v4 than on v3 | v4 follows the source more closely | audition both; the ear decides; keep the model on the row |
A newly generated line no longer matches a locked one | the model "may shift over time" | never regenerate a locked line; use the lock, and re-audition new lines |
The performer's slip survives a conversion | the input decides the output | re-record the input |
A sound is the wrong size or material | a vague prompt | name what struck what, with force, scale and space (25.7) |
A sound effect carries music or a voice | the prompt invited it | end with "No music, no voices"; make the event alone |
The retry ladder, smallest lever first. (1) Play the take word by word and name the failing word. (2) Correct the punctuation or the phrase length. (3) Change stability by one step. (4) Mark the word. (5) Respell it for this voice only, with the reviewer's approval. (6) Add IPA on v4, or a dictionary alias. (7) Audition the other model on the same voice. (8) Recast the voice. (9) A person who signed, or a conversion of a native performance (21.3). Measure after every rung: the lock is the timing truth.
23.9 Limits, prices and rights
Limits (checked 30 Sep 2026).
Item | Limit |
Characters in a request | v4 10,000 (about ten minutes); v3 5,000 (about five); eleven_multilingual_v2 10,000; eleven_flash_v2_5 40,000 |
Text to Dialogue | at most 10 voices; total text at or below 2,000 characters for reliable generation; previous_text and future_text at most 100 characters each |
Voice Changer | 5 minutes and 50 MB; 1,000 characters a minute |
Voice Design | preview text 100 to 1,000 characters; three previews, charged once; a saved voice takes a voice slot |
Instant clone | 1 to 2 minutes of clean speech recommended; over 3 minutes adds little (the v4 announcement says a sample of 10 seconds is enough; use the documentation) |
Professional clone | 30 minutes at least, 2 to 3 hours best; your own voice only; fine-tuning usually 3 to 6 hours |
Sound effects | 30 s a generation; prompt 450 characters |
Plans and credits (monthly billing, checked 30 Sep 2026). One credit pool serves every product: speech costs 1 credit a character, Voice Changer 1,000 a minute, Eleven Music about 900 a minute (approximate, the pricing FAQ; 27.9).
Plan | Price | Credits | Adds |
Free | $0 | 10,000 | no commercial licence; music has its own terms (27.9) |
Starter | $6 | 30,000 | commercial licence, Instant cloning, music commercial use |
Creator | 22(11 first month) | 121,000 | Professional cloning |
Pro | $99 | 600,000 | 44.1 kHz PCM through the API |
Scale | $299 | 1,800,000 | 3 seats, 3 Professional clones |
Business | $990 | 6,000,000 | 10 seats, 10 Professional clones |
Enterprise | custom | custom | custom terms |
The API's 192 kbps MP3 is listed as Creator and above in the reference and under Pro on the pricing page. API speech is priced per 1,000 characters: v4 $0.08, on offer at $0.022 until 12 Oct; v4 Turbo $0.04; v3 $0.08; Multilingual $0.08. Voice Changer is $0.12 a minute. A v4 promotion runs to 12 October, and the pages disagree about it (a banner says "3x credits included on Creator+", the pricing page "up to 2x"); read the cost the app shows before you generate. Sound effects have three prices that do not agree: 200 credits a generation (pricing FAQ), 40 credits a second when a duration is set (overview), $0.12 a minute (API). Budget the highest and read the app.
Rights.
Generate on a paid plan. "All paid plans include a commercial license, provided you're not using Beta Services." Content made outside a paid subscription, before or after, cannot be used commercially, and content made with Beta Services cannot be used "in any production environment". Confirm on your plan that the model you use, v4 or v3, carries no alpha or beta label before any client use: one app FAQ heading still reads "Eleven v3 (Alpha)".
A voice needs a person who signed. An Instant clone asks you to confirm "that you have the right and consent to clone the voice". A Professional clone can only be made of your own voice: "Even with their consent, you cannot clone someone else's voice". The actor builds it in her account and shares it. Clones cannot be exported. The use policy bars replicating another person's voice "without consent or legal right". Never clone a library voice or a public clip.
The release covers what the maker takes. ElevenLabs licenses your content, voice included, to provide and improve its services and develop new ones, and "will not commercialize your voice on a standalone basis without your permission". Put that in the actor's release (Chapter 34).
A library voice is for casting and screening. Its owner may set a notice period from 30 days to 2 years; disable_at_unix shows the date; ElevenLabs may remove any voice at its discretion; earlier outputs "will continue to exist and remain available". The Voice Library terms bar creating and sharing a voice model based on another person's voice. The licence texts say nothing on advertising use. A library voice enters a final only on the conditions of 21.6.
A sound effect is not exclusive. Under the Sound Effects Terms (12 Feb 2026), sound-effect outputs may be sublicensed to third parties, including other users. The Disable control on the sound-effects page stops that for new use only. Do not use a generated effect as a signature sound; record it.
You keep the audio you make on a paid plan, subject to the terms. The client is told in writing which files are generated, by which model, on what date (Chapter 34).
What may ship to a client. A take made on a paid plan, on a model that is not a beta service, in the voice of a person who signed or a library voice that meets 21.6, with its take-ledger row, its consent record and its release on file. A conversion of a performer's recording, with the performer's consent. A sound effect that is not a signature sound, with its row. A designed voice or a test line, never as a final.
23.10 Version notes
What Eleven v4 changed (28 and 29 Sep 2026). More expressive delivery and tags followed more reliably; cloning that follows the source voice more closely, so a v4 clone can sound different from its v3 version; Professional clones fully supported; IPA in slashes; Stability and Similarity only, with no Style or Speed slider and no SSML; and cross-language speech that is fluent in the target language instead of carrying the source accent. The maker says v4 "will keep evolving" and its behaviour "may shift over time", and asks you to re-test periodically. Voice Design voices work on v4 but "may not be as performative". For cloning it advises training audio in a single speaking style. v3 remains available, with Professional clones "not fully optimized".
Announced. A toggle for the cross-language accent is "a research project", with no timeline. ElevenLabs is working on a Director's Mode for finer control. Nothing else is dated.
Re-check when the next model ships: the models page and its language lists (the help-centre page adds Javanese, Kamba, Latvian and Zulu for v4, and the FAQ omits them); whether Style and Speed return; the stability labels; the character limits; the API defaults; the beta status; the price and the promotion; the Terms of Use (31 Mar 2026), the use policy (17 Aug 2026) and the Sound Effects Terms (12 Feb 2026).
Open questions that change what you do
How v4 and v3 read Egyptian, Najdi and Hejazi. No maker page says. Audition both on every new Arabic voice, with a native ear, and write the model on every take.
Whether IPA works for Arabic, and whether a dictionary's phoneme rule applies on v4 or v3. Not documented. Audition it on the failing word; use an alias for a recurring brand.
Whether v4, or v3, is a beta service on your plan. If it is, its output cannot ship. The models and help pages carry no such label, but one app FAQ heading still reads "Eleven v3 (Alpha)". Read the label before client use.
Whether a library voice may appear in an advertisement. The licence texts are silent. Ask the client's legal side in writing (21.6).
What Voice Changer defaults to, and whether it reads Egyptian. The default model is not stated and Egyptian is not listed. Name the model and hear a test.
What the cross-language accent rule does to a non-Arabic voice reading an Egyptian line. Not stated. Cast a native voice.
What a sound effect costs, and how short it can be. Three prices and two minimum durations are published.
What to remember
The voice decides the dialect and the text is the score. Cast the voice before you write a tag; the language setting names a language, never a dialect.
Name the model on every take: eleven_v4 or eleven_v3. The API defaults are older or different. For each new Arabic voice, audition v4 against v3 and let a native ear choose.
One bracket at the head, describing the voice: quality and physical state, then attitude, then address. Punctuation is the score. One breath group per take. Neither model reads SSML break tags.
Repair only the word that fails, by the ladder: a mark, a respelling for this voice, IPA on v4, a dictionary alias, another voice.
Change one variable at a time; three takes per round; record the model, voice, settings, seed and engine text on every take.
Text to Dialogue is a rhythm pass. Name the model, keep the text under 2,000 characters, and remake every line alone before it is locked.
Voice Changer keeps the performer's dialect and cadence; name eleven_multilingual_sts_v2. Voice Design is an audition, and a clone is only of a person who signed.
A sound effect is one physical event described in 450 characters; make each event alone, keep all four variations, and place it in the edit. It is not exclusive.
Generate on a paid plan, confirm that the model is not a beta service, and never regenerate a locked line because a newer model exists.




Comments