top of page

Chapter 24 · Sync Lipsync 3

Writer: Yasser Ashour
Yasser Ashour
12 minutes ago
16 min read

Part 4 · Voice and speaking faces

Sync Lipsync 3 takes a finished video and your audio and redraws the mouth, and the face around it, to match the audio. It has no prompt, so all your direction goes into the talking plate and the take. This chapter teaches what the pass reads, how to prepare the take and plate, the setting never to leave at its default, the checks, and three complete plates.

In this chapter

  • What Sync Lipsync 3 is, and when another route is right

  • What the pass reads, and where the direction lives without a prompt

  • Templates, and three examples at the standard

  • Settings, checks, failures, and limits, prices and rights

Before you start. Chapter 22 decides whether the exact take must ship and teaches the plate and its timing; Chapters 21 and 23 make the take; Chapters 16 and 17, the plate.

24.1 What it is

Sync Lipsync 3 is Higgsfield's name for sync.so's lip-sync model, sync-3; the names mean one model (checked 30 Sep 2026). A finished video and an audio file go in; the same video comes back with the mouth, and the face around it, redrawn to the audio. sync-3 launched on 6 April 2026, and the maker's changelog (latest entry 31 August 2026) lists nothing newer.

Reach for it when the exact take must ship (Route A of Chapter 22): a locked recording spoken by a face in a talking plate, or a line changed after picture lock. Reach for something else when a new performance is acceptable, or English speech in a described voice is wanted (Seedance 2.5, 17.3); when you need the plate (Kling, 16.3, or Seedance, 17.3); when two faces share the frame and one must speak (singles and a cut, 22.3, because the connector cannot choose a face); or when a spoken clip needs a changed performance (react-1, 24.10).

How it is reached. On Higgsfield it is the video-to-video model sync_so. sync.so also runs it in its own Studio and API, which expose more controls (24.6).

24.2 How it reads what you give it

The pass has three levers and no prompt.

Input

What it owns

It must not carry

Video (the talking plate)

everything the audience sees: eyes, head, breath, hands, light, the frame's size and rate

the words; a second face; anything crossing the mouth

Audio (the take)

the words, the dialect, the pronunciation and the timing

music, a second speaker, a first sound with no head silence

sync_mode

what happens when the two lengths differ

—

The pass starts video and audio together at the first frame, with no offset and no timing per word, so the speech lands where the take's own head puts it (22.4). It redraws the mouth for as long as the audio lasts.

It redraws more than the lips. The maker documents that sync-3 works by "generating all frames at once rather than stitching together isolated segments", from "a larger spatial window than any previous model, giving it a much wider field of view around the face", and that it "Preserves the original speaker's cadence and emotional expression". Which build Higgsfield runs is not stated, so treat everything outside the mouth as the plate's and check the whole face against the plate afterwards (24.7).

fig24-1

Figure 24.1 — The pass has three levers and no prompt. It redraws the mouth for the length of the take and, the maker documents, the face around it; the rest comes from the plate, so the acting must already be in the plate and the whole face is checked against it afterwards.

Not on the connector: a prompt, the resolution, the aspect, a seed, a still image as the video, text for the maker's own speech, and a choice of face or speaker; keep one face in every shot. What comes back: the plate's frame size and rate, cut to the take's length, and your take as AAC at 48 kHz mono, near-identical to the upload, not bit-identical.

24.3 The prompt, in this tool's dialect

There is no prompt: Higgsfield exposes no prompt field, and a cost check that includes one is rejected. What a prompt would carry lives elsewhere. The face, the acting, the business and the landing go in the plate's prompt, written for Kling (16.3) or Seedance (17.3) on the seven elements of Chapter 14. The speech's timing lives in the take's head silence and the plate's length (22.4), the words in the take (Chapters 21 and 23), and how much the mouth is asked to do in the shot's exposure and framing (22.3).

What it responds to is motion in the plate. sync.so advises that the model performs best when the character appears to be talking naturally, and says sync-3 "can open silent lips to match audio, though results are generic rather than speaker-style matched". So prompt the plate to talk quietly with no words, for about as long as the line. Nothing you type reaches the pass.

Every plate prompt for this tool contains, beyond Chapter 14's anatomy: a measured business and the start of the talk in seconds; "talks quietly", with no words and no "No dialogue" line, which would tell the plate model not to speak; the face fully visible and lit, nothing crossing the mouth; a landing held to the end; and the hygiene line. There is no sound line in the prompt: the pass discards the plate's own sound and lays your take, so the plate's sound is a setting, set by what has passed. Kling 3.0: on. The one plate that passed (12 Sep 2026) ran with sound on; a sound-off plate would save 2.5 credits and is untried. Seedance 2.5: off (generate_audio). The plates that passed (12 and 17 Sep 2026) were silent, the file then has no audio stream, and the price is the same. Wan 3.0: off, as Seedance; no Wan plate has been made.

24.4 Templates

Written for this book · template; not run

Template: the finished-grade talking plate

A tripod-locked [shot size] holds the composition for the full [N] seconds, with no drift; all the movement belongs to [Name] and the room. [Business, 0 to T seconds: one or two measured actions, mouth closed.] Then, from a little after [T] seconds and for about [S] seconds, [Name] talks quietly toward [target], [physical state]: lips and jaw move in an easy speaking rhythm, the head shifts a few degrees with each phrase, the eyes stay on [target], the shoulders lift with each breath. [The room's life: one or two things that move on their own.] Keep [Name]'s face fully visible and evenly lit, both lips clear, with no hand, sleeve or prop crossing the mouth, and keep [prop and room] exactly as in the frame. [He/She] lands on [one small exhale, mouth closed, eyes on ...] and holds there to the end. No captions, no subtitles, no on-screen text.

Written for this book · template; not run

Template: the test-grade talking plate

[Shot size and camera, who and where.] [Name] [the business], then talks quietly toward [the target] for about [S] seconds, [state], with small natural head movements, then [the landing]. Face stays fully visible and well lit throughout, nothing in front of the mouth. No captions, no subtitles, no on-screen text.

Written for this book · template; not run

Template: the pass record, written before the pass

Shot <id> · Line <id>
Plate: <model, mode, start frame, duration, sound, size and frames returned> · talk window <start–end, measured>
Take: <file and version> · <head> + <speech> + <tail> = <total> s · sum against the plate <plate length − total>, inside 0.3 to 0.5 s
Pass: input_video = <plate job or file> · input_audio = <confirmed upload> · sync_mode = cut_off · expect <size, fps, seconds>
Cost: <plate + pass> · captions timed to the take

24.5 Examples at the standard

Each example gives the plate prompt and the record of its pass; Chapter 22 draws the clock of the first (Figure 22.3).

Written for this book · telecom and tech · Kling 3.0 standard plate, start frame, sound on, 5 s, 16:9, then Sync Lipsync 3 with cut_off · Egyptian line, exact take · not run

Example 24.1 — "Can you see me now?": a father on a fibre video call

A tripod-locked medium close-up holds the composition for the full five seconds, with no drift; all the movement belongs to Adel and the room. In the first second Adel leans back into the sofa about ten centimetres in one slow, settling movement and lifts his phone to chest height, his mouth closed. Then, from a little after the first second and for about three seconds, he talks quietly toward the phone, loose and unhurried: his lips and jaw move in an easy speaking rhythm, his head shifts a few degrees with each phrase, his eyes stay on the screen, and his shoulders lift with each breath. A ceiling fan turns slowly and the phone screen brightens softly as the call picture sharpens. Keep his face fully visible and evenly lit, both lips clear, with no hand, sleeve or phone crossing the mouth, and keep the phone, the sofa and the room exactly as in the frame. He lands on one small exhale, his mouth closed and his eyes on the screen, the phone steady at chest height, and holds there to the end. No captions, no subtitles, no on-screen text.
Shot 12 · Line T-04 «شايفاني كويس دلوقتي؟ الصورة نضيفة أهي!»
Plate: Kling 3.0 standard · start frame 12_keyframe_v3 · 5 s · 16:9 · sound on · returns 1284 × 716, 24 fps, 5.04 s · talk window measured 1.2–4.2 s
Take: T-04_lock_v02 · 1.2 s head + 3.0 s speech + 0.4 s tail = 4.6 s · sum against the plate 5.04 − 4.6 = 0.44 s
Pass: input_video = the plate's job · input_audio = the MP3 · sync_mode = cut_off · expect 1284 × 716, 24 fps, about 4.6 s
Cost: 8.75 + 18.15 = 26.9 credits · captions timed to the take
  • Why it works: the plate is written to a stated take. The business ends at 1.2 s where the talk begins, the talk is as long as the speech, and the take's head is padded to the measured talk onset.

  • What to copy: the phone below the mouth, its screen a glow with no legible text; protection naming the mouth; a landing and a hold. Sound is on, as 24.3 sets for Kling.

Written for this book · beauty and personal care · Seedance 2.5 plate, 720p, start frame, generate_audio off, 5 s, 16:9, then Sync Lipsync 3 with cut_off · Saudi (Najdi) line, exact take · not run

Example 24.2 — a serum testimonial: a Najdi voice, a bottle kept clear of the chin

Image 1 is the first frame, exactly as it is. The camera is locked on a medium close-up; all the movement belongs to the woman and the room. 0-1.4 s: she lifts the bottle to chest height in one settling movement of about twenty centimetres and looks at it, then at the lens, her mouth closed. 1.4-4.2 s: she talks quietly to the lens, as if answering a friend; her lips and jaw move in an easy speaking rhythm, her head tilts a few degrees with each phrase, her eyes stay on the lens, and the bottle stays at chest height, clear of her chin. 4.2-5 s: she lets out a small breath, her lips close into the beginning of a smile, and she lowers the bottle a few centimetres against her chest. A sheer curtain moves faintly. Keep her face fully visible and evenly lit, both lips clear, with no hand, sleeve, headscarf edge or bottle crossing the mouth, and keep the bottle's plain body and the room exactly as in the frame. The clip lands on her with the bottle at her chest and the smile just begun, and holds for the last half second. No subtitles, captions or on-screen text.
Shot 07 · Line T-11 «صراحة ما توقعت أشوف فرق بهالسرعة.»
Plate: Seedance 2.5 · references mode, start frame 07_keyframe_v2 · 720p · 16:9 · 5 s · generate_audio off · returns 1280 × 720, 24 fps, 5.04 s · talk window measured 1.4–4.2 s
Take: T-11_lock_v03, a Najdi voice actor's recording under a signed release, line and take approved by a native Najdi reviewer · 1.4 s head + 2.8 s speech + 0.4 s tail = 4.6 s · sum against the plate 0.44 s
Pass: input_video = the plate's job · input_audio = the MP3 · sync_mode = cut_off · expect 1280 × 720, 24 fps, about 4.6 s
Cost: 35 + 18.15 = 53.15 credits · captions timed to the take
  • Why it works: the bottle is parked at chest height, clear of the chin, so nothing crosses the mouth; a 1.4 s head gives the business room. Seedance is the fallback plate when Kling's face fails twice or neighbouring shots are Seedance clips.

  • What to copy: ordered beats with whole-second stamps, the form the Seedance maker documents; a plain label, so the brand is composited in post; sound off, as 24.3 sets for Seedance. The line and take need a reader and a listener of the Najdi variety (Chapter 21).

Written for this book · automotive and luxury · Kling 3.0 pro plate, start frame, sound on, 6 s, 16:9, then Sync Lipsync 3 with cut_off · Egyptian line, exact take · not run

Example 24.3 — the dealer at the showroom: a longer plate at delivery size

A tripod-locked medium close-up holds the composition for the full six seconds, with no drift; all the movement belongs to the dealer and the showroom. In the first two seconds the dealer takes half a pace toward the lens, about thirty centimetres, and rests his right palm flat on the roof of the car beside him at shoulder height, his mouth closed. Then, for about three seconds from just after the second second, he talks warmly toward the lens: his lips and jaw move in an easy speaking rhythm, his head tilts a few degrees with each phrase, his eyes stay on the lens, and his palm stays on the roof. Beyond the glass wall behind him, out-of-focus traffic slides steadily from left to right. Keep his face fully visible and evenly lit, both lips clear, with no hand, sleeve or reflection crossing the mouth, and keep the car's body and the showroom exactly as in the frame. He closes his mouth into a small smile, gives one nod, and holds with his palm on the roof to the end. No captions, no subtitles, no on-screen text.
Shot 21 · Line T-19 «العربية دي مستنياك من الصبح يا باشا.»
Plate: Kling 3.0 pro · start frame 21_keyframe_v1 · 6 s · 16:9 · sound on · returns about 1912 × 1080, 24 fps, 6.04 s · quote it first (pro 5 s with sound off was 7.5 credits; pro with sound on is not quoted) · talk window asked for 2.0–5.0 s, to be measured
Take: T-19_lock_v01 · 2.0 s head + 3.0 s speech + 0.6 s tail = 5.6 s · sum against the plate 6.04 − 5.6 = 0.44 s
Pass: input_video = the plate's job · input_audio = the MP3 · sync_mode = cut_off · expect about 1912 × 1080, 24 fps, about 5.6 s (134 or 135 frames)
Cost: the plate as quoted + the pass (18.15 for a 4.6 s take; about 22 if the price grows with length)
Length flag: a plate over 5 s and a take over about 4.7 s have not been run; try a scratch copy of the shot first
  • Why it works: the plate is made at the delivery raster, because the result keeps the plate's frame size; the palm on the car roof is a parked hand, and the traffic behind glass shows no second face.

  • What to copy: a longer head (2.0 s) on a six-second plate, a 0.6 s tail for the nod, and the sum written out before a credit is spent. The car's badge and any legal line are composited in post.

24.6 Settings

On Higgsfield (the connector's contract, read 27 Sep 2026):

Setting

Values

Use

input_video

a completed job id, or an upload

the plate's finished job, passed straight in; an upload only for a trimmed plate

input_audio

a confirmed upload only

the prepared take, as MP3

sync_mode

bounce (default) · loop · cut_off · silence · remap

cut_off

folder_id

a folder

the project folder

The cost check is rejected ("extra inputs not permitted"), so the price first appears in your transactions after the job.

sync_mode decides what happens when the plate and the take differ in length; sync.so advises keeping the two "close to each other".

Value

What Higgsfield says it does

Use

cut_off

"truncates to the shorter input"

always

bounce (default)

"ping-pongs the video": forwards, then backwards

never: it reverses the actor's motion

loop

"repeats it"

never: the acting visibly repeats

silence

"pads the audio"

a scratch test only

remap

"retimes the video"

never without a creative decision

The default differs by surface: sync.so's Studio picks bounce or cut_off by which input is longer; Higgsfield defaults to bounce either way. When the take outlasts the plate, bounce plays the plate backwards; cut_off clips the line, and you hear it.

Prepare the take the same way every time: clean single-speaker speech with no music (the maker: "Avoid audio with music, background noise, or multiple simultaneous speakers"); a WAV at 48 kHz mono, about −20 LUFS, peaks at −3 dBTP, at least 200 ms of silence at the head and 400 ms at the tail (more at the head when the plate's business comes first); uploaded as an MP3 made from it, and confirmed.

Prepare the plate as in 22.4, at 24 fps (the maker accepts 24, 25 or 30 constant) and the delivery size. Seedance's maker delivers 1080p as 10-bit HEVC, while sync.so asks for 8-bit SDR BT.709: probe and conform such a plate first (Chapter 30).

On sync.so's own surfaces (checked 30 Sep 2026) you can also choose the speaker; sync-3 "manages expressiveness natively", with no temperature to set.

24.7 Checks before you sign

The pass changed the mouth, and possibly the face around it, for the length of the take: check that region and stretch closely, and the rest against the plate. Chapter 22 (22.7) gives the order.

Before you sign a pass

  • Probe the file: the frame size is the plate's; the length is the take's, in frames (112 frames, 4.667 s, for a 4.649 s take); 24 fps; one audio stream, AAC 48 kHz mono.

  • Normal playback at 100 %, with the sound, before anything else.

  • Frame-step at 100 %: the first visible articulation against the first audible sound; the lips meet on every ب and م; the last consonant; no movement in a pause; the mouth closed in the hold.

  • At 200 %: teeth and mouth corners, the skin and jaw line around the redrawn region, light and grain matching the plate, no blur and no pasted-on look.

  • Side by side with the plate: the eyes, head, breath and hands are the plate's, and the face still matches the master.

  • The audio, decoded and compared with the upload for content and timing, then the native ear; then the first and last frames, and again in the final export.

A mouth-movement score cannot tell a passing pass from a failing one; the eye and the ear decide.

24.8 Failures and fixes

Symptom

Cause

Smallest fix

A word sounds wrong

the take

re-lock the take with the spelling repaired (Chapter 21); no more video credits

Sync right, face dead

the plate

re-plate with more business and a talking mouth (22.4)

The first sound is clipped

the audio window

restore 200 ms of head silence

Lips start before or after the sound

the take's head does not match the plate's talk onset

pad the head to the measured onset, or trim the plate (22.4)

The mouth moves in a pause or after the last word

the plate talked longer than the speech

shorten the plate's talk; end the take after its tail

Motion reversed, repeated or sped up

bounce (the default), loop or remap

set cut_off

The line is clipped at the end

the take outlasts the plate

a plate 0.3 to 0.5 s longer than the padded take

Head and jaw fight the new line

the approved clip spoke the old line

a talking plate of the same shot

The mouth is blurred or pasted on

the face is small, turned or occluded

a tighter, nearer-frontal plate; else redesign the shot (22.3)

The retry ladder, smallest lever first. (1) Fix the file, not the pass: the take's spelling, padding and confirmation. (2) Line up the windows: pad the head to the plate's talk onset, or trim the plate. (3) Re-plate on the same model with one change. (4) Re-plate on the other model. (5) Redesign the shot (22.3), record to picture (22.5), or, where a new performance is acceptable, change to Route B in writing. Two failed passes on one plate mean a new plate; two failed plates mean a new shot.

24.9 Limits, prices and rights

Item

Value

Dated

Price on Higgsfield

18.15 credits per pass, for a plate of about 5 s and a take of about 4.6 s (about 3.9 a second of take; growth with length not measured); no cost check

charged 12 Sep 2026

Price at the maker

$0.107–0.133 a second in the docs; pricing page $0.1334 (Hobbyist, Creator), $0.1266 (Growth), $0.1066 (Scale); plans $5, $19, $49 and $249 a month

checked 30 Sep 2026

Duration

maker: free 20 s, Hobbyist 1 min, Creator 5 min, Growth 10 min, Scale 30 min; a free account gets one sync-3 generation a month, up to 15 s, and the pricing page lists "No watermark" from Creator up. Higgsfield states no cap; plates of about 5 s and takes under about 4.7 s are the established sizes

checked 30 Sep 2026

Resolution and files

480p minimum, 1080p recommended, 4096 × 2160 maximum, 4K native output at the maker; Higgsfield returns the plate's size. MP4, MOV, WebM, AVI; 24, 25 or 30 fps constant; at least 10 Mbps; 8-bit SDR BT.709; audio WAV or MP3, one speaker, no music; direct upload 20 MB

checked 30 Sep 2026; formats page updated 22 May 2026

Languages

95+; nothing said about Egyptian or Saudi mouth shapes

checked 30 Sep 2026

Rights and terms. Higgsfield's terms (updated 26 Jul 2026, read 30 Sep) allow commercial use on paid plans, and its providers' policies ban the use of a likeness or a voice without consent. sync.so's terms (last revised 6 Jun 2024) require that you hold the rights, privacy and publicity included, in what you upload, and are silent on commercial use of the output. What may ship to a client: a pass on a plate whose face is generated or released, with a take from a person who signed (Chapters 21 and 34).

24.10 Version notes

Current. sync-3 is the maker's default model ("Default model for all users") and the only one Higgsfield offers as sync_so (checked 30 Sep 2026). The maker announced image-to-video on sync-3 on 9 June 2026 (changelog, 14 July). It generates all frames at once from a larger window around the face, detects obstructions automatically, handles profile and over-the-shoulder angles, and outputs 4K.

Older models. The maker's models page (1 June 2026) still lists lipsync-2 and lipsync-2-pro and advises "For the best quality, start with sync-3", at a higher price; Higgsfield's sync_so is sync-3 only.

A separate model. react-1 adjusts a performance (lips, face or head modes, with emotion prompts) on clips of up to 15 s, at $0.167 a second. It is not on Higgsfield's connector as read on 28 Sep 2026.

Announced: nothing newer than sync-3. Re-check when the next version ships: the model id on Higgsfield; the sync_mode default; whether a prompt or a speaker choice is exposed; the price per pass and whether a cost check works; the sizes and lengths returned.

Open questions

  • Which sync-3 build Higgsfield runs, and whether it has the maker's wider face window: not settleable from outside. Watch the face around the mouth on every pass.

  • Sync-3 on Egyptian and Saudi mouth shapes: the maker says nothing about dialects; only a native eye and ear settle it.

  • Not yet run: a take head of a second or more, room tone in the pad, a trimmed plate joined to its own business frames, a sound-off Kling plate, a plate over 5 s, a take over about 4.7 s, silence, a 1080p plate, an upscale after the pass, a profile, a beard or a hand near the mouth. Test on a scratch clip before a client shot depends on any of them.

What to remember

  1. Sync Lipsync 3 (current on 30 Sep 2026) redraws the mouth and the face around it to your audio. It has no prompt: the plate carries the acting and the take carries the words.

  2. Set sync_mode to cut_off on every pass, with a plate 0.3 to 0.5 s longer than the padded take. Never leave the default.

  3. Prepare the take the same way each time: clean single-speaker speech, 48 kHz mono, at least 200 ms of head and 400 ms of tail, uploaded and confirmed.

  4. The pass starts plate and take together: measure the plate's talk window and pad the take's head until the speech starts where the talk starts.

  5. The plate talks quietly with no words, keeps one face fully visible and nothing across the mouth, and lands on a held state.

  6. Fix a word in the take, a dead face in the plate, and neither by running the same pass again.

  7. Check at 100 % and 200 %, frame by frame at the onsets and closures, side by side with the plate, then by ear, then in the final export.

Comments


bottom of page