Chapter 18 · Wan 3.0
Part 3 · Motion
Wan 3.0 is Alibaba's video model and the economy choice in this book: words alone, references of pictures, footage and sound, or a first frame, for up to thirty seconds with its own sound. This chapter teaches what each input owns, how to bind, place and protect a reference, how to write an English speaking shot for a new performance, and what to check before you sign a clip.
In this chapter
What it takes in, and the two families of input that never share a job
The maker's prompt order, references to video, footage, and the maker's edit and extension
The economy speaking route for a new performance
Templates, three examples, settings, checks, failures, limits, prices and rights
Before you start. Chapter 13 chooses the operation. Chapter 14 teaches the seven elements of a moving shot, Chapter 10 the anchors, and Chapter 22 the speaking-shot contract. This chapter maps them to one tool and does not repeat them.
18.1 What it is
Wan 3.0 is a video model from Alibaba. One model takes text, a first frame, a first and last frame, references of images, clips and audio, an edit and an extension, and makes picture and sound together for up to 30 seconds at 30 frames a second. Alibaba's Model Studio lists it as wan3.0-video and, faster with the same capabilities, wan3.0-video-prime. It was announced on 13 Aug 2026 and reached Higgsfield on 24 Aug as wan3_0 and wan3_0_prime (Chapter 33). It is the current Wan: Alibaba's guides and API reference were last updated on 28 Sep and list no later model (checked 30 Sep 2026). The API reference still says "Currently in preview", so read its limits and prices again before a client job.
The jobs it takes.
A plate, or a cheap probe of one camera sentence, from words alone. Text alone cannot hold a face.
A shot composed from your anchors by number with no start frame, or one that borrows a camera path, a timing or an action from footage.
The economy new performance: an English line in a chosen voice. Compare it with Seedance on the same line before a client sees it (22.8).
The maker's edit and extension of a clip, which the connection may not offer (18.10).
The jobs it does not take.
A camera move measured from a signed frame: Kling (Chapter 16). A locked first frame with a voice: Seedance (Chapter 17).
The exact recorded take: a wordless plate, then Sync Lipsync 3 (Chapters 22 and 24). Egyptian and Saudi dialogue: the maker lists no dialogue languages, so the route is Seedance's (22.6, 22.8).
Titles, subtitles, logos and Arabic in the frame: post (Chapter 30). The same take twice: there is no seed.
How it is reached. Name the id in every request. The host may answer an image-driven request with a preset suggestion instead of a job: decline it and resubmit.
18.2 How it reads what you give it
Wan routes a job by the type of each input and the intent of your words. It counts Image N, Video N and Audio N separately, in upload order.
Input | What it owns | How you bind it |
Words | the whole picture when nothing is attached; otherwise the action and the sound | the prompt (up to 20,000 characters) |
Start frame, end frame | frame 0 and the landing; the result takes the aspect | their slots; the words describe the travel |
Image references | who, what and where; never the frame | "Image N", a name, one job, a fence |
Video references | a camera path, a timing, an action; or the clip to edit or extend | "Video N" and the channel taken |
Audio references | a voice's timbre, or the rhythm a picture follows | "Audio N", once per speaker |

Figure 18.1 — A job carries boundary frames or references, never both.
The two families never share a job. The maker: "First-and-last-frame mode and all-reference mode are mutually exclusive; once enabled, reference material input is no longer supported". The API refuses the pair, and so has the connection. A start frame takes no audio input either, so a shot with a voice cannot start from a locked first frame.
Numbers follow upload order, and must match it, or the wrong material is cited. The same file can be cited for its look and, elsewhere, its voice.
A reference guides the picture, does not fix the frame, and passes everything in it. Shot size, camera position and who stands where exist only if you write them. Give each file one job and a fence ("take nothing of its framing").
The maker also documents keyframe and storyboard references, which have not been run here; Seedance carries both (17.2).
18.3 The prompt, in this tool's dialect
Chapter 14 teaches the seven elements of a moving shot. This section gives the order Wan wants them in, its syntax and what it answers to.
The order
The maker's formula is "Prompt = [Overall description] + [Reference material citation: Image N / Video N / Audio N] + [Shot N (start-end seconds): Subject + Scene + Motion + Aesthetic control]", then dialogue, sound, style and mood, and negatives; skip any slot you do not need. A finished prompt runs:
The overall line: subject, mood and, if it matters, pace ("pacing shifts from steady restraint to explosive release").
"One continuous shot." The output may be one shot or several; the maker gives "Generate single shot" and "One continuous shot" for no split.
On a reference job, the manifest (10.5) and the locks (below), then the shot: size and height, one camera instruction, the subject and place, the action in order.
The world's life, the protection, the endpoint and hold, the pace, the exclusions and the sound.
A sentence is a probe; a finished prompt is complete. The look is words (light, colour, fine restrained grain); delivery grain is added in post.
Camera and time
Camera words are plain: "Use plain language, e.g. push in, pull out, orbit, handheld follow". A still camera is "Fixed shot, camera static, position unchanged": declare it, put it first in the shot, and give a move its pace and where it settles. The maker's examples use lens sizes and degrees but document no distances; this book keeps lens numbers out (Chapter 5), and a distance or an angle is a repair after two wrong moves (Chapter 14).
Stamps follow the formula, "Shot N (start-end seconds)", and "each shot connects end-to-end with no gaps or overlaps, 2–5 seconds per segment"; the bracket style is free, so pick one. One shot stays the default (13.5), since several bake the cut into the prompt.
Bind, place, lock, protect
The maker's reference sentence binds and places without describing: "The child in Image 1 stands in the entryway of the modern apartment in Image 2." With no first frame, the words do four jobs. Bind each file to a number, a name, one job and a fence. Place the subject: shot size, camera height, locked or one move, who is in frame. Lock each subject ("must strictly match Image 1") and count the people. Protect by name what the motion could break, after the action; write it as a keep line. Describe in full only what no file carries.
Footage as a reference
A clip lends a camera path, a timing, an action or an effect: "Replace the skateboarding man in Video 1 with the woman in Image 1, keeping the skateboarding action, trajectory, and scene unchanged". It passes everything in it, so name the one channel taken and send the rest to "take nothing of its people, place or light". Use footage you own (18.9).
Sound
The switch and the words both decide it. A described action, material or weather brings its own sound; name an effect as its own sentence when it is tied to a beat. Music has three levels: leave it to the model, specify it, or exclude it with the maker's "No BGM, generate only ambient and action sounds". "No dialogue" must be written, or the model decides whether to include dialogue.
The speaking route: a new performance
Each project decides how a line is delivered (Chapter 22). On the exact-take contract Wan makes at most the wordless plate, sound off (Chapter 24). The pass discards the plate's own sound and lays the take, so plate sound follows what has passed: Kling 3.0 on, Seedance 2.5 off, Wan 3.0 off (24.3). On the new-performance contract it is the economy route, for English only.
The keyframe is an image reference, not the start frame (Figure 18.1), so the framing is written and the mouth is signed on the keyframe: lips together, relaxed, any gesture the words trigger not yet made.
The voice is a sample bound once: "Voice timbre reference Audio 1", once per character. It is a voice you may use in a final (Chapter 34), 1 to 15 seconds, on a different sentence from the line.
The line is written once: "character + speaking verb + colon + quoted content", the language named, then Lip sync, then "says only this line". The business is timed by order, and the mouth is written.
Buy seconds for a breath, the line and a hold. With a recorded take as the audio reference, the clip's sound off and the take laid in, a frontal 4.6 s line with one pause held its mouth in both runs, and a three-quarter 4.65 s line with two pauses drifted in its second half (connection, 12 Sep 2026).
Check the returned track before use: whether it is the sample's own speech or your line in the sample's voice is not yet known (18.10).
Editing and extending, in the maker's terms
The maker documents both; whether the connection offers them is not confirmed (18.10). An edit is named by its verb (add, delete, replace, change to): "Edit the video: change xxx in video 1 to yyy, keep the rest of the frame unchanged". The maker finds that "modifying one item at a time is the most stable", and a stamp such as "4-6s" pins a segment. List what stays, the sound included. An extension carries an intent word ("extend", "continue"); restate the character, clothes and scene, then add the next beat only. Input and output together may not pass 30 seconds.
On Wan the direction words read the other way from Seedance. The maker's "extend backward" starts from the last frame and generates what comes after; "extend forward" ends at the first frame and generates what came before. On Seedance 2.5 backward means before (17.3). Write what happens instead: "Continue from the last frame of Video 1", or "Before the first frame of Video 1".
Exclusions. The connection has no negative field, so they live in the prose: three or fewer, each earned, with five as the ceiling (the maker: "Write only what you do not want to appear"), then the closing line.
18.4 Templates
Written for this book · template; not run · finished grade · Wan 3.0, text to video
Template — text to video
[Overall description, in a sentence]. One continuous shot.
Shot 1 (0-[T] seconds): [Shot size and camera height.] [Camera: fixed shot, camera static, position unchanged; or one move in plain words, with its pace and where it settles.] [Subject, place, time of day and light.] [The one thing that moves, in order, at a stated pace.] [What the world does around it: two or three small systems.] Keep [the named things the action could break]. The shot ends on [a composition you can cut on] and holds. [Sound.] [Up to three earned exclusions.] No BGM, generate only ambient and action sounds. No overlaid text, subtitles, logos, or interface graphics.Written for this book · template; not run · finished grade · Wan 3.0, references to video, with a Video line when the shot borrows from footage
Template — references to video
One continuous shot. [Shot size and camera height; locked, or the one move and where it settles.] Image 1 is [name]: take [face, hair, build, the named clothes and the one prop] only; [name] must strictly match Image 1; take nothing of its framing or background. Image 2 is [the place]: take [the room, its fixtures and its light]; take no people from it. [Video 1 gives only [the camera's path and speed / the action]: take nothing of its people, place or light.] Exactly [one / two] people, each appearing once: [who is in frame at the first moment, where, doing what]. [The action, in order, with its end.] [What the room does around them.] Keep [the named things the action could break]. The shot ends on [a composition you can cut on] and holds. [Sound.] [Up to three earned exclusions.] No BGM, generate only ambient and action sounds. No overlaid text, subtitles, logos, or interface graphics.Written for this book · template; not run · finished grade · Wan 3.0, an English speaking shot, new performance
Template — a speaking shot, new performance, English
One continuous shot. [Shot size and camera height; fixed shot, camera static, position unchanged.] Image 1 is [name]: take [face, hair, build, clothes, room and light]; [name] must strictly match Image 1; take nothing of its framing. Exactly one person, [name], [where, mouth closed and relaxed] from the first moment. Voice timbre reference Audio 1: take its timbre and manner only, not its words. [Name] says in English, [delivery in a few words]: "[the line]" Lip sync. [Business by order: "as she finishes the first sentence …", never a quotation of the line.] Lips and jaw move with the words and are still in the pauses and after the line. [Name] says only this line. After the line [she] holds, mouth closed, [where the eyes rest], to the end. Keep [identity, wardrobe, room, light] exactly as in Image 1. [Room sound.] No BGM. No other voices. No overlaid text, subtitles, logos, or interface graphics.Written for this book · template; not run · test grade · Wan 3.0, any operation, at 480p
Template — the test grade
[Overall line.] One continuous shot. Shot 1 (0-5 seconds): [Shot size; fixed shot, camera static, position unchanged.] [One action with its end.] [Sound.] No overlaid text, subtitles, logos, or interface graphics.18.5 Examples at the standard
Three finished-grade clips: a plate from words, a shot from anchors, and a line to camera.
Written for this book · food and drink · Wan 3.0, text to video, nothing attached, 6 s, 720p, 16:9, audio on · finished grade; not run
Example 18.1 — a baker draws six loaves from the oven at first light
A Cairo street baker draws round loaves of aish baladi from a wood-fired oven at first light, warm and unhurried. One continuous shot.
Shot 1 (0-6 seconds): A medium shot at counter height, on a tripod: fixed shot, camera static, position unchanged, framing the baker from the chest down across a worn steel counter and the oven's mouth. Flour-dusted forearms slide a long wooden peel out of the oven and turn it over the counter at an even pace; six blistered brown loaves slide onto a wooden tray and settle side by side. Steam lifts from the loaves into the oven's glow, flour dust drifts down through the beam, and the fire flickers on white wall tiles behind. Warm amber from the oven on the loaves, cool blue dawn from the street door. Keep exactly six loaves, round with a blistered crust, and the peel and tray at one size. The shot ends on the six loaves on the tray, the steam thinning, and holds. The scrape of the peel, loaves landing on the tray, the low roar of the fire. No dialogue. No BGM, generate only ambient and action sounds. No overlaid text, subtitles, logos, or interface graphics.Copy the frame that hides the face: words alone cannot hold a face, so the baker is framed from the chest down.
Copy the count and the small systems: "exactly six loaves" stops the tray gaining or losing bread; steam, dust and firelight keep a fixed frame alive.
Written for this book · automotive and luxury · Wan 3.0, two image references (a stand-in flacon, an empty terrace plate), 6 s, 720p, 16:9, audio on · finished grade; not run
Example 18.2 — a flacon on a Nile terrace at dusk, composed from two anchors
The last sun of the day moves across a fragrance flacon standing alone on a stone ledge above the Nile. One continuous shot.
Shot 1 (0-6 seconds): A close shot at the height of the ledge; the camera pushes in on a dolly, slowly and evenly, from the whole ledge to a close-up of the flacon's shoulder, and settles. Image 1 is the flacon: take its shape, proportions, glass, amber liquid and gold cap only; it must strictly match Image 1; its label area is blank and stays blank; take nothing of its background or light. Image 2 is the terrace: take the stone ledge, the balustrade and the river beyond; take no people from it. Exactly one flacon, upright at the centre of the ledge from the first moment. A sheer curtain at the edge of the frame lifts once and falls; far behind, a felucca crosses the river, out of focus; one warm highlight slides down the glass and gathers in the amber. Cool blue dusk from the open sky. Keep the flacon's shape, cap and liquid line exactly as in Image 1, and the ledge and skyline as in Image 2. The shot ends on the close-up of the shoulder, the light resting on the glass, and holds. Soft wind, the faint slap of water. No dialogue. No BGM, generate only ambient and action sounds. No overlaid text, subtitles, logos, or interface graphics.Copy the manifest and the placing: one job and one fence a file, the flacon locked with a count; a reference does not fix the frame, so the size, the height and the one move are written.
Copy the blank label: the wordmark is composited in post from the client's artwork.
Written for this book · beauty and personal care · Wan 3.0, the speaking keyframe as an image reference and an 8 s sample of a licensed English voice as the audio reference, 5 s, 720p, 16:9, audio on · English project, new-performance contract · finished grade; not run
Example 18.3 — a line to camera at the bathroom mirror, in the chosen English voice
One continuous shot. A medium close-up at eye level, the camera locked: fixed shot, camera static, position unchanged. Image 1 is Salma in her bathroom in the morning light: take her face, hair, build, the white cotton robe, the bathroom and its light; Salma must strictly match Image 1; take nothing of its framing. Exactly one person, Salma, stands at the vanity from the first moment, facing the lens, her mouth closed and relaxed. Voice timbre reference Audio 1: take its timbre and manner only, not its words. She draws one small breath and looks into the lens. Salma says in English, warmly and evenly: "Ten minutes. Then I am out of the door." Lip sync. As she finishes the first sentence her fingertips touch the edge of the sink. Her lips and jaw move with the words and are still in the pause and after the line. Salma says only this line. After the line she smiles slightly, mouth closed, and holds to the end, her eyes on the lens. Keep her identity, robe, room and light exactly as in Image 1. Soft room tone, a tap running behind her. No BGM. No other voices. No overlaid text, subtitles, logos, or interface graphics.Copy the routing: the keyframe is an image reference, so the framing is written and the voice can enter; a start frame would refuse the audio.
Copy the split of jobs: the frame carries her, the robe and the room; the sample the voice; the words the language, the delivery, the business and the mouth.
Copy the seconds: a breath, about 3 seconds of line, a hold.
18.6 Settings
The connection's contract was read on 27 Sep 2026; the maker's limits on 30 Sep 2026.
Setting | Values | Use |
Model | wan3_0 · wan3_0_prime | name it every time; Prime is faster and costs about 1.5 to 1.7 times as much |
Duration | 2 to 30 s; default 5; -1 is billed as 10 s | 5 to 6 s for most clips; never -1 |
Resolution | 480p · 720p · 1080p; default 720p | 480p to test a camera sentence; 720p to judge and usually keep; 1080p only when the delivery needs it |
Sound (generate_audio) | on · off; default on | off for a wordless plate (24.3); on for speech and room sound. The same price |
Aspect | auto · 16:9 · 9:16 · 1:1 · 4:3 · 3:4; none stated | set it every time; there is no 21:9, so matte in post |
Not offered | a seed, a negative field, a frame rate, a lip-sync flag, a language; enable_thinking is off and untested | fence in words; conform the rate in post |
Defaults that bite. Sound is on by default, so a plate for Sync Lipsync 3 is switched off (24.3). -1 costs a 10-second clip. Aspect has no default.
What comes back. A 720p job returns 1280 × 720 at 30 fps (ignore the job record's 1344 × 768), so conform to the sequence's rate on purpose (Chapter 30). With sound off there is no audio stream; with sound on and a voice reference the file carries its own track.
On the maker's own service the default resolution is 1080P, 21:9 is offered, and there are a seed (results still differ), a watermark switch (off by default) and prompt_extend (on by default: a language model rewrites the prompt).
18.7 Checks before you sign
On the returned file
Probe it: raster, frame rate (30), length, audio stream (Chapter 30).
The opening: everyone who should be there is, the count matches the cast, nothing is duplicated.
Faces at 200 % against the master or the in-look anchor; the costume; the product's shape and marks.
Leaks: no reference's backdrop, light or pose; no invented lettering.
The framing is the one written, not the keyframe's.
The move, at the start, the middle and the end: one dominant camera action, real parallax.
The ending: a composition you can cut on, then a live hold.
Hands, closures and labels at 100 % and 200 %.
Sound by ear: what owns the room; no music if excluded; no speech where "No dialogue" was written.
A speaking shot, heard word by word: the mouth moves with the words, is still in the pauses, and the track is compared with the sample.
Record the verdict with the inputs in upload order, the prompt as sent, the job number and the price (Chapters 8 and 33); acceptance is the director's (Chapter 29).
18.8 Failures and fixes
You see | Why | Smallest fix |
People appear in an empty plate | the model fills the space | "No people", and what makes emptiness plausible |
The camera moves, or a cut you did not ask for | no fixed-shot sentence; no single-shot line | the fixed-shot sentence first inside the shot; "One continuous shot" on the first line |
The framing differs from the keyframe | a reference does not fix the frame | write the shot size and camera position; if it must match, a Seedance or Kling start frame |
A reference's backdrop or light leaks in | no fence | "take nothing of its framing or background" |
The job is rejected | a start or end frame went in with a reference | remove the frame, or drop the references |
The mouth drifts in the second half | a three-quarter view, or two pauses | a frontal keyframe, one pause, a shorter line |
Speech where there should be none, or extra words | "No dialogue" or "says only this line" missing | write it; sound off for a wordless plate |
An extension goes the wrong way | direction words read the other way from Seedance | "Continue from the last frame of Video 1", or "Before the first frame of Video 1" |
The retry ladder is the one in 17.8, smallest lever first. On Wan, two failed retries mean the fix is upstream: rebuild the anchor, use a start frame on Seedance or Kling, or composite in post. Never a third wording of the same prompt.
18.9 Limits, prices and rights
The maker's limits (checked 30 Sep 2026).
What | Limit |
One generation | 2 to 30 s at 30 fps, MP4; with a video input, input plus output no more than 30 s |
Images | up to 10; 240 to 8000 px a side; ratio no more than 8:1; up to 20 MB each |
Clips | up to 5, 15 s in all; MP4 or MOV; 1 to 15 s each; 16 fps or more; up to 100 MB each |
Audio | up to 5, 15 s in all; WAV or MP3; 1 to 15 s each; up to 15 MB each |
Service | 20 references in all; a first or last frame excludes every one; the task and result link last 24 hours; 5 requests a second, 5 running at once |
Prices on the connection (about five US cents a credit, Higgsfield's planning figure, Aug 2026).
Job | Wan 3.0 | Prime | When |
5 s, 480p | 5 | 7.5 | quoted 27 Sep |
5 s, 720p, references or none, sound on or off | 8.75 | 15 | quoted 27 Sep |
5 s, 1080p | 17.5 | 30 | quoted 27 Sep |
6 s, 720p / 1080p; 10 s or -1, 720p | 10.5 / 21; 17.5 | not known | quoted 12 and 27 Sep |
With a video reference | not known | not known | never priced |
Prices at the maker (Alibaba's pricing page, updated 28 Sep 2026): in Beijing, Prime costs 0.45, 0.9 and 1.8 yuan a second at 480P, 720P and 1080P, and the standard model 0.3, 0.6 and 1.2, marked "limited-time 30% off"; other regions differ. A video input's seconds are billed with the output's; a failed task is not billed. Before every paid run, ask the platform for the exact price at no cost (Chapter 34).
Rights and terms.
Higgsfield's terms on outputs, on training from uploads and on what stays off the platform (read 30 Sep 2026) are in 34.5 and 34.6.
Alibaba's Model Studio service terms (read 28 Sep 2026) require AI-generated content to be labelled and forbid removing the service's labels. The acceptable-use policy that governs Wan on the connection's route has not been read.
What may ship. An accepted clip may ship, through Chapter 8's four gates: LOCK, SPEND, RIGHTS and ACCEPT.
18.10 Version notes
Current. Wan 3.0, announced 13 Aug 2026; the guides and the API reference updated 28 Sep 2026 (checked 30 Sep 2026). The API reference calls it a preview. Nothing later is listed.
Re-check when the next version ships, or the preview ends. The identifiers; the reference and length limits; the credit rate and the yuan prices; what the connection offers; the audio behaviour with a voice reference; the terms.
Not yet known
Whether sound on returns your sample's own speech or your line in the sample's voice. Run one 480p job with a sample on a different sentence and compare the track with the upload before a shot depends on it.
Whether the connection carries a clip as a video reference, an edit or an extension, and what it costs. Ask for the price at no cost with the clip attached, then run one 480p job.
More than one image reference (two people, or a person and a product), a line longer than about five seconds, and two speakers have not been run.
Whether the preview label lets limits, prices or output change without notice. Re-read the pages before a client job.
What to remember
A job carries boundary frames or references, never both; a voice is a reference, so a speaking shot on Wan is a reference shot.
A reference guides the picture and does not fix the frame: write the shot size, the camera position and who is where, give every file one job and a fence, and protect by name what the motion could break.
Open with the overall line, then "One continuous shot"; declare the camera as design, "Fixed shot, camera static, position unchanged".
Set the sound switch and write the sound: "No dialogue" must be written, and music is left, specified or excluded in words.
Wan's extension words read the other way from Seedance's: write what happens.
Wan is the economy route: test a camera sentence at 480p, compare a speaking line with Seedance on the same line, and never use it for Arabic.




Comments