top of page

Chapter 18 · Wan 3.0

Writer: Yasser Ashour
Yasser Ashour
11 minutes ago
18 min read

Part 3 · Motion

Wan 3.0 is Alibaba's video model and the economy choice in this book: words alone, references of pictures, footage and sound, or a first frame, for up to thirty seconds with its own sound. This chapter teaches what each input owns, how to bind, place and protect a reference, how to write an English speaking shot for a new performance, and what to check before you sign a clip.

In this chapter

  • What it takes in, and the two families of input that never share a job

  • The maker's prompt order, references to video, footage, and the maker's edit and extension

  • The economy speaking route for a new performance

  • Templates, three examples, settings, checks, failures, limits, prices and rights

Before you start. Chapter 13 chooses the operation. Chapter 14 teaches the seven elements of a moving shot, Chapter 10 the anchors, and Chapter 22 the speaking-shot contract. This chapter maps them to one tool and does not repeat them.

18.1 What it is

Wan 3.0 is a video model from Alibaba. One model takes text, a first frame, a first and last frame, references of images, clips and audio, an edit and an extension, and makes picture and sound together for up to 30 seconds at 30 frames a second. Alibaba's Model Studio lists it as wan3.0-video and, faster with the same capabilities, wan3.0-video-prime. It was announced on 13 Aug 2026 and reached Higgsfield on 24 Aug as wan3_0 and wan3_0_prime (Chapter 33). It is the current Wan: Alibaba's guides and API reference were last updated on 28 Sep and list no later model (checked 30 Sep 2026). The API reference still says "Currently in preview", so read its limits and prices again before a client job.

The jobs it takes.

  • A plate, or a cheap probe of one camera sentence, from words alone. Text alone cannot hold a face.

  • A shot composed from your anchors by number with no start frame, or one that borrows a camera path, a timing or an action from footage.

  • The economy new performance: an English line in a chosen voice. Compare it with Seedance on the same line before a client sees it (22.8).

  • The maker's edit and extension of a clip, which the connection may not offer (18.10).

The jobs it does not take.

  • A camera move measured from a signed frame: Kling (Chapter 16). A locked first frame with a voice: Seedance (Chapter 17).

  • The exact recorded take: a wordless plate, then Sync Lipsync 3 (Chapters 22 and 24). Egyptian and Saudi dialogue: the maker lists no dialogue languages, so the route is Seedance's (22.6, 22.8).

  • Titles, subtitles, logos and Arabic in the frame: post (Chapter 30). The same take twice: there is no seed.

How it is reached. Name the id in every request. The host may answer an image-driven request with a preset suggestion instead of a job: decline it and resubmit.

18.2 How it reads what you give it

Wan routes a job by the type of each input and the intent of your words. It counts Image N, Video N and Audio N separately, in upload order.

Input

What it owns

How you bind it

Words

the whole picture when nothing is attached; otherwise the action and the sound

the prompt (up to 20,000 characters)

Start frame, end frame

frame 0 and the landing; the result takes the aspect

their slots; the words describe the travel

Image references

who, what and where; never the frame

"Image N", a name, one job, a fence

Video references

a camera path, a timing, an action; or the clip to edit or extend

"Video N" and the channel taken

Audio references

a voice's timbre, or the rhythm a picture follows

"Audio N", once per speaker

fig18-1

Figure 18.1 — A job carries boundary frames or references, never both.

The two families never share a job. The maker: "First-and-last-frame mode and all-reference mode are mutually exclusive; once enabled, reference material input is no longer supported". The API refuses the pair, and so has the connection. A start frame takes no audio input either, so a shot with a voice cannot start from a locked first frame.

Numbers follow upload order, and must match it, or the wrong material is cited. The same file can be cited for its look and, elsewhere, its voice.

A reference guides the picture, does not fix the frame, and passes everything in it. Shot size, camera position and who stands where exist only if you write them. Give each file one job and a fence ("take nothing of its framing").

The maker also documents keyframe and storyboard references, which have not been run here; Seedance carries both (17.2).

18.3 The prompt, in this tool's dialect

Chapter 14 teaches the seven elements of a moving shot. This section gives the order Wan wants them in, its syntax and what it answers to.

The order

The maker's formula is "Prompt = [Overall description] + [Reference material citation: Image N / Video N / Audio N] + [Shot N (start-end seconds): Subject + Scene + Motion + Aesthetic control]", then dialogue, sound, style and mood, and negatives; skip any slot you do not need. A finished prompt runs:

  1. The overall line: subject, mood and, if it matters, pace ("pacing shifts from steady restraint to explosive release").

  2. "One continuous shot." The output may be one shot or several; the maker gives "Generate single shot" and "One continuous shot" for no split.

  3. On a reference job, the manifest (10.5) and the locks (below), then the shot: size and height, one camera instruction, the subject and place, the action in order.

  4. The world's life, the protection, the endpoint and hold, the pace, the exclusions and the sound.

A sentence is a probe; a finished prompt is complete. The look is words (light, colour, fine restrained grain); delivery grain is added in post.

Camera and time

Camera words are plain: "Use plain language, e.g. push in, pull out, orbit, handheld follow". A still camera is "Fixed shot, camera static, position unchanged": declare it, put it first in the shot, and give a move its pace and where it settles. The maker's examples use lens sizes and degrees but document no distances; this book keeps lens numbers out (Chapter 5), and a distance or an angle is a repair after two wrong moves (Chapter 14).

Stamps follow the formula, "Shot N (start-end seconds)", and "each shot connects end-to-end with no gaps or overlaps, 2–5 seconds per segment"; the bracket style is free, so pick one. One shot stays the default (13.5), since several bake the cut into the prompt.

Bind, place, lock, protect

The maker's reference sentence binds and places without describing: "The child in Image 1 stands in the entryway of the modern apartment in Image 2." With no first frame, the words do four jobs. Bind each file to a number, a name, one job and a fence. Place the subject: shot size, camera height, locked or one move, who is in frame. Lock each subject ("must strictly match Image 1") and count the people. Protect by name what the motion could break, after the action; write it as a keep line. Describe in full only what no file carries.

Footage as a reference

A clip lends a camera path, a timing, an action or an effect: "Replace the skateboarding man in Video 1 with the woman in Image 1, keeping the skateboarding action, trajectory, and scene unchanged". It passes everything in it, so name the one channel taken and send the rest to "take nothing of its people, place or light". Use footage you own (18.9).

Sound

The switch and the words both decide it. A described action, material or weather brings its own sound; name an effect as its own sentence when it is tied to a beat. Music has three levels: leave it to the model, specify it, or exclude it with the maker's "No BGM, generate only ambient and action sounds". "No dialogue" must be written, or the model decides whether to include dialogue.

The speaking route: a new performance

Each project decides how a line is delivered (Chapter 22). On the exact-take contract Wan makes at most the wordless plate, sound off (Chapter 24). The pass discards the plate's own sound and lays the take, so plate sound follows what has passed: Kling 3.0 on, Seedance 2.5 off, Wan 3.0 off (24.3). On the new-performance contract it is the economy route, for English only.

  • The keyframe is an image reference, not the start frame (Figure 18.1), so the framing is written and the mouth is signed on the keyframe: lips together, relaxed, any gesture the words trigger not yet made.

  • The voice is a sample bound once: "Voice timbre reference Audio 1", once per character. It is a voice you may use in a final (Chapter 34), 1 to 15 seconds, on a different sentence from the line.

  • The line is written once: "character + speaking verb + colon + quoted content", the language named, then Lip sync, then "says only this line". The business is timed by order, and the mouth is written.

  • Buy seconds for a breath, the line and a hold. With a recorded take as the audio reference, the clip's sound off and the take laid in, a frontal 4.6 s line with one pause held its mouth in both runs, and a three-quarter 4.65 s line with two pauses drifted in its second half (connection, 12 Sep 2026).

  • Check the returned track before use: whether it is the sample's own speech or your line in the sample's voice is not yet known (18.10).

Editing and extending, in the maker's terms

The maker documents both; whether the connection offers them is not confirmed (18.10). An edit is named by its verb (add, delete, replace, change to): "Edit the video: change xxx in video 1 to yyy, keep the rest of the frame unchanged". The maker finds that "modifying one item at a time is the most stable", and a stamp such as "4-6s" pins a segment. List what stays, the sound included. An extension carries an intent word ("extend", "continue"); restate the character, clothes and scene, then add the next beat only. Input and output together may not pass 30 seconds.

On Wan the direction words read the other way from Seedance. The maker's "extend backward" starts from the last frame and generates what comes after; "extend forward" ends at the first frame and generates what came before. On Seedance 2.5 backward means before (17.3). Write what happens instead: "Continue from the last frame of Video 1", or "Before the first frame of Video 1".

Exclusions. The connection has no negative field, so they live in the prose: three or fewer, each earned, with five as the ceiling (the maker: "Write only what you do not want to appear"), then the closing line.

18.4 Templates

Written for this book · template; not run · finished grade · Wan 3.0, text to video

Template — text to video

[Overall description, in a sentence]. One continuous shot.
Shot 1 (0-[T] seconds): [Shot size and camera height.] [Camera: fixed shot, camera static, position unchanged; or one move in plain words, with its pace and where it settles.] [Subject, place, time of day and light.] [The one thing that moves, in order, at a stated pace.] [What the world does around it: two or three small systems.] Keep [the named things the action could break]. The shot ends on [a composition you can cut on] and holds. [Sound.] [Up to three earned exclusions.] No BGM, generate only ambient and action sounds. No overlaid text, subtitles, logos, or interface graphics.

Written for this book · template; not run · finished grade · Wan 3.0, references to video, with a Video line when the shot borrows from footage

Template — references to video

One continuous shot. [Shot size and camera height; locked, or the one move and where it settles.] Image 1 is [name]: take [face, hair, build, the named clothes and the one prop] only; [name] must strictly match Image 1; take nothing of its framing or background. Image 2 is [the place]: take [the room, its fixtures and its light]; take no people from it. [Video 1 gives only [the camera's path and speed / the action]: take nothing of its people, place or light.] Exactly [one / two] people, each appearing once: [who is in frame at the first moment, where, doing what]. [The action, in order, with its end.] [What the room does around them.] Keep [the named things the action could break]. The shot ends on [a composition you can cut on] and holds. [Sound.] [Up to three earned exclusions.] No BGM, generate only ambient and action sounds. No overlaid text, subtitles, logos, or interface graphics.

Written for this book · template; not run · finished grade · Wan 3.0, an English speaking shot, new performance

Template — a speaking shot, new performance, English

One continuous shot. [Shot size and camera height; fixed shot, camera static, position unchanged.] Image 1 is [name]: take [face, hair, build, clothes, room and light]; [name] must strictly match Image 1; take nothing of its framing. Exactly one person, [name], [where, mouth closed and relaxed] from the first moment. Voice timbre reference Audio 1: take its timbre and manner only, not its words. [Name] says in English, [delivery in a few words]: "[the line]" Lip sync. [Business by order: "as she finishes the first sentence …", never a quotation of the line.] Lips and jaw move with the words and are still in the pauses and after the line. [Name] says only this line. After the line [she] holds, mouth closed, [where the eyes rest], to the end. Keep [identity, wardrobe, room, light] exactly as in Image 1. [Room sound.] No BGM. No other voices. No overlaid text, subtitles, logos, or interface graphics.

Written for this book · template; not run · test grade · Wan 3.0, any operation, at 480p

Template — the test grade

[Overall line.] One continuous shot. Shot 1 (0-5 seconds): [Shot size; fixed shot, camera static, position unchanged.] [One action with its end.] [Sound.] No overlaid text, subtitles, logos, or interface graphics.

18.5 Examples at the standard

Three finished-grade clips: a plate from words, a shot from anchors, and a line to camera.

Written for this book · food and drink · Wan 3.0, text to video, nothing attached, 6 s, 720p, 16:9, audio on · finished grade; not run

Example 18.1 — a baker draws six loaves from the oven at first light

A Cairo street baker draws round loaves of aish baladi from a wood-fired oven at first light, warm and unhurried. One continuous shot.
Shot 1 (0-6 seconds): A medium shot at counter height, on a tripod: fixed shot, camera static, position unchanged, framing the baker from the chest down across a worn steel counter and the oven's mouth. Flour-dusted forearms slide a long wooden peel out of the oven and turn it over the counter at an even pace; six blistered brown loaves slide onto a wooden tray and settle side by side. Steam lifts from the loaves into the oven's glow, flour dust drifts down through the beam, and the fire flickers on white wall tiles behind. Warm amber from the oven on the loaves, cool blue dawn from the street door. Keep exactly six loaves, round with a blistered crust, and the peel and tray at one size. The shot ends on the six loaves on the tray, the steam thinning, and holds. The scrape of the peel, loaves landing on the tray, the low roar of the fire. No dialogue. No BGM, generate only ambient and action sounds. No overlaid text, subtitles, logos, or interface graphics.
  • Copy the frame that hides the face: words alone cannot hold a face, so the baker is framed from the chest down.

  • Copy the count and the small systems: "exactly six loaves" stops the tray gaining or losing bread; steam, dust and firelight keep a fixed frame alive.

Written for this book · automotive and luxury · Wan 3.0, two image references (a stand-in flacon, an empty terrace plate), 6 s, 720p, 16:9, audio on · finished grade; not run

Example 18.2 — a flacon on a Nile terrace at dusk, composed from two anchors

The last sun of the day moves across a fragrance flacon standing alone on a stone ledge above the Nile. One continuous shot.
Shot 1 (0-6 seconds): A close shot at the height of the ledge; the camera pushes in on a dolly, slowly and evenly, from the whole ledge to a close-up of the flacon's shoulder, and settles. Image 1 is the flacon: take its shape, proportions, glass, amber liquid and gold cap only; it must strictly match Image 1; its label area is blank and stays blank; take nothing of its background or light. Image 2 is the terrace: take the stone ledge, the balustrade and the river beyond; take no people from it. Exactly one flacon, upright at the centre of the ledge from the first moment. A sheer curtain at the edge of the frame lifts once and falls; far behind, a felucca crosses the river, out of focus; one warm highlight slides down the glass and gathers in the amber. Cool blue dusk from the open sky. Keep the flacon's shape, cap and liquid line exactly as in Image 1, and the ledge and skyline as in Image 2. The shot ends on the close-up of the shoulder, the light resting on the glass, and holds. Soft wind, the faint slap of water. No dialogue. No BGM, generate only ambient and action sounds. No overlaid text, subtitles, logos, or interface graphics.
  • Copy the manifest and the placing: one job and one fence a file, the flacon locked with a count; a reference does not fix the frame, so the size, the height and the one move are written.

  • Copy the blank label: the wordmark is composited in post from the client's artwork.

Written for this book · beauty and personal care · Wan 3.0, the speaking keyframe as an image reference and an 8 s sample of a licensed English voice as the audio reference, 5 s, 720p, 16:9, audio on · English project, new-performance contract · finished grade; not run

Example 18.3 — a line to camera at the bathroom mirror, in the chosen English voice

One continuous shot. A medium close-up at eye level, the camera locked: fixed shot, camera static, position unchanged. Image 1 is Salma in her bathroom in the morning light: take her face, hair, build, the white cotton robe, the bathroom and its light; Salma must strictly match Image 1; take nothing of its framing. Exactly one person, Salma, stands at the vanity from the first moment, facing the lens, her mouth closed and relaxed. Voice timbre reference Audio 1: take its timbre and manner only, not its words. She draws one small breath and looks into the lens. Salma says in English, warmly and evenly: "Ten minutes. Then I am out of the door." Lip sync. As she finishes the first sentence her fingertips touch the edge of the sink. Her lips and jaw move with the words and are still in the pause and after the line. Salma says only this line. After the line she smiles slightly, mouth closed, and holds to the end, her eyes on the lens. Keep her identity, robe, room and light exactly as in Image 1. Soft room tone, a tap running behind her. No BGM. No other voices. No overlaid text, subtitles, logos, or interface graphics.
  • Copy the routing: the keyframe is an image reference, so the framing is written and the voice can enter; a start frame would refuse the audio.

  • Copy the split of jobs: the frame carries her, the robe and the room; the sample the voice; the words the language, the delivery, the business and the mouth.

  • Copy the seconds: a breath, about 3 seconds of line, a hold.

18.6 Settings

The connection's contract was read on 27 Sep 2026; the maker's limits on 30 Sep 2026.

Setting

Values

Use

Model

wan3_0 · wan3_0_prime

name it every time; Prime is faster and costs about 1.5 to 1.7 times as much

Duration

2 to 30 s; default 5; -1 is billed as 10 s

5 to 6 s for most clips; never -1

Resolution

480p · 720p · 1080p; default 720p

480p to test a camera sentence; 720p to judge and usually keep; 1080p only when the delivery needs it

Sound (generate_audio)

on · off; default on

off for a wordless plate (24.3); on for speech and room sound. The same price

Aspect

auto · 16:9 · 9:16 · 1:1 · 4:3 · 3:4; none stated

set it every time; there is no 21:9, so matte in post

Not offered

a seed, a negative field, a frame rate, a lip-sync flag, a language; enable_thinking is off and untested

fence in words; conform the rate in post

Defaults that bite. Sound is on by default, so a plate for Sync Lipsync 3 is switched off (24.3). -1 costs a 10-second clip. Aspect has no default.

What comes back. A 720p job returns 1280 × 720 at 30 fps (ignore the job record's 1344 × 768), so conform to the sequence's rate on purpose (Chapter 30). With sound off there is no audio stream; with sound on and a voice reference the file carries its own track.

On the maker's own service the default resolution is 1080P, 21:9 is offered, and there are a seed (results still differ), a watermark switch (off by default) and prompt_extend (on by default: a language model rewrites the prompt).

18.7 Checks before you sign

On the returned file

  • Probe it: raster, frame rate (30), length, audio stream (Chapter 30).

  • The opening: everyone who should be there is, the count matches the cast, nothing is duplicated.

  • Faces at 200 % against the master or the in-look anchor; the costume; the product's shape and marks.

  • Leaks: no reference's backdrop, light or pose; no invented lettering.

  • The framing is the one written, not the keyframe's.

  • The move, at the start, the middle and the end: one dominant camera action, real parallax.

  • The ending: a composition you can cut on, then a live hold.

  • Hands, closures and labels at 100 % and 200 %.

  • Sound by ear: what owns the room; no music if excluded; no speech where "No dialogue" was written.

  • A speaking shot, heard word by word: the mouth moves with the words, is still in the pauses, and the track is compared with the sample.

Record the verdict with the inputs in upload order, the prompt as sent, the job number and the price (Chapters 8 and 33); acceptance is the director's (Chapter 29).

18.8 Failures and fixes

You see

Why

Smallest fix

People appear in an empty plate

the model fills the space

"No people", and what makes emptiness plausible

The camera moves, or a cut you did not ask for

no fixed-shot sentence; no single-shot line

the fixed-shot sentence first inside the shot; "One continuous shot" on the first line

The framing differs from the keyframe

a reference does not fix the frame

write the shot size and camera position; if it must match, a Seedance or Kling start frame

A reference's backdrop or light leaks in

no fence

"take nothing of its framing or background"

The job is rejected

a start or end frame went in with a reference

remove the frame, or drop the references

The mouth drifts in the second half

a three-quarter view, or two pauses

a frontal keyframe, one pause, a shorter line

Speech where there should be none, or extra words

"No dialogue" or "says only this line" missing

write it; sound off for a wordless plate

An extension goes the wrong way

direction words read the other way from Seedance

"Continue from the last frame of Video 1", or "Before the first frame of Video 1"

The retry ladder is the one in 17.8, smallest lever first. On Wan, two failed retries mean the fix is upstream: rebuild the anchor, use a start frame on Seedance or Kling, or composite in post. Never a third wording of the same prompt.

18.9 Limits, prices and rights

The maker's limits (checked 30 Sep 2026).

What

Limit

One generation

2 to 30 s at 30 fps, MP4; with a video input, input plus output no more than 30 s

Images

up to 10; 240 to 8000 px a side; ratio no more than 8:1; up to 20 MB each

Clips

up to 5, 15 s in all; MP4 or MOV; 1 to 15 s each; 16 fps or more; up to 100 MB each

Audio

up to 5, 15 s in all; WAV or MP3; 1 to 15 s each; up to 15 MB each

Service

20 references in all; a first or last frame excludes every one; the task and result link last 24 hours; 5 requests a second, 5 running at once

Prices on the connection (about five US cents a credit, Higgsfield's planning figure, Aug 2026).

Job

Wan 3.0

Prime

When

5 s, 480p

5

7.5

quoted 27 Sep

5 s, 720p, references or none, sound on or off

8.75

15

quoted 27 Sep

5 s, 1080p

17.5

30

quoted 27 Sep

6 s, 720p / 1080p; 10 s or -1, 720p

10.5 / 21; 17.5

not known

quoted 12 and 27 Sep

With a video reference

not known

not known

never priced

Prices at the maker (Alibaba's pricing page, updated 28 Sep 2026): in Beijing, Prime costs 0.45, 0.9 and 1.8 yuan a second at 480P, 720P and 1080P, and the standard model 0.3, 0.6 and 1.2, marked "limited-time 30% off"; other regions differ. A video input's seconds are billed with the output's; a failed task is not billed. Before every paid run, ask the platform for the exact price at no cost (Chapter 34).

Rights and terms.

  • Higgsfield's terms on outputs, on training from uploads and on what stays off the platform (read 30 Sep 2026) are in 34.5 and 34.6.

  • Alibaba's Model Studio service terms (read 28 Sep 2026) require AI-generated content to be labelled and forbid removing the service's labels. The acceptable-use policy that governs Wan on the connection's route has not been read.

  • What may ship. An accepted clip may ship, through Chapter 8's four gates: LOCK, SPEND, RIGHTS and ACCEPT.

18.10 Version notes

Current. Wan 3.0, announced 13 Aug 2026; the guides and the API reference updated 28 Sep 2026 (checked 30 Sep 2026). The API reference calls it a preview. Nothing later is listed.

Re-check when the next version ships, or the preview ends. The identifiers; the reference and length limits; the credit rate and the yuan prices; what the connection offers; the audio behaviour with a voice reference; the terms.

Not yet known

  • Whether sound on returns your sample's own speech or your line in the sample's voice. Run one 480p job with a sample on a different sentence and compare the track with the upload before a shot depends on it.

  • Whether the connection carries a clip as a video reference, an edit or an extension, and what it costs. Ask for the price at no cost with the clip attached, then run one 480p job.

  • More than one image reference (two people, or a person and a product), a line longer than about five seconds, and two speakers have not been run.

  • Whether the preview label lets limits, prices or output change without notice. Re-read the pages before a client job.

What to remember

  1. A job carries boundary frames or references, never both; a voice is a reference, so a speaking shot on Wan is a reference shot.

  2. A reference guides the picture and does not fix the frame: write the shot size, the camera position and who is where, give every file one job and a fence, and protect by name what the motion could break.

  3. Open with the overall line, then "One continuous shot"; declare the camera as design, "Fixed shot, camera static, position unchanged".

  4. Set the sound switch and write the sound: "No dialogue" must be written, and music is left, specified or excluded in words.

  5. Wan's extension words read the other way from Seedance's: write what happens.

  6. Wan is the economy route: test a camera sentence at 480p, compare a speaking line with Seedance on the same line, and never use it for Arabic.

Comments


bottom of page