top of page

Reference A · Quick cards

Writer: Yasser Ashour
Yasser Ashour
12 minutes ago
27 min read

Reference

One card for each tool, in book order: version and date checked, what to reach for it for, the order of its prompt, the settings to set every time, limits, price, the one mistake never to make, and the chapter to read. Every fact is copied from the tool's own section.

In this reference: thirteen cards. The top row gives the current version and the date it was checked. Every limit and price carries the date its section gives; where a section says "checked 30 Sep 2026 unless dated otherwise", so do the cards. Ask for the price of the exact request on the day. A Higgsfield credit is worth about five US cents; Suno and ElevenLabs credits are separate. Where a card and its chapter differ, the chapter is right.

A.0 Which card

Card

Tool

Version and date checked

Chapter

A.1

Nano Banana Pro

Gemini 3 Pro Image, 30 Sep 2026

11

A.2

GPT Image 2.5

Flare and Sunburst, snapshots of 8 Sep 2026, 30 Sep 2026

12

A.3

Kling 3.0, and 4.0 as announced

VIDEO 3.0, 30 Sep 2026; 4.0 announced 28 Sep 2026

16

A.4

Seedance 2.5

dreamina-seedance-2-5-260628, 30 Sep 2026

17

A.5

Wan 3.0

wan3.0-video and Prime, 30 Sep 2026

18

A.6

Veo 3.1 and MiniMax H3

Veo 3.1 and H3, 30 Sep 2026

19

A.7

ElevenLabs v4 and v3

eleven_v4 and eleven_v3, 30 Sep 2026

23

A.8

Sync Lipsync 3

sync-3, 30 Sep 2026

24

A.9

Suno

v6, 30 Sep 2026

26

A.10

Eleven Music

Music 2.5, 30 Sep 2026

27

A.11

Firefly sound effects, Stable Audio 3 and Lyria

Firefly Audio Model, Stable Audio 3.0, Lyria 3.5, 30 Sep 2026

28

A.12

Premiere Pro and the upscalers

Premiere 26.5, 30 Sep 2026

32

A.13

Higgsfield

read by date, 30 Sep 2026

33

A.1 Nano Banana Pro

Version and date

Google's Gemini 3 Pro Image (model id gemini-3-pro-image), released November 2025, generally available 28 May 2026. Google's pages show no newer Pro image model (checked 30 Sep 2026). Nano Banana 2 (gemini-3.1-flash-image) is a different, Flash-class model.

Reach for it for

The stills camera: hero frames, the master and anchors, keyframes the video models animate, plates of rooms and products, treatment boards, and repairs of a frame that is 80 % right. Best with faces, rooms, products in rooms and several references blended. For a transparent cut-out or a layout-led frame use GPT Image 2.5 (A.2); brand marks come from the real artwork in post; nothing here moves.

Prompt order

The twelve elements of Chapter 9 as one paragraph of plain prose: operation and medium → subject and frozen instant → identity clause → wardrobe and props → frame inventory and edges → environment → must-show anchors → camera placement → light → atmosphere → finish → at most five exclusions, then "no overlaid text, subtitles, logos, or interface graphics". An anchored frame runs manifest → locked identity clause → change contract → the frame.

Set every time

Model nano_banana_pro, by name resolution 2k as the baseline (1k costs the same; 4k when a crop or a large delivery needs the pixels) aspect from the delivery aspect, ten ratios, no automatic choice variants 1 for a finished prompt pictures in the order identity, state or body, frame or plate, the one object, style: two to five, one job each. Aspect and resolution are settings, never words.

Limits

1K, 2K and 4K in ten aspects; 16:9 is 1376 × 768, 2752 × 1536 and 5504 × 3072 14 pictures in all: up to 6 objects, 5 characters and 3 style references (Google); Higgsfield states no maximum input cap 65,536 tokens, output cap 32,768 variants 1 to 4 no word limit; working bands 400 to 600 words for a hero, 150 to 300 for an insert, 40 to 150 for an edit no mask, seed, exclusion field or transparent background; thinking always on SynthID on every output (checked 30 Sep 2026).

Price

On Higgsfield (checked 27 Sep 2026): 1k and 2k cost 2 credits, 4k costs 4, and two references at 2k also cost 2; each variant is charged; about $0.10 for a 2k frame and $0.20 for a 4k. At Google (Gemini API paid tier, pricing page updated 24 Sep 2026): $0.134 per 1K or 2K image, $0.24 per 4K; batch $0.067 and $0.12.

The one mistake

Sending a request that names no model. A request with no model runs GPT Image 2.5 (checked 27 Sep 2026), and the job history records nano_banana_2 for a job submitted as Nano Banana Pro, so log both labels.

Read

Chapter 11: 11.3 the prompt, 11.6 settings, 11.9 limits, prices and rights.

A.2 GPT Image 2.5

Version and date

OpenAI's image model, current on 30 Sep 2026, in two models with snapshots dated 8 Sep 2026: Flare, the small model, optimized for speed, and Sunburst, the base model, optimized for quality and the one OpenAI points to where editing precision matters most. GPT Image 2 is the previous generation and is not used.

Reach for it for

An element on a transparent background (the only stills route with a real alpha channel); a frame led by text or layout, such as a key visual, a poster study, a diagram or a packaging flat; a bounded edit of a frame with no signed face; fast exploration. A signed identity carried through several references, and the photographic keyframe, stay on Nano Banana Pro (A.1). Type that ships is set in post. The host offers no mask.

Prompt order

A photographic frame: the twelve elements of Chapter 9, one paragraph. A layout, a diagram or a text-heavy brief: four labels in order, Scene → Subject → Details → Constraints. An edit: open with the change in physical terms → its consequences (a shadow, a contact) → what stays, region by region; one change per edit, the preserve list repeated. A transparent element: describe only the element, with no ground, and repeat that the transparent background is preserved.

Set every time

All four: variant (Flare for exploration, layouts and elements; Sunburst for an edit where precision matters; an unwritten variant falls back to Flare) quality high (medium for layout drafts, low only for a first glance; the default is low) resolution 2k (4k for a hero you will crop or print; the default is 1k) aspect, always set (fifteen are listed; no 1.85:1, make 16:9 and crop) background transparent for an element, opaque otherwise one picture slot, named in the prompt.

Limits

One picture slot on the host, no stated maximum complex prompts may take up to 2 minutes (OpenAI) 1k, 2k or 4k with fifteen aspects not available: seed, exclusions field, mask, output format at 2k and 16:9 the host returned 2688 × 1520 in September 2026 (Nano Banana Pro returns 2752 × 1536) a blocked job is not retried unchanged (checked 30 Sep 2026).

Price

Higgsfield, Flare, credits per frame (read 27 Sep 2026): low at 1k 0.25 (the defaults); medium at 2k 1; high 1.5 at 1k, 2.75 at 2k, 4.25 at 4k; xhigh at 2k 4.5; max at 2k 9. Sunburst at high and 2k quoted the same 2.75. About 14 cents for a high 2k frame. The same shape quoted 5.5, 3 and 2.75 credits between 11 and 27 Sep. OpenAI's own service: USD 8 per million image tokens in (2 cached) and 30 out (checked 30 Sep 2026).

The one mistake

Running on the host's defaults, which are the lowest quality tier at the smallest size (low at 1k) and the Flare variant. Write all four settings on every job, and the settings line into the run record.

Read

Chapter 12: 12.3 the prompt, 12.6 settings, 12.9 limits, prices and rights.

A.3 Kling 3.0, and 4.0 as announced

Version and date

Kling VIDEO 3.0, Kuaishou's video model (the maker's guide is dated 6 Feb 2026), with a native 4K mode added on 23 Apr 2026, and Turbo and stronger video editing in the Omni variant added on 17 Jun 2026 (checked 30 Sep 2026). On the Higgsfield connector: kling3_0, kling3_0_turbo and kling_video_edit (contract last read 28 Sep 2026).

Reach for it for

One signed still, moved: the camera is the shot (a dolly, a track, a crane, a small push whose size matters); a product or hand insert; lettering that must survive a gentle move; a quiet performance with no words; the wordless plate the exact-take speaking route needs. For more than one liquid or cloth system, a crowd near the camera, or a face with no fixed frame, use Seedance 2.5 (A.4). Egyptian or Saudi dialogue is never Kling's.

Prompt order

One flowing paragraph of prose, 120 to 220 words as a drafting expectation. When the move is the shot, open on the camera: the verb, the rig with its weight, the destination, one magnitude and a speed profile that ends. When the shot is an action, open on who does what and declare the camera next. Then: the world's life → protection by name → the endpoint as a composition and its hold → the closing line. Each moving system gets one approximate magnitude, its duration and its visible result.

Set every time

Model kling3_0, named mode std (pro for a hero shot at 1080p; 4k only after a probe) duration 3 to 15 whole seconds, default 5 aspect 16:9, 9:16 or 1:1, no default sound off for every clip whose sound you design (the default is on), on for a talking plate start frame: the signed keyframe end frame only when the landing is the point.

Limits

3 to 15 whole seconds (five has run often, ten once, fifteen never) one start frame, one end frame, a prompt and a sound switch; no references, no audio, no video, except Omni Edit's source clip speech from text in Chinese, English, Japanese, Korean and Spanish returned: std 1284 × 716, pro 1912 × 1080, 24 frames a second, 121 frames for 5 s (probes of 12 to 26 Sep 2026); the 4k mode has never been probed.

Price

Higgsfield credits, about five US cents each: std, 5 s, sound off, start frame 6.25 (charged 12, 26 and 27 Sep 2026); std, 10 s, sound off 12.5; std, 5 s, sound on 8.75 (charged 12 Sep); pro, 5 s, sound off 7.5; 4k, 5 s, sound off 30 (quote of 21 Sep); Turbo, 5 s, 7.5 at 720p and 16:9, 10 at 1080p and 9:16. A std plate, sound on, plus a Sync Lipsync 3 pass: about 26.9. Kling's own site is priced in its own credits per second, a different currency.

The one mistake

Leaving the sound switch on its default. Sound is on unless you turn it off, and a generated mix cannot be taken apart: an unset switch is a decision you did not make.

Read

Chapter 16: 16.3 the prompt, 16.6 settings, 16.9 limits, prices and rights; 16.1 and 16.10 for 4.0.

Kling 4.0, as announced (28 Sep 2026): not available to use

  • Kling's release note says Kling 4.0 will launch officially in October. Its first member, Kling 4.0 Flash, became available on 28 September to a limited group of users. The connector's model list, last read on 28 Sep, has no 4.0. Kling gives no price.

  • What Kling states for 4.0: clips of 3 to 30 seconds (Flash 3 to 20); 720p, 1080p and 4K (Flash 720p); 10-bit HDR at 1080p and 4K marked as coming soon; up to 10 keyframe images; up to 15 reference items in all; a prompt of up to 8,000 tokens; stereo audio, with dialogue in Chinese, English, Japanese, Korean, Spanish, Portuguese, German, French and more. Arabic is on neither list.

  • Every setting on this card is 3.0's until the connector lists 4.0 and a first job has been tested (16.10).

A.4 Seedance 2.5

Version and date

ByteDance's Seedance 2.5, dreamina-seedance-2-5-260628 on BytePlus: announced 31 Jul 2026, on Higgsfield since 6 Aug, prompt guide updated 28 Sep. No later model is listed (checked 30 Sep 2026). On the Higgsfield connection as seedance_2_5 (contract read 28 Sep 2026).

Reach for it for

A shot composed from your anchors with no start frame (references to video); a shot from a start frame, or from a start frame and an end frame; a shot that borrows a camera path or timing from footage; the edit or extension of an accepted clip; a speaking shot when a new performance is acceptable; a long take of up to 30 seconds. A camera move whose size must be measured is Kling's (A.3). Egyptian and Saudi lines are never typed for it.

Prompt order

On a reference job: the manifest and the locks (bind, place, lock, protect) → the subject and the event → one declared camera instruction → the world's life, protection by name, the endpoint and hold, the pace, up to three earned exclusions and the sound → the closing line "No overlaid text, subtitles, logos, or interface graphics." Time in whole seconds: "0-3 seconds ... 3-7 seconds". The maker's sound notation: () music, <> sound effects, {} dialogue, never 【】 (whether the connection passes it is not yet known).

Set every time

Model seedance_2_5, named mode: t2v is the default; omni_reference whenever anything is attached, a start frame included; video_edit; video_extension duration 4 to 30 s, default 5 (5 to 6 s for most clips) resolution 480p to prove a route, 720p (the default) to judge, 1080p only when the delivery needs it aspect, set every time sound on or off, default on (off for a plate) bitrate standard variants one.

Limits

The maker's (checked 30 Sep 2026): one generation 4 to 30 s 50 assets in a job: 30 images, 10 videos, 10 audios images 300 to 6000 px a side, ratio 0.4 to 2.5, under 30 MB videos 2 to 30 s each, 30 s in all, 24 to 60 fps, up to 200 MB, MP4 or MOV; an edit source 4 to 30 s audio WAV or MP3, 2 to 30 s each, 30 s in all, up to 15 MB output 480p and 720p (8-bit), 1080p (10-bit HEVC). Returned: 720p is 1280 × 720 at 24 fps, and a 5 s request returns 121 frames, 5.04 s. No seed, negative field or mask.

Price

Higgsfield credits, about five US cents each: 5 s, 720p 35 (charged 20 and 26 to 27 Sep); 6 s, 720p 42 (charged 20 Sep); 5 s, 480p, one picture attached 15 (charged 27 Sep); 5 s, 1080p 45 to 60 (quoted 11 and 20 Sep). The rate is 7 credits a second at 720p. Edit, extension, video input and audio-only jobs: not known. At the maker, a 5 s clip at 16:9 costs USD 0.514 at 480p, 1.156 at 720p and 2.843 at 1080p (BytePlus pricing, updated 28 Sep 2026).

The one mistake

Leaving the mode on its default. t2v is text to video, so a job with a start frame, a picture, a clip or a sound attached must be set to omni_reference. The job history does not store the mode, so write it in the run record.

Read

Chapter 17: 17.3 the prompt, 17.6 settings, 17.9 limits, prices and rights.

A.5 Wan 3.0

Version and date

Alibaba's Wan 3.0: wan3.0-video and, faster, wan3.0-video-prime in Model Studio; announced 13 Aug 2026; on Higgsfield since 24 Aug as wan3_0 and wan3_0_prime. The guides and API reference were last updated on 28 Sep and list no later model; the API reference still says "Currently in preview" (checked 30 Sep 2026). Connection contract read 27 Sep 2026.

Reach for it for

A plate, or a cheap probe of one camera sentence, from words alone (text alone cannot hold a face); a shot composed from your anchors by number with no start frame, or one that borrows a camera path, a timing or an action from footage; the economy new performance, an English line in a chosen voice, compared with Seedance on the same line before a client sees it. A measured camera move is Kling's (A.3), a locked first frame with a voice is Seedance's (A.4). Never Egyptian or Saudi dialogue.

Prompt order

The overall line (subject, mood, pace) → "One continuous shot." → on a reference job the manifest and the locks → the shot: size and height, one camera instruction, subject and place, the action in order → the world's life, the protection, the endpoint and hold, the pace, up to three exclusions and the sound. A still camera is "Fixed shot, camera static, position unchanged". "No dialogue" must be written.

Set every time

Model wan3_0 or wan3_0_prime, named (Prime is faster and costs about 1.5 to 1.7 times as much) duration 2 to 30 s, default 5, 5 to 6 s for most clips resolution 480p to test a camera sentence, 720p (the default) to judge and usually keep, 1080p only when the delivery needs it sound on or off, default on (off for a wordless plate) aspect, set every time; there is no 21:9.

Limits

The maker's (checked 30 Sep 2026): 2 to 30 s at 30 fps, MP4; with a video input, input plus output no more than 30 s images up to 10 (240 to 8000 px, ratio no more than 8:1, 20 MB each) clips up to 5, 15 s in all audio up to 5, 15 s in all 20 references in all; a first or last frame excludes every one. Returned: 720p is 1280 × 720 at 30 fps. No seed, negative field or lip-sync flag on the connection.

Price

Higgsfield credits, about five US cents each, Wan 3.0 / Prime (quoted 27 Sep 2026): 5 s, 480p 5 / 7.5; 5 s, 720p 8.75 / 15; 5 s, 1080p 17.5 / 30. With a video reference: not known. At the maker (Alibaba's pricing page, updated 28 Sep 2026), in Beijing, Prime costs 0.45, 0.9 and 1.8 yuan a second at 480P, 720P and 1080P, and the standard model 0.3, 0.6 and 1.2, marked "limited-time 30% off".

The one mistake

Sending a first or last frame together with references. A job carries boundary frames or references, never both, and a voice is a reference, so a speaking shot on Wan is a reference shot with the keyframe as an image reference.

Read

Chapter 18: 18.3 the prompt, 18.6 settings, 18.9 limits, prices and rights.

A.6 Veo 3.1 and MiniMax H3

Version and date

Veo 3.1 is Google's current Veo (checked 30 Sep 2026): a preview on the Gemini API, in a fast variant (the default), a slower preview variant and Veo 3.1 Lite. Gemini Omni Flash 1.1 (generally available 27 Aug 2026) is Google's default video model; Chapter 20 has it. H3 (Hailuo 3.0) is MiniMax's video model, announced 31 Jul 2026, with H3 Max as the fast variant; no newer model (checked 30 Sep 2026).

Reach for it for

Veo: a sound-led English action, a visible impact whose sound carries the idea (a cap twisted shut, a can opened), or a premium wordless plate whose track is stripped afterwards. H3: a 2K picture from a keyframe; a plate for delivery at 1080p or better in one job; mixed references (images, footage, audio) in one pass. Check either on a scratch clip before a client relies on it. Egyptian and Saudi dialogue is never generated here; H3 is not a speech route and not for a plate that must arrive silent.

Prompt order

The seven elements of Chapter 14: camera with its rig physics → the subject's action → the world's life → what is protected by name → the endpoint as a cuttable composition with its hold → the speed profile → earned exclusions. Nothing the start frame carries is described again. Close with a sound line naming each sound at its moment and the ambience beneath it, because the track cannot be switched off. H3 puts camera brackets such as [pan], [zoom] or [static] after the description they apply to, and ends with a "Sound:" line.

Set every time

Veo: duration 4, 6 or 8 s (the default is 8, so set 4) quality basic (default), high or ultra fast (default) or preview aspect 16:9 or 9:16. H3: duration 4 to 15 s, default 5 2K only (H3 Max: 480p or 768p, default 768p) aspect batch 1 to 4, default 1, each video billed. Write every setting on every job, and ask for the price of the exact shape first.

Limits

Veo 3.1 (Google, checked 30 Sep 2026): 24 fps; 720p by default, 1080p and 4K only at 8 s (no 4K on Lite); one video per request; extension adds 7 s up to 20 times, on Veo clips at 720p; videos kept for 2 days; SynthID on every video. Not exposed on the connection: references, extension, seed, negative prompt, language, audio input. H3 (MiniMax, checked 30 Sep 2026): 768P or 2K on the maker's API, 2K only on the connection; 24 fps (the maker's page of 11 Sep 2026; probe the file). Not exposed: seed, negative prompt, frame rate, sound switch.

Price

Google's price per second: Veo 3.1 USD 0.40, Fast 0.10 at 720p, Lite 0.05. Higgsfield (27 Sep 2026), credits: Veo fast, basic, 4 s / 8 s 16 / 32; preview, basic 40 / 80; Lite 6 / 12. H3, 2K, 5 s / 10 s / 15 s 10 / 20 / 30; 5 s, batch of four 40; H3 Max, 5 s, 768p / 480p 12.5 / 7.5. MiniMax's price is USD 0.13 a second at 2K and 0.08 at 768P. Its terms of service could not be read as text on the day.

The one mistake

Leaving Veo's duration at its default of 8 seconds when the idea is a sound-led event that needs 4. The track is always on, and "No music, no voice" is a request, not a switch: keep the clip at 4 seconds and check the file.

Read

Chapter 19: 19.3 the prompt, 19.6 settings, 19.9 limits, prices and rights; Chapter 20 for Omni.

A.7 ElevenLabs v4 and v3

Version and date

Eleven v4 (eleven_v4) is the current speech model, "Our most emotionally rich, expressive speech synthesis model" (models page, checked 30 Sep 2026): 90+ languages, 10,000 characters a request. Eleven v3 (eleven_v3) is the "previous generation" model: 70+ languages, 5,000 characters. Both stay, and neither is preferred in advance for Arabic. v4 Turbo is not used. Sound effects run on eleven_text_to_sound_v2. Not reached through Higgsfield.

Reach for it for

Text to Speech on v4 or v3 (dialogue, voice-over, scratch); Text to Dialogue to hear a scene's rhythm; Voice Changer to carry a native performance into another voice; Voice Design for an audition; an Instant or Professional clone of a consenting person; one sound event or a short bed. A client final's voice is an actor's recording or the consenting clone of that person. A sung line goes to a singer.

Prompt order

The voice and the model, chosen outside the box → one bracket at the head: voice quality and physical state, then attitude, then address → the canonical words in the spoken dialect, plain spelling → beat marks (ellipsis, full stop, em dash, [pause]) → a repair on the failing word only → one breath group per take. A sound effect: object, material, action, force, surface, perspective, space, attack, decay and refusals, within 450 characters.

Set every time

Model eleven_v4 or eleven_v3, named on every take (the API default is eleven_multilingual_v2; Text to Dialogue defaults to eleven_v3); for each new Arabic voice, run v4 and v3 and let a native ear choose language ar or en, a language and never a dialect Stability 0.5 Similarity 0.75 on v4 Style 0 and Speed 1.0 (neither on v4) seed, one per round, written down numbers written as words test output mp3_44100_128, lock source WAV or PCM Voice Changer model eleven_multilingual_sts_v2.

Limits

Checked 30 Sep 2026. A request: v4 10,000 characters (about ten minutes), v3 5,000 (about five) neither model supports SSML break tags Text to Dialogue: at most 10 voices, total text at or below 2,000 characters Voice Changer 5 minutes and 50 MB Voice Design preview text 100 to 1,000 characters Instant clone 1 to 2 minutes of clean speech recommended Professional clone 30 minutes at least, and your own voice only sound effects 30 s a generation.

Price

Monthly billing, one credit pool for every product (checked 30 Sep 2026): Free $0, 10,000 credits, no commercial licence; Starter $6, 30,000; Creator 22(11 first month), 121,000; Pro $99, 600,000; Scale $299, 1,800,000; Business $990, 6,000,000. Speech costs 1 credit a character, Voice Changer 1,000 a minute. API speech per 1,000 characters: v4 $0.08 (on offer at $0.022 until 12 Oct), v3 $0.08. Sound effects have three prices that do not agree; budget the highest and read the app.

The one mistake

Expecting a dialect from the language setting or from the text. ar names a language, never Egyptian, Najdi or Hejazi: the voice carries the accent, so cast the voice before you write a tag.

Read

Chapter 23: 23.3 the prompt, 23.6 settings, 23.9 limits, prices and rights; Chapter 21 for the craft of a line.

A.8 Sync Lipsync 3

Version and date

Higgsfield's name for sync.so's lip-sync model sync-3, launched 6 April 2026 and labelled "Default model for all users"; the maker's changelog (latest entry 31 August 2026) lists nothing newer (checked 30 Sep 2026). On Higgsfield it is the video-to-video model sync_so (contract read 27 Sep 2026).

Reach for it for

The exact take must ship (Route A of Chapter 22): a locked recording spoken by a face in a talking plate, or a line changed after picture lock. When a new performance is acceptable, use Seedance 2.5 (A.4); when two faces share the frame and one must speak, cut singles, because the connector cannot choose a face.

Prompt order

There is no prompt: Higgsfield exposes no prompt field, and a cost check that includes one is rejected. The direction lives in the plate's prompt, written for Kling or Seedance on the seven elements of Chapter 14: a measured business and the start of the talk in seconds → "talks quietly", no words and no "No dialogue" line → the face fully visible and lit, nothing crossing the mouth → a landing held to the end → the hygiene line. No sound line: the pass discards the plate's sound.

Set every time

sync_mode cut_off, always input_video the plate's finished job id input_audio the prepared take as an MP3, confirmed folder_id the project folder. The take: clean single-speaker speech with no music, WAV at 48 kHz mono, about −20 LUFS, peaks at −3 dBTP, at least 200 ms of silence at the head and 400 ms at the tail. The plate: 24 fps, at the delivery size.

Limits

Higgsfield states no cap; plates of about 5 s and takes under about 4.7 s are the established sizes. The maker (checked 30 Sep 2026): 480p minimum, 1080p recommended, 4096 × 2160 maximum MP4, MOV, WebM, AVI 24, 25 or 30 fps constant at least 10 Mbps 8-bit SDR BT.709 audio WAV or MP3, one speaker, no music direct upload 20 MB 95+ languages, and nothing said about Egyptian or Saudi mouth shapes.

Price

Higgsfield: 18.15 credits per pass, for a plate of about 5 s and a take of about 4.6 s; no cost check is possible (charged 12 Sep 2026). At the maker (checked 30 Sep 2026): $0.107 to 0.133 a second in the docs; the pricing page shows $0.1334 (Hobbyist, Creator), $0.1266 (Growth) and $0.1066 (Scale), with plans at $5, $19, $49 and $249 a month.

The one mistake

Leaving sync_mode at the default. Higgsfield defaults to bounce, which plays the plate forwards and then backwards and so reverses the actor's motion; cut_off truncates to the shorter input. Set it on every pass, with a plate 0.3 to 0.5 s longer than the padded take.

Read

Chapter 24: 24.3 the prompt, 24.6 settings, 24.9 limits, prices and rights; Chapter 22 for the route and the plate's timing.

A.9 Suno

Version and date

Suno v6, released 9 Sep 2026, in three variants: v6 ("our most advanced model"), v6-wild ("built for experimentation") and v6-mini ("a faster, more efficient way to create"). v6 and v6-wild need a Pro or Premier plan; v6-mini is open to every user. Up to eight minutes a piece; every generation returns two songs and costs 10 credits (checked 30 Sep 2026). On suno.com, on your own paid account; not on Higgsfield.

Reach for it for

A score cue that meets picture; a temp track under a rough cut; a rhythm reference that sets cut points before picture exists; a sketch to find a direction in minutes. An exclusive score, guaranteed originality or timing exact to the bar goes to a composer or a licensed library; timed sections for online delivery go to Eleven Music (A.10); a sung line in Egyptian or Saudi that must be right goes to a singer.

Prompt order

Custom mode. The Style field is one directed paragraph under about 950 characters, most important first: form and function → the figure and what it does → roles → arc, entry and ending → harmonic identity → performance and room → speech space → tempo, last and approximate. The Lyrics field holds the running order: [Instrumental], one tag for each picture section, [End]. Refusals go in Exclude Styles, each with a positive twin in the text.

Set every time

Plan Pro or Premier from the first generation model v6 for picture cues (v6-wild only to explore, v6-mini for sketches), written on every candidate Custom mode Instrumental on Weirdness low, start at 25 to 30 Style Influence high, start at 80 to 85 Variety 0 Max Mode off for a 30 second cue Exclude Styles one to five sounds a neutral Title download the candidate you choose, WAV on Pro and Premier. Run the brief twice as written before changing a word.

Limits

Up to eight minutes per generation on v6 Suno publishes no field lengths; a third-party reading gives 1,000 characters for Style and for Exclude Styles, 5,000 for Lyrics and 100 for Title, so trust the counter on your screen stems up to 12 with Auto Split uploads: the pricing page lists up to 8 minutes on the free plan and 30 on Premier, the help article 60 seconds free and 8 minutes paid, so test the limit on the day (checked 30 Sep 2026).

Price

Suno's pricing page (checked 30 Sep 2026; it shows no date): Free $0, 50 credits a day, no downloads, v6-mini only, no commercial rights; Pro $8 a month billed yearly, 2,500 credits a month, 20 downloads; Premier $24 a month billed yearly, 10,000 credits, 60 downloads, with Studio and uploads up to 30 minutes. Auto Split costs 50 credits; Split from Mix and Advanced Split 10 per extraction. Terms revised 10 Aug 2026, effective 3 Sep 2026: commercial use needs Pro or Premier and a permitted download, and nothing warrants copyright.

The one mistake

Briefing a mood, a genre or an artist ("trailer music", "epic orchestral", "in the style of"). Each names the average of a genre, and Suno delivers the average. Brief one figure and what it does in this cue.

Read

Chapter 26: 26.3 the prompt, 26.6 settings, 26.9 limits, prices and rights.

A.10 Eleven Music

Version and date

ElevenLabs' music model, reached in the ElevenCreative app and through the API. Its current model is Music 2.5, "our most advanced music model", music_v2_5 (changelog of 14 Sep 2026); music_v2 and music_v1 remain available. The API still defaults to music_v1 (documentation checked 30 Sep 2026).

Reach for it for

A cue that needs sections of known length, such as a 30 second online spot with a change at second 24, when the delivery is online. For film, television or radio choose a composer, a licensed library, or Suno on a paid plan (A.9); a single sound event belongs to a sound-effect generator (A.11).

Prompt order

The brief of Chapter 25 written as roles and behaviour, then split into chunks measured from the picture, with the theme in the first chunk. A chunk's text is a section name in square brackets with inline directions in braces: [Intro] {claps alone, then the riff}. "Instrumental" and the vocal refusals go in every chunk. Styles are written in English. Never an artist, a band or a lyric.

Set every time

model_id music_v2_5, named in every request mode: a plan for timed sections, a prompt for a temp cue, one per request store_for_inpainting true, at generation (the default is false) output_format named (the default is auto; MP3 at 192 kbps needs the Creator plan or above, PCM at 44.1 kHz the Pro plan or above) force_instrumental and music_length_ms in prompt mode only seed optional, not with a prompt.

Limits

A plan holds up to 30 chunks of 3 to 120 seconds each and runs from 3 seconds to 10 minutes; the overview page says a track may run to 5 minutes, so test a long cue the prompt limit is 4,100 characters nothing shorter than three seconds, and there is no hit-point field (checked 30 Sep 2026).

Price

Music uses the account's credits, priced by the length of the track and the number of variants; hover over the remaining-credits display before generating. The pricing page (30 Sep 2026) shows Free $0 with 10,000 credits a month, Starter $6 with 30,000, Creator $22 (first month $11) with 121,000, Pro $99 with 600,000, and Scale and Business above that; the pricing FAQ puts music at approximately 900 credits a minute, and the API bills $0.15 a minute. Free is a trial: attribution required and no downloads.

The one mistake

Using a self-serve plan's music beyond online. The model-specific terms (updated 26 May 2026) read: "All online and offline commercial use permitted, except film, TV, radio, & Studio Games"; only Enterprise Music lifts that. Music 2.5 is not named in those terms; its id puts it in the music_v2 family they cover: ask ElevenLabs in writing before any final beyond online.

Read

Chapter 27: 27.3 the prompt, 27.6 settings, 27.9 limits, prices and rights.

A.11 Firefly sound effects, Stable Audio 3 and Lyria

Checked 30 Sep 2026

Adobe Firefly sound effects

Stable Audio 3

Google Lyria

Version

The Firefly Audio Model, no number published; generally available since 20 Aug 2026; its video editor is still labelled beta

Stable Audio 3.0, released 20 May 2026: Small SFX, Small, Medium, Large

Lyria 3.5 (lyria-3.5, stable) for songs of a couple of minutes; Lyria 3 Clip (lyria-3-clip-preview, preview) for 30-second clips

Reach for it for

Effects and short beds, when the timing of a performed action matters more than words can say

Effects and music; a recording you hold that needs improving, or a bed of an exact length on your own machine

An instrumental sketch to audition beside Suno; music as clips, loops and songs

Prompt order

Short and direct ("lion roaring"): material and object → action → quality → distance → space. One sound per generation, layered on tracks. English only

Open with TrackType: SFX, → what makes the sound, how it is triggered and how long it lasts, where the microphone is and what the room is like → end "No music, no voices."

The eight parts of the brief as one paragraph with instruments, BPM, key, mood and structure → timestamps ([0:00 - 0:10] Intro: …) or section tags → "Instrumental only, no vocals." Each refusal becomes a positive twin

Set every time

Duration slider: a maximum up to 30 s, the action plus about a second of decay voice guide on the start frame four variations a generation WAV for audio-only, 48 kHz

Model small-sfx (up to 120 s, on a CPU), medium (up to 380 s, Nvidia GPU with Flash Attention 2), large (API only) the exact duration a seed, fixed and recorded

Clip to audition a brief, lyria-3.5 for the longer cut MP3 by default, WAV on Lyria 3.5; probe the file

Limits

Four variations a generation, each up to 30 s

small-sfx 120 s; medium and large 380 s; no intelligible vocals

Clip 30 s; 3.5 about a couple of minutes; single-turn; eight languages listed on Google Cloud, Arabic not among them

Price

10 generative credits a generation (Adobe, 10 Sep 2026); free daily generations with an Adobe account

Local: your own machine; the large API price was not read

Paid tier only: $0.08 a song on 3.5, $0.04 a Clip (Gemini pricing, 24 Sep 2026)

The one mistake

Putting several sounds into one generation

Shipping before the licensee is settled: free unless "you or your organization generate over USD $1M" of annual revenue

Briefing an Arabic vocal

Read

Chapter 28: 28.3, 28.6, 28.9; ElevenLabs effects are in Chapter 23

A.12 Premiere Pro and the upscalers

Version and date

Adobe Premiere Pro, release 26.5 of 10 Sep 2026 (checked 30 Sep 2026). Version 27 is in beta: Apple silicon Mac (M1 or later) on macOS 15 or later, Intel Macs stay on 26.x, no release date. Topaz Video desktop 1.7.1 (18 Sep 2026). On the platform: upscale_video with Topaz or ByteDance as the engine, and fps_boost (contracts read 27 to 28 Sep 2026, all untested). FLUX Video Upscale from Black Forest Labs, used directly.

Reach for it for

Building, judging, finishing and exporting the cut: the assistant plans and checks, Premiere executes and confirms. Generative Extend for a handle of up to two seconds; captions, with Arabic built from approved copy (Speech to Text lists eighteen languages and no Arabic). Upscale only after picture lock, if the delivery needs it, and only the intervals the cut uses. A handle longer than two seconds, or a change of action, is a new generation.

Prompt order

Premiere takes no prompt. Its order: sequence settings before media → media probed before it is placed → the Text Engine set on a caption track before the Arabic goes in → captions on their own track above the picture → generated frames and transforms listed as you make them → the export last, then probed. The upscale requests are settings: upscale_video with provider: topaz takes video_id, resolution 1080p or 2160p, aspect_ratio.

Set every time

Sequence timebase and aspect from the delivery plan Text Engine per caption track, "South Asian and Middle Eastern", before typing Arabic (test on your build) right-to-left typing on for an Arabic title caption export chosen per version XML through File > Import (Open Project refuses .xml). Topaz 1080p (default) or 2160p. ByteDance: width and height from the probe, resolution 2k (default), fps 24 (default), preset aigc for generated footage; fps above 30 doubles the cost. FLUX: creativity 0 for faces and products, upscale_factor the smallest that reaches the delivery size.

Limits

Checked 30 Sep 2026. Generative Extend: up to two seconds of video and ten seconds of audio; a video clip of at least two seconds; 360p to 4K UHD since 26.5; one reviewer reports it generates at 30 fps at most and outputs SDR 8-bit, so try it on an HDR or 60 fps clip. Topaz Video takes input at or under 4K and outputs up to 4K at 2×, 3× or 4×. FLUX Video Upscale: source MP4 of at most 20 seconds, 50 MB and 2560 × 1440, at least 480p; output 1080p, 2K or 4K.

Price

Premiere: single app US22.99amonthontheannualplanbilledmonthly, with25generativecreditsamonth; CreativeCloudProUS34.99 for the first three months and US69.99after, with4, 000amonth(USprices).TopazPersonalUS39 a month with 25 monthly video cloud credits and unlimited local rendering. Platform upscalers: no quote before the job, the price shows in the ledger; failed jobs are refunded. FLUX: precise US0.07(aboutUS0.14 a second at 1080p), creative US0.10(aboutUS0.20).

The one mistake

Generating media inside Premiere's AI Assistant or Generative Media Tool. No request there names its engine, and the tool works outside the run record and the spend packet. Missing coverage is generated on the platform, not in Premiere.

Read

Chapter 32: 32.3 the prompt, 32.6 settings, 32.9 limits, prices and rights; Chapter 30 for the finishing order.

A.13 Higgsfield

Version and date

Higgsfield has no version number, so it is read by date (checked 30 Sep 2026): its changelog's latest entry is 27 Sep 2026, its terms are of 26 Jul 2026, its help pages of August 2026. The model contracts were last read on the connector on 27 and 28 Sep 2026; the developer API launched on 16 Sep 2026. It hosts image, video and audio models from several makers behind one account and one credit balance, and makes none of the models this book uses.

Reach for it for

Running the stack from one account: stills, clips and lip-sync passes, a price check before every spend, jobs chained by number, a history of image and video jobs and a ledger of every charge. For a voice use ElevenLabs (Higgsfield has no Arabic voice), for music Suno; probing, cutting and finishing are your own tools. A feature only the web app shows is its own route with its own lock, test and spend approval.

Prompt order

The order of a request: admit the route and read the contract → name the model → set the inputs in the order the model wants (identity first on stills; on Seedance and Wan, the order of first appearance), each with one job → type every setting the contract exposes → write the prompt, with duration, aspect and resolution as settings and never prompt text → quote the exact request with its media attached → submit, and write the job number down.

Set every time

The exact model id the aspect, typed everywhere the resolution the sound (Kling 3.0 on or off, default on; Seedance 2.5 generate_audio, on by default) Seedance mode omni_reference whenever pictures or audio go in lip-sync sync_mode: cut_off count 1 unless you mean several jobs get_cost: true to quote leave out use_unlim declined_preset_id only on the retry after a preset suggestion.

Limits

Last read on the connector on 27 and 28 Sep 2026. Batch tools take 1 to 12 requests, with count fixed at 1 and no quote inside waiting takes 1 to 12 job numbers and at most 15 seconds a call uploads take 1 to 20 files show_generation_by_ids takes up to 60 jobs rate-limit errors above about three jobs on one model (12 Sep); four in one batch accepted (27 Sep) the connector states no result-link lifetime.

Price

A credit is worth about five US cents (Higgsfield's blog, 6 Aug 2026) and has no cash value. Everything generated through the connector deducts credits, whatever the plan. The Ultimate plan holds 1,200 credits a month, granted at the reset around the 11th at 02:00 UTC, and the unused balance is removed then. Platform readings: the image upscale (ByteDance), 2752 × 1536 to 4k, 2 credits (quoted 27 Sep); Sync Lipsync 3, 18.15 credits (charged 12 Sep). Each model's prices are on its own card.

The one mistake

Resending after a timeout, a rate-limit error or a partial failure before you know what happened. A wait that runs out does not stop the job: read the ledger and the job history first, and resend only what is missing. "Completed" is a status, not a success.

Read

Chapter 33: 33.3 the request, 33.6 settings, 33.9 limits, prices and rights; Chapter 34 for money, rights and terms.

Comments


bottom of page