Chapter 28 · Sound-effect and music generators
Part 5 · Music and sound
Three makers fill the gaps that a recording, a library or Suno leaves. Adobe Firefly makes a sound effect that follows your voice against the picture. Stability AI's Stable Audio 3 makes an effect from words, remakes a recording you made, or redoes one region of a sound. Google's Lyria makes a music clip or a song from a written brief. This chapter teaches each under the same ten headings; the craft of choosing, describing and placing a sound is Chapter 25.
In this chapter
What each tool is, its version, and how it reads text, voice, audio and images
The prompt in its dialect, templates and four examples at the standard
Settings, checks, repairs, and limits, prices and terms, dated 30 Sep 2026
Before you start. Chapter 25 decides what a sound or a cue must do and where it comes from. ElevenLabs sound effects are in Chapter 23, Suno in Chapter 26 and Eleven Music in Chapter 27. Chapter 34 covers rights across the whole job.
28.1 What it is
(checked 30 Sep 2026)
Adobe Firefly sound effects | Stable Audio 3 | Google Lyria | |
Current version | the Firefly Audio Model, no number published; generally available since 20 Aug 2026; its video editor is still labelled beta | Stable Audio 3.0, released 20 May 2026: Small SFX, Small, Medium, Large | Lyria 3.5 (lyria-3.5, stable) for songs of a couple of minutes; Lyria 3 Clip (lyria-3-clip-preview, preview) for 30-second clips |
Makes | effects and short beds, optionally timed by your voice | effects and music; remakes your recording or one region | music: clips, loops and songs, instrumental or sung |
Reach for it when | the timing of a performed action matters more than words can say | you hold a recording to improve, or need a bed of an exact length on your own machine | you want an instrumental sketch to audition beside Suno |
Reach for another when | words will do (ElevenLabs, Chapter 23) or the prop can be recorded (25.8) | the recording already sounds right: it ships | a cue must follow sections and bars (Suno, Chapter 26; Eleven Music, Chapter 27) |
Reached through | Adobe's Firefly on the web: the Audio module, or the video editor's Generate media panel | model weights on Hugging Face, run locally (Small on a CPU, Medium on an Nvidia GPU); Large through Stability's API | Google AI Studio and the Gemini API; Google Cloud (its Gemini Enterprise Agent Platform pages) |
All three are reached at the makers' own sites and interfaces, not through the Higgsfield connector (Chapter 33). Adobe also announced Generate Music and Generate Speech on 20 Aug 2026; this chapter does not cover them.
28.2 How it reads what you give it
Input | Adobe Firefly | Stable Audio 3 | Google Lyria |
Text | English only; owns the material and the action | owns the source, the action, the microphone and the room; opens with a TrackType tag | owns genre, instruments, mood, tempo, key and structure; lyrics and section tags |
Voice or audio | your voice, recorded against the picture or uploaded, up to 30 s or the clip's length; owns timing and energy; never becomes speech | init audio, your recording; owns timing and attack, and a noise level says how far the result may move | none |
Picture | the clip you time against, with the playhead on the start frame | none | up to 10 images may be attached; Google's example asks for their mood and colours |
A region | none | inpainting: mask start and end seconds, or several regions | none |
How each binds. In Firefly the playhead frame starts the effect, the voice and the text combine in one generation, and each generation returns four variations. In Stable Audio the prompt, the init audio and its noise level, or the mask, go in one call, and a seed repeats a setup. In Lyria one request carries the text (and images); it is single-turn, so a result is not refined by a follow-up.
28.3 The prompt, in this tool's dialect
25.7 gives the anatomy of an effect: object, material, action, surface, force, distance, space, attack, decay, and the refusals. Each tool wants it cut differently.
Firefly: short and direct. Adobe's guide writes "lion roaring", not "the sound of a lion roaring"; adjectives give the quality, verbs the action; one sound per generation, layered on tracks; an ambience is best described broadly ("room tone"). Order: material and object, action, quality, distance, space. Timing is your voice, not the words. The help pages describe no field for refusals.
Stable Audio: three questions. The maker's guide asks what makes the sound, how it is triggered and how long it lasts, and where the microphone is and what the room is like. Open with TrackType: SFX,, set a short duration for a one-shot, and end with "No music, no voices." (the guide documents no exclusion field, so the refusal is written into the prompt). Its models "don't output intelligible vocals".
Lyria: a brief, then a map. Write the eight parts of the brief (25.3) as one paragraph with instruments, BPM, key, mood and structure, then timestamps ([0:00 - 0:10] Intro: …) or section tags, and "Instrumental only, no vocals." Turn each refusal into a positive twin: Google's music-generation page documents no field for exclusions. Name no artist; the safety filter blocks specific artist voices.
What each ignores. Firefly ignores a language other than English. Stable Audio ignores another tool's controls (timing, influence). Lyria treats a time as a request: results "may vary between calls, even with the same prompt".
28.4 Templates
Paste only what follows a label such as Prompt; the Record lines and the voice guide are never pasted.
Written for this book · Adobe Firefly · sound effect, finished and test grade · template; not run
FINISHED
Prompt: [material and object] [action verb], [force or quality], [close / medium / far], [space]
Voice guide (performed, not typed): [what you perform, over which frames; the release on its own frame]
Record: Firefly sound effects · clip [name] · playhead [frame] · duration max [s] · variation [1–4] · [date]
TEST
[object] [verb], [quality] — no guide, duration 2 to 5 s, judge all four variationsWritten for this book · Stable Audio 3 · sound effect, finished and test grade · template; not run
FINISHED
TrackType: SFX, [object] of [material] [action] on [surface], [force], [microphone and room], [attack], [decay]. No music, no voices.
Record: [small-sfx / medium / large] · duration [s] · seed [n] · init audio [file / none] · noise level [ ] · mask [s–s / none]
TEST
TrackType: SFX, [object] [action], [room] — duration 3 s, fixed seedWritten for this book · Google Lyria · music cue, finished and test grade · template; not run
FINISHED
[Length] instrumental cue for [job]. Instrumental only, no vocals. [Figure]; [roles]; [harmony, room]; [speech space]; about [n] BPM, [key].
[0:00 - 0:0x] [Section]: [what happens]
[0:0x - 0:xx] [Section]: [what happens]
[0:xx - end] Ending: [behaviour: held chord that decays, no fade]
Record: [lyria-3-clip-preview / lyria-3.5] · [MP3 / WAV] · run [1, 2] · [date]
TEST
A 30-second instrumental [genre] cue, [three instruments], [mood], about [n] BPM.28.5 Examples at the standard
Written for this book · beauty and personal care · Adobe Firefly sound effects, Audio module or video editor, duration maximum 2 s, voice guide over the turn · the lid of Example 25.2, rows 2+14 to 3+02 · written for this book; not run
Example 28.1 — a jar lid turned and released, timed by voice
Prompt: Frosted glass jar, metal screw lid unscrewing a quarter turn, fine thread friction ending in one soft release, very close, dry small studio
Voice guide (performed, not typed): a dry "sss-sss-sss", slowing, over the quarter turn from 2+14, ending in one short "tk" on 3+02, the frame the lid lifts free
Record: Firefly sound effects · clip jar_lid_macro_v2 · playhead 2+14 · duration max 2 s · variation [1–4]Why it works
The voice carries the timing, so the words carry only material and action, in the short direct form Adobe's guide asks for.
The turn and the release are one action in a fixed order, so one file is right; pick a variation by its release, then place the transient (25.10).
Written for this book · food and drink · Stable Audio 3 small-sfx, run locally, 30 s, fixed seed · a street-food griddle under a 30-second scene · written for this book; not run
Example 28.2 — a griddle sizzle bed, one file for the scene
TrackType: SFX, thin strips of marinated beef and sliced onion cooking on a hot flat steel griddle, already at a steady fine crackle with small irregular pops from the first second, close microphone, dry tiled street-food kitchen, even level for the whole length. No music, no voices.
Record: small-sfx · duration 30 s · seed 4172 · init audio none · mask noneWhy it works
A bed is a place, and a place is a list. It starts already at a steady sizzle, so the first hiss and the spatula stay their own events on their own frames.
Duration is the scene's length and the seed is recorded, so a good bed can be made again identically and needs no loop seam.
Written for this book · automotive and luxury · two files: a set-recorded door remade in Stable Audio 3 small-sfx, and an engine start timed by voice in Adobe Firefly · written for this book; not run
Example 28.3 — a saloon door and its engine start
DOOR — Stable Audio 3 (audio-to-audio)
TrackType: SFX, a heavy luxury saloon door closing firmly, a deep dense low thud with a short soft latch click, no rattle, close microphone, quiet concrete car park, short dry decay. No music, no voices.
Record: small-sfx · duration the recording's length · init audio door_set_take3.wav · noise level 0.4 · seed 88 · original kept
ENGINE START — Adobe Firefly
Prompt: Petrol engine starting and settling to a steady idle, exterior, close, concrete car park
Voice guide (performed, not typed): a low "rrm-rrrm" that catches on the ignition frame, rises for two beats, then settles to a flat hum
Record: Firefly sound effects · playhead on the ignition frame · duration max 4 s · variation [1–4]Why it works
The recording carries the timing and a lowish noise level lets it stay; the words name only what is missing, weight and no rattle. 0.4 is the low end of the maker's own style-transfer examples (0.4 to 0.5).
The engine's ignition and rise are performed, so its energy follows the picture.
Door and engine are separate files, each placed on its own frame. If the remake does not improve the recording, the recording ships.
Written for this book · telecom and tech · Google Lyria 3 Clip (`lyria-3-clip-preview`), MP3, 30 s, run twice as written · a launch film under a close voice-over · written for this book; not run
Example 28.4 — a 30-second instrumental cue
A 30-second instrumental cue for a telecom launch film. Instrumental only, no vocals. Bright, precise and human: a clean electric piano plays a rising three-note figure, then repeats it under a soft round bass and light finger-snap percussion. About 104 BPM, G major, steady pulse, dry and close.
[0:00 - 0:08] Intro: the three-note figure alone on electric piano.
[0:08 - 0:22] Groove: soft bass and finger snaps enter and the figure repeats; the middle register stays open for a voice-over.
[0:22 - 0:30] Ending: the figure once more, then one held major chord that decays to silence, no fade.Why it works
One small figure, and what it does in each section (25.3); the times ask for order and proportion, and the timeline holds the exact times.
Every refusal is a positive twin ("dry and close", "the middle register stays open"), and the ending is a behaviour, so the cue does not fade or run on.
28.6 Settings
(checked 30 Sep 2026)
Tool | Setting | Value to use, and why |
Firefly | Duration slider | a maximum, up to 30 s: the action plus about a second of decay |
Firefly | Voice guide | playhead on the start frame; record while the picture plays silent; up to 30 s or the clip's length |
Firefly | Variations | four per generation; audition all and choose by the release. A new prompt or duration makes new media |
Firefly | Download | WAV for audio-only files, MP4 for video with audio; 48 kHz |
Stable Audio 3 | Model | small-sfx for effects, up to 120 s, on a CPU (Apple Silicon and CPU runtimes exist); medium up to 380 s, Nvidia GPU with Flash Attention 2; large, API only |
Stable Audio 3 | Duration | the exact length needed; short for a one-shot; it is variable, so a bed needs no loop |
Stable Audio 3 | Seed | fix and record it: without one, no two runs match |
Stable Audio 3 | Noise level | with init audio: 0.4 to 0.5 are the maker's own style-transfer values; a higher level filters out rhythm and melody, so the attack moves |
Lyria | Model | Clip to audition a brief; lyria-3.5 for the longer cut |
Lyria | Output | MP3 by default; WAV on Lyria 3.5. Probe the file: the docs give 44.1 kHz and, on one page, 48 kHz for Clip |
28.7 Checks before you sign
Listen on headphones and on speakers, at delivery volume, with the picture.
The transient is on the contact frame, and the release (or the last note) lands where the sheet says (25.10).
One event, not several; the object's size and material, not a bigger one; the tail is what you asked for.
No music, voices or unintelligible vocal texture in an effect or a bed; a bed is even for its whole length.
Beside a recording, the generated file is better for what was missing; if not, the recording ships.
A cue ends as briefed, and the middle register is open for the voice.
The raw file and the settings are on the row, with the rights line (34.10) filled.
28.8 Failures and fixes
Tool | You hear | Cause | Smallest fix |
Firefly | contacts off the action | the guide was performed watching the face | record the guide again, watching the object; keep the words |
Firefly | the wrong material, or an ambience like drumming on pots | the words | rewrite the words; keep the guide |
Firefly | several sounds mixed | more than one sound in the prompt | one sound per generation, layered on tracks |
Stable Audio | several taps or a long tail | duration too long, or "impact" without a scale | a shorter duration, "once", a size |
Stable Audio | the attack moves or doubles | noise level too high | bring it back towards the recording, or keep the recording |
Stable Audio | a region repeats the original | mask too small | widen the mask |
Lyria | a generic bed | a vague prompt | instruments, BPM, key, mood and structure |
Lyria | voices appear | a hint of singing in the brief | "Instrumental only, no vocals", remove any sung wording, audition another run |
Lyria | blocked prompt | the safety filter (an artist's voice, copyrighted lyrics) | describe the sound, name no one |
The retry ladder, smallest lever first: place or trim in the edit; run again with the same settings; change one phrase; change the guide or the noise level; change the route (record it, or take it from a library). After one batch and one change of wording without the sound, change the route.
28.9 Limits, prices and rights
(checked 30 Sep 2026)
Adobe Firefly sound effects | Stable Audio 3 | Google Lyria | |
Limits | English prompts; four variations per generation, each up to 30 s; guide up to 30 s | small-sfx 120 s; medium and large 380 s; no intelligible vocals | Clip 30 s; 3.5 about a couple of minutes; single-turn; eight languages listed on Google Cloud, Arabic not among them |
Price | 10 generative credits a generation (Adobe, 10 Sep 2026); free daily generations with an Adobe account; a credit's price depends on the plan, which was not read | local: your own machine; the large API price was not read | paid tier only: $0.08 a song on 3.5, $0.04 a Clip (Gemini pricing, 24 Sep 2026) |
Rights | Adobe's FAQ (13 May 2026): outputs can be used commercially, beta features too unless the product says otherwise; the product page allows royalty-free commercial use of audio "developed using the commercially released version of Adobe Firefly" and calls the tools "safe for commercial use"; no contract clause on audio outputs was read | Stability's licence page: "as between you and Stability, you own outputs"; free unless "you or your organization generate over USD $1M" of annual revenue; then an Enterprise licence, which offers indemnification (blog, 20 May 2026) | Gemini API terms (28 Apr 2026): "Google won't claim ownership"; Google "may generate the same or similar content for others"; paid services do not use your prompts or responses to improve Google's products |
Marks | Adobe may attach Content Credentials to certain exports | none read | a SynthID watermark on all audio |
May ship to a client | after the terms are read and the rights line written | once the licensee is settled (below) | with the client told the content may not be exclusive |
Write on the row the tool, model, plan, date and the terms line that covered the file (25.15).
28.10 Version notes
Firefly. Adobe's audio tools became generally available on 20 Aug 2026; the help pages are of 18 and 19 Aug 2026 (Audio module) and 16 Jun 2026 (video editor). Re-check the model's name and number, whether the video editor leaves beta, the credit cost, and any clause on audio outputs.
Stable Audio. 3.0 is the current family (20 May 2026); Stability says its next generation is already being built. Re-check the licence, the large API price and any new model.
Lyria. 3.5 is stable (Sep 2026). Clip is a preview: Google's pricing page lists both Lyria 3 previews, Clip and Pro, as legacy, while its music-generation page still documents Clip. Re-check whether Clip stays, the sample rate and the language list.
Open questions
Firefly's contract terms for audio. The FAQ allows commercial use; the terms page for generative AI could not be read. Read Adobe's terms before a client final.
Whose revenue counts in the Community License. The licence page says "you or your organization" and, elsewhere, "your organization, or the organization(s) performing the commercial use". If your production company, your agency or your client earns over USD 1M a year, ask Stability, or a lawyer, before a client final.
Lyria and Arabic. Arabic is not on the eight-language list read, so do not brief an Arabic vocal here. Egyptian and Saudi speech comes from a voice actor or ElevenLabs (Chapter 21).
The sample rate of a Clip file. Two pages give 44.1 and 48 kHz. Probe the file.




Comments