top of page

Chapter 20 · Other video tools worth knowing

Writer: Yasser Ashour
Yasser Ashour
2 hours ago
7 min read

Part 3 · Motion

Eight video tools sit outside the routes this book teaches. Each has a current version and a job it might one day take from a commercial director. None is a route yet. This chapter is not a recommendation to spend; it says what each tool does, so that a client's question finds you knowing where to look.

In this chapter

  • What each tool does, and its current version (checked 30 Sep 2026)

  • The job it might matter for

  • What would have to be read before it became a route

Before you start. Chapter 13 chooses the route from the tools taught in Chapters 16 to 19. Chapter 34 covers money and rights.

20.1 How to read an entry

Each entry gives the maker's current version, what the maker says the tool does, when it might matter, and what stands between it and a route. Prices and limits are the maker's own; the connection (Chapter 33) may expose fewer inputs, other prices, or none. It listed four of these on 28 Sep 2026: Gemini Omni Flash 1.1, FLUX 3 Video with FLUX Video Edit, Grok Imagine Video 1.5 (as "Grok Video 1.5") and Depth Anything Video; what it exposes of each has not been read. The rest are reached only through their makers.

20.2 Runway Act-Two

Version. Act-Two is Runway's current performance model; its changelog and research page list no successor (checked 30 Sep 2026). It runs in Runway's app (Standard plan or higher) and on its API.

What it does. You give it a driving video of a person acting and a character, as an image or a video. It carries the performer's movement, speech timing and expression onto the character: up to 30 seconds, 24 fps, at 1280 × 720 (16:9) and five other aspects, for 5 credits a second with a 3-second minimum. Runway asks for one subject in each input.

When it might matter. A mascot or stylised brand character that must act with a real actor's timing, such as an Egyptian or Saudi performance already recorded.

Before it could be a route. Runway shows no Arabic performance, and whether the actor's own voice can ship untouched is not confirmed. A character built from a signed photograph raises the rights questions of Chapter 34.

20.3 Runway Aleph 2.0

Version. Aleph 2.0, launched on 21 May 2026 in Runway's Edit Studio, is the current edit model; Gen-4 Aleph stopped working on the developer API on 30 Jul 2026. Runway's own generator is Gen-4.5, 2 to 10 seconds (checked 30 Sep 2026).

What it does. It edits footage you already have, up to 30 seconds at 1080p, changing what you ask for and leaving the rest as shot. You edit one frame the way you want and it carries the change through the clip, and across the shots of a multi-cut video. It also runs in a panel inside Premiere Pro and After Effects (from 8 Sep 2026).

When it might matter. Versions of finished footage: another colourway, a seasonal background, a stray car removed from a locked plate.

Before it could be a route. Chapter 13's map already names tools that edit a clip; Aleph 2.0 would have to beat one on a real product, and its price has not been read.

20.4 Luma Ray3.2

Version. Ray3.2 is Luma's newest model, in Luma's app and on its API (checked 30 Sep 2026).

What it does. One model runs text to video, image to video, Modify Video and Reframe. Image to video takes up to 16 keyframes in one clip, in native 5 or 10 second clips up to 1080p. Modify Video remakes existing footage (a new wall, lighting, wardrobe or product) at 1080p for up to 20 seconds at 24 fps, 15 at 30 and 7 at 60, and keeps the source audio. Generated clips have no sound. Output can be 16-bit HDR or EXR frames (ACES2065-1), at two and three times the credits.

When it might matter. Where the grade is the point: EXR frames a colourist can open; a keyframed move that hits several beats; a filmed action in a changed world.

Before it could be a route. Its results on a real product and a real face have not been checked. Luma's page gives the Reframe length limit as both 12 and 16 seconds.

20.5 LTX-2.5 and LTX-2.3

Version. LTX-2.5, in fast and pro variants, is the current model; LTX-2.3 stays beside it on LTX's API (checked 30 Sep 2026).

What it does. Text, image and audio to video, portrait or landscape, up to 4K (3840 × 2160) at 24, 25, 48 or 50 fps. Clips run 6 to 20 seconds on fast at 720p or 1080p and 24 or 25 fps, and 6 to 10 seconds elsewhere. Sound is generated unless switched off; a last frame can be pinned; in audio to video the clip's length follows the audio supplied. Only LTX-2.3 offers retake and extend.

When it might matter. A 4K plate at 48 fps in a single job; a clip built around a track already recorded, where the picture must follow the sound.

Before it could be a route. Whether audio to video keeps the supplied audio unchanged is not documented, and its Arabic is untested. It overlaps the speech routes of Chapter 22.

20.6 Gemini Omni Flash 1.1

Version. Generally available since 27 Aug 2026 on the paid tier of the Gemini API, and now named by Google its default video model; Veo 3.1 stays for extension, last-frame control and older pipelines (checked 30 Sep 2026; Chapter 19).

What it does. It takes text, images, audio and video and returns video with sound: text to video, image to video, first and last frame, subject references, and conversational editing, in which each follow-up changes the last result and keeps what you did not mention. Extension adds 3 to 10 seconds, up to 40 in all. Uploaded videos for editing must be 10 seconds or shorter, and are unavailable in the EEA, Switzerland and the UK. Output is 720p by default, with 1080p and 4K upscaled, at about USD 0.10 a second at 720p. Every file carries a SynthID watermark. Google says English is fully supported and other languages are not evaluated.

When it might matter. It is the likeliest of the eight to become a route: an English clip with its sound, refined by instruction ("change the lighting; keep everything else the same") instead of re-rolled.

Before it could be a route. Its contract on the connection has not been read. Egyptian and Saudi speech is not a route on it, as on Veo.

20.7 FLUX 3 Video and FLUX Video Edit [fast]

Version. FLUX 3 Video is Black Forest Labs' video model, and the maker calls FLUX 3 a preview. FLUX Video Edit [fast] is a separate editing endpoint with no preview label (checked 30 Sep 2026).

What it does. FLUX 3 makes video with synchronized sound from text, pinned keyframes, or a clip to continue: up to 20 seconds, up to 4K (3840 × 2176 at 16:9), 24 fps, with several scenes in one generation and readable in-scene text. The maker claims multilingual speech with strong lip-sync. Text or image to video costs USD 0.17 a second in HD and 0.80 in UHD. Reference images and videos are announced, not yet offered. Video Edit changes an existing clip from a text instruction: sources up to 15 seconds, output at most 720p, USD 0.03 a second, source audio kept. It takes no mask or reference.

When it might matter. A long single take; typography inside a scene; a cheap edit while a look is still being found.

Before it could be a route. A preview model is not one for a client's shot, and its Arabic speech is unproven.

20.8 Grok Imagine Video 1.5

Version. xAI's model left preview on 16 Jun 2026 and is generally available on xAI's API; the connection lists it as "Grok Video 1.5" (checked 30 Sep 2026).

What it does. Image to video and text to video (the model makes a first frame, then animates it), with native 1080p; reference to video, capped at 720p, with reference images and a preset voice; a pinned last frame; and up to four keyframes at exact moments. Sound effects, ambience and dialogue come in the same pass. It costs USD 0.08 a second at 480p, 0.14 at 720p and 0.25 at 1080p.

When it might matter. Cheap 480p looks at a keyframed idea before a costly render elsewhere; English sound that lands on the action.

Before it could be a route. Which inputs the connection passes, and xAI's terms for commercial work, have not been read; nothing is known of its Arabic.

20.9 Depth Anything Video

Version. The connection lists "Depth Anything Video" without a version. The research model of that name, Video Depth Anything from ByteDance (CVPR 2025), shows no newer version on its project page (checked 30 Sep 2026); that the listing runs it is not confirmed.

What it does. It makes a depth map for every frame of a clip, steady from frame to frame, on videos of any length. It makes no picture, only a grey-scale map of near and far.

When it might matter. In post: a depth-driven defocus, fog or relight behind a subject and in front of a background; a matte helper where an edge is hard to key.

Before it could be a route. Its output format and resolution on the connection are unread. Chapter 30 would take it up if it earned a place in finishing.

Summary

  • None of the eight is a route in this book. Chapter 13's map is the guide until one is.

  • Read first: Gemini Omni Flash 1.1, Google's default video model now; Runway Act-Two, for an actor's timing; Luma Ray3.2, for EXR frames.

  • A tool becomes a route when its inputs, limits and terms are read on the connection you will use, a scratch job passes the checks of its stage chapter, and the director decides.

Comments


bottom of page