Click to Play Episode
Sora is shut down, Google runs two video models, Kling 3 does lip-synced dialogue, and open-weight MiniMax H3 is what you can actually fine-tune. What a usable clip costs, which models do native audio, how reference consistency works, and why the unit of work is the shot.
Companion show: for weekly, hands-on coverage of the AI video pipeline, from a first usable clip to scenes that cut together, listen to AI Video Generation.
Second of three episodes on AI media generation, covering Veo, Kling, Runway and MiniMax H3. Four questions: what a usable clip costs, which models generate sound and dialogue natively, how character and shot consistency work now, and where open-weight video fits for a programmer. Ends with the shot-to-scene mental model.
Native audio is now the baseline at the frontier: Veo 3.1, Kling 3.0, MiniMax H3 and LTX-2.5 sample audio and frames from one model, so lip movement and sound effects land on the right frame. Clips grew from four or five seconds to eight to fifteen, with a few models advertising thirty. Every serious product ships a reference-conditioning feature (Google "ingredients", Kling "elements", Runway references) that holds a character or object across clips. Image-to-video, not text-to-video, is the professional path: lock the first frame with an image model (see AI Image Generation and Editing), then ask the video model to move it. The old "storyteller vs animator" split resolved in favor of the animators.
Google now runs two video models in two places. Veo 3.1 is the developer baseline on the Gemini API and Vertex, in Quality, Fast and Lite tiers, all with native audio; clips are 4, 6 or 8 seconds, and 1080p/4K are upscales of the 8-second clip. It accepts up to three reference images, first-and-last-frame interpolation, and extend in 7-second steps up to 20 times; the Ingredients to Video update added identity consistency, native vertical and 4K upscaling. At Google I/O 2026 Google announced Gemini Omni; Gemini Omni Flash replaced Veo inside the Gemini app and Flow, taking text, image, video or audio as input and supporting conversational video-to-video editing. Per Flow's model matrix it currently tops out at 10 seconds and 720p. All output carries SynthID; the detector portal is still waitlisted.
OpenAI launched Sora 2 on September 30, 2025 with native audio and a free iOS app; it then hit a copyright reckoning over opt-out character use, SAG-AFTRA and Bryan Cranston pushback on likeness, and a court order barring the word "Cameo". In March 2026 OpenAI announced a two-stage shutdown: app and web closed April 26, 2026, API closes September 24, 2026. NBC's reporting attributes it to reallocating compute to coding, reasoning and enterprise; Sora continues only as internal world-model research. It stays in the episode as the case study of a strong model without a business.
Kuaishou's Kling 3.0 launched globally in March 2026 as a unified image, video and audio model: up to 15 seconds per shot, native 4K, and per the Kling Omni audio guide lip-synced dialogue in five languages with sound effects and ambience generated in the same pass. The control surface is the point: an elements library built from images or short reference video with per-element voice binding, multi-shot generation with continuity, motion transfer from a reference video, motion brush, six-axis camera control, extend and retake. Sold as a credit-based consumer app with commercial rights on paid tiers, plus a first-party API fronted in the West by fal and Replicate. fal's three per-second prices (audio off, audio on, voice control) make the cost of joint audio-video sampling visible.
Runway Gen-4.5 shipped December 1, 2025, briefly topped the Artificial Analysis leaderboard, and was candid about causal reasoning, object permanence and "success bias" failures. Runway's differentiator is editing: Aleph is video-to-video (new angles, relighting, add/remove objects, restyle), and per the Runway API changelog Aleph 2.0 takes 2-30 second inputs with up to five keyframes; Act-Two transfers a filmed performance onto a character. Gen-4.5's release notes do not claim native dialogue or effects. The Runway API now also resells ByteDance Seedance 2.5 (30-second clips, large reference budgets, audio) and Wan 3, and studio deals with Adobe, AMC Networks and Lionsgate anchor the enterprise story.
As of this recording the Artificial Analysis text-to-video arena (blind pairwise human preference) has none of Veo 3.1, Kling 3.0 or Gen-4.5 in its top five: Wan 3.0, Gemini Omni Flash, fal's post-trained MiniMax H3 Max, MiniMax H3, then Seedance 2.0. The headline products win on control, distribution and enterprise fit, not the taste test.
Naming matters for Wan: Wan 2.2 is Apache-2.0 open weight (14B MoE needing an 80GB GPU, or a 5B model for a 24GB card) and its GGUF quantizations still trend on Hugging Face; Wan 2.5, 2.7 and 3.0 have no published weights on the Wan-AI Hugging Face org or GitHub and are served as APIs. The open-weight center of gravity is MiniMax H3: 33B, native stereo audio, up to 2K, 4-15 seconds, a community license permitting commercial use, official ComfyUI workflows, and a LoRA and step-distillation ecosystem. LTX-2.5 (19B, native audio, community license free under $10M revenue) and HunyuanVideo-1.5 (8.3B, 14GB with offload, no audio) round out the runnable set; MAGI-2 is a preview with no confirmed license. Elsewhere: Luma shipped Ray3, Ray3 Modify and Ray3.14; Pika pivoted to effects, an agent and MCP on Pika 2.5; ByteDance's Seedance 2.0 and 2.5 ride Dreamina and CapCut distribution; Grok Imagine is a priced API video model outside the top ten; Higgsfield is an aggregator and creative suite, as are fal, Replicate and OpenArt on the developer side.
Every consistency feature is conditioning under a different name: text, reference images, first frame, last frame, and reference video are slots the denoiser attends to. A first frame is the strongest condition, which is why image-to-video wins. Reference characters (Veo's three images, Kling elements with voice binding, Seedance's dozens of references) fight identity drift, still the main failure mode. Start and end frames bound a camera move and let shots hand off to each other. Explicit camera controls beat prompt text. Video-to-video (Aleph, Luma Modify, Omni Flash, Kling) means fixing a nearly right shot instead of regenerating. Extend compounds drift, so use it to finish a shot, not build a scene. Open models add LoRAs: Musubi Tuner trains adapters for HunyuanVideo and Wan 2.x, and the tooling lags each new frontier open release by months.
Native audio means one sampling process produces waveform and frames, conditioned on each other; Veo 3.1 prices everything as video with audio, Kling 3.0 exposes it as a paid toggle, H3 and LTX-2.5 do it in open weights, Gen-4.5 and Wan 2.2 do not. Post-hoc remains a valid choice: MMAudio generates synchronized sound from finished video with an explicit alignment module, and ElevenLabs sound effects generate timed effects from text. Native dialogue holds for a line or two; longer talking heads still favor performance-driven tools like Act-Two. Voice and music proper are in The AI Media Pipeline.
As of this recording, from Gemini API pricing: Veo 3.1 Quality about $0.40/s with audio (720p/1080p), Fast about $0.10/s, Lite about $0.05/s; Gemini Omni Flash is billed per token, working out to roughly $0.10/s of 720p. From Runway API pricing: Gen-4.5 $0.12/s, Aleph 2 $0.28/s with a minimum. From fal: Kling v3 about $0.08/s silent and $0.13/s with audio, Wan 2.5 $0.05/s; Grok Imagine video $0.05-0.08/s. An 8-second Veo Quality shot with audio is a bit over $3; Kling or Gen-4.5 about $1. No vendor publishes success rates; budgeting four generations per usable shot puts a frontier clip with audio at $3-13 and a Fast or open-weight clip under $1. Iterate on the cheap tier, render on the expensive one.
Every model generates a shot: one continuous take, one camera, one action, 4-15 seconds. A scene is three to eight shots cut together, and continuity is your job: same references in every shot, first and last frames handing off, the same elements or LoRA, one audio bed over the cut. Storyboard as shots, lock first frames with an image model, iterate cheap, render expensive, fix with video-to-video, assemble in an editor. Assembly, voice, music, ComfyUI and driving it from code are the next episode.
A note on how this one was made. I generated this episode with Gnothi, my tool that takes a topic, researches it, writes the chapters, and reads them aloud. This episode is the overview. The dedicated Gnothi show on AI video generation goes further, from a first usable clip to scenes that cut together. O C devel dot com slash video. That's O C D E V E L dot com, slash video.
This is the second of three episodes on AI media generation. The last one, AI Image Generation and Editing in twenty twenty-six, covered the still image and why editing replaced generating as the core task. This one is video. The next one, The AI Media Pipeline, is about the glue: voice, music, ComfyUI, the APIs, and how you assemble finished pieces. So if you're wondering where the text to speech and the music models went, they're in the next episode. Here I want to answer four questions. What does a usable clip cost. Which models do sound and dialogue natively. How do you keep a character and a shot consistent from one clip to the next. And where does open weight video fit for a programmer who wants to run things locally. I'll end with the mental model I keep coming back to: the unit of work in this field is the shot, not the film.
Let me start with what changed. A lot did, and some of it isn't what anyone expected a year ago.
The first shift is native audio. A year ago, Google's Veo three had just shipped sound in the same generation as the picture, and everyone else was silent video you had to score afterward. As of this recording, native audio is the baseline at the frontier. Veo three point one generates dialogue, sound effects, and ambience with every clip. Kling three does it. The open weight models from MiniMax and Lightricks do it. The models aren't bolting audio on as a second pass. They're sampling the audio and the video together from one model, which is why the lip movement matches the words and the footstep lands when the foot does. I'll come back to how that works.
The second shift is length. The atom of generated video used to be four or five seconds. Now the common frontier clip is eight to fifteen seconds, and a few models advertise thirty seconds in one go. That sounds like a small change, but eight seconds is a real shot. It's long enough for a line of dialogue, a camera move, and a reaction. Fifteen seconds is long enough that you start needing the model to plan, and you can see where the models that plan well pull ahead.
The third shift is consistency tooling. A year ago, keeping the same character across two clips was a prompt engineering trick. Now every serious product has a way to hand the model reference images, or a reference video, and say this person, this product, this room, keep it. Google calls its version ingredients. Kling calls its version elements. Runway calls them references. The names differ, the mechanism is the same, and it's the thing that turned video generation from a slot machine into something you can direct.
The fourth shift is that image to video became the professional path. Text to video is the demo. Image to video is the workflow. You generate or photograph your first frame, you get it exactly the way you want with an image model, which is what the previous episode was about, and then you ask the video model to move it. That gives you control over composition, lighting, and identity before any motion is involved, and it makes the video model's job smaller. If you take one thing from this episode for your own pipeline, take that.
The old framing was storytellers versus animators. Sora was the storyteller, the model you gave a paragraph of narrative and it interpreted the scene. Kling was the animator, the model you gave a picture and it moved it. That framing still helps, but the market answered it, and the answer is that the animators won. The storyteller product is gone, and I'll explain what happened to it. The tools that survived are the ones that let you start from an image and control what happens next.
So that's what changed. Now the question everyone asks, which is which model. I'm going to go through the four names in the title. One of them is there because you'll search for it, not because you can use it.
Let me start with Google, because Google's story got confusing this year and I want to untangle it. There are now two Google video models and they live in different places. Veo three point one is the model you get through the Gemini API and through Vertex, and it's the one developers build on. It comes in three tiers, which Google calls Quality, Fast, and Lite, and all three generate audio natively. Clips are four, six, or eight seconds, and the higher resolutions, ten eighty and four K, are upscales on the eight second clip, not native generation. That distinction matters when a marketing page says four K. The model is drawing at seven twenty and a separate step is enlarging it.
Veo three point one is the model that got the consistency features first. You can hand it up to three reference images and it will hold the person, the object, or the style across the clip. You can give it a first frame and a last frame and it will generate the motion between them, which is the trick for a controlled camera move or a controlled transformation. And you can extend a clip you already generated, seven seconds at a time, up to twenty times, so in principle you can chain your way to a couple of minutes of continuous footage. In practice the extend feature is how you get past the eight second wall when a shot needs to breathe, and it's also where drift creeps in, because each extension is a new generation conditioned on the last few frames of the previous one.
Then, at Google I O in May of twenty twenty-six, Google announced a new family it calls Gemini Omni, and the first model, Gemini Omni Flash, replaced Veo as the video model inside the consumer Gemini app and inside Flow, which is Google's filmmaking front end. Omni is a different idea. It's a multimodal model that takes text, images, video, or audio as input and produces video, and it supports conversational editing, where you generate a clip, then say make the sky overcast and keep everything else, and it edits rather than regenerating. As of this recording Omni Flash tops out at ten seconds and seven twenty resolution in Flow, which is lower than Veo's upscaled output, but it does video to video editing, which Veo doesn't. The developer version is on the Gemini API as well, and it's billed differently, which I'll get to in the cost section, because it says something about where this is going.
So if you're a developer, as of this recording, Veo three point one is your baseline on Google, and Omni Flash is the thing to watch. If you're a consumer in the Gemini app, you're already using Omni whether you noticed or not. Everything Google generates carries an invisible watermark called SynthID, which survives cropping and recompression. The catch is that the detector isn't public. Google runs a verification portal for journalists and media organizations with a waitlist. So the watermark is in your file and the ability to check for it isn't in your hands.
Now Sora, and this is the one in the title you can't use. OpenAI launched Sora two at the end of September twenty twenty-five, with native synchronized audio, better physics, and a free iOS app that hit number one on the App Store. It had a feature that let you upload your own face and voice and drop yourself into other people's videos, which OpenAI called a cameo until a court told them they couldn't use that word because the celebrity video company Cameo had sued. The app launched with an opt out policy for copyrighted characters, so the feeds filled with SpongeBob and Pikachu within days, and OpenAI reversed to opt in with revenue sharing. SAG AFTRA and individual actors, Bryan Cranston most publicly, pushed back on likeness use, and OpenAI tightened the guardrails.
Then in March of twenty twenty-six OpenAI announced it was shutting Sora down. The consumer app and website closed in late April. The API shuts down at the end of September. The reporting says the reason was compute. Downloads had peaked in November and fallen by two thirds by February, in app revenue was small, and OpenAI wanted the GPUs for coding, reasoning, and enterprise products ahead of a possible public offering. The Sora research effort continues internally as a world model project aimed at robotics and what they call the physical economy, but there's no OpenAI consumer video product anymore, and ChatGPT doesn't render video.
I keep Sora in this episode for two reasons. One is that you'll search for it, and you deserve a straight answer. The other is that it's the cleanest case study I know of what happens when the frontier is expensive and the business isn't there. Sora two was a very good model. It was arguably the best storyteller in the sense I used earlier. It didn't matter. The compute went where the revenue was, and video generation was not it, at least not for OpenAI. If you built on the Sora API, you're migrating this month, and the migration targets are the rest of this episode.
Now Kling, from Kuaishou, which is the Chinese short video company that competes with Douyin. Kling was the animator in that framing, the model that took an image and produced the most convincing motion, especially human motion. Kling three point zero was announced in February twenty twenty-six and launched globally in March. It's a unified model that does image, video, and audio in one architecture, and the audio is generated in the same forward pass as the frames. That means lip synced dialogue, and as of this recording Kling advertises lip sync in five languages, English, Chinese, Japanese, Korean, and Spanish, plus sound effects and ambience. Clips run up to fifteen seconds, which is the longest of the headline four, and it advertises native four K, which unlike Google's is claimed as native rather than upscaled. I've seen that claim reported with different frame rates by different outlets, so take the details as marketing until you test them.
What makes Kling interesting for a producer is the control surface. The elements feature lets you register a character from images or from a short reference video, bind a specific voice to that character, and then address that element by name in later prompts. The multi shot feature lets you generate several connected shots in one session with continuity of character, lighting, and space. There's motion transfer, where you give it a reference video of a person moving and it applies that motion to your character. There's a motion brush for painting where things should go, camera control on six axes, and the usual extend and retake. It's a lot of buttons, and the buttons are the point. Kling is what the animator side of the market grew into once the models got good enough to be worth directing.
Kling is sold as a consumer app on web, mobile, and desktop with a credit based ladder of plans, and commercial use rights come with the paid tiers. For developers there's a first party API, and it's fronted by fal and Replicate in the West, which matters because a lot of Western teams can't or will not sign up directly with a Chinese vendor. One thing to flag, because it's a clean mechanism lesson. On fal, Kling three has three per second prices for the same clip: audio off, audio on, and audio on with voice control. Audio on costs about fifty percent more than silent. That surcharge is the compute of joint audio video sampling made visible on an invoice. Native audio isn't free, it's just bundled into the model instead of into your pipeline.
Now Runway. Runway is the company that has been in this the longest of the four, and they have positioned themselves as the tool studios use rather than the app consumers play with. The current generation model is Gen four point five, released December first, twenty twenty-five, alongside Gen four rather than replacing it. When it launched it briefly topped the Artificial Analysis text to video leaderboard, ahead of Veo three, Kling two point five, and Sora two Pro. That lead didn't last, which I'll come back to, but the model is solid, and Runway was unusually candid in its release notes about what the model gets wrong: causal reasoning, object permanence, and what they call success bias, which is the tendency of a model to show the action succeeding because the training data is mostly footage of things working.
Runway's differentiator is the editing, not the generator. Aleph is their video to video model. You give it an existing clip and a prompt, and it changes the camera angle, relights the scene, adds or removes an object, restyles the footage, or extends the scene, all conditioned on the input video rather than starting from nothing. Aleph two point zero shipped in June of twenty twenty-six and takes input clips from two to thirty seconds and up to five keyframe images to guide the edit. Act Two is their performance capture model: you film yourself doing the performance, provide a character reference image, and it transfers your expressions, head motion, and body language onto the character. Whether Act Two is on the current API I couldn't confirm as of this recording, so treat that as a web app feature.
Two more things about Runway. First, Gen four point five doesn't, as far as the release notes say, generate dialogue or sound effects natively. Audio on the Runway platform comes from separate tools, or from the third party models they now serve through the same API, which is the second thing. Runway's API in twenty twenty-six resells other labs' models. Seedance two point five from ByteDance is on there, and it's the one that does thirty second clips with a very large reference budget, up to thirty images and ten videos and ten audio clips as references in one generation, with audio. Wan three from Alibaba is on there too. So Runway is becoming a front end with its own models plus the field's, and their studio deals with Lionsgate, AMC Networks, and Adobe, which put the Runway models inside Premiere and Photoshop, are the commercial reason to believe that will stick.
So that's the headline four, and here's where they sit. As of this recording, the Artificial Analysis video arena, which ranks models by blind human preference between pairs of clips, has none of Veo three point one, Kling three, or Gen four point five in its top five. Sora is gone. The top of the board is Alibaba's Wan three, then Google's Omni Flash, then two versions of MiniMax's H three, then ByteDance Seedance two. The headline products are the ones with the best control surfaces, the best distribution, and the best enterprise stories. They aren't currently the ones winning the taste test. That gap between product and model is the story of this year, and it's why the next section matters.
Now the special mentions, and this year they're less special than the name implies.
Wan is Alibaba's video family, and I need to be precise here because the naming is confusing and a lot of the internet gets it wrong. Wan two point two, released in twenty twenty-five, is open weight under the Apache two license, with a fourteen billion parameter mixture of experts model that wants an eighty gigabyte GPU, and a five billion parameter model that runs on a twenty-four gigabyte consumer card. That's the Wan you can download, and quantized versions of it are still among the most downloaded video models on Hugging Face. Wan two point five, two point seven, and three point zero are the models on the leaderboard, and as of this recording none of them has published weights. They're API models served through Alibaba and through fal, Replicate, and Runway. There's a Hugging Face page marketing Wan three as open, with no license and no weights, pointing at a third party website, and I'd treat that as noise. So when someone says Wan is the best open model, the accurate version is that Wan two point two is a good open model and Wan three is a good closed one from the same lab.
The open weight center of gravity as of this recording is MiniMax H three. MiniMax is the Shanghai lab whose consumer brand is Hailuo, and H three is their flagship: thirty-three billion parameters, native stereo audio at thirty-two kilohertz, up to two K resolution, clips from four to fifteen seconds. The license is a community license that permits commercial use, with an application step for companies in the US, EU, UK, and South Korea. It runs in ComfyUI with official workflow templates, on fal, and in the standard serving stacks. The recommended deployment is four GPUs, so this isn't a laptop model, but it's the model the fine tuning community has moved to. If you look at what is trending in video on Hugging Face, it's H three LoRAs, H three step distillations, and H three ComfyUI packages, the way it was Wan two point one LoRAs a year ago. The number three model on the leaderboard is fal's own post trained version of H three, sitting above MiniMax's original, which is a thing that can only happen with open weights.
Lightricks LTX two point five is the other open model with native audio. Nineteen billion parameters, a built in audio decoder, a community license that's free under ten million dollars in annual revenue, and it's designed to run with aggressive offloading in the twenty to thirty-two gigabyte range. Tencent's HunyuanVideo one point five is the small one, at about eight billion parameters, runs at fourteen gigabytes with offload, no audio, and it has been lapped: it isn't on the leaderboard and not trending. If you want something that fits a single consumer card, it and the five billion parameter Wan two point two are still the answers. MAGI two from Sand AI is a preview of a one hundred fourteen billion parameter mixture of experts model that does audio and video with singing and dialogue, and I couldn't confirm a license or weights, so I mention it only because it's on the leaderboard.
MiniMax as a consumer product is still Hailuo, and Hailuo two point three is on the aggregators. Luma shipped Ray three in September twenty twenty-five, which they market as a reasoning video model, then a Modify video editing model, then Ray three point one four in January of twenty twenty-six, faster and cheaper with native ten eighty. Luma is shipping steadily and isn't in the leaderboard top ten. Pika is alive but has pivoted: the current model is Pika two point five and the site leads with effects, an agent, and an MCP server rather than with a frontier model. That's a company that read the leaderboard and decided not to compete on it.
Seedance from ByteDance is the sleeper. Seedance two point zero sits fifth on the leaderboard under the Dreamina brand, which is ByteDance's creative app, and Seedance two point five is the thirty second model I mentioned on Runway, also on fal and Higgsfield. It accepts image, video, and audio references, and I'd expect it to matter more by the time you hear this, because ByteDance has the same distribution advantage through CapCut that Google has through YouTube.
Grok Imagine from xAI is a real product in the API sense: there's a video model with a per second price, and Replicate resells it. It isn't in the leaderboard top ten and its audio comes from separate voice models rather than from the video model, as best I can tell. If you use Grok for coding, that's a different product and it lives in the vibe coding episode. Midjourney has a video model, and I couldn't verify its current version or feature set as of this recording, so the safe statement is that it exists and it animates Midjourney images in the Midjourney style. And Higgsfield, which people ask about, isn't a lab. It's an aggregator and creative suite that fronts Seedance, Kling, Veo, and the image models, with its own effects and camera tools layered on top. That's a reasonable thing to be, and it's the same thing fal, Replicate, OpenArt, and Krea are on the developer side. The aggregator layer is where I'd start if I were building today, because the model that wins next quarter will show up there without a rewrite.
So that's the roster. Now the mechanism under all of it, which is how you keep things consistent.
Every consistency feature in this field is the same trick with a different name: conditioning. A diffusion video model starts from noise and denoises toward something that matches its conditions. Text is one condition. A reference image is another. A first frame is a very strong condition on the first few frames. A last frame is a strong condition on the last few. A reference video is a condition on motion. The product features are just different slots for conditions, and the reason image to video is the professional path is that a first frame is the strongest condition you can give. It fixes composition, lighting, and identity before the model has to invent anything.
Reference characters are the next slot. Veo takes up to three reference images and holds the subject across the clip. Kling's elements go further, letting you register a character from images or a short video, bind a voice, and refer to the element by name in a later prompt. Seedance two point five takes dozens of references in one generation. The mechanism is that the model attends to the reference embeddings while denoising, so the face in frame ninety looks like the face in the reference rather than the face in frame one after it has drifted. It isn't perfect, and drift is still the failure mode you'll fight most: the face slowly becomes a generic version of itself, the jacket changes color, the extra finger shows up in the third extension.
Start and end frames are how you control a camera move or a transformation. You give the model where the shot begins and where it ends and it interpolates the motion. This is the feature to reach for when you need a shot to land in a specific composition, because it turns a vague motion prompt into a bounded problem. It's also the feature that makes multi shot continuity possible by hand: the last frame of shot one becomes the first frame of shot two, and the model has to respect it.
Camera control ranges from text, where you write slow dolly in and hope, to explicit controls. Kling advertises six axis camera control, and several products ship camera presets, orbit, crane, push in, that map to trained motion conditions. Where the models support it, explicit camera control beats prompt text, because prompt text about cameras is where models most often ignore you.
Video to video is the editing layer, and it's where the field is heading. Runway's Aleph takes a finished clip and changes the angle, the lighting, the objects, or the style. Luma's Modify does the same. Gemini Omni Flash does it conversationally. Kling has a video to video mode. What this means for a workflow is that you stop regenerating from scratch when the shot is nearly right. You fix the shot. That's the same shift the previous episode described for images, where edit this replaced generate this, and it's arriving in video about a year later.
Extend is the last control, and the one to be careful with. Veo extends seven seconds at a time up to twenty times. Kling and most others have some version of it. Every extension is a fresh generation conditioned on the tail of the last one, so the identity and physics drift compound. My rule is that extend is for finishing a shot that needed a bit more room, not for building a scene. Scenes are built from shots, which is the closing section.
For the open weight models, there's one more control that closed models can't offer: a LoRA. You fine tune a small adapter on a few dozen clips or images of a character, a product, or a style, and then the model holds that identity without any references at inference time. The tooling exists, Musubi Tuner and its peers train LoRAs for HunyuanVideo and Wan two point x, and the community is doing the same for H three. But as of this recording Musubi doesn't support the newest frontier open model, which is a pattern you should expect: a new open model ships, and the fine tuning ecosystem catches up a few months later. If you need a LoRA workflow today, use the model the tooling already supports.
That's control. Now audio, because it changed the shape of the work more than any single feature.
Native audio means the model generates the waveform and the frames from one sampling process, so the sound is conditioned on the picture and the picture on the sound. That's why lip sync works in Veo and Kling without a separate step, and why a door slam lands on the frame the door closes. Veo three point one does this in all three tiers and Google prices it as video with audio, there's no silent option. Kling three does it as a toggle, which is why fal can show you the surcharge. MiniMax H three and LTX two point five do it in open weights. Gemini Omni Flash does it, though I couldn't confirm it's on quite the same footing as Veo. Runway's Gen four point five, as far as its release notes say, doesn't, and Wan two point two doesn't: the Wan speech to video model is driven by audio you supply rather than generating it.
The alternative is post hoc audio, and it still has a place. MMAudio is the open reference for generating synchronized sound from a finished video: it looks at the frames with a vision encoder, aligns with a synchronization module at twenty-five frames per second, and produces sound that matches. ElevenLabs sells sound effects from text with timing control. If you're on a silent model, or you want the sound design to be a separate creative decision from the picture, the post hoc path is a choice, not a compromise. Native audio is convenient. Separate audio is controllable. The next episode covers the voice and music side of that in depth.
Lip synced dialogue is where the quality gap still shows. Native dialogue from Veo and Kling is good for a line or two. For a talking head that needs to hold for a minute, the dedicated performance tools, Runway's Act Two, and the avatar and lip sync products, still do a better job because they're conditioned on a real performance rather than inventing one. I couldn't source a rigorous comparison as of this recording, so that's a judgment call, not a benchmark.
Now cost, and I want to be careful here, because prices change and I don't want you to quote this episode at your boss. So as of this recording, in round numbers.
Google's Veo three point one Quality tier costs about forty cents per second of video with audio at seven twenty or ten eighty, and more at four K. The Fast tier is about a quarter of that, around ten cents a second. The Lite tier is around five cents a second. Runway's Gen four point five is about twelve cents a second, Aleph two about twenty-eight cents a second with a minimum charge. Kling three through fal is about eight cents a second silent and about thirteen with audio. Wan two point five on fal is around five cents a second, and Grok Imagine is in the same range. Gemini Omni Flash is the odd one out, and I think the telling one: Google bills it per token, at a rate that works out to roughly ten cents a second of seven twenty video. Video pricing is migrating into the same meter as text.
What those numbers mean in practice is that an eight second Veo Quality shot with audio is a bit over three dollars, and the same shot on Kling or Gen four point five is about a dollar. Then multiply by your hit rate. Nobody publishes success rates, so here's a working number: for a shot with a specific composition and a character, budget four generations per usable clip. That puts a usable frontier shot with audio somewhere between three and thirteen dollars depending on the model, and a usable Fast or open weight shot under a dollar. The lesson is that iteration happens on the cheap tier and the final render on the expensive one. If your pipeline doesn't separate those two steps you're paying frontier prices for drafts.
The other cost is the subscription credit model. Every consumer product sells credits, a clip costs credits proportional to length and resolution, and the free tiers watermark. I couldn't verify the current caps and watermark policies per vendor, and they change monthly, so I'll just say: read the credit table before you plan a project, and prefer the API when you can, because the API price is the one you can plan on.
So that's money. Let me close with the mental model.
The unit of work is the shot. Every model I described generates a shot: one continuous take, four to fifteen seconds, one camera, one action. Nothing generates a scene. A scene is three to eight shots that cut together, and the continuity between them is your job, done with the tools from the control section: the same reference images in every shot, first and last frames that hand off to each other, the same registered elements in Kling, the same LoRA in an open model, and a consistent audio bed laid over the cut so the native sound from each shot doesn't fight the one next to it.
If you think in shots, the workflow falls out. Storyboard the scene as shots. Lock the first frame of each shot with an image model, where you have full editing control. Generate each shot on a cheap tier until the motion is right, then render it on the tier you can afford. Use extend to finish a shot, never to build a scene. Fix nearly right shots with video to video instead of regenerating. Assemble in an editor. That last step, assembly, along with voice, music, ComfyUI, and driving all of this from code, is the next episode, The AI Media Pipeline.
Quick recap. Native audio and reference based consistency are the changes that matter. Google has two models, Veo three point one for developers and Gemini Omni Flash in the app, both watermarked with SynthID. Sora is shut down, with the API closing at the end of September. Kling three is the animator grown up, with elements, multi shot, and lip sync. Runway Gen four point five is a good generator attached to the best editing model, Aleph, and an API that now sells other labs' models too. The leaderboard is topped by closed Wan three and Google Omni, while open weight MiniMax H three and LTX two point five are what you can actually run and fine tune. Start from an image, iterate cheap, render expensive, and think in shots.
Next episode: The AI Media Pipeline, voice, music, ComfyUI, APIs, and finishing.