Google Opens Gemini Omni Flash Video Generation to Developers at $0.10 Per Second

Google has officially released its Gemini Omni Flash model to developers and enterprise customers through the Gemini API and Google AI Studio, as reported by VentureBeat. The rollout on June 30, 2026 marks the model's transition from consumer preview — first shown at Google I/O in May — into a production-ready tool for businesses building video workflows at scale.
The model generates 720p video at $0.10 per second, matching the Veo 3.1 Fast tier but dramatically undercutting standard Veo 3.1, which costs $0.40 per second. That 75% price reduction puts a ten-second clip at roughly $1 — a meaningful shift for developers and marketing teams that previously found AI video cost-prohibitive at scale.
Conversational Editing Changes the Production Workflow
The headline capability isn't the generation itself — it's iterative, conversational editing. Built on Google's new interactions API, a stateful interface designed for multi-turn tasks, Omni Flash lets users describe changes in plain language and have the model apply them to an existing clip without regenerating from scratch. A marketer can relight a product shot, reframe a scene, or swap wardrobe details through a sequence of natural-language instructions, preserving what already worked.
This collapses what was previously a five-tool pipeline — separate LLMs for scripting, image generators, video models, lip-sync tools, and voice generators — into a single model that accepts text, images, and video as inputs and returns a finished clip with synchronized audio.
Enterprise Features: Brand Assets, Text Insertion, and Deepfake Guardrails
For enterprise teams, Omni Flash supports multimodal reference inputs: feed it a product photo or brand logo alongside a text prompt and the model incorporates the real asset into the generated scene rather than inventing a generic placeholder. The model also handles in-scene text and signage — useful for localized training videos or ads that need logos placed contextually.
Google has built explicit limits against misuse. The model refuses to lip-sync a still photo of a person to an audio track — a direct measure against deepfakes. It will, however, translate existing video speech into other languages, useful for organizations with multilingual teams. Every output carries Google's SynthID watermark, and the company is extending C2PA Content Credentials across its generative video tools.
Current constraints are worth noting: clips cap at 10 seconds, and the model only outputs at 720p — no 1080p or 4K option. Teams needing longer content must stitch clips together, and premium brand work requiring high-resolution output will need to look at the Veo 3.1 tiers instead.
Competitive Context: Sora Gone, Runway Watching
The API launch comes as OpenAI has discontinued Sora, its own text-to-video product, and Runway remains the leading independent alternative in the enterprise AI video market. Google's pricing and the conversational editing angle represent a direct play for the enterprise segment those competitors targeted.
For independent developers and small teams especially, the $0.10/second rate at the API level means video generation is now economically viable for prototype workflows and internal tooling that wouldn't justify Runway's subscription costs or the complexity of stitching together multiple point tools.
Gemini Omni Flash is available now via the Gemini API and Google AI Studio. Developers can access the model card and technical specifications through Google DeepMind's published documentation.
Originally reported by VentureBeat. Read the original article for additional details.
View original source