Google launches Gemini Omni, the I/O 2026 release that wasn't about agents
Gemini Omni Flash, Google's new video-generation model, shipped May 19 at I/O 2026 alongside Gemini 3.5. Conversational editing is the differentiator; ten-second clips and SynthID watermarks are the current shape.
Google opened I/O 2026 with two model launches, not one. Gemini 3.5 took the agentic-and-coding headline, and Gemini Omni Flash, a new multimodal video generation model, shipped alongside it on May 19. Omni is the first model in a family Google describes as one that “can create anything from any input, starting with video,” and the launch positions Google squarely against OpenAI’s Sora, Runway, and Pika in the video generation tier where the company had been lagging a year ago.
What shipped
Gemini Omni Flash accepts images, audio, video, and text as input and produces video as output. Current clips run ten seconds, which Google characterizes as a deployment choice rather than a model limit. The distinctive product surface is conversational editing: once a clip is generated, a user can ask Omni to “make the background a rainy Tokyo street” or “give him a leather jacket” and the model maintains character consistency, physics, and scene continuity across multiple turns. Every generated video carries an imperceptible SynthID watermark.
The rollout is wider than most frontier model launches. Gemini Omni Flash is available globally to Google AI Plus, Pro, and Ultra subscribers through the Gemini app and Google Flow. It is also free on YouTube Shorts and YouTube Create, which puts the model in front of consumers who have not signed up for an AI subscription. APIs for developers and enterprise customers are dated “in the coming weeks,” meaning the developer surface is still ahead of the consumer one.
Koray Kavukcuoglu, CTO of Google DeepMind, framed the model around the editing loop: “Gemini Omni gives you an easier way to edit video, with natural language.” That framing matters. Most video generation models so far have been one-shot, prompt-in, clip-out; Omni’s pitch is that the next clip is a conversation, not a new prompt. For a working developer or content team, that’s the difference between treating video as a generated artifact and treating it as an editable session.
Where this lands in the market
The strategic frame for the I/O 2026 announcements is now clearer than it was last week. Gemini 3.5 was Google’s argument that frontier-model capability still matters and that Google can compete at the top tier on agentic and coding work. Gemini Omni is the parallel argument for generative video, where OpenAI’s Sora launch in 2024 had given OpenAI an early lead and Runway and Pika had owned the indie creative tier. Omni Flash entering the market for free on YouTube Shorts is a distribution move competitors cannot match.
For developers, the relevant question is what Omni does to the build-or-buy calculus on creative tooling. A team that has been piecing together Runway, ElevenLabs, and a stitching pipeline now has a single-vendor option that handles iterative edits in one conversation. The trade-off is the Google-stack dependence and the ten-second clip cap, which makes longer-form work a stitching problem again. The “starting with video” phrasing in the announcement also suggests Omni is positioning to expand into image and audio outputs, at which point Google is competing across the full generative-media stack rather than just one slice of it.
The “riskiest feature” framing from TechTimes points at the avatar capability, which lets users generate videos using their own voice and likeness. Google has held this feature back from the broad launch, which reads as an acknowledgment that voice and likeness generation is where the next set of safety and policy battles is going to land.
What’s worth watching
Three signals are worth tracking over the next 90 days:
- API pricing when Omni reaches developers. Free distribution on YouTube Shorts is a consumer-acquisition move, not the eventual developer pricing. Where Omni lands relative to Sora, Runway, and Pika per-clip and per-second will decide whether teams move work onto it.
- The clip-length lift. Ten seconds is enough for short-form social but not for ads, explainers, or product demos. Whether the cap moves to thirty seconds or a minute within 90 days will signal whether the constraint was capacity or a model limitation Google prefers not to advertise.
- The avatar surface. If Google ships the voice and likeness feature broadly, the safety scaffolding around it (consent flows, detection, watermark robustness) becomes the new bar for the category. If Google holds it back, expect OpenAI or Meta to test the boundary first.
The strategic frame to hold: Gemini Omni is the half of I/O 2026 that did not get the attention Gemini 3.5 did, but for content teams and creative tools it may be the more consequential release. Stackmaven’s follow-up coverage will land around August 19.
- Google: Introducing Gemini Omni (May 19, 2026) blog.google
- TechCrunch: Google's Gemini Omni turns images, audio, and text into video techcrunch.com
- Google: 9 demos of Gemini Omni and Gemini 3.5 in action (May 29, 2026) blog.google
- TechTimes: Google Launches Gemini Omni Video Model, but Holds Back Its Riskiest Feature www.techtimes.com