-
Google has unveiled Gemini Omni, a major new family of generative AI models designed to create content from virtually any input. The first model, Omni Flash, specializes in generating AI videos using text, photos, videos, and audio. According to Google's blog post, the long-term vision for Omni is to "create anything from any input" — hence the "Omni" name.
Google positions Omni Flash as a video counterpart to its popular Nano Banana image generation model, which has already generated over 50 billion images since its launch. With Omni Flash, users can generate clips up to 10 seconds long with both video and audio. The company is already working on extending this duration.
Google already offers Veo, a text-to-video generation model. However, Omni Flash goes beyond by accepting video as input and incorporating far more world knowledge. According to Koray Kavukcuoglu, CTO of Google DeepMind, Omni Flash has "a lot" more world knowledge than Veo thanks to Gemini's extensive training data.
Gemini Omni Flash will be available starting Tuesday in the Gemini app, Google Flow, and YouTube Shorts. This integration makes AI video generation accessible directly within popular Google platforms.
One notable feature is the ability to insert a user's likeness into videos. While this may raise privacy concerns, Nicole Brichtova, who leads the Omni product team, notes that Google has seen strong adoption of similar features with Nano Banana image generation. Users have already generated billions of images using their own likeness.
Google envisions Omni evolving beyond video generation to handle any type of content creation. The company is actively working on extending clip length beyond 10 seconds and adding more capabilities. This positions Gemini Omni as a central pillar of Google's AI strategy going forward.
Comment