English

Model ReleasesGoogleGemini Omni

Google Introduces Gemini Omni for Natural Language Video Editing and Multimodal Generation

This article is a translation. Read the Japanese original

Google has announced "Gemini Omni," an AI video editing model. The company positions Gemini Omni as the video version of "Nano Banana." Users can iteratively edit the aesthetics, actions, and effects of a video through natural language dialogue. Each edit builds upon the previous state, allowing the scene to progress while maintaining consistency and coherence.

Gemini Omni combines an intuitive understanding of physics with Gemini's knowledge of historical, scientific, and cultural contexts. According to Google, this bridges the gap between photorealism and meaningful storytelling. It is capable of converting any reference—such as images, text, video, or audio—into a single integrated output.

Specific examples provided include prompts where a mirror ripples like liquid when a person touches it, or a person's arm transforms into a reflective mirror material. Interactive edits are also possible, such as a toy animal making its characteristic sound when touched by a finger. Additionally, effects such as 3D architectural structures, the sun, or airplanes appearing on a palm can be created based on reference images.

Furthermore, users can replace specific characters or objects, change camera angles, or swap the environment with that of another image. Google claims that by combining Gemini's reasoning capabilities with creativity, Gemini Omni brings a leap forward in world understanding, multimodality, and editing capabilities.


Source: Gemini Omni (HN 323pt, 146 comments) (HN Search (backfill), 2026-05-20)