Create, edit and star in videos with two Google Vids updates

Google has announced a significant expansion of its AI-powered video creation suite, Google Vids, introducing two major features designed to streamline the production of professional-grade content: Gemini Omni and Personal Avatars. These updates represent a pivot toward a more conversational and accessible video editing experience, allowing users to generate, refine, and appear in video content using natural language prompts and digital likenesses. By integrating the latest advancements in multimodal AI, Google aims to eliminate the traditional barriers to video production—such as technical editing skills, expensive equipment, and time-consuming filming schedules—positioning Google Vids as a central tool for corporate communication, marketing, and training.
The Integration of Gemini Omni: Conversational Video Editing
At the core of this update is Gemini Omni, a multimodal AI model that allows users to interact with the video creation process through "everyday language." While previous iterations of AI video tools focused primarily on text-to-video generation, Gemini Omni introduces a more sophisticated "chat-to-edit" workflow. This feature enables users to treat the AI as a collaborative editor rather than a static generator.
The system supports multimodal inputs, meaning a user can provide a text prompt alongside reference images, such as a brand logo, a rough sketch of a storyboard, or a specific photograph, to guide the AI’s output. This ensures that the generated content aligns more closely with a specific creative vision or corporate identity. Once a draft is created, the "chat-to-edit" functionality allows for iterative refinements. Instead of manually navigating a timeline or adjusting keyframes, users can simply type instructions like "change the background to a modern office setting," "fix the lighting to make it warmer," or "add a cinematic filter to this segment."
This shift toward step-by-step editing is intended to solve one of the primary frustrations with generative AI: the "all-or-nothing" nature of early models. By allowing for granular, language-based adjustments, Google Vids provides a level of control that bridges the gap between automated generation and manual professional editing.
Digital Twins in the Workplace: The Rise of Personal Avatars
Perhaps the most provocative feature included in the update is the introduction of Personal Avatars. This tool allows users to create a high-fidelity digital version of themselves to "star" in videos without ever having to step in front of a camera. To set up an avatar, a user must upload a selfie and a short voice recording to train the model. Once the avatar is generated, the user can simply type a script, and the digital twin will deliver the message with synchronized lip movements, facial expressions, and a voice that mimics the original recording.
This feature targets the growing demand for personalized video messaging in professional environments. In a corporate landscape increasingly defined by remote work and asynchronous communication, personalized video updates are often more engaging than text-based emails but are frequently avoided due to the time required for setup and recording. Personal Avatars allow executives, managers, and educators to provide a "human" face to their communications at scale.
Google has implemented strict guardrails for this technology. Personal Avatars are currently restricted to users aged 18 and older in specific regions. Furthermore, the system is designed to prevent the creation of unauthorized deepfakes; avatars are strictly linked to the user’s Google Account and are restricted to the account holder’s own likeness.
Chronology of Google’s AI Video Evolution
The rollout of Gemini Omni and Personal Avatars is the latest milestone in a rapid development cycle for Google’s video efforts. To understand the significance of these updates, it is necessary to view them within the context of Google’s broader AI roadmap:
- April 2024: Google first unveiled Google Vids at its Cloud Next conference, describing it as an AI-powered video creation app for work that sits alongside Docs, Sheets, and Slides.
- May 2024: Google introduced Veo, its most capable generative video model, designed to compete with industry rivals like OpenAI’s Sora and Runway’s Gen-3.
- February 2025: The company integrated Veo 3.1 into Google Vids, significantly improving the visual fidelity and temporal consistency of generated clips.
- Present Day: The introduction of Gemini Omni and Personal Avatars shifts the focus from "generation" to "personalization and refinement," making the tool more viable for daily business use cases.
This timeline reflects an industry-wide race to capture the enterprise market. While competitors have focused on high-end cinematic generation, Google’s strategy emphasizes integration within the Workspace ecosystem, leveraging its existing user base of millions of businesses.
Market Context and Supporting Data
The demand for accessible video tools is supported by shifting trends in workplace productivity and content consumption. According to recent industry reports, video is expected to account for over 80% of all internet traffic by 2025. In the corporate sector, a 2024 survey of HR and internal communications professionals found that employees are 75% more likely to watch a video than read a long-form document or email.
However, the "production gap" remains a significant hurdle. Traditional video production can cost anywhere from $1,000 to $5,000 per finished minute for professional-grade content. By automating the editing and "starring" roles, Google Vids aims to reduce these costs to the price of a Workspace subscription.

The generative AI market as a whole is projected to add trillions of dollars in value to the global economy. Gartner predicts that by 2026, 30% of outbound marketing messages from large organizations will be synthetically generated, up from less than 2% in 2022. Google’s latest updates are a direct play for this burgeoning market, competing against specialized startups like Synthesia and HeyGen, which have previously dominated the AI avatar space.
Transparency and Ethical Safeguards
As the line between real and synthetic media continues to blur, Google has emphasized its commitment to "Responsible AI." A critical component of the new Google Vids updates is the integration of SynthID. Developed by Google DeepMind, SynthID is an invisible digital watermark embedded directly into the pixels of generated video clips.
Unlike traditional watermarks, SynthID is designed to be undetectable to the human eye but identifiable by specialized software, even if the video is cropped, compressed, or edited. This provides a layer of content transparency, allowing viewers or platforms to verify that a video was created with AI.
"We want you to feel confident sharing what you create," stated Justin Luk, Product Manager at Google, in the official announcement. "To ensure content transparency, every generated clip includes an invisible SynthID digital watermark. This allows people to verify that a video was created with AI, so you can explore your creativity and share your AI videos responsibly."
These safeguards are part of a broader industry push toward provenance standards, such as those championed by the Coalition for Content Provenance and Authenticity (C2PA). By embedding these features at the product level, Google is attempting to mitigate the risks associated with misinformation and deepfakes while still enabling the creative benefits of the technology.
Broader Implications for the Future of Work
The implications of Gemini Omni and Personal Avatars extend beyond simple convenience. They signal a shift in the "democratization" of media production. In the same way that desktop publishing in the 1980s allowed anyone to become a layout artist, and digital photography in the 2000s turned everyone into a photographer, generative AI is turning office workers into video producers.
For training and development (L&D) departments, this means the ability to create localized, multi-language training videos in minutes. For sales teams, it allows for the creation of personalized video pitches for hundreds of prospects simultaneously. For internal communications, it enables a more personal touch in global organizations where face-to-face interaction is rare.
However, the rise of digital avatars also raises questions about the "authenticity" of professional interactions. As digital twins become indistinguishable from their human counterparts, the value of "live" presence may increase, or conversely, the definition of presence may evolve to include digital representation.
Access and Availability
Gemini Omni and the Personal Avatar features are currently available to a specific subset of Google’s user base. This includes subscribers of Google AI Pro and Ultra plans, as well as Google Workspace business customers. By gating these features behind premium tiers, Google is positioning them as professional-grade tools rather than consumer novelties.
As the technology matures, it is expected that these features will see wider rollout, potentially integrating more deeply with other Workspace apps—such as an avatar that can "attend" a Meet call or a Gemini Omni feature that can turn a Google Doc into a full video presentation with a single click.
With these updates, Google Vids has transitioned from a promising experiment to a robust platform for the future of enterprise communication. By combining the conversational intelligence of Gemini with the visual potential of digital avatars, Google is betting that the future of the office is not just digital, but increasingly cinematic.







