Google I/O 2024: Introducing Veo and Imagen 3 generative AI tools
- Document
- 14 May 2024
- Event
- 14 May 2024
- Retrieved
- 16 September 2026
The shot
On 14 May 2024, at Google I/O, Google introduced Veo alongside the image model Imagen 3. The announcement post, credited to VP Eli Collins and Senior Research Director Douglas Eck, states Veo 'generates high-quality 1080p resolution videos in a wide range of cinematic and visual styles that can go beyond a minute.' The post adds that, 'starting today, Veo is available to select creators in private preview in VideoFX' by joining a waitlist, and highlights an early collaboration with filmmaker Donald Glover's studio Gilga. It names a lineage of prior Google video research the model builds on, including DVD-GAN, Imagen-Video, Phenaki, WALT, VideoPoet and Lumiere.
What the documents show
The 1080p, over-a-minute claim is Google's own stated capability, not an outcome any independent benchmark confirms in this post. Google DeepMind's current Veo model page, as retrieved on 16 September 2026, documents native audio generation and precise camera controls as features of Veo 3.1 specifically, describing options such as 'move back' and 'zoom in.' Neither term, audio or camera control, appears in the May 2024 post. The two documents together show a model family whose current, most-discussed capabilities were absent from the version Google actually announced at I/O.
The workflow
For a production in May 2024, Veo was reachable only through a waitlist-gated private preview inside VideoFX, not a general product, subscription or API. The Donald Glover collaboration illustrates the vendor's selected-partner testing model rather than a self-serve entry point available to any filmmaker. The announcement describes generation output and stylistic understanding only; it offers no editing, extension or camera-direction controls at this stage, meaning any resulting clip still needed a human continuity pass before it could sit inside an edit, exactly as any other pre-visualisation source would.
What the tool does not change
Google's own description of an 'advanced understanding of natural language and visual semantics' documents how the model interprets a prompt; it does not substitute for a director framing what a shot needs to accomplish in a scene. A waitlist-only rollout also meant no production could treat Veo as a dependable pipeline stage in May 2024, whatever its stated capabilities. Nothing in the announcement claims general availability, and no independent evaluation is cited in the post itself.
- Which named version of a model, not just which vendor, is a production actually being offered access to.
- Has a capability advertised today, such as audio or camera control, been confirmed as present in the version under discussion.
- Is a demonstration built from a private, curated collaboration, or from output any account holder could reproduce.
Read side by side, the two documents mark Veo's actual starting point in May 2024: a 1080p, waitlist-gated preview without audio generation or camera control, a baseline worth keeping distinct from the audio-and-camera-equipped later version described in Google's currently live documentation.
Sources & reading trail
Google's own announcement stating Veo's launch capabilities (1080p, over a minute) and its private-preview waitlist access model.
Source published: 14 May 2024 · Retrieved: 16 September 2026
Google DeepMind's current model page, describing audio generation and camera controls as features of Veo 3.1, distinguishing them from the original announcement.
Source published: Not established · Retrieved: 16 September 2026
Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.