
The shot
On 29 September 2022, Meta AI Research published a blog post and an accompanying paper introducing Make-A-Video, a system that turns a written prompt into a few seconds of video without training on a dataset of videos paired with captions. The paper, credited to thirteen Meta researchers including Uriel Singer and Yaniv Taigman, frames the release alongside sample clips on a project page: unicorns painted in watercolor, a teddy bear washing dishes, a dog in a superhero outfit. These are lab-selected demonstrations, not footage from any production.
What the documents show
Meta's own account of the method is specific: the system, in the company's words, 'learns what the world looks like from paired text-image data and how the world moves from video footage with no associated text.' An existing text-to-image model supplies the visual vocabulary, while separate spatiotemporal modules, trained on unlabeled video, supply motion. The paper describes a pipeline of decoding, frame interpolation, and super-resolution stages that extend a short low-resolution clip into a longer, sharper one. The project page adds comparative figures the company generated itself, claiming '3x Better representation of text input' and '3x Higher quality' against a prior system, based on Meta's own user studies. Because these numbers come from the paper's authors and were never run by an outside evaluator against a fixed benchmark, this is a promotional demonstration paired with a vendor-reported comparison, not an independent test.
The workflow
No production workflow existed around Make-A-Video at release. Meta describes a 'demo experience' shared alongside the research, not a licensed tool, and the project page says the team was still testing before any broader release. The documented pipeline previews a sequence later commercial models would ship as a product: describe a scene in text, generate a short clip, then run interpolation and upscaling passes for smoothness and resolution. Meta also states the system can animate a still image or generate variations of an existing video, functions that would resurface, packaged for actual use, in later tools from other vendors.
What the tool does not change
Nothing in Meta's materials claims the system handles story structure, blocking, performance, or continuity between shots; the paper stays scoped to a generation method. The company states it applies source-data filtering and 'mandatory watermarking' to outputs, a disclosure measure, not a substitute for editorial review of what a clip depicts. Sound, dialogue, and where a generated clip belongs in a larger sequence stay outside what either document describes.
- Is the clip in front of you a lab-selected sample or output from a tool anyone else can run?
- Does a quoted benchmark number come from the model's own authors or from a separate evaluator?
- What continuity, sound, and story decisions has no cited document claimed the system makes?
Make-A-Video is best read as a marker of method rather than of production readiness. It documents a training recipe, pairing labeled images with unlabeled video, that later systems extended, but Meta's own framing keeps the outputs research samples with self-reported comparisons. A record of a training strategy is not evidence that any production has used it.
Sources & reading trail
Meta's own announcement of the method, its training data split between images and unlabeled video, and the framing of outputs as research demonstrations.
Source published: 29 September 2022 · Retrieved: 16 September 2026
The technical paper documenting the spatiotemporal decomposition pipeline and the authors' own state-of-the-art claims.
Source published: 29 September 2022 · Retrieved: 16 September 2026
Meta's project page stating its own comparative benchmark figures and its watermarking and pre-release testing measures.
Source published: Not established · Retrieved: 16 September 2026
Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.