RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 198 retrospective records ↗
Screen Method

The archive / Video generation

Video generation / From the archive · 16 November 2023 event · prepared 16 September 2026

Emu Video split generation into an image step and a video step

Meta's Emu Video paper describes a two-step image-then-video method and reports preference results Meta gathered itself.

arxiv.orgprimary record

Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Document
17 November 2023
Event
16 November 2023
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The shot

Meta's AI research division announced Emu Video on 16 November 2023, in a research blog post titled "Introducing Emu Video and Emu Edit, our latest generative AI research milestones," alongside a separate image-editing research paper. The accompanying technical paper, posted to arXiv the following day, describes a text-to-video method that factorizes generation into two diffusion steps: first producing an image conditioned on a text prompt, then producing a video conditioned on both the text and that generated image. Meta states the resulting model generates 512-by-512, four-second videos at 16 frames per second using just two diffusion models, where it says prior work needed a "deep cascade" of as many as five.

What the documents show

Both documents report the same core preference-study numbers: Meta states its generations were "preferred over Make-A-Video by 96% of respondents based on quality and by 85% of respondents based on faithfulness to the text prompt," and the paper extends the comparison, reporting its videos strongly preferred over Google's Imagen Video by 81%, Nvidia's PYOCO by 90%, and outperforming commercial systems including RunwayML's Gen-2 and Pika Labs. These are Meta's own human-evaluation results, run and reported by the paper's authors; this record could not locate an independent replication of these specific percentages. The blog post explicitly frames Emu Video as building on Meta's earlier Emu image model, one dated data point in a publication sequence this site otherwise covers only at its earlier and later ends.

The workflow

Emu Video is documented here as a research paper, not a released tool: neither source offers a public download, an API, or a license. The practical fact the record establishes is architectural rather than operational: the factorized, image-then-video approach the paper describes reduces the number of models a pipeline needs, a design choice later research and products can adopt or diverge from, not a step a production can run today. A human decision is unavoidable at the level of tracking which vendors eventually build shippable tools on similar factorized methods.

What the tool does not change

A vendor's own preference-study percentages, however specific, describe results under conditions Meta selected and reported; they are not a substitute for a production's own side-by-side comparison of tools it can actually license. This is an editorial caution the two documents do not state themselves.

  • Did Meta ever release Emu Video as a public product, API, or open model following this research announcement?
  • What comparison would a production need to run today to test whether Emu Video's stated advantages carried into any shipped Meta tool?
  • How does the factorized image-then-video method described here relate to Meta's later Movie Gen research?

Emu Video is a documented step in Meta's video-generation research, useful for tracing method and stated results, but its own materials stop well short of describing a tool a filmmaker could use.

Sources & reading trail

Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning ↗

States the paper's own factorized image-then-video method and Meta's self-reported human-evaluation preference results against Imagen Video, PYOCO, Make-A-Video, Gen-2, and Pika Labs.

Source published: 17 November 2023 · Retrieved: 16 September 2026

Introducing Emu Video and Emu Edit, our latest generative AI research milestones ↗

Meta's own announcement, dated 16 November 2023, describing Emu Video as research built on the Emu image model and stating its preference-study percentages against Make-A-Video; cited via Wayback Machine archive since the live page no longer resolves.

Source published: 16 November 2023 · Retrieved: 16 September 2026

Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.