RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 198 retrospective records ↗
Screen Method

The archive / Evidence & limits

Evidence & limits / From the archive · 5 June 2024 event · prepared 16 September 2026

The best model still failed physics most of the time

VideoPhy's own 2024 test found generated video followed captions and physics together only 39.6% of the time, with stated limits.

arxiv.orgprimary record

VideoPhy: Evaluating Physical Commonsense for Video Generation

Document
5 June 2024
Event
5 June 2024
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The shot

On 5 June 2024, researchers from the University of California, Los Angeles and Google Research posted 'VideoPhy: Evaluating Physical Commonsense for Video Generation' to arXiv, revising it on 3 October 2024 to add two more tested models. The paper, by Hritik Bansal, Zongyu Lin, Tianyi Xie, Zeshun Zong, Chenfanfu Jiang, Yizhou Sun, Kai-Wei Chang and Aditya Grover of UCLA with Michal Yarom and Yonatan Bitton of Google Research, introduces a benchmark of 688 curated prompts testing whether generated video depicts material interactions — solid-solid, solid-fluid, fluid-fluid — according to real-world physics.

What the documents show

The paper reports its own result plainly: across the models tested — Zeroscope, LaVIE, VideoCrafter2, OpenSora, CogVideoX and a Stable Video Diffusion pipeline as open models, and Runway's Gen-2, Pika, Lumiere and Dream Machine as closed ones — 'the best performing model, CogVideoX-5B, generates videos that adhere to the caption and physical laws for 39.6% of the instances,' based on the authors' own human evaluation. That is the authors' self-reported result for the specific models and prompts tested on the dates stated, not an independent audit, and not a shared test set with the separately run Physics-IQ benchmark. The project's GitHub repository, maintained by the lead author, lists the same 39.6% figure and hosts training instances for a proposed automatic evaluator, VideoCon-Physics, built to score newly released models without repeating full human annotation.

The workflow

The benchmark's method — curating prompts through a 'three-stage data curation pipeline,' generating video from each tested model, then scoring outputs by human annotators for both caption adherence and physical plausibility — describes a research evaluation, not a production tool. A citation to this benchmark is a citation to the authors' scored snapshot of specific model versions available in mid-to-late 2024; the paper's own limitations note that testing 'an exhaustive list of models' is 'financially and computationally challenging,' so newer or unlisted models were simply not included.

What the tool does not change

The authors state their own limitations directly, and they belong with any citation: annotations were gathered largely from Amazon Mechanical Turk workers based in the US and Canada, so, in the paper's words, 'the human annotations in this work do not represent the diverse demographics around the globe' and 'reflect the perceptual biases of the annotators from Western cultures.' A later, differently designed benchmark, VideoPhy-2, appeared in 2025 and should not be treated as the same test or scores as the paper described here.

  • Which specific model version and generation date does a cited VideoPhy score refer to, given how quickly text-to-video models are updated?
  • Does a claim distinguish VideoPhy's material-interaction prompts from Physics-IQ's separate object-permanence tests, rather than treating 'passed a physics benchmark' as one claim?
  • Whose cultural assumptions did the human annotators bring to scoring 'physical commonsense,' and is that population representative of the intended audience?

VideoPhy's 39.6% figure is a specific, dated, self-reported measurement of specific models, not a permanent verdict on video generation's grasp of physics. The paper's own caveats about model coverage and annotator demographics are part of what that number means.

Sources & reading trail

VideoPhy: Evaluating Physical Commonsense for Video Generation ↗

The paper's own dataset construction, tested models, headline 39.6% result, and stated limitations on model coverage and annotator demographics.

Source published: 5 June 2024 · Retrieved: 16 September 2026

VideoPhy Dataset and Benchmark (GitHub Repository) ↗

The authors' own maintained repository confirming the reported leaderboard figure and describing the VideoCon-Physics auto-evaluator and later VideoPhy-2 follow-up.

Source published: Not established · Retrieved: 16 September 2026

Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.