
The shot
A domino chain gets interrupted mid-fall by a rubber duck: does a generated continuation know the far side of the chain should still tumble? That is the kind of test posed by 'Do generative video models understand physical principles?', posted to arXiv on 14 January 2025 by researchers including a Google DeepMind team, alongside the Physics-IQ project page hosting the benchmark and code.
What the documents show
The paper states the authors filmed 396 real videos, 66 physical scenarios shot from three camera angles in two takes each, at 3840x2160 resolution with static cameras, covering solid mechanics, fluid dynamics, optics, thermodynamics and magnetism. Each tested model, given a conditioning frame, must generate a plausible continuation, scored against the real footage on where action happens, when it happens, how much of it happens, and how it happens, combined into a single Physics-IQ score normalized so that the natural variance between two real takes of the same scenario equals 100 percent. Tested across Sora, Runway Gen 3, Pika 1.0, Lumiere, Stable Video Diffusion and VideoPoet, the paper reports that 'the best model' scored only '29.5%' on that scale, which the authors state is evidence that 'physical understanding is severely limited, and unrelated to visual realism.' This is the authors' own independent test: a repeatable method, named model versions, and comparative results, run by a lab that itself makes at least one of the compared products.
The workflow
As the paper documents it, a model under test receives only the conditioning frame or frames, generates a five-second continuation, and that output is compared frame-by-frame against the real filmed continuation using the four component metrics before aggregation into the single score. A visual-effects supervisor evaluating a model for physically grounded shots, water, collisions, falling objects, could read the per-scenario breakdown rather than the aggregate score to see which physical categories a candidate model handles better or worse.
What the tool does not change
The paper states plainly that 'visual realism doesn't imply physical understanding,' meaning a shot can look convincing while still violating the physical law the scenario is testing. Deciding whether a specific generated shot's physics reads as acceptable to an audience, versus what this benchmark's frame-by-frame comparison measures, remains a creative and technical judgment the paper does not make for a production.
- Does your intended shot involve a physical interaction similar to one of the paper's 66 tested scenarios?
- Are you judging a model on visual realism alone, when the paper explicitly separates that from physical accuracy?
- Which of the four component metrics, location, timing, amount or manner of motion, matters most for your specific shot?
A 29.5 percent score against real-world physical variance is a stated, reproducible result for the six models the authors tested in January 2025; it is not evidence about any model version released after that testing window closed.
Sources & reading trail
States the Physics-IQ dataset design, the four-metric scoring method, and the reported 29.5 percent top score.
Source published: 14 January 2025 · Retrieved: 16 September 2026
Hosts the benchmark project page confirming the dataset and code release alongside the paper.
Source published: Not established · Retrieved: 16 September 2026
Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.