RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 198 retrospective records ↗
Screen Method

The archive / Evidence & limits

Evidence & limits / Craft note · Note note · prepared 16 September 2026

A fair model comparison starts with the settings menu

An editorial method for comparing video models by fixing prompts, seeds and settings before judging results.

Visual for this record: A fair model comparison starts with the settings menu
Visual published by d3phaj0sisr2ct.cloudfront.net, shown for identification of the record. Credit: d3phaj0sisr2ct.cloudfront.net · source page ↗ Rights: owner-review-pending.

The shot

Comparing two video-generation models fairly is a documented, controllable process, not a matter of judging whichever clip looks best. Google's own Veo product page, retrieved 16 September 2026, shows one version of the alternative: Google reports that 'Veo 3.1 performs best on overall preference' against competing models on benchmarks it names, using human raters Google itself recruited. That is a vendor's own reported comparison. An independent alternative exists in the academic literature: the VBench paper, published 29 November 2023, scores video models against sixteen separate, named dimensions such as motion smoothness and subject consistency, using fixed prompts released with the paper rather than a verdict chosen by one vendor.

What the documents show

Google's page states its methodology with a disclosed gap: a footnote concedes it 'was unable to compare image to video with Sora 2 Pro because it currently does not support realistic human images,' meaning even a company's own benchmark carries stated limits on what was actually tested. VBench, by contrast, is an independent test: a fixed, published prompt set and named evaluation dimensions any outside party can rerun against a new model, rather than a result only the maker can reproduce. Separately, Runway's own help-center article documents a mechanism a comparison needs to control: a generation seed, described as a 'visual fingerprint' that can be fixed or randomized and that carries over between generations once set.

The workflow

Building a comparison this site can stand behind means fixing every variable the vendor documentation identifies as adjustable before generating a single clip: the same text prompt and, where supported, the same reference image, submitted to each model at a stated, named version; a fixed rather than randomized seed, recorded per Runway's own instructions; and identical resolution, aspect ratio, and duration settings across both tools. A test kit then applies that fixed setup across a small set of varied categories: a walking figure interacting with an object, two people exchanging dialogue, a moving camera, a physical collision, and one difficult edge case, recording every setting so the test could be rerun against a newer model version.

What the tool does not change

Neither Google's page nor Runway's documentation claims that controlling a seed or a prompt eliminates a model's limitations; Google's own footnote about Sora 2 shows a vendor's stated methodology has edges it discloses rather than hides. Fixing settings controls for one unfairness, an evaluator unintentionally giving one model an easier prompt, but it does not replace a rubric an evaluator still applies by human judgment.

  • Was the seed, model version, and prompt held constant across every compared clip, and is that documented?
  • Is a reported result reproducible by an outside party, or only by the vendor that ran it?
  • Does a benchmark's own documentation disclose what it did not test, the way Google's Sora 2 footnote does?

This is an editorial method for this site's own comparisons, built from vendor documentation about what can be controlled and an independent benchmark's example of fixed, published categories; it is not a claim that any specific test has already been run under it.

Sources & reading trail

Using seed numbers ↗

Runway's own documentation defines a generation seed as a 'visual fingerprint' and explains how to fix or randomize it across generations.

Source published: Not established · Retrieved: 16 September 2026

Veo ↗

Google's own comparisons of Veo against competitors on named benchmarks, run and reported by Google using human raters, and a disclosed testing gap.

Source published: Not established · Retrieved: 16 September 2026

VBench: Comprehensive Benchmark Suite for Video Generative Models ↗

An independently published method scoring video models against sixteen fixed-prompt dimensions rather than a single vendor-chosen comparison.

Source published: 29 November 2023 · Retrieved: 16 September 2026

Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.