RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 198 retrospective records ↗
Screen Method

The archive / Evidence & limits

Evidence & limits / Craft note · Note note · prepared 16 September 2026

Vendor benchmark claims disclose wildly different amounts of method

An editorial checklist for reading vendor comparison claims, built from what named pages disclose or omit.

Visual for this record: Vendor benchmark claims disclose wildly different amounts of method
Visual published by starrocks.io, shown for identification of the record. Credit: starrocks.io · source page ↗ Rights: owner-review-pending.

The shot

This is an editorial reading guide, not a new test: how to read the comparative claims vendors publish on their own product pages. Two contrasting examples, both retrieved on 16 September 2026, are Google DeepMind's Veo page, which discloses granular test conditions for its head-to-head claims, and OpenAI's Sora 2 launch announcement, which makes a comparative claim with no disclosed test methodology at all.

What the documents show

The Veo page states its 'Scene Extension' claim was tested across '80 diverse examples,' at '720x1280 resolution,' with Veo clips '8 seconds long' against competitor clips of a different stated length, shown 'without sound' except for one metric, and marked 'Last updated October 2025.' That is an unusually complete disclosure of resolution, clip count, clip length and audio condition, though it still omits rater count and recruitment. By contrast, the Sora 2 announcement states only, in prose, that the model is 'more physically accurate, realistic, and more controllable than prior systems,' illustrated with a basketball example, without naming a sample size, a rater pool, or a scoring method at all. Independent efforts such as Artificial Analysis's published methodology show what a fuller disclosure looks like from outside a vendor: a stated aggregation method, per-model vote counts, and confidence intervals, still without disclosing voter recruitment. None of these three pages is independently audited; each discloses a different slice of its own method.

The workflow

A practical reading checklist, built from these three pages: note whether resolution and clip length are stated at all; note whether the claim gives a sample size or is purely qualitative; note whether the comparison controls for audio, since Veo's own page treats sound as a separate condition; and note whether the source is the vendor's own team, an outside leaderboard, or an academic benchmark, since each carries a different evidentiary weight under this site's own labels.

What the tool does not change

None of this checklist substitutes for testing footage relevant to a specific production. A vendor page that discloses granular test conditions, like Veo's, is still the vendor's own internal test, not an outside audit; a vendor page with no disclosed methodology at all, like the Sora 2 announcement, is a claim to be read as marketing until a production tests it directly.

  • Does the claim name a sample size, resolution, and clip length, or only a qualitative comparison?
  • Is the source the vendor's own team, an independent arena, or a peer-reviewed benchmark?
  • Would the disclosed test conditions, where any exist, resemble the footage and format your production actually needs?

Reading a benchmark claim carefully means separating what is disclosed from what is asserted; the amount of disclosure varies enormously even among vendors making similar claims, and no cited page here claims to replace a production's own test.

Sources & reading trail

Veo (Google DeepMind model page) ↗

Shows a vendor benchmark claim with disclosed resolution, clip length, example count and audio condition.

Source published: Not established · Retrieved: 16 September 2026

Sora 2 is here ↗

Shows a vendor comparative claim made in prose with no disclosed sample size or test methodology.

Source published: 30 September 2025 · Retrieved: 16 September 2026

Video Generation Methodology (Artificial Analysis) ↗

Illustrates a non-vendor methodology disclosure standard, naming its aggregation method while still omitting voter recruitment.

Source published: Not established · Retrieved: 16 September 2026

Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.