RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 198 retrospective records ↗
Screen Method

The archive / Evidence & limits

Evidence & limits / Craft note · Note note · prepared 16 September 2026

An independent arena ranks video models by blind human votes

Artificial Analysis discloses its Elo methodology and vote counts but not how its voters are recruited.

Visual published with the cited source for this record: An independent arena ranks video models by blind human votes
Visual published with the cited source, shown for identification of the record. Credit: artificialanalysis.ai · source page ↗ Rights: owner-review-pending.

The shot

Rather than a vendor scoring its own model, an independent analysis group runs a public arena: Artificial Analysis's text-to-video leaderboard, as retrieved on 16 September 2026, shows people two videos generated from the same prompt by different, unlabeled models and records which one they prefer.

What the documents show

The site's own video methodology page states that 'users compare two videos generated from the same prompt by different models and select the one they prefer,' with results aggregated using 'Bradley-Terry Maximum Likelihood Estimation' and rescaled 'to an Elo-like range for readability,' recomputed hourly as new votes arrive. The methodology page states explicitly that Elo is 'reported separately for each modality' and that 'arena matchups only pair outputs from the same modality,' meaning a silent clip is never compared against an audio-enabled one, and video editing is never compared against generation. The leaderboard page itself lists per-model vote counts ranging from roughly 3,000 to more than 20,000, each score shown with a 95 percent confidence interval. This is an independent test in the sense used here: a repeatable method with named model versions and comparative results, run by a third party rather than a vendor. What neither page discloses is how voters are recruited or screened, or any demographic or professional-experience information about who is doing the judging.

The workflow

As the methodology page describes it, a visitor is shown a blind pair, votes for a preferred clip, and that single vote feeds the hourly Bradley-Terry recalculation across every model in that modality's pool. A production team can read the resulting ranking as a rough, continuously updated popularity signal for a specific prompt style and modality, not as a test of the specific footage or use case that team needs.

What the tool does not change

A blind vote from an anonymous visitor is not the same judgment a working editor, colorist or client would make about a shot's fitness for a specific scene, and the site does not claim otherwise. Because the confidence intervals and vote counts vary by model, a narrow Elo gap between two models may not represent a reliable preference at all; the page's own disclosed vote counts are the only way to check that before treating a ranking as decisive.

  • How many votes support the specific models you are comparing, and how wide is the stated confidence interval?
  • Are you comparing scores across modalities the arena itself says are never matched against each other?
  • Does an anonymous voter's blind preference tell you anything about how a shot performs in your specific edit?

An arena-style leaderboard is a genuine independent signal precisely because a vendor cannot set its own test conditions, but the site's own silence on voter recruitment means the population doing the judging is, by its own account, undisclosed.

Sources & reading trail

Video Generation Methodology (Artificial Analysis) ↗

States the blind-comparison voting method, the Bradley-Terry Elo aggregation, and the same-modality pairing rule.

Source published: Not established · Retrieved: 16 September 2026

Text to Video Leaderboard (Artificial Analysis) ↗

Shows per-model Elo scores, confidence intervals, and vote counts for the current text-to-video ranking.

Source published: Not established · Retrieved: 16 September 2026

Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.