
The shot
Before EvalCrafter, a team comparing text-to-video models mostly had two borrowed image- and video-quality numbers to lean on. On 17 October 2023, researchers at Tencent AI Lab with City University of Hong Kong, the University of Macau and the Chinese University of Hong Kong posted the EvalCrafter paper, later shown at CVPR 2024, proposing a framework built around 700 prompts drawn from an analysis of real-world user requests and constructed with the help of a large language model, then scored with 17 separate objective metrics spanning visual quality, content quality, motion quality and text-video alignment.
What the documents show
The paper states that a per-model score is produced by fitting coefficients that align the 17 objective metrics to scores collected in a human user study, which the authors report correlates with human opinion more closely than simply averaging the raw metrics. The accompanying project leaderboard, as retrieved on 16 September 2026, still lists only 2023-era models such as VideoCrafter, Gen-2, PikaLab and Zeroscope, evidence of the specific comparison window this is scoped to. That leaderboard is an independent test result for those model versions; it says nothing about any model released since. The paper's own Limitation section, numbered 5.3, states three named gaps: 700 prompts cannot represent every real-world request; evaluating general motion quality is 'hard' with tools available at the time; and the human alignment labels came from 'fewer human annotators,' which the authors say may introduce bias into the fitted score.
The workflow
As the repository documents it, a team generates its own model's output against the published 700-prompt list, runs the 17 metrics locally, and can request the current prompt set from the project to join the public leaderboard. Used this way, EvalCrafter gives a producer comparing candidate models a dated snapshot for the specific prompt style, camera motion and subject types the benchmark encodes, not a general verdict.
What the tool does not change
The authors' own admission that annotator count was limited means the 'human alignment' step encodes a small group's taste, not a production's client or audience. Whether the prompt style in the 700-prompt set resembles a specific shoot's needs, and whether 2023-era leaderboard rankings say anything about a model released years later, are judgments EvalCrafter leaves to whoever is running the comparison.
- Does the 700-prompt set resemble the kind of prompt your production would actually write?
- How recent are the model versions on the leaderboard you are reading, and does that date match the model you plan to use?
- What did the paper's own annotator-count limitation do to the reliability of its human-alignment score?
EvalCrafter's contribution was a repeatable, multi-metric methodology at a moment when the category had almost none; reading its leaderboard as a live ranking of current models, rather than the dated comparison the authors published, extends the claim past what the paper supports.
Sources & reading trail
States the 700-prompt methodology, 17 objective metrics, human-alignment fitting, and the paper's own stated limitations.
Source published: 17 October 2023 · Retrieved: 16 September 2026
Shows the leaderboard is populated with 2023-era model versions, evidencing the benchmark's comparison window.
Source published: Not established · Retrieved: 16 September 2026
Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.