HunyuanVideo: A Systematic Framework For Large Video Generative Models
- Document
- 3 December 2024
- Event
- 3 December 2024
- Retrieved
- 16 September 2026
The shot
On 3 December 2024, Tencent researchers posted 'HunyuanVideo: A Systematic Framework For Large Video Generative Models' to arXiv, later revised through a sixth version by March 2025. The paper reports a video generative model of 'over 13 billion parameters,' calling it 'the largest among all open-source models' at the time. The accompanying code repository dates the practical release separately, stating 'Dec 3, 2024: We release the inference code and model weights of HunyuanVideo,' and lists supported configurations up to '720px1280px129f,' a 720p recommended resolution generating 129 frames per clip.
What the documents show
The paper states that 'according to evaluations by professionals, HunyuanVideo outperforms previous state-of-the-art models, including Runway Gen-3, Luma 1.6, and three top-performing Chinese video generative models,' a claim the repository elaborates as a human evaluation of '1,533 text prompts' judged by 'more than 60 professional evaluators' across text alignment, motion quality and visual quality. This is Tencent's own reported comparison; neither document shows the named competitor systems' own scores on their own evaluations, only Tencent's judges rating outputs against each other. The two documents agree on the parameter count and the professional-evaluation method, giving the claim more procedural detail than a bare marketing statement would.
The workflow
A team adopting HunyuanVideo works from open weights and code rather than a hosted product, choosing between the 720p and a lower '544px960px129f,' or 540p, configuration depending on available compute. The 129-frame ceiling stated in the repository sets a fixed clip length a pipeline must plan around, whether by generating multiple segments or accepting a short single take; nothing in either document describes an automated way to extend a generation beyond that frame count.
What the tool does not change
A professional-evaluator panel judging prompt alignment and motion quality is a structured preference study, not a substitute for a supervising editor deciding whether a specific generated take serves a specific scene. The paper's own framing is that it aims to 'bridge the gap between closed-source and open-source communities,' a stated goal about ecosystem access, not a claim that the released weights match a closed competitor's product on every task a production might need.
- Were the named comparison models tested at their own recommended settings, or only as Tencent's evaluators configured them.
- Does a 129-frame ceiling require a production to plan multi-segment generation for anything longer than roughly five seconds.
- Is 'professional evaluator' defined in the source with enough detail to judge how the panel was selected.
Tencent's own paper and repository document a specific, dated open-weights release with a stated evaluation method; the comparison itself remains the developer's self-reported result, worth citing as exactly that rather than as an independently confirmed ranking.
Sources & reading trail
Tencent's technical paper stating the parameter count and reporting its own professional-evaluator comparison against named competitor models.
Source published: 3 December 2024 · Retrieved: 16 September 2026
Tencent's code repository dating the practical release and detailing the evaluation's prompt count and evaluator panel size, plus supported resolutions.
Source published: Not established · Retrieved: 16 September 2026
Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.