RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 198 retrospective records ↗
Screen Method

The archive / Video generation

Video generation / From the archive · 6 August 2024 event · prepared 16 September 2026

CogVideoX splits its open license between two model sizes

Zhipu AI's two open CogVideoX tiers carry different licenses, and only the smaller one is unrestricted for commercial use.

github.comprimary record

CogVideo & CogVideoX (repository README)

Document
6 August 2024
Event
6 August 2024
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The shot

Zhipu AI's CogVideoX entered its repository's own release log on 6 August 2024, when the smaller CogVideoX-2B model and an accompanying 3D Causal VAE were open-sourced; the larger CogVideoX-5B followed on 27 August 2024. A technical report, posted to arXiv on 12 August 2024, describes the model as a text-to-video diffusion system built around a 3D variational autoencoder and an "expert transformer" with expert adaptive LayerNorm, and states it can generate ten-second, 768-by-1360-pixel video at 16 frames per second.

What the documents show

The paper's claim of "state-of-the-art performance across both multiple machine metrics and human evaluations" is Zhipu AI's own reported result. The repository's release notes add a fact the paper does not: the two model sizes do not carry the same license. CogVideoX-2B's license "has been changed to the Apache 2.0 License," the notes state, while the larger CogVideoX-5B is released under a separate CogVideoX License. That document states it "allows you to freely use all open-source models in this repository for academic research," but that "users who wish to use the models for commercial purposes must register and obtain a basic commercial license," capped at "1 million visits per month" before a further license is required, and it separately restricts military or illegal use.

The workflow

For a production or tool-builder, the practical fact the license text settles is that downloading CogVideoX-5B's weights is not the same as clearing it for a monetized release: registration through the vendor's own commercial-license form is a required step the code repository does not perform automatically. CogVideoX-2B, by contrast, carries no such registration step under Apache 2.0. A human decision remains about which tier's cost, quality, and licensing terms fit a given release, and about confirming current registration requirements directly with Zhipu AI before shipping.

What the tool does not change

The license's own restriction language, including its bar on military or illegal use and its reference to Chinese law for dispute resolution, is not legal advice and does not substitute for a production's own counsel reviewing the current document, since the license itself notes it is "subject to update to a more comprehensive version." The paper's benchmark numbers likewise remain Zhipu AI's self-reported comparison until an independent evaluation states otherwise.

  • Which CogVideoX model size does a planned use actually require, and does that size carry the Apache 2.0 license or the separate CogVideoX License?
  • Has the production registered for a commercial license where the model's own terms require one?
  • What does the current, as-retrieved version of the license state, given the license text's own note that it may be updated?

CogVideoX shows that "open-source" is not one fixed condition across a single project: two model sizes from the same repository can carry two different sets of obligations, and only the license text itself, not the paper or the marketing framing, settles which applies.

Sources & reading trail

CogVideo & CogVideoX (repository README) ↗

States the 6 August 2024 open-source release of CogVideoX-2B and the differing Apache 2.0 vs custom CogVideoX License terms for the 2B and 5B tiers.

Source published: Not established · Retrieved: 16 September 2026

CogVideoX-5b LICENSE (The CogVideoX License) ↗

Gives the CogVideoX License's own text: free for academic research, a required registered commercial license for commercial use, a 1-million-visits-per-month cap, and restrictions against military or illegal use.

Source published: Not established · Retrieved: 16 September 2026

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer ↗

States the model's 3D VAE and expert transformer architecture, the 10-second/768x1360/16fps generation capability, and the paper's own state-of-the-art performance claim.

Source published: 12 August 2024 · Retrieved: 16 September 2026

Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.