RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 198 retrospective records ↗
Screen Method

The archive / Video generation

Video generation / Craft note · Living document · prepared 16 September 2026

Vidu's own paper claims parity with Sora, not proof of it

Shengshu Technology's Vidu paper reports performance on par with Sora using the developer's own comparison, not a third party's.

arxiv.orgprimary record

Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Document
7 May 2024
Event
no single event
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The shot

Shengshu Technology, working with Tsinghua University researchers, published a technical paper on 7 May 2024 introducing Vidu, described as "a high-performance text-to-video generator that is capable of producing 1080p videos up to 16 seconds in a single generation." The paper's own listing names a project page at shengshu-ai.com, tying the model to that company. As retrieved on 16 September 2026, Vidu's current product site shows the same name now attached to a broader, actively updated consumer platform, including a real-time interactive model and social-media-oriented tools, which the 2024 paper does not describe.

What the documents show

The paper states Vidu is "capable of generating both realistic and imaginative videos, as well as understanding some professional photography techniques, on par with Sora," which is Shengshu Technology's own comparative claim rather than an independent test against OpenAI's system; this record located no third-party evaluation confirming that specific comparison. The paper is explicit about the boundary of its own evidence, describing additional capabilities including canny-to-video generation and subject-driven generation as "initial experiments" with results the authors call merely "promising," language that itself signals these are not the paper's confirmed core claims. The current product site's description of tools like real-time generation and multi-image reference input represents later development the 2024 paper does not cover, and neither document states whether they share the original architecture.

The workflow

The paper documents a research architecture rather than a production workflow: a diffusion model built on a U-ViT backbone chosen, the authors state, for its scalability and ability to handle long videos. The product site shows this line evolved into a commercial tool with paid credits and enterprise API access, the practical entry point for a production today rather than the research paper itself. A human decision remains in verifying, against current documentation rather than the 2024 paper, exactly which model version and licensing terms govern the tool a production would actually use.

What the tool does not change

A developer's own "on par with Sora" comparison, made in a research paper's abstract, is not the same evidentiary weight as an independent side-by-side test; a production considering Vidu for a shot should treat that comparison as Shengshu Technology's claim about its 2024 research system, not as a current, verified statement about the actively updated commercial product the same name now describes.

  • Does the currently marketed Vidu product still rest on the architecture the 2024 paper describes, or has it changed since?
  • What would an independent comparison against Sora, or against another named competitor, need to test to confirm the paper's claim?
  • Which of the paper's initial experiments, such as canny-to-video generation, became stable features in the current product?

Vidu's research paper and its current product share a name and a company, but the specific claims in one do not automatically carry over to the other, and only the current, as-retrieved documentation settles what the commercial tool now does.

Sources & reading trail

Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models ↗

States Shengshu Technology's own architecture (a U-ViT backbone diffusion model), its stated 1080p/16-second generation capability, and its own comparison against Sora.

Source published: 7 May 2024 · Retrieved: 16 September 2026

Vidu AI Video Generator (product site) ↗

As retrieved 16 September 2026, shows the current Vidu product line, including a real-time interactive model and consumer-facing tools, distinct from the system the 2024 research paper describes.

Source published: Not established · Retrieved: 16 September 2026

Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.