Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
- Document
- 7 May 2024
- Event
- no single event
- Retrieved
- 16 September 2026
The shot
Shengshu Technology, working with Tsinghua University researchers, published a technical paper on 7 May 2024 introducing Vidu, described as "a high-performance text-to-video generator that is capable of producing 1080p videos up to 16 seconds in a single generation." The paper's own listing names a project page at shengshu-ai.com, tying the model to that company. As retrieved on 16 September 2026, Vidu's current product site shows the same name now attached to a broader, actively updated consumer platform, including a real-time interactive model and social-media-oriented tools, which the 2024 paper does not describe.
What the documents show
The paper states Vidu is "capable of generating both realistic and imaginative videos, as well as understanding some professional photography techniques, on par with Sora," which is Shengshu Technology's own comparative claim rather than an independent test against OpenAI's system; this record located no third-party evaluation confirming that specific comparison. The paper is explicit about the boundary of its own evidence, describing additional capabilities including canny-to-video generation and subject-driven generation as "initial experiments" with results the authors call merely "promising," language that itself signals these are not the paper's confirmed core claims. The current product site's description of tools like real-time generation and multi-image reference input represents later development the 2024 paper does not cover, and neither document states whether they share the original architecture.
The workflow
The paper documents a research architecture rather than a production workflow: a diffusion model built on a U-ViT backbone chosen, the authors state, for its scalability and ability to handle long videos. The product site shows this line evolved into a commercial tool with paid credits and enterprise API access, the practical entry point for a production today rather than the research paper itself. A human decision remains in verifying, against current documentation rather than the 2024 paper, exactly which model version and licensing terms govern the tool a production would actually use.
What the tool does not change
A developer's own "on par with Sora" comparison, made in a research paper's abstract, is not the same evidentiary weight as an independent side-by-side test; a production considering Vidu for a shot should treat that comparison as Shengshu Technology's claim about its 2024 research system, not as a current, verified statement about the actively updated commercial product the same name now describes.
- Does the currently marketed Vidu product still rest on the architecture the 2024 paper describes, or has it changed since?
- What would an independent comparison against Sora, or against another named competitor, need to test to confirm the paper's claim?
- Which of the paper's initial experiments, such as canny-to-video generation, became stable features in the current product?
Vidu's research paper and its current product share a name and a company, but the specific claims in one do not automatically carry over to the other, and only the current, as-retrieved documentation settles what the commercial tool now does.
Sources & reading trail
States Shengshu Technology's own architecture (a U-ViT backbone diffusion model), its stated 1080p/16-second generation capability, and its own comparison against Sora.
Source published: 7 May 2024 · Retrieved: 16 September 2026
As retrieved 16 September 2026, shows the current Vidu product line, including a real-time interactive model and consumer-facing tools, distinct from the system the 2024 research paper describes.
Source published: Not established · Retrieved: 16 September 2026
Documentation, agreements and rulings establish the note; the workflow reading is Screen Method editorial analysis. This retrospective draft does not imply the site published on the event date.