
Wan 3.0 Video from Alibaba, a 30-second clip straight from your slide deck
Alibaba released Wan 3.0, the first model that turns a slide deck, spreadsheet, or PDF directly into video. Let's break down what it does, how it differs from the competition, and how to work with it in the unitool.ai interface.
Video Straight From a Slide Deck
Every video generator over the past couple of years shared one trait. You describe a scene in words or attach an image, and the model draws a clip. Alibaba's Wan 3.0 adds something none of its competitors currently offer. The model reads documents.
We're talking about ordinary work files, presentations, spreadsheets, text documents, PDFs, and even links to web pages. You upload a café's brand guide, add a single sentence describing what you want, and get a finished commercial back.
Public beta opened on August 6, 2026, and on August 24 Alibaba officially released the model to everyone. The second major change is length. A clip now runs up to 30 seconds in a single generation, double the previous Wan 2.7.
Thirty seconds in one take instead of stitched clips
Most generators output a few seconds, which gives you a shot rather than a story. Alibaba framed the shift this way, the model used to generate a frame, now it tells a complete story.
The difference isn't just the number. Thirty seconds leaves room for narrative pacing, continuous camera moves, and complex techniques like one-take sequences. There's no need to stitch the clip together from short pieces, which means no visible seams.
Two supporting features make this practical. Smart duration recommendation analyzes your prompt and suggests an appropriate length, so a simple product beat doesn't get padded to thirty seconds and an ambitious idea doesn't get cut short. Video extension lets you append a continuation to an existing clip so the story keeps developing.
Documents become video
This is the release's signature feature, and nothing else on the market matches it. Beyond text, images, audio, and video, Wan 3.0 accepts files in doc, xls, ppt, pdf, txt, and md formats, plus key, pages, and numbers, meaning Apple's office formats. The limit is simple, one file or link per request, up to 100 megabytes and 50 pages.
Here's what that produces in practice,
A product deck becomes a commercial or brand film
A training deck becomes video courseware
Spreadsheet data becomes animated charts
A management report becomes a narrated video briefing
This works because the model treats your prompt as the control center. It explains how to interpret the uploaded file, what to pull from it, and how to present it. That's why wildly different source material turns into coherent video rather than a slideshow with animation.
Lifelike faces and stable characters
AI-generated people have a recognizable problem, they all look like the same glossy person. Alibaba tackled exactly that. The model was trained to produce diverse, believable faces, render features and skin with more care, keep emotional expression natural and restrained, and link facial micro-expressions to body language. Even in group scenes, different characters carry their own shades of emotion.
The other half of the work concerns consistency, which matters most in production. In reference-to-video mode, the model holds four things stable at once,
Characters, including facial features, hairstyle, hair color, build, clothing, and accessories
Props, their appearance from multiple angles, structure, logos, and materials
Space, meaning character blocking and camera perspective in correct relation to each other
Style, where cinematic tone renders accurately and multi-style projects avoid bleeding between looks
One more capability carried over from Wan 2.7 and still works. A finished video can be edited by changing visuals, plot, and dialogue without regenerating from scratch. A revision stays a revision instead of becoming a do-over.
How Wan 3.0 Video works in the unitool.ai interface
Now for the practical side. On our site the model runs in three modes, and the right one is selected automatically based on what you attach to your request.
Text to video. Prompt only, the scene is built from scratch.
First frame to video. You attach a photo and the model animates that exact frame. You can add a last frame too, and the model will build a transition between them. Image requirements are jpeg, png, webp, or bmp, sides from 240 to 8000 pixels, an aspect ratio no steeper than eight to one, and a file size up to 20 megabytes. Transparency in png isn't supported.
References to video. Up to ten images, with characters and style holding steady through the whole clip. In the prompt you refer to them as Image1, Image2, and so on.

An important restriction, the first frame and reference images can't be used at the same time, you have to pick one scenario. The available settings are resolution at 480p, 720p, or 1080p, aspect ratio, duration from 2 to 30 seconds, and a separate audio toggle. The model scores the clip with speech, ambient sound, and music according to the prompt, and if you don't want audio, you simply switch the toggle off.
Reference video and reference audio
Beyond images, the model accepts finished clips and audio tracks as reference material.
Video limits work like this, up to five clips in mp4 or mov format, each 1 to 15 seconds long, no more than 15 seconds in total, up to 100 megabytes each, with sides between 240 and 4096 pixels. In the prompt you refer to them as Video1, Video2.
Audio follows a similar scheme, up to five wav or mp3 files, each 1 to 15 seconds, no more than 15 seconds total, up to 15 megabytes. Referenced in the prompt as Audio1, Audio2. There's a catch here, audio can't be your only attachment, you need to pair it with an image or video.
An important billing note. When you use reference video, both the input clip seconds and the output seconds are charged, and their combined total can't exceed 30 seconds. Worth keeping in mind when planning your budget.
What generation costs
Alibaba charges per second of finished video, with the rate depending on resolution. That works out to 0.05 dollars per second at 480p, 0.10 dollars at 720p, and 0.20 dollars at 1080p. So a full thirty-second clip at maximum quality costs 6 dollars, while a 480p draft runs just 1.5 dollars.
That leads to a simple working habit that applies to our interface too. Run drafts at 480p or 720p while you dial in the wording and composition, then produce the final version at 1080p. The per-second price gap between lowest and highest quality is fourfold, so the savings during the drafting stage add up quickly.
In the unitool.ai interface, cost is calculated in tokens and shown right below the input field before you launch the generation, so there are no surprises. A five-second clip at 720p, for example, costs 61 tokens.
What Alibaba admits still needs work
The company openly named two areas that still need improvement.
The first is audio quality. Sound is generated, but its texture isn't yet at the level the developers want.
The second, and more significant for anyone working in Russian, is the accuracy of on-screen text rendering. If you need readable signage, captions, or logos inside the video, the result may disappoint. Here it's worth either budgeting for several attempts or adding text afterward during editing.
Alibaba describes both as areas of active refinement, so future versions should improve on them.
From Document to Finished Clip
Wan 3.0 solves a problem nobody had solved before it. Turning a work file, whether a deck, a spreadsheet, or a report, directly into video with no scriptwriter and no editor. Add thirty seconds in one take, stable characters from references, and edits to a finished clip without regenerating it.
That said, Alibaba named the weak spots honestly, audio and on-screen text. So for clips with readable lettering, other models currently do better, while for long coherent scenes or video built from an office document, Wan 3.0 has almost no competition.
As always, a universal model doesn't exist. The only way to know which one handles your specific task is hands-on, running one idea through several generators and comparing the results with your own eyes.
That's exactly what unitool.ai is for. Wan 3.0 Video is already available alongside dozens of other video, image, and audio models, under one subscription with no separate sign-up on each platform. Start with a 480p draft, dial in your wording, then produce the final version at maximum quality.