Create Wan 3.0 Video Online

A Wan 3.0 generator that takes more than a prompt: start and end frames, up to 10 image, 5 video and 5 audio references, a document, or a public webpage. Render 2 to 30 seconds at 480P, 720P, or 1080P with native audio.

The Wan 3.0 model

One Wan 3.0 model, five ways into the shot.

Wan 3.0 is Alibaba's all-in-one video model, and wan3.0-video reads text, images, video, audio, documents, and webpages as a single context. Give it whichever material you already have, and it returns one consistent take with native sound.

PromptDescribe the subject, action, setting, style, and camera movement in words.
FramesPin the opening image, and optionally the closing one. First frame required; last frame optional.
ReferencesGuide identity, motion, atmosphere, or sound with media. Up to 10 images, 5 videos, and 5 audio files.
Context · Document OR WebpageUse one document or public webpage as story context. Private and sign-in links are not supported.
Wan 3.0
One model
Generated Video
2–30s · 480P / 720P / 1080P · Native audio
4 routesWorkflow families
2–30sOutput duration
480P–1080PResolution
Native audioSound in one pass
Wan 3.0 tutorial

How to use the Wan 3.0 video generator

Three steps: pick the input route, write the prompt and name your references, then set duration and resolution.

  1. Choose

    Pick your input route

    Text to video for a prompt-only shot. Frames when the opening and closing images are fixed. Reference to video when an image, a clip, or a track should carry identity, motion, or rhythm. Document or webpage when the brief already exists somewhere and you would rather not retype it.

    TextFramesReference
  2. Write

    Write the prompt and name your references

    Keep the same order every time: subject, action, setting, camera movement, lighting, style, ending state, audio. Then say what each attached file is for. Naming references by number — Image1, Video1, Audio1 — is the part of a Wan 3.0 prompt guide that actually changes the result.

    "Keep Image1's face. Follow Video1's walk. Cut to a high angle at 3s."
  3. Generate

    Set duration and resolution, then generate

    Choose 2 to 30 seconds, a resolution, and an aspect ratio — or leave it adaptive and let Wan 3.0 match your input. Native audio is on by default and can be switched off. Check the credit cost, submit, and leave if you want: the task keeps running and the result waits for you on this page.

    Wan 3.0 · 480P / 720P / 1080P2–30s6 ratios · Adaptive
Use cases

What Wan 3.0 video generation is actually for

01Text to video

A 30-Second Short, In One Pass

Wan 3.0 generates up to 30 seconds natively, so a prompt can carry camera moves, a line of dialogue, and native sound design without stitching five-second clips together. Long enough for a trailer, a brand story, or a scene with a real beginning, middle, and end — photoreal or cartoon, the prompt decides.

02Image to video

One Reference Image, Any World

Upload a character still and name it Image1 in the prompt. Wan 3.0 holds the face, hair, wardrobe, and proportions, then puts that person somewhere new and moves them. One image is enough to keep identity stable across the whole clip — no LoRA, no training run.

03Reference to video

One Take, Every Angle

Feed a continuous action clip as Video1 and keep the performance exactly as shot. Change only the camera: close-up, high angle, side track, over-the-shoulder, insert. This is the closest thing to a Wan 3.0 edit — coverage that used to need a second unit, generated from one reference video.

Wan 3.0 samples

One prompt in,
one finished take out

Every clip below came from text alone — no reference image, no source footage. A one-take dance, a 2D cartoon, hand-painted animation, a four-shot scene. Each Wan 3.0 prompt is one you can adapt; the structure is what carries the result.

See every input route
Wan 3.0 · 1080P00:08

"One continuous shot, dynamic follow mixed with an orbiting perspective — a Dunhuang dancer in a warm golden light beam inside a dark stone cavern"

Lip sync00:05

"Frontal medium close-up, she speaks straight to camera, brows slightly furrowed, a faint handheld breathing sensation"

2D cartoon00:06

"Japanese sports anime 2D animation — a boxer throws continuous punches, the gloves visibly rebounding off the heavy bag"

Hand-drawn00:06

"Hand-drawn animation, oil painting brushstrokes, cold palette — a green beam lifts a cow off an American farm"

Multi-shot00:10

"3D animated fantasy in four shots — a giant transparent plant rises from a glowing night lake"

Ultra-wide00:08

"Ultra-widescreen plush universe — a fabric spaceship flees a deep-blue plush whale past knitted yarn planets"

FAQ

Wan 3.0 questions, answered

Prompting, references, duration limits, editing, credit cost, and the difference between wan3.0-video and Prime — the details worth knowing before you generate. Can't find what you're looking for?

Wan 3.0 is Alibaba's all-in-one video model, published on the Kie model market as wan3.0-video. It reads text, images, video, audio, documents, and webpages as one context and returns a single 2–30 second take at 480P, 720P, or 1080P, with a native audio track on by default. Videoify runs Wan 3.0 AI as a hosted service through an upstream API provider, so there are no model weights to download, no environment to configure, and no GPU of your own.

Wan 3.0 online

Generate your next
Wan 3.0 video

Bring a prompt, a pair of frames, a reference set, a document, or a link. Pick 2 to 30 seconds at up to 1080P and let Wan 3.0 render it with sound.

Account and credits required · Cost shown before generation