A Wan 3.0 generator that takes more than a prompt: start and end frames, up to 10 image, 5 video and 5 audio references, a document, or a public webpage. Render 2 to 30 seconds at 480P, 720P, or 1080P with native audio.
The Wan 3.0 model
One Wan 3.0 model, five ways into the shot.
Wan 3.0 is Alibaba's all-in-one video model, and wan3.0-video reads text, images, video, audio, documents, and webpages as a single context. Give it whichever material you already have, and it returns one consistent take with native sound.
PromptDescribe the subject, action, setting, style, and camera movement in words.
FramesPin the opening image, and optionally the closing one. First frame required; last frame optional.
ReferencesGuide identity, motion, atmosphere, or sound with media. Up to 10 images, 5 videos, and 5 audio files.
Context · Document OR WebpageUse one document or public webpage as story context. Private and sign-in links are not supported.
Wan 3.0
One model
Wan 3.0
One model
Generated Video
2–30s · 480P / 720P / 1080P · Native audio
4 routesWorkflow families
2–30sOutput duration
480P–1080PResolution
Native audioSound in one pass
Wan 3.0 tutorial
How to use the Wan 3.0 video generator
Three steps: pick the input route, write the prompt and name your references, then set duration and resolution.
01Choose
Pick your input route
Text to video for a prompt-only shot. Frames when the opening and closing images are fixed. Reference to video when an image, a clip, or a track should carry identity, motion, or rhythm. Document or webpage when the brief already exists somewhere and you would rather not retype it.
TextFramesReference
02Write
Write the prompt and name your references
Keep the same order every time: subject, action, setting, camera movement, lighting, style, ending state, audio. Then say what each attached file is for. Naming references by number — Image1, Video1, Audio1 — is the part of a Wan 3.0 prompt guide that actually changes the result.
"Keep Image1's face. Follow Video1's walk. Cut to a high angle at 3s."
03Generate
Set duration and resolution, then generate
Choose 2 to 30 seconds, a resolution, and an aspect ratio — or leave it adaptive and let Wan 3.0 match your input. Native audio is on by default and can be switched off. Check the credit cost, submit, and leave if you want: the task keeps running and the result waits for you on this page.
Wan 3.0 generates up to 30 seconds natively, so a prompt can carry camera moves, a line of dialogue, and native sound design without stitching five-second clips together. Long enough for a trailer, a brand story, or a scene with a real beginning, middle, and end — photoreal or cartoon, the prompt decides.
02Image to video
One Reference Image, Any World
Upload a character still and name it Image1 in the prompt. Wan 3.0 holds the face, hair, wardrobe, and proportions, then puts that person somewhere new and moves them. One image is enough to keep identity stable across the whole clip — no LoRA, no training run.
03Reference to video
One Take, Every Angle
Feed a continuous action clip as Video1 and keep the performance exactly as shot. Change only the camera: close-up, high angle, side track, over-the-shoulder, insert. This is the closest thing to a Wan 3.0 edit — coverage that used to need a second unit, generated from one reference video.
Wan 3.0 samples
One prompt in, one finished take out
Every clip below came from text alone — no reference image, no source footage. A one-take dance, a 2D cartoon, hand-painted animation, a four-shot scene. Each Wan 3.0 prompt is one you can adapt; the structure is what carries the result.
"One continuous shot, dynamic follow mixed with an orbiting perspective — a Dunhuang dancer in a warm golden light beam inside a dark stone cavern"
Lip sync00:05
"Frontal medium close-up, she speaks straight to camera, brows slightly furrowed, a faint handheld breathing sensation"
2D cartoon00:06
"Japanese sports anime 2D animation — a boxer throws continuous punches, the gloves visibly rebounding off the heavy bag"
Hand-drawn00:06
"Hand-drawn animation, oil painting brushstrokes, cold palette — a green beam lifts a cow off an American farm"
Multi-shot00:10
"3D animated fantasy in four shots — a giant transparent plant rises from a glowing night lake"
Ultra-wide00:08
"Ultra-widescreen plush universe — a fabric spaceship flees a deep-blue plush whale past knitted yarn planets"
FAQ
Wan 3.0 questions, answered
Prompting, references, duration limits, editing, credit cost, and the difference between wan3.0-video and Prime — the details worth knowing before you generate. Can't find what you're looking for?
Wan 3.0 is Alibaba's all-in-one video model, published on the Kie model market as wan3.0-video. It reads text, images, video, audio, documents, and webpages as one context and returns a single 2–30 second take at 480P, 720P, or 1080P, with a native audio track on by default. Videoify runs Wan 3.0 AI as a hosted service through an upstream API provider, so there are no model weights to download, no environment to configure, and no GPU of your own.
Open the generator at the top of this page, pick an input route, write your prompt, then set duration, resolution, and aspect ratio. The credit cost appears before you submit. Wan 3.0 costs 2 credits per second at 480P, 4 at 720P, and 8 at 1080P, so a 5-second 720P clip is 20 credits. When you attach a reference video, the seconds of that input are billed at the same rate as the output, because the model processes both. You can compose a prompt while signed out and the draft is restored after you sign in, but generating needs an account and enough credits. A failed generation returns your credits automatically.
Describe one shot in this order: subject, action, setting, camera movement, lighting, style, ending state, and the audio you want to hear. Say what changes during the clip instead of describing a frame that already exists. State the style outright — photoreal, 2D cartoon, hand-drawn, anamorphic — because a Wan 3.0 cartoon comes back flat and outlined only if you ask for it in those words. Close with what you do not want: no subtitles, no extra people, no on-screen text. Prompts accept up to 20,000 characters in English or Chinese, but a tight paragraph beats a stack of adjectives, and anything past ten seconds is better written as a short beat-by-beat timeline than as more adjectives.
In reference to video, files are numbered in the order you attach them and you address them in the prompt by name: Image1 and Image2 for identity, wardrobe, or location, Video1 for motion and timing, Audio1 for rhythm and pace. That naming is the whole trick — "keep Image1's face, follow Video1's walk" is far more reliable than "use the references." It is also how Wan 3.0 music to dance works: attach a track as Audio1 and ask the dancer to hit its accents, and the choreography lands on the beat instead of drifting. Limits are 10 images, 5 video clips totalling 15 seconds, and 5 audio clips totalling 15 seconds. Audio works best alongside an image or video rather than on its own, and reference media cannot be combined with a first or last frame.
Yes, up to a 30-second ceiling. With no video input, one generation produces 2 to 30 whole seconds. To extend footage you already have, switch to reference to video, attach the clip as Video1, and describe what happens next — Wan 3.0 continues from it rather than starting over. The catch is that reference video length plus output length must stay within 30 seconds, so a 10-second input leaves you 20 seconds of new material. For anything longer, work in passes: reuse the same reference images so the character and setting stay recognizable, generate each stretch on its own, and cut them together in an editor.
Within limits. A Wan 3.0 edit means describing a change over a reference clip — replace the subject, swap the background, remove an accessory, shift the visual style, or re-frame the camera while the performance stays put — and the model generates a new take that follows your instruction. What it is not is a timeline. There is no inpainting brush, no keyframe editor, no object-level repair, and nothing guarantees that every pixel outside your described change survives untouched. Use a dedicated editor for cuts, captions, transitions, and exact frame work.
Kie publishes two tiers of the same model: wan3.0-video, the standard one, and wan3.0-video-prime, a high-speed variant with the same inputs, the same 30-second ceiling, and the same 480P, 720P, and 1080P options. The stated difference is generation speed, not capability. Videoify runs the standard wan3.0-video here, which also means there is no separate "Prime reference" quality switch to find: on either tier, quality comes from the references you supply. Use clean, unobstructed, well-lit source images at a resolution close to your output, give each reference exactly one job, name each one in the prompt, and remove anything that contradicts another file. Fewer and better references at 1080P will do more for the result than any change of tier.
Right here in the browser — the generator at the top of this page is the whole tool, with nothing to install and no API key to manage. Videos export without a watermark. Two layers govern what you can make. Videoify's acceptable use terms rule out illegal activity, infringement of intellectual-property, privacy, or publicity rights, sexual content involving minors, non-consensual intimate content, impersonation, harassment, and misleading synthetic media released without a disclosure the law requires. Separately, the upstream provider running the model applies its own filters, which can reject a prompt or a reference file that our terms would technically allow — most often around real public figures, graphic violence, and explicit material. We do not publish a filter list and the boundary can shift, so when a prompt is refused, rewrite the sensitive part rather than trying to route around it.
Wan 3.0 online
Generate your next Wan 3.0 video
Bring a prompt, a pair of frames, a reference set, a document, or a link. Pick 2 to 30 seconds at up to 1080P and let Wan 3.0 render it with sound.