Wan 3.0
All-in-one video generation with native audio and up to 30-second clips
Try Wan 3.0About
Wan 3.0 is Alibaba's all-in-one video generation model. It automatically supports text-to-video, first-frame image-to-video, and multi-reference video generation based on your inputs. It creates synchronized audio, supports up to 10 reference images, and can generate clips up to 30 seconds at 720p or 1080p. Wan 2.7 generation and editing modes remain available for legacy workflows.
How to prompt
- Describe motion and pacing: "slow motion", "time-lapse", "smooth tracking shot", "handheld camera"
- For image-to-video, provide a strong first frame — the model animates forward from it
- Use multiple reference images and refer to them as "Image 1", "Image 2", and so on in the prompt
- Describe dialogue, ambience, music, and sound effects directly in the prompt to use Wan 3.0's native audio generation
- Longer durations (15–30s) work best with a clear sequence of simple actions
- Specify aspect ratio to match your platform: 16:9 for YouTube, 9:16 for TikTok/Reels
Best for
Social media video, product animation, character-consistent clips, native audio, commercials, and longer-form generation
Good to know
Wan 3.0 is currently a preview model, so availability and output behavior may change. First-frame and multi-reference modes cannot be combined in one request. Legacy Wan 2.7 modes do not generate native audio.
Ready to create with Wan 3.0?
Open Lyvia