← All models
AlibabaVideo

Wan 3.0

All-in-one video generation with native audio and up to 30-second clips

Try Wan 3.0

About

Wan 3.0 is Alibaba's all-in-one video generation model. It automatically supports text-to-video, first-frame image-to-video, and multi-reference video generation based on your inputs. It creates synchronized audio, supports up to 10 reference images, and can generate clips up to 30 seconds at 720p or 1080p. Wan 2.7 generation and editing modes remain available for legacy workflows.

CinematicCreativeCommercialStylized

How to prompt

  • Describe motion and pacing: "slow motion", "time-lapse", "smooth tracking shot", "handheld camera"
  • For image-to-video, provide a strong first frame — the model animates forward from it
  • Use multiple reference images and refer to them as "Image 1", "Image 2", and so on in the prompt
  • Describe dialogue, ambience, music, and sound effects directly in the prompt to use Wan 3.0's native audio generation
  • Longer durations (15–30s) work best with a clear sequence of simple actions
  • Specify aspect ratio to match your platform: 16:9 for YouTube, 9:16 for TikTok/Reels

Best for

Social media video, product animation, character-consistent clips, native audio, commercials, and longer-form generation

Good to know

Wan 3.0 is currently a preview model, so availability and output behavior may change. First-frame and multi-reference modes cannot be combined in one request. Legacy Wan 2.7 modes do not generate native audio.

Ready to create with Wan 3.0?

Open Lyvia