Wan3.0

Wan3.0 is Alibaba's higher-end all-in-one multimodal video model for longer, reference-heavy, production-oriented creation. It combines text-to-video, first- and last-frame image-to-video, reference-driven video generation, localized video editing, and temporal extension in one model, with support for image, video, audio, document, and webpage inputs. It is especially well suited to branded storytelling, product demos, explainers, music-led sequences, and other commercial workflows that need stronger character or product consistency, more impactful audiovisual motion, and tighter control over complex multi-input prompts.

Complete technical specification for integration
Ready-to-use code snippets for common workflows
Step-by-step tutorials for advanced use cases
Image inputs with Wan 3.0 How to use images with Wan 3.0's video generation: pin an image as the first frame, morph between two frames, and carry subject or product identity across new scenes with reference images.
Prompting Wan 3.0 How to write prompts for Wan 3.0 that get you the shot: how the model reads layered directives, camera language, style and register range, directing native audio, and the dimensions and duration rules.
Video and audio references with Wan 3.0 How to use video and audio references with Wan 3.0: carrying a subject across new scenes with referenceVideos, driving video from an audio source with referenceAudios, and combining multiple reference types in one call.