A prompt becomes a shot
The text-to-video model supports landscape, portrait, square, 4:3, and 3:4 framing.
Video generation with audio
Get to know Wan 2.7: video generation with audio and reference-based tools for shaping your visual story.
The Wan3.video workspace currently offers W3.0 and W3.0 Pro. This page is a model guide.

Wan 2.7
Start with a prompt or work from references. Choose the specific Wan 2.7 model that matches the material you have.
The text-to-video model supports landscape, portrait, square, 4:3, and 3:4 framing.
Text-to-video can generate matching audio or use an audio file you provide.
The dedicated reference-to-video model accepts reference images, videos, and audio to guide a new scene.
Write a scene from scratch, or prepare a reference image when the model supports it. Decide what should move and what should stay consistent.
Name the subject, action, setting, lighting, and camera movement. Start with one clear moment before attempting a more complex sequence.
Check motion, composition, and subject consistency. Change one instruction at a time so you can see what improves the result.
The workspace currently offers W3.0 and W3.0 Pro, not a selectable Wan 2.7 model. The create button opens that workspace; the links below lead to official Wan 2.7 documentation.
Alibaba Cloud documents 720P and 1080P output, with 2–15 seconds for its text-to-video model. Check the specific model and service for other workflows.
Yes, the dedicated reference-to-video model supports image and video references. Its input requirements differ from the text-to-video model.
Describe the subject and one clear action, then add camera movement and lighting. When sound matters, describe the atmosphere or dialogue you want. Review the result before adding more detail.
Model details: official documentation. Options vary by model and service.
Ready to create online? Open the Wan 3.0 workspace, choose your settings, and build your first shot.