# Wan 2.2 > Wan 2.2 is Alibaba's open-weights AI video model, released in July 2025 under Apache 2.0, with a mixture-of-experts design: one expert lays out each clip, another refines the detail. In Kenerate AI's studio you can run Wan 2.2 and Wan 2.2 Ultra Fast for text or image to video, 5 or 8 seconds at 480p or 720p, no GPU. URL: https://kenerateai.com/wan-2-2 Publisher: Kenerate AI (https://kenerateai.com) Last updated: 7 October 2026 ## Key facts - Made by: Alibaba (Wan team), July 2025 - Licence: Apache 2.0, open weights - Modes here: Text to video, image to video - Length · resolution: 5 or 8 s · 480p or 720p - Aspect ratios: 16:9 and 9:16 - Sound: No (silent clips) - Versions here: Wan 2.2, Wan 2.2 Ultra Fast - Wan 2.2 (Alibaba): Text, or image (+ end frame) → 5 or 8 s · 480p or 720p. Open-model quality, quick clips. - Wan 2.2 Ultra Fast (Alibaba): Text, or image (+ end frame) → 5 or 8 s · 480p or 720p. Fast, low-cost drafts. - Wan 2.2 Animate (Alibaba): Photo + a reference video → Your character, moving. Motion transfer. - Wan 3 (Alibaba): Text or image (+ end frame) → 2–30 s · up to 1080p · sound. Longer clips with sound. ## Steps 1. Open the video studio — Pick Wan 2.2, or Wan 2.2 Ultra Fast for quicker drafts. 2. Start from text or an image — Image to video takes a start image and an optional end frame; text to video needs only the prompt. 3. Write the motion — A prompt is required: one action, one camera move and the pace. In image to video, leave the look to the image. 4. Pick 5 or 8 s and a resolution — Draft at 480p, keep the take you like at 720p; choose 16:9 or 9:16 for text to video. ## In depth ### What Wan 2.2 is Wan 2.2 is the video model Alibaba's Wan team open-sourced on 28 July 2025, code and weights, under Apache 2.0 [1]. Its 14B text-to-video and image-to-video models use a mixture-of-experts design: a high-noise expert lays out each clip early in generation and a low-noise expert refines the detail, about 27B parameters in total with 14B active per step [1][2]. Compared with Wan 2.1 it was trained on 65.6% more images and 83.2% more videos, with aesthetic data labelled for lighting, composition, contrast and colour tone, which is where its cinematic look comes from [1]. ### The Wan 2.2 family "Wan 2.2" (also written wan2.2 or wan.2.2) covers several open models from the same release line [1]: - T2V-A14B: text to video at 480p or 720p, mixture of experts. - I2V-A14B: image to video at 480p or 720p, mixture of experts. - TI2V-5B: a smaller model for text and image to video at 720p and 24 fps [3]. - S2V-14B (August 2025): speech to video. - Animate-14B (September 2025): character animation and replacement, on Kenerate as Wan 2.2 Animate. ### How Kenerate runs Wan 2.2 In the studio, Wan 2.2 and Wan 2.2 Ultra Fast are served through WaveSpeed AI's hosted Wan 2.2 endpoints [5], one per mode and resolution. Both take text, or a start image plus an optional end frame, and return a silent 5 or 8 second clip at 480p or 720p, in 16:9 or 9:16 for text to video. Every clip on this page came from that same Wan 2.2 image-to-video or text-to-video endpoint at 480p, unedited. ### Wan 2.2 online or on your own GPU Locally, the TI2V-5B model runs on a consumer GPU such as an RTX 4090, with a 5-second 720p clip in under 9 minutes per the official repo [1], and ComfyUI's native templates fit it into 8 GB of VRAM with offloading [4]. The 14B models need at least 80 GB of VRAM on one GPU [1]. Run it locally if you want the open weights, custom workflows and no per-clip cost; run it online if you'd rather have a clip in minutes from a browser. References: [1] Wan 2.2 on GitHub (Wan-Video): https://github.com/Wan-Video/Wan2.2 [2] Wan2.2-I2V-A14B on Hugging Face: https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B [3] Wan2.2-TI2V-5B on Hugging Face: https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B [4] ComfyUI docs — Wan2.2 native workflows: https://docs.comfy.org/tutorials/video/wan/wan2_2 [5] WaveSpeed AI — Wan 2.2 image to video: https://wavespeed.ai/models/wavespeed-ai/wan-2.2/i2v-480p ## FAQ Q: Is Wan 2.2 free? A: The model is: Alibaba released Wan 2.2's code and weights under Apache 2.0, so you can download and run it on your own GPU at no cost. Running it online on Kenerate AI needs a paid plan, because video isn't part of the free plan. Wan 2.2 Ultra Fast is the studio's lowest-cost video model. Q: Can I use Wan 2.2 online without a GPU? A: Yes. In Kenerate's video studio, pick Wan 2.2, add a prompt (and a start image if you want image to video) and choose 5 or 8 seconds at 480p or 720p. It runs on hosted GPUs, so there's no ComfyUI, no model files and no VRAM limit to worry about. Q: How does Wan 2.2 image to video work? A: You give it a start image and a motion prompt, which is required, and Wan 2.2 animates the image for 5 or 8 seconds; the clip keeps the image's framing. You can add a second image as the end frame. Describe the action, camera move and pace rather than re-describing what's already in the picture. Q: Can Wan 2.2 do text to video? A: Yes. With no image, Wan 2.2 makes a clip from the prompt alone, in 16:9 or 9:16, at 480p or 720p, 5 or 8 seconds long. The paper-boat clip on this page was made that way at 480p. Q: What's the difference between Wan 2.2 and Wan 2.2 Ultra Fast? A: Both do text and image to video at 5 or 8 seconds, 480p or 720p, 16:9 or 9:16. Ultra Fast runs on WaveSpeed AI's accelerated Wan 2.2 endpoints, so it suits quick, low-cost drafts; switch to standard Wan 2.2 for the take you keep. Q: Does Wan 2.2 make sound? A: No. Wan 2.2's text and image-to-video models make silent clips, and so do both versions in the studio. Alibaba's separate Wan 2.2 S2V model does speech-to-video, but it isn't in the studio. For a clip with sound, use Wan 3. Q: What hardware does Wan 2.2 need locally? A: It depends on the model. The official repo says the TI2V-5B model makes 720p video at 24 fps on a consumer GPU such as an RTX 4090, and ComfyUI's docs say it fits 8 GB of VRAM with offloading. The 14B models need at least 80 GB of VRAM on a single GPU. Q: How is Wan 2.2 different from Wan 2.5, 2.7 and Wan 3? A: Wan 2.1 and 2.2 have open weights; Wan 2.5, 2.6, 2.7 and Wan 3 run only as hosted models. The newer ones add longer clips, 1080p and, on Wan 3, sound and up to 30 seconds. Wan 2.2 stays the quick, open option for 5–8 second clips. ## Related pages - Wan AI: https://kenerateai.com/wan-ai-video-generator - Wan 2.2 Animate: https://kenerateai.com/wan-2-2-animate - Wan 2.2 NSFW (18+): https://kenerateai.com/wan-2-2-nsfw - Wan Image: https://kenerateai.com/model/wan - Image to video AI: https://kenerateai.com/image-to-video-ai - AI video models: https://kenerateai.com/ai-video-models ## Sources - Kenerate video studio: https://kenerateai.com/app/video - Wan 2.2 on GitHub (Wan-Video): https://github.com/Wan-Video/Wan2.2 - Wan2.2-I2V-A14B on Hugging Face: https://huggingface.co/Wan-AI/Wan2.2-I2V-A14B - Wan2.2-TI2V-5B on Hugging Face: https://huggingface.co/Wan-AI/Wan2.2-TI2V-5B - ComfyUI docs — Wan2.2 native workflows: https://docs.comfy.org/tutorials/video/wan/wan2_2 - WaveSpeed AI — Wan 2.2 image to video: https://wavespeed.ai/models/wavespeed-ai/wan-2.2/i2v-480p