Wan 2.5 Prompts

Alibaba's video model built on tight audio-visual sync — every visual beat lands with its sound, in one generation.

video

About Wan 2.5

Wan 2.5 is Alibaba's video generation model, and its defining discipline is audio-visual sync. Sound is not layered on after the fact — it is generated locked to the picture, so a door slams the frame it closes, a spoken line lands on the lip movement, and footsteps hit the ground exactly when the feet do.

That tightness is what makes Wan the practical choice for content where audio timing is the whole point: spoken-word clips, product demos with narration, musical moments, and any scene where a sound out of step by a beat would break the illusion. The picture side holds up too, but sync is the axis the model is built around.

Prompting Wan well therefore means writing sound and picture as one script, not two. A prompt that pairs each visual beat with its audio cue produces clips that feel edited on purpose; a prompt that only describes visuals leaves the soundtrack to chance. The gallery below collects tested Wan prompts with real outputs — study how the working ones interleave the two tracks.

How to write Wan 2.5 prompts

  1. 1Write paired beats: every visual event gets its sound in the same sentence — "the match flares with a sharp hiss", "her heels click across the marble lobby". Pairing is what the sync engine locks onto.
  2. 2Use "as" and "when" to weld sound to frame: "as the shutter rolls down, the street noise cuts to a muffled hum". Temporal connectives are sync instructions in disguise.
  3. 3Put dialogue and lyrics in quotes so the audio lands on the delivery — quoted lines give the model an exact utterance to sync lips and timing against.
  4. 4Layer the mix deliberately: name a foreground sound, then the ambient bed behind it — "close-mic'd pencil scratching, library silence underneath". Two named layers beat one vague atmosphere.
  5. 5Describe the character of each sound, not just its presence: "muffled through the wall", "echoing in the stairwell", "close and dry". Texture words steer the render of the audio itself.
  6. 6Give the clip one audio focal point. A single sharp sync moment — a slam, a beat drop, a spoken line — reads tighter than a soundscape trying to sync everything at once.

Frequently asked questions

What is Wan 2.5 best at?

Tight audio-visual sync — clips where sound and picture are generated locked together, so impacts, speech, and music land on the exact frame. It is the strong choice whenever audio timing is the point of the clip, from narrated demos to musical moments.

How do I write Wan 2.5 prompts with sound?

Script sound and picture together: pair every visual beat with its audio cue in the same sentence, use temporal words like "as" and "when" to weld them, and put any dialogue in quotes. The tested prompts on this page show how working prompts interleave the two tracks.

Does Wan 2.5 generate audio together with the video?

Yes — audio is generated in the same pass, synchronized to the picture rather than layered on afterwards. That is why prompts that specify what should be heard, and exactly when, outperform prompts that describe visuals alone.

Where can I try these Wan prompts?

Wan is available through Alibaba's Tongyi Wanxiang platform, and developers can access it via Alibaba Cloud Model Studio. Copy any prompt from this page, paste it there, and adapt the sound-picture pairings to your own scene.