MiniMax H3 Prompts
MiniMax's multimodal video model — start from text, an image, a clip or a soundtrack, and get video with native stereo sound.
About MiniMax H3
MiniMax H3 is the flagship of MiniMax's Hailuo video line, and its defining trait is that it is genuinely multimodal on the way in and complete on the way out. You can start a generation from a text description, a reference image, an existing video clip, an audio track, or a combination of them — and the result arrives with native stereo sound generated alongside the picture, not dubbed on afterwards.
That combination makes it the model to reach for when the clip has to feel finished: ad spots where the ambience sells the product, story scenes where footsteps and room tone carry the mood, image-to-video shots where a still photograph needs to come alive with believable motion and sound. It inherits the physical realism the Hailuo line is known for, and follows camera direction — dolly, orbit, handheld — with unusual obedience.
If you have used Hailuo 02, H3 will feel familiar but less forgiving of vagueness. The older line rewards a short scene description and fills in the rest; H3 gives you six or seven independent levers — subject, action, setting, camera, light, sound, and the input modality itself — and hands each unpulled lever back to its own defaults. More control is only an advantage if you actually spend it.
The practical consequence is that a MiniMax H3 prompt is closer to a shot list than to a caption. A prompt that reads "a woman walking in the rain, cinematic" gets you a competent, anonymous clip. A prompt that names the lens, the direction of the walk, the camera's move relative to her, the quality of the light, and what the rain sounds like on what surface gets you a shot you could actually cut into something.
Because H3 accepts so many kinds of input, the prompt is what tells it which job to do — and vague prompts waste exactly the control that makes it strong. The gallery above collects tested MiniMax H3 prompts with their real generated videos, so you can copy a structure that already works and swap in your own subject, motion, and soundscape.
Everything on this page is free to copy. There is no sign-up, no paywall, and no watermark on the prompt text — take the structure, change the nouns, and run it wherever you have access to the model.
How to write MiniMax H3 prompts
- 1Structure the prompt as subject → action → setting → camera → light and mood → sound. H3 reads all six layers; leaving one out means the model decides it for you.
- 2Write it as a shot, not as a caption. "Cinematic" and "high quality" are the two least useful words you can spend — replace them with the specific thing you mean: the lens, the grade, the time of day, the weather.
- 3Use real camera vocabulary: "slow dolly-in", "handheld tracking shot", "orbit around the subject". H3 follows named camera moves far more precisely than vague phrases like "dynamic shot".
- 4For slow motion, name it as a property of the footage rather than of the subject: "filmed in slow motion, water droplets hanging in the air" works; "she moves slowly" just makes her walk slowly at normal speed. If you want a speed change mid-clip, say where it happens — "real time, then ramps into slow motion as the glass hits the floor".
- 5Direct the audio explicitly — sound is native, so it listens: "rain on a tin roof, distant thunder, no music" or "upbeat synth track builds under the voiceover". Unprompted audio defaults to generic ambience.
- 6Name the surface, not just the sound. "Footsteps" is generic; "boots on wet gravel" gives the model the material to synthesize. The same trick works for everything from "rain" to "engine".
- 7If you want no sound at all, say so — "no music, no dialogue, ambient room tone only". Silence is a direction H3 will follow, but only if you give it.
- 8When starting from a reference image, describe only the motion and the camera — the image already fixes subject, style, and framing, and re-describing them invites drift.
- 9Give each generation one continuous action beat. A clip that tries to cover three story moments blurs all of them; chain separate generations when the story needs more beats.
- 10Freeze your look in a reusable clause — film stock, color grade, lens character — and keep it identical across a campaign so every clip cuts together.
- 11To hold a character across shots, feed the same reference image into each generation rather than re-describing the person. Text descriptions of faces drift between runs; an image does not.
- 12When a result misses, change one layer and regenerate rather than rewriting the whole prompt. Rewriting everything makes it impossible to learn which clause was actually responsible.
Frequently asked questions
How do I write a good MiniMax H3 prompt?
Write it as a shot list, in this order: subject, action, setting, camera move, light and mood, sound. H3 reads each of those as a separate instruction, so every layer you leave blank is a layer the model fills in with its own default. Keep one continuous action per generation, use real camera vocabulary for the move, and name the audio explicitly. Every prompt in the gallery above follows that shape — the fastest way to learn it is to copy one and swap in your own subject.
Is MiniMax H3 free to use?
The prompts on this page are free — copy any of them with no account and no paywall. Access to the model itself runs through MiniMax's Hailuo AI app and the MiniMax open platform API, and their current plans and free-tier terms are set on MiniMax's side, so check there for what a run costs today.
What is MiniMax H3 best at?
Clips that need to arrive finished: video and stereo audio generated together, from whatever input you have — text, a still image, an existing clip, or a soundtrack. Its physical realism and precise camera-following make it a strong pick for ads, cinematic single takes, and image-to-video work.
Does MiniMax H3 generate sound?
Yes — stereo sound is generated natively with the picture, not added afterwards. You get the most out of it by directing the mix in the prompt: name the ambience, the effects, and whether you want music. If you say nothing, you get generic ambience.
How do I get slow motion out of MiniMax H3?
Describe slow motion as a property of the footage, not of the subject. "Filmed in slow motion" or "high-speed camera, slow-motion playback" changes how the clip is shot; "she moves slowly" only makes the subject sluggish at normal speed. Pair it with a detail that only reads in slow motion — spray, dust, fabric, hair — so the effect has something to show, and name the moment if you want the speed to change mid-shot.
Can I start from an image or an audio track instead of text?
Yes. H3 is multimodal on input: a reference image can define the subject and framing, a clip can be extended or restyled, and an audio track can drive the scene. When an image carries the look, keep the text prompt focused on motion and camera only.
How is MiniMax H3 different from Hailuo 02?
They are two generations of the same Hailuo video line. Hailuo 02 takes a short scene description and fills in most of the decisions for you. H3 is multimodal on input, generates its stereo audio natively with the picture, and exposes far more of those decisions as things you control — which means it rewards a detailed prompt and punishes a lazy one more than the older line did.
Why did MiniMax H3 ignore part of my prompt?
Usually because the prompt asked for more than one thing at once. H3 handles a single continuous action beat well; a prompt covering three story moments gets averaged into a blur. The other common cause is re-describing what a reference image already fixed — when the image says one thing and the text says another, the result drifts. Cut the prompt to one beat, change one layer at a time, and regenerate.
Where can I use these MiniMax H3 prompts?
H3 is available through MiniMax's Hailuo AI app, and developers can access it via the MiniMax open platform API. Copy any prompt from this page, paste it there, and swap in your own subject, camera move, or soundscape.
Related models
Hailuo 02
MiniMax's Hailuo video model with standout physical realism.
Speech 2.5
MiniMax's multilingual voice synthesis model.
Seedance 2.0
ByteDance's #1-ranked video model — multi-shot direction with native lip-sync and spatial audio.
Seedance 2.5
Seedance 2.5 is a next-gen AI video model that generates continuous, synchronized 30-second 4K clips from a single prompt.