by Google

Veo 3.1 Prompts

Google's cinematic text-to-video model — direct shots like a filmmaker, with sound generated in the same pass.

video
{ "duration": "10s", "aspect_ratio": "16:9", "fps": 30, "style": "photorealistic premium tech commercial, cinematic macro photography", "scene": "A 3-inch-tall adult male presenter in black streetwear with subtle neon-cyan accents interacts with gigantic premium black wireless headphones on a white marble studio desk.", "action": [ "0-3s: The miniature presenter confidently runs along the curved matte-black aluminum headband. Ultra-low macro tracking camera follows directly behind him.", "3-6s: He reaches an oversized knurled metal volume dial, grips it with both hands and forcefully rotates it. A subtle translucent sonic shockwave expands from the headphones.", "6-8s: He jumps from the headband onto the giant plush black leather memory-foam ear cushion, which visibly compresses on impact.", "8-10s: He stands on the cushion, smiles at the camera and gives a quick thumbs-up as the camera rapidly cranes upward and pulls back, revealing the full headphones, laptop and coffee cup on the desk." ], "camera": "Extreme macro shallow depth of field, smooth tracking, controlled whip-pan to the dial, dynamic follow-through on the jump, then rapid crane-up pullback into a wide hero shot.", "lighting": "Soft studio lighting with subtle cyan and electric-magenta rim lights, realistic metal, leather and marble reflections.", "audio": { "music": "NONE", "dialogue": "NONE", "voiceover": "NONE", "effects": "Realistic sneaker squeaks, mechanical dial clicks, deep rotation sound, subtle sonic whoosh and soft leather cushion impact." }, "quality": "8K-quality detail, realistic human anatomy, physically accurate materials and physics, stable character and product design.", "negative": "no music, no dialogue, no text, no logos, no watermark, no extra people, no floating, no teleportation, no clipping, no distorted hands, no changing scale, no character or clothing changes, no cartoon CGI, no excessive motion blur, no camera shake." }
First Frame Image Prompt A realistic, eye-level street view of a busy New York City intersection during the day. In the background, a massive billboard covers the side of a tall, classic brown brick building with visible windows. The billboard features a high-fashion, hyper-realistic image of a handsome man with short, tousled brown hair lying on his side against a backdrop of a light blue sky with fluffy white clouds. He is wearing a white button-down shirt with blue detailing and dark denim jeans. To his right on the billboard is a dark amber glass jar with a black lid and a white label. In the foreground, a bright yellow NYC taxi is driving past, and several pedestrians in casual clothing are walking across a crosswalk. There is a green street sign and scaffolding below the billboard. Natural sunlight, photorealistic style. Image-to-Video Prompt Theme: Faux out-of-home (FOOH) CGI advertisement where a billboard model comes to life and throws a product into the real world Visuals: Busy NYC street corner, massive billboard on a brick building showing a fashion model lying against a cloudy sky backdrop with a scented candle jar beside him; yellow taxi, pedestrians, green street sign, scaffolding, photorealistic urban daylight Camera: Static POV from a pedestrian on the street, slight rack focus shift at the end toward the caught candle in the foreground Style: Photorealistic, cinematic, hyper-realistic CGI blending 2D billboard illusion with 3D pop-out effect Action + Sound Design: The man on the billboard remains still for a brief moment as street traffic moves naturally in the foreground — ambient city sounds, car engines, distant chatter He seamlessly transitions from a static 2D image to a moving 3D figure, sitting up smoothly and crossing his legs left over right, revealing polished black leather boots, while smiling confidently toward the camera He reaches down with his left hand, picks up the dark amber candle jar beside him, and holds it up as if preparing to offer it He tosses the candle forward, out of the billboard's frame, creating a 3D pop-out illusion as the jar flies toward the viewer A real human hand reaches up from the bottom of the frame and catches the candle; the camera pulls focus to the jar in the foreground, revealing the white label reading "SCENTED CANDLE," "PRESS," "No 5," and "KEY NOTES: Ginger, Verveine, Vanilla," while the billboard blurs softly in the background Ambient street traffic continues naturally throughout the entire sequence.

About Veo 3.1

Veo 3.1 is Google DeepMind's flagship video generation model. It turns a written scene description into short cinematic clips, and unlike most video models it generates native audio — dialogue, ambient sound, and effects — synchronized with the footage in a single pass.

Its standout strength is direction control. Veo understands film language: shot types, lens choices, camera moves like dolly-ins and crane shots, lighting setups, and mood. Prompts that read like a shot list from a screenplay consistently outperform loose descriptions.

That is exactly why prompt quality matters more on Veo than on most models. The gallery below collects tested Veo prompts with their actual outputs, so you can copy a working structure instead of rediscovering it from scratch.

How to write Veo 3.1 prompts

  1. 1Structure prompts like a shot brief: subject → action → setting → camera → lighting → mood. Veo rewards film vocabulary over adjective piles.
  2. 2Name the camera work explicitly — "slow dolly-in", "handheld tracking shot", "aerial crane down". One camera instruction per clip keeps motion coherent.
  3. 3Direct the audio too: describe dialogue in quotes, then ambient sound ("rain on tin roof, distant thunder"). Unprompted audio tends toward generic ambience.
  4. 4Anchor the look with a lighting and grade reference: "golden hour backlight, shallow depth of field, filmic teal-orange grade".
  5. 5Keep one scene per prompt. Veo excels at a single continuous shot; cramming multiple cuts into one prompt fragments the motion.
  6. 6Iterate by changing one variable at a time — camera, light, or action — so you can tell which phrase actually moved the result.

Frequently asked questions

What is Veo 3.1 best at?

Cinematic single-shot clips with realistic motion and synchronized native audio. It is the strongest choice when you need footage that feels directed — deliberate camera moves, coherent lighting, and sound that matches the scene.

How do I write good Veo 3.1 prompts?

Write like a director, not a tagger. Describe one scene with explicit camera work, lighting, and audio cues in natural sentences. The tested prompts on this page are a faster starting point than writing from zero — copy one, then swap in your subject.

Does Veo 3.1 generate sound and dialogue?

Yes — native audio is Veo's signature feature. It generates speech, ambient noise, and sound effects synchronized with the video. Put spoken lines in quotation marks and describe the soundscape explicitly for best results.

Where can I try these Veo prompts?

Veo is available through Google's Gemini app and Flow, and via the Gemini API for developers. Copy any prompt from this page, paste it there, and adjust the subject or camera language to fit your scene.