AI Video Prompts
Cinematic prompts for AI video generation — Seedance, MiniMax H3, Kling, Hailuo, Runway and more.
Explore AI video prompts with real generated clips. Every prompt is tested, tagged by model and style, and ready to copy into your video generator.
All categories3D RenderAbstractAdvertisingAerialAnimalsAnimationAnimeArchitectureBlack & WhiteBrightCartoonCharacterCinematicCityscapeCyberpunkDarkDigital ArtDramaticDynamicEtherealFantasyFashionFoodGame AssetsGothicHistoricalHolidayHorrorLandscapeMinimalistMotion GraphicsMusic VideoMysteriousMythologyNatureNeonPeacefulPixel ArtPortraitRealisticSci-FiShort FilmSlow MotionSpaceSportsSteampunkSurrealTimelapseTrailerVehiclesVFXVintageWhimsical
30 seconds | 16:9 | live-action supernatural horror drama
Fictional dramatization. Documentary-style visual realism.
SETTING
An isolated old farmhouse surrounded by dense forest and fields.
Time: late afternoon transitioning into night.
Weather: overcast, cold and slightly windy.
Aged wooden exterior, neglected rooms, narrow staircase, old family photographs, dusty furniture, dim corridors and a locked basement door.
Natural environmental activity, distant wind, creaking wood and subtle household sounds.
CHARACTERS
A young couple in their early 30s and their 8-year-old daughter.
Maintain consistent faces, clothing, hairstyle, physical appearance and character identities throughout the sequence.
Human behaviour remains grounded and restrained.
STORY
00–04s — ARRIVAL
Wide observational shot establishing the isolated farmhouse.
The family carries boxes from their vehicle toward the front door.
The daughter stops and looks toward an upstairs window.
The parents continue walking, unaware.
Natural wind moves the trees and clothing.
04–08s — THE HOUSE
Handheld perspective follows the family through the dusty interior.
The father explores the living room while the mother opens curtains.
The daughter quietly looks toward the staircase.
A faint creaking sound comes from upstairs.
She asks, "Is someone up there?"
Her parents dismiss it as the old house settling.
08–12s — THE FIRST SIGN
The family moves through the hallway.
The daughter suddenly stops.
At the far end of the corridor, a bedroom door slowly moves open by itself.
The parents turn toward it.
The camera operator hesitates and slowly approaches.
The room appears completely empty.
Avoid exaggerated supernatural movement.
12–16s — NIGHT
Cut to the farmhouse after dark.
The family is preparing for bed.
The daughter suddenly looks toward the hallway.
A faint silhouette appears at the far end for only a moment before disappearing when the mother looks.
The daughter quietly says, "She's there again."
The parents exchange an uneasy glance.
16–20s — THE FOOTSTEPS
Handheld camera follows the father through the dark hallway.
Slow footsteps can be heard above him.
He looks toward the staircase.
The footsteps stop.
He cautiously climbs the stairs.
The camera follows behind with natural handheld movement.
A bedroom door at the end of the hallway is slightly open.
20–24s — THE LOCKED ROOM
The father discovers the daughter standing outside a locked basement door.
She is completely still.
She quietly says, "She told me not to open it."
The father looks at the door.
The handle slowly moves from the other side.
The camera operator instinctively steps backward.
24–28s — THE BASEMENT
The door suddenly opens slightly.
Cold darkness fills the doorway.
The mother grabs her daughter and pulls her backward.
A faint female voice comes from inside the basement.
The camera struggles to maintain focus in the darkness.
No visible monster.
No graphic violence.
28–30s — FINAL REVEAL
Camera remains in the hallway.
The family backs away from the basement.
The daughter suddenly looks past her parents toward the upstairs landing.
A pale female figure is standing silently at the top of the stairs.
Nobody notices her.
The camera slowly turns toward the figure.
The image abruptly cuts to black.
CAMERA
Grounded supernatural-documentary cinematography.
Mix:
wide environmental coverage
medium handheld following shots
brief close reaction shots
slow observational hallway perspectives
Natural handheld instability.
Realistic operator movement.
Plausible camera positions.
Autofocus hunting in darkness.
Small framing corrections.
Natural motion blur.
No impossible camera movement.
No floating camera.
No excessive slow motion.
No polished Hollywood horror cinematography.
LIGHTING
Lighting originates from practical and environmental sources.
Natural overcast daylight during arrival.
Warm household lamps and weak practical lighting after dark.
Natural exposure changes when moving between illuminated rooms and dark hallways.
Deep but detailed shadows.
Realistic highlight roll-off.
No artificial rim lighting.
No supernatural glow.
PHYSICS
Real gravity and momentum.
Feet maintain contact with floors and stairs.
Furniture has believable weight.
Doors move through physical hinges.
Curtains respond to airflow.
Dust reacts naturally to footsteps.
Clothing responds to movement and wind.
Wood, doors and floorboards behave realistically.
No floating objects.
No impossible movement.
No exaggerated supernatural physics.
HUMAN PERFORMANCE
Restrained, believable fear.
Quiet breathing.
Small eye movements.
Hesitation before entering dark spaces.
Natural glances toward unexplained sounds.
Parents instinctively protect their daughter.
The daughter behaves frightened but not theatrically possessed.
No screaming without cause.
No exaggerated facial expressions.
No heroic behavior.
No theatrical horror acting.
AUDIO
Natural farmhouse ambience.
Wind through trees.
Distant insects.
Wood creaking.
Footsteps on wooden floors.
Floorboards reacting under weight.
Quiet household sounds.
Distant birds during daylight.
Low nighttime ambience.
Soft breathing.
Subtle door movement.
Faint footsteps upstairs.
A barely audible female voice from the basement.
No music.
No jump-scare sound effects.
No oversized Hollywood trailer effects.
No artificial bass drops.
VISUAL CHARACTER
Photorealistic live-action footage.
Observational supernatural-documentary character.
Natural skin texture.
Realistic fabric and materials.
Authentic old farmhouse textures.
Subtle sensor noise.
Natural motion blur.
Imperfect handheld framing.
Practical lighting.
Atmospheric darkness.
The supernatural presence should feel physically present within an otherwise completely realistic environment.
Avoid:
video-game aesthetics
CGI appearance
excessive HDR
oversaturated colors
floating ghosts
Scene 1: Lavender is imprisoned in a dungeon cell, standing with his back against the wall. Heavy chains bind his arms raised high above his head, exposing both armpits. Lavender looks angry and defiant, yanking at the chains, glancing around furiously, and roaring in Japanese that they will regret this. Scene 2: Two mechanical arms slowly extend from the wall behind Lavender, each ending in a hand wearing white gloves. Lavender notices the movement behind him, looking confused and uneasy, and asks in Japanese: "What now?" Scene 3: The two white-gloved hands approach Lavender's exposed armpits. One hand reaches into one armpit, the other into the other armpit. They begin tickling Lavender's armpits with gentle, fluid, continuous finger movements. Comical tickling sound effects play. Lavender immediately breaks into a grin and lets out muffled laughter, because he is extremely ticklish. (Comical tickling sound effects) Scene 4: Close-up shot of Lavender's upper body. The two gloved hands continue tickling his exposed armpits. Lavender squeezes his eyes shut, grinning widely and letting out muffled laughter, twisting his body slightly to escape the gloved hands, but he cannot get away, and the gloved hands keep tickling Lavender's armpits. He says in Japanese in a giggling tone: "Stop!" then lets out soft laughter again. (Comical tickling sound effects continue) Scene 5: After the two mechanical arms continue tickling his armpits for a while, they suddenly stop and pull back slightly, but still remain close to Lavender. Lavender giggles a few times, then regains his composure a little, hoping the tickling is over, and sighs, saying in Japanese with relief: "Finally stopped." Scene 6: The mechanical arms' hands move down from Lavender's armpits to his stomach, and Lavender watches them move with a hint of worry. The two white-gloved hands begin tickling Lavender's stomach with gentle, fluid, continuous finger movements. Lavender starts giggling softly again, saying in his distinctive Japanese laughing voice: "Oh, don't tickle there!" then starts giggling again, the laughter growing louder. (Comical tickling sound effects continue) Scene 6: Full frontal view of Lavender in the cell. The two white-gloved hands continue tickling his stomach, and he giggles, twisting his upper body slightly, but unable to escape. He grins with his eyes closed, struggling weakly against the chains, while the gloved hands continue tickling his stomach. There are two mechanical arms in total, and exactly two white-gloved hands. They extend from behind Lavender. First, the two hands tickle Lavender's exposed armpits, then move down to tickle his stomach. When tickling the armpits, focus on the hollow of the armpits, and do not tickle the ribs, chest, or sides. The finger movements should be gentle, fluid, and continuous. Use comical, funny tickling sound effects in all tickling scenes. The voice is in Japanese.
Create a 20-second cinematic vlog featuring a cheerful young Japanese woman exploring a hidden little town in Japan.
Character & Outfit: A cheerful young Japanese woman wearing a cute pastel short summer dress with a delicate floral pattern, white ankle socks, and adorable sneakers. Soft hairstyle, natural makeup, bright expressive eyes, warm smile, youthful and wholesome summer-vlog aesthetic.
0–5s — Discovering Japan: She walks through a beautiful hidden Japanese alley decorated with colorful lanterns, flowers, traditional wooden houses, and tiny cafés. She turns the camera toward herself, smiles excitedly and says: “Hi guys! I found the cutest little place in Japan!”
5–9s — The Cat Appears: She sits outside a cozy flower-filled café, holding a cute drink. Suddenly, a tiny fluffy cat slowly walks toward her and stops in front of her. She looks surprised and smiles: “Oh… hello, little one. Did you come to see me?” The cat gently rubs against her leg. She giggles and kneels down.
9–13s — First Bonding Moment: She offers her hand. The cat curiously sniffs it, then gently places one tiny paw on her hand. She smiles warmly and says: “Aww… I think we’re friends now.” She gently pets its head while the cat happily stays beside her.
13–17s — Cute Friendship: The cat curls up next to her on the bench. She gently strokes its fluffy fur and laughs softly: “You’re way too cute!” The cat looks up at her and gives a tiny playful meow. She looks at the camera and whispers: “Okay… you’re coming with me.”
17–20s — Sweet Ending: Golden sunset beside a peaceful Japanese river. She sits with the cat beside her, gently pets its head, looks into the camera with a warm smile and says: “Best day ever. See you tomorrow!” She giggles and waves goodbye while the cat looks toward the camera.
Visual & Audio Style: Ultra-realistic cinematic vlog, natural English dialogue, realistic Japanese female voice, accurate lip-sync, adorable facial expressions, playful reactions, realistic fluffy cat fur and movements, authentic Japanese scenery, lanterns, flowers, cozy café, peaceful river, warm golden-hour lighting, cinematic depth of field, smooth handheld vlog camera, gentle close-ups, natural transitions, soft ambient café and street sounds, wholesome kawaii atmosphere, photorealistic 4K.
Ultra cinematic wide shot: A vast field of burning embers stretches into darkness. Suddenly, an enormous wall of fire races across the ground, sending billions of glowing sparks into the air. The flames twist into a gigantic rotating fire vortex hundreds of feet high. Molten embers, smoke, and liquid fire spiral together, naturally forging the word "IGNITE" from blazing flames. The letters burn brighter and brighter before erupting into a colossal shockwave that fills the screen with fire and incandescent particles. Hollywood fire simulation, ultra-realistic VFX, blockbuster title reveal.
Cinematic shot, ARRI Alexa 65, 20mm wide-angle lens, symmetrical composition. The camera slowly pushes in and slowly orbits about 10 degrees.
A grand wide-angle architectural photograph, showcasing a massive white celestial marble staircase, high above the clouds, leading to heaven.
Perfect central symmetry, extreme vertical one-point perspective, steep converging lines from bottom to top.
Foreground: wide white stone steps, low-angle side light casting rhythmic diagonal shadows;
2 Chinese immortal female figures seen from behind, the one on the left wearing a large crimson robe, the one on the right wearing a white robe, they walk slowly in the distance, appearing tiny against the magnificent architecture. The figures occupy 5% of the frame height, appearing as extremely small specks of color, walking slowly forward. Their bodies are still, their garments and long hair catching the high-altitude breeze from the same direction, flowing backward. There is an empty distance dozens of times their body length around them. Shown from behind, with a calm and elegant bearing. They ascend from the very bottom.
Midground: carved marble railings on both sides of the upper-middle section of the stairs, adorned with huge auspicious cloud sculptures; in the distant left of the sea of clouds floats a small Chinese temple pavilion.
Background: at the top of the endless staircase, a magnificent and enormous golden-roofed celestial palace stands atop thick clouds.
A vast ocean of white clouds, crisp azure sky, sacred atmosphere, bright morning light, grand scale contrast,
the sea of clouds churns rapidly, drifting around the pathways and architecture, evoking a serene and solemn mood, like a Chinese fairyland;
epic Eastern aesthetics, solemn and ethereal, with extreme proportional contrast.
A hyper-realistic handheld phone video at a normal city crosswalk during the day. People wait for the light, cars pass, and the ambience is completely ordinary and mundane. In the middle of the scene, a low distant whale call is heard from above. Several people, including the camera operator, look up. High in the sky, partially obscured by clouds, a colossal blue whale slowly swims overhead as though the atmosphere were an ocean. It is enormous, larger than several city blocks, and moves gracefully through the clouds with realistic weight and motion. The surrounding buildings make its impossible scale instantly clear. More of its body becomes visible as it glides across the sky, with fins, tail, and subtle cloud interaction. The pedestrians below stand frozen, looking up in disbelief. The scene should feel like an ordinary urban phone recording interrupted by one vast, surreal, awe-inspiring event that is presented as if it were really happening.
A sparkling, super-cute gal switches between adorable faces and hand poses while lip-syncing to the rhythm of the song. Make it a 15-second MV hook with that rowdy gal energy. Add flashy effects that match the music and action.
integrated_multimodal_description: [Shot 1] Handheld cellphone footage from a passenger seat on a commercial airplane, filmed through the window during daylight above a thick cloud layer. Mild turbulence, realistic cabin ambience, totally normal amateur footage at first: the wing, clouds, and slight shaking from the aircraft. From 00:00.000 to 00:03.500, nothing unusual happens. At 00:03.500 the cloud deck far below begins to bulge and part in a disturbing way. At 00:05.000 a colossal dark tentacle, thicker than a skyscraper, slowly rises out of the clouds and reaches upward toward the plane. The filmer’s voice cracks, [English, terrified] Oh my god— what is that?!. At 00:07.500 the tentacle suddenly accelerates toward the aircraft. At 00:09.000 it smashes violently through the fuselage and into the cabin, causing a burst of debris, wind, screaming passengers, and structural tearing. But instead of attacking further, the end of the gigantic tentacle is revealed to be holding an enormous paper takeout bag. From 00:10.000 to 00:14.500 the camera shakes wildly amid the damaged cabin as the tentacle awkwardly extends the bag deeper into the aisle like a delivery driver completing a dropoff. A calm, matter-of-fact voice from somewhere in the clouds says, [English, calm] DoorDash for Kevin?. The contrast should be completely deadpan and unexpected, with no comedic music or visual cue beforehand.
overall_soundscape: Realistic airplane cabin noise, muffled engine hum, mild turbulence, rising passenger panic, violent impact, metal tearing, rushing wind, screams, rattling debris, frantic breathing from the filmer, followed by the absurdly calm delivery line from outside.
non_diegetic_music: none
Create a 30-second ultra-realistic video that feels like a genuine forgotten home-video recording from the early 2000s. It must NOT look cinematic, staged, polished, or AI-generated.
MAIN SUBJECT:
Young Korean woman, early 20s, natural everyday appearance. Faded charcoal-grey sleeveless crop top, loose high-waisted light-wash jeans, black canvas sneakers, black cord necklace. Black wavy hair tied in a messy side ponytail with wispy bangs. Minimal makeup, realistic pores and skin texture, subtle imperfections. Warm, approachable personality. Maintain EXACTLY the same identity, face, hairstyle, clothing, proportions, and appearance throughout.
LOCATION:
Authentic quiet Korean residential neighborhood on a warm late afternoon. Narrow concrete lanes, low-rise houses, small gates, potted plants, old bicycles, utility poles, tangled overhead wires, parked scooters, laundry hanging from balconies, mature trees, uneven concrete, faded walls. No shops, signs, advertisements, cafés, crowds, or tourist landmarks.
CAMERA:
Early-2000s consumer MiniDV camcorder recorded casually by her friend. Heavy natural handheld shake, imperfect framing, occasional accidental cropping, autofocus hunting, lens breathing, exposure pumping when entering shade, slight rolling shutter, motion blur during quick movement, mild digital compression, subtle sensor noise, faded consumer-camera colors, soft contrast. No stabilization. No cinematic camera movement. No gimbal. No modern HDR. No artificial film look.
00:00–00:04
The recording begins abruptly while the camera is already moving. She is walking out through a small metal gate carrying a reusable grocery bag. She notices the camera and gives a slightly confused smile, as if she didn't realize her friend had started recording. The operator walks backward awkwardly and briefly loses her face from frame.
00:04–00:08
They turn into a narrow residential alley. She suddenly notices a small delivery box left near a neighbor's gate. She picks it up and looks around, trying to figure out whose it is. The camera moves closer too quickly, autofocus briefly locks onto the box instead of her face.
00:08–00:12
A middle-aged Korean neighbor appears from a doorway and gestures that the package is hers. The woman laughs softly, hands it over, and gives a small polite bow. The interaction feels spontaneous and unscripted. The camera shakes slightly as the person filming laughs behind the camera.
00:12–00:16
She continues walking and stops beside an old vending-style outdoor water dispenser near a residential entrance. She takes a small sip from her reusable bottle, wipes her mouth with the back of her hand, then notices the camera still pointed at her and playfully covers the lens for a moment.
00:16–00:21
The camera drops slightly as her hand moves away, revealing her walking ahead beneath large trees. Sunlight flickers naturally across the pavement and her clothing as leaves move in the breeze. A bicycle passes slowly in the distant background. She casually kicks a small fallen leaf forward while walking.
00:21–00:25
She reaches a tiny neighborhood convenience-style residential rest area with a concrete bench, but there are NO commercial signs or branding. She sits down, places the grocery bag beside her, and checks her phone. She suddenly looks toward something off-camera and smiles naturally, reacting to her friend rather than posing.
00:25–00:28
The friend walks closer. She looks directly into the lens and says casually, "Why are you still filming me?" with a small laugh. Her expression should feel completely spontaneous, not performed.
00:28–00:30
She stands and starts walking toward the camera. The operator instinctively backs away, causing noticeable shaky footage and brief focus hunting. She reaches toward the camera as if trying to stop the recording.
The image abruptly cuts to black while her hand is still approaching the lens.
AUDIO:
ONLY authentic location sound. No music. No cinematic sound design. No narration.
Include distant birds, leaves moving, footsteps on concrete, fabric rustling, a faint scooter passing, bicycle wheel sounds, neighborhood voices far away, subtle traffic in the distance, the neighbor's brief greeting, her natural laughter, and the actual sound of the camcorder operator moving and breathing. Dialogue should sound naturally recorded through a cheap built-in microphone, slightly compressed and imperfect.
REALISM REQUIREMENTS:
The entire video must feel accidentally captured by a real person, not directed for a film. Avoid beautiful composition, dramatic lighting, perfect framing, slow motion, smooth transitions, excessive depth of field, exaggerated facial expressions, perfect skin, artificial bokeh, cinematic color grading, CGI-looking environments, or overly clean textures.
Small imperfections are essential: missed focus, exposure fluctuations, awkward framing, slight camera shake, realistic motion blur, occasional clipped highlights, compression artifacts, imperfect timing, natural pauses, and genuine micro-expressions.
Create a 15-second ultra-realistic cinematic pizza-making sequence featuring the girl from the uploaded reference image as the exact character reference. Preserve her facial identity, green-hazel eyes, long jet-black hair, skin tone, facial structure, and overall appearance consistently throughout every shot. Dress her in a stylish black apron over her white fitted T-shirt, with her hair tied into a practical low ponytail.
The video opens with an intense close-up of her holding a freshly baked pizza slice toward the camera. She takes a confident bite as molten cheese stretches dramatically from the slice, steam rising around her face.
She immediately turns toward the counter and begins making another pizza. She spins a round piece of dough high above her head, catches it perfectly, and slaps it onto the counter as flour bursts into the air. The camera follows the movement with a fast whip-pan.
She rapidly spreads bright-red tomato sauce across the dough, then throws fresh mozzarella, ricotta, and pepperoni onto the pizza with energetic precision. Use fast macro cuts of sauce swirling, cheese landing, and pepperoni bouncing onto the dough.
She slides the topped pizza into a blazing stone oven. Extreme macro shots show the crust rapidly puffing into huge leopard-charred bubbles, pepperoni curling into crispy cups, cheese bubbling and melting, and heat distortion shimmering around the pizza.
She pulls the finished pizza from the oven, scatters fresh basil over the bubbling surface, then lifts one slice high. A huge molten cheese pull stretches between the slice and pizza as she gives the camera a confident satisfied look.
Finish with a dramatic close-up of the golden blistered crust, melted cheese, crispy pepperoni, fresh basil, and steam filling the frame.
Style: Premium cinematic food commercial, female chef protagonist, aggressive handheld camera, whip pans, smooth tracking, extreme macro food photography, dramatic blue-and-warm-amber lighting, realistic dough and cheese physics, natural steam, oven heat distortion, detailed skin and hair, shallow depth of field, cinematic film grain, photorealistic, ultra-detailed, 4K HDR, 24fps, 16:9.
Audio: Natural diegetic sounds only — dough slapping, flour scattering, sauce spreading, toppings landing, oven roar, bubbling cheese, crust crackling, basil flutter, and final bite crunch. No music, dialogue, subtitles, logos, text, or watermark.
Create a 15-second ultra-photorealistic live-action medieval war sequence set during the 12th-century Crusades, depicting a fictionalized battlefield confrontation between Crusader forces and the army of Salahuddin (Saladin). The entire scene must feel grounded, documentary-like, raw, and physically realistic, as if authentic historical footage had somehow been captured with a primitive period camera.
Environment: A medieval settlement and battlefield in the Levant during the Crusades, with weathered stone buildings, fortified walls, dirt roads, wooden structures, tents, farmland, horses, carts, siege equipment, scattered shields and weapons, drifting smoke, dust in the air, damaged structures, and a tense wartime atmosphere. Harsh overcast afternoon light mixed with dusty sunlight, natural atmospheric haze, and realistic environmental wear.
Characters: Crusader soldiers wearing historically inspired 12th-century mail armor, surcoats, helmets, shields, boots, and period-accurate equipment. Salahuddin's soldiers wear historically inspired Ayyubid-era Middle Eastern military clothing, chainmail, helmets, shields, and period-accurate equipment. Natural faces, realistic skin texture, sweat, dirt, fatigue, fear, and believable body movements. Horses and animals must move according to realistic biomechanics.
0–3s — Establishing Shot:
Wide handheld shot of a dusty medieval settlement suddenly filled with smoke and confusion. Crusader soldiers and Salahuddin's forces move rapidly between stone buildings and defensive positions. Horses pull wooden carts while civilians rush toward safer areas. Dust and smoke drift naturally through the battlefield.
3–6s — Tension:
Camera moves through the battlefield at shoulder height, following several Crusader soldiers as distant battle cries, horns, arrows, and metal impacts are heard. They immediately react and take cover behind a stone wall and overturned wooden cart. Across the battlefield, Salahuddin's soldiers advance cautiously through the dust.
6–10s — Combat:
Fast handheld tracking shot as Crusader and Ayyubid soldiers clash between cover. Shields absorb impacts, swords collide with realistic weight and momentum, arrows strike wooden structures and shields, and small pieces of stone, wood, dust, and debris fall naturally from nearby impacts. Weapon movement, armor weight, recoil from physical impacts, momentum, balance, and body weight must be physically accurate. Keep the violence realistic and restrained.
10–13s — Human Moment:
Camera briefly focuses on a wounded soldier being helped behind cover by another soldier. Their breathing, facial expressions, fear, exhaustion, body language, armor weight, and movement should feel natural and unscripted. In the background, the battle continues through smoke and dust.
13–15s — Final Shot:
Camera pulls back into a wide shot of the medieval battlefield as smoke slowly moves through the settlement. Crusader forces remain behind defensive positions while Salahuddin's army advances in the distant background. Horses, damaged stone buildings, shields, banners, and scattered battlefield equipment fill the frame. The scene ends with an authentic, tense historical-documentary feeling.
Visual Style: Ultra-photorealistic live-action, historically grounded 12th-century Levantine environment, vintage 35mm film texture, subtle film grain, natural imperfections, realistic exposure, handheld documentary cinematography, muted historical color palette, realistic smoke and dust, natural shadows, accurate depth of field, physically convincing medieval materials and armor.
Physics: Strictly obey real-world gravity, momentum, inertia, friction, weight, collision physics, horse biomechanics, armor movement, shield impacts, sword momentum, arrow trajectories, and human biomechanics. No exaggerated explosions, impossible movements, superhero behavior, or choreographed-looking combat.
Negative Prompt: modern buildings, modern vehicles, firearms, smartphones, modern clothing, modern weapons, futuristic technology, fantasy armor, fantasy creatures, CGI appearance, video-game graphics, superhero action, excessive explosions, excessive blood, gore, impossible physics, unrealistic sword movement, weightless armor, floating weapons, horses with incorrect anatomy, distorted faces, extra limbs, floating objects, plastic skin, artificial-looking environments.
CORE MESSAGE: "What if today, you could use the break to find your focus?" — a KitKat break creates a moment to pause, reset, and return to what matters with renewed focus. Feel like a premium global chocolate ad with emotional family storytelling — NOT a generic product commercial, NOT a fitness ad, NOT a children's cartoon. Warm, aspirational, elegant, believable.
SCENE 1 — MORNING BEGINS (0:00–0:03)
Luxury family home, soft morning light. Camera moves through the elegant kitchen (marble island, warm wood, pendant lighting, garden view). Mother (Ref 1) prepares breakfast in elegant cream/beige homewear; son (Ref 3) enters sleepily; father (Ref 2) preps coffee in background. Red KitKat (Ref 5) sits naturally on the island; camera pushes toward it — integrated, not a catalog shot.
VO (warm female): "Every day brings a new challenge…"
SCENE 2 — WAKE UP, CHAMPION (0:03–0:06)
Mother wakes son, opens curtains. MOTHER: "Good morning, champion." BOY: "Five more minutes…" Natural, warm. Mother stays in casual homewear — no office clothes.
SCENE 3 — BREAKFAST / KITKAT INTRO (0:06–0:09)
Family at the island. Son notices the KitKat, picks it up, looks at it (doesn't eat yet — introduction moment), places it by his school bag.
VO: "Sometimes, all you need is a moment…"
SCENE 4 — READY FOR SCHOOL (0:09–0:12)
Son dressed in navy blazer/uniform. Mother straightens his collar; father nods. KitKat visible in his open bag. He heads out the front door into daylight.
VO: "...to reset."
SCENE 5 — SCHOOL / THE CHALLENGE (0:12–0:16)
Realistic classroom. Son looks focused but slightly overwhelmed at his desk, tries to work through a problem, briefly distracted. Natural classroom ambience, no music yet.
SCENE 6 — THE KITKAT BREAK (0:16–0:20) (key product moment)
Bell rings. Son takes a breath, pulls the exact KitKat (Ref 5) from his bag. Close-ups: opens wrapper, snaps off a finger, bites, pauses — visibly relaxes and refocuses.
VO (slow, deliberate): "What if today, you could use the break to find your focus?"
Music lifts gently. Only one bite — no repeat consumption.
SCENE 7 — BACK TO FOCUS (0:20–0:22)
Match cut to him at his desk, calm and writing confidently; teacher notices.
VO: "Pause. Reset. Go again."
SCENE 8 — FOOTBALL FIELD (0:22–0:26)
Match cut to his foot kicking a ball. Realistic child-level play — receives, passes, shoots, scores; teammates celebrate. School bag with KitKat visible on the sideline (no second KitKat eaten — visual continuity only).
VO: "And get back to what matters." Music turns energetic; football ambience.
SCENE 9 — FATHER AND SON (0:26–0:28)
Home, cream sofa. KitKat nearby on coffee table. FATHER: "Proud of you, kid." BOY: "Thanks, Dad." Warm, intimate two-shot.
SCENE 10 — FINAL KITKAT HERO SHOT (0:28–0:30)
Premium product shot on the marble island: unopened KitKat, a few wafer fingers, chocolate crumbs, warm natural light, home softly visible behind.
VO: "Have a break, have a KitKat." ON-SCREEN TEXT: same line. End clean on product.
Create a 30-second authentic 1980s Japanese game-show broadcast sequence inspired by classic Japanese obstacle-course television. Realistic live-action footage, broadcast-camera aesthetics, natural handheld movement, analog film grain, slightly washed colors, chaotic energy, practical stunt physics, and no modern visual artifacts. No reference images provided; generate everything from this description.
0–5s — Wide Establishing Shot:
A young Japanese man in his late 20s, wearing a red tracksuit and white safety helmet, sprints onto a massive horizontal rotating log above a deep muddy pit. The log spins continuously beneath him while several giant padded wrecking balls swing across the path. The camera tracks him with imperfect handheld broadcast movement.
5–10s — Medium Action Shot:
The first wrecking ball swings toward him. He ducks desperately, rapidly shuffling his feet to maintain balance as the rotating log accelerates. His arms flail naturally, and muddy water splashes below where other contestants have fallen.
10–15s — Close Action Shot:
A second padded wrecking ball strikes his shoulder, sending him spinning off balance. He drops onto his knees and desperately hugs the rotating log as his body slides around its wet wooden surface. Strong motion blur and realistic momentum.
15–20s — Dynamic Tracking Shot:
He ends up hanging underneath the rotating log, gripping it with both hands while his legs kick through the air. He struggles to crawl toward the opposite platform as the log continuously rotates. His grip gradually weakens and his fingers slip.
20–25s — Major Impact:
He finally loses his grip and falls into the deep muddy pit. He hits the water with a huge, physically realistic explosion of brown mud and water. The camera reacts with chaotic handheld movement as the splash briefly fills the frame.
25–30s — Comedic Ending Shot:
He resurfaces completely covered in mud, helmet crooked and face stunned. He slowly raises one defeated fist toward the camera while the rotating log continues spinning behind him and a wrecking ball swings overhead. Maintain the authentic chaotic game-show atmosphere.
Visual requirements: High-fidelity realistic live action, accurate rotating-log physics, believable human movement and impacts, realistic mud and water simulation, practical-stunt appearance, coherent character identity throughout, 1980s Japanese television broadcast grain, subtle analog noise, imperfect exposure, natural motion blur, period-accurate camera movement, no CGI-looking surfaces, no modern cameras, no futuristic elements, no text overlays, no subtitles.
Ultra-realistic summer café scene with subtle frozen-time VFX. One continuous 20-second shot. Natural handheld camera with a smooth backward track and one gentle side move. Bright daylight, warm terrace atmosphere, realistic shadows, soft film grain, no dramatic action styling.
One adult man seated at an outdoor café table, wearing casual summer clothes. One adult waitress carrying a glass of beer on a tray. A few adult customers remain in the background.
Sunny café terrace with small tables, umbrellas, chairs and a quiet city street.
Photorealistic pause-time effect, simple body motion, stable anatomy, realistic glass handling, coherent background extras, natural summer atmosphere, seamless camera continuity, no fall, no injury, no direct body repositioning.
0-5s: Medium-wide tracking shot - The man waits calmly at his table. The waitress approaches with a cold beer on a tray.
5-8s: Gentle push-in - Her foot catches slightly on the edge of a chair. She loses balance for a moment. The beer glass begins sliding off the tray.
8-11s: Close-up - The man snaps his fingers. Time pauses instantly. The waitress becomes motionless in a tilted but stable pose. The beer glass and a few droplets remain suspended close to the tray.
11-16s: Slow side move - The man stands, calmly takes the suspended beer glass, places it on his table, then moves the chair slightly out of the waitress's path. He does not touch her.
16-20s: Push-in to medium close-up - He sits back down and snaps again. Time resumes. The waitress regains her footing naturally and looks confused when she sees the beer already on the table. The man smiles and says in English: "Perfect timing."
15 second vertical comedy clip, 9:16, photorealistic, shot like a real phone video that turns into a fast fashion montage. Cut driven, handheld, no music.
MAYA. 27, olive skin, long chestnut brown wavy hair with sun lightened strands, dark brown eyes, gold hoops, layered gold necklaces, thin bracelets. Wears a fitted ivory sleeveless square neck top, high waisted colorful print resort trousers, thin tan belt, barefoot. Very expressive face.
LENA. 28, tall, golden skin, long dark espresso brown glossy waves, hazel brown eyes, gold earrings, calm and confident. Same face, same hair, same body, same age in every shot. Only her clothes change.
SETTING: a white motor yacht anchored in a calm turquoise Mediterranean bay. White sailboats in the distance, low rocky green coastline on the same side the whole time, cloudless blue sky, bright late morning sun from the same direction in every shot. White deck, steel railings, cream cushions, coiled ropes. Light sea breeze moves hair and fabric.
STYLE: natural sunlight, real skin texture with pores, correct hands and fingers, real fabric movement, clear turquoise water, no beauty filter, no plastic skin, no heavy HDR.
CAMERA: handheld with small natural shake. Fast whip pans and hard match cuts between outfits. Normal lens, no slow motion, no floating drone moves, no heavy blur.
0.0 to 1.6s Maya stands on the deck, medium close up, sea behind her. She stares off to the right, eyebrows up, mouth open, and leans forward. Maya says, flat and shocked: {No. Absolutely not.} Do not show what she is looking at.
1.6 to 3.0s Camera pushes in. Maya covers her mouth, then drops her hand and points hard to the right. Maya says: {You said ONE outfit!} On the word "outfit" the camera whip pans right along her arm.
3.0 to 4.6s The pan lands on Lena at the bow, wearing an oversized open white linen shirt over a black one piece swimsuit, black sunglasses, flat sandals, big woven tote. Beside her sit one large suitcase, a garment bag and two tote bags. Medium full shot, railing leading toward her. Lena says calmly: {It is one outfit.} Small pause, she lifts one finger: {For now.}
4.6 to 6.0s Hard match cut, same body position. Lena now wears a cobalt blue swimsuit with a white sheer sarong tied at her waist, hair damp at the temples, barefoot. She takes two steps toward camera and points at the water. Lena says: {Swim.}
6.0 to 7.4s A white towel sweeps past the lens, and behind it Lena now sits at a small deck table with fruit and sparkling water, wearing a coral orange halter dress, sunglasses on her head, woven handbag. She lifts the sunglasses slightly and smiles. Lena says: {Lunch.}
7.4 to 8.8s Hard match cut. Lena stands at the railing in an emerald green satin slip dress with gold statement earrings, same daylight. Wind lifts her hair, she turns her head toward Maya. Lena says: {Sunset.}
8.8 to 10.3s Fast whip pan across the cabin. Lena walks past camera in an ivory cropped blazer, matching wide leg trousers, champagne silk camisole, gold heels, small metallic clutch. Camera tracks back one second. Lena says: {Dinner.}
10.3 to 11.8s Hard match cut. Medium close up of Lena in a silver metallic evening mini dress with crystal earrings. She adjusts one earring, completely serious. Lena says: {Emergency.}
11.8 to 13.2s Smash cut to a tight close up of Maya, honestly confused, looking toward Lena. Maya asks: {What emergency?}
13.2 to 15.0s Cut back to Lena, medium shot, silver dress catching the sun, turquoise water, sailboats and green coast behind her. She slides her sunglasses on, gives a tiny smile to camera. Lena says: {Bad photos.} Maya laughs loudly off screen. Lena walks out of frame. Hard end, no fade.
AUDIO: only real sound, meaning ocean, wind, boat movement, footsteps, fabric, bracelets, glass, laughter. Dialogue language: English, natural conversational delivery, never theatrical, never both women talking at once. No music at any point. No narration. No subtitles.
Last frame: empty yacht deck at the bow, turquoise bay, sailboats and green coastline behind, bright sunlight, gentle water movement. No on screen text, no captions, no logo, no watermark.
Negative prompt: music, score, subtitles, on screen text, watermarks, logos, brand names, extra people, two Lenas, two Mayas, changing face, changing hair, changing age, mixed or merged outfits, clothes morphing on the body without a cut, magical transformation, slow motion, drone shots, warped hands, extra fingers, plastic skin, beauty filter, oversaturated colors, moving coastline, changing sun direction, horizontal framing.
Use exactly 2 uploaded image assets.
image1 = The adult female lead is the sole and highest-priority character identity reference. Strictly maintain her face, facial feature proportions, skin tone, hairstyle, body type, sense of age, overall temperament, clothing, and accessories.
image2 = A Korean adult male wearing glasses is the sole and highest-priority character identity reference. Strictly maintain his face, facial feature proportions, skin tone, hairstyle, glasses, body type, sense of age, overall temperament, clothing, and accessories.
Generate a 30-second, 16:9 landscape, 4K, 24fps, hyper-realistic live-action film-grade action comedy short. The scene is a high-rise modern top-floor office / luxury apartment-style space at night, with floor-to-ceiling windows overlooking the city nightscape. The interior includes an open living room, a long hallway, modern furniture, and mixed warm-cool lighting. No anime feel, no game CG feel, no cheap VFX, no slow motion, no pseudo slow motion. Action must be realistic, fast, clear, and cinematic.
【Premise】
This is a pre-choreographed, mutually consensual action comedy performance between two adult actors. The overall tone is light, playful, and competitive — no fear, threat, or coercion. It is not a life-or-death fight, not a vulgar skit, but a high-energy action comedy where "the female lead keeps attacking, and the male lead keeps dodging while disrupting her rhythm with playful cheek kisses."
【Core Dynamic】
The female lead is on the offensive the entire time. The male lead almost never truly fights back — he only dodges, repositions, moves in close, and backs away. His main "scoring method" is not punches or kicks, but using momentary gaps in the action rhythm to suddenly step in, deliver a quick, clear, playful cheek kiss, then immediately pull back. There are 5 cheek kisses total. Each one must be a clear comedic beat. After each kiss, the female lead becomes more annoyed and more serious, continuing the chase.
【Action Principles】
Female lead: Continuously presses forward, using jabs, hooks, elbow strikes, knee strikes, side kicks, high kicks, spinning kicks, turning attacks, and relentless pursuit. The further it goes, the more serious and aggressive she becomes.
Male lead: Almost never throws a punch, never delivers real kick or punch counters, no heavy throws or slams. His main movements are only: dodging, retreating, sidestepping, ducking, leaning back, gliding, repositioning, redirecting attack paths, suddenly stepping in, cheek kissing, and immediately backing off.
The structure must keep escalating repeatedly:
Female lead attacks continuously → male lead dodges continuously → male lead seizes a tiny gap and suddenly delivers a cheek kiss → female lead gets angrier → she attacks again even more fiercely.
This structure repeats 5 times.
【Camera Rules】
Action segments primarily use medium shots, medium-wide shots, wide shots, close follow cam, handheld feel, lateral tracking, and slight breathing-like camera shake, clearly conveying spatial movement.
Before each cheek kiss, the camera first follows the rhythm of both performers' movements, capturing the male lead suddenly entering the female lead's close personal space.
When each kiss happens, do not cut away suddenly, do not switch camera angles. Must use a fast, smooth camera push-in / dolly-in within the same continuous shot:
Action medium shot → male lead suddenly steps in → camera rapidly pushes in to medium close-up / close-up → clearly showing the brief cheek kiss, the male lead's slightly smug expression, and the female lead's momentary shock or irritation → camera naturally pulls back as the two separate and continues following the chase.
Each kiss lasts approximately 0.3–0.6 seconds, must be clearly visible but very brief. No stopping to strike a pose.
【Male Lead Expression & Personality】
The male lead maintains a composed, playful, slightly cheeky prankster vibe throughout. He is not angry and does not want to actually overpower the female lead; his enjoyment comes from continuously avoiding her attacks and then seizing a momentary opening to deliver a cheek kiss.
After each successful kiss, he must show a very brief, natural, slightly smug expression: a subtle upward curl of the mouth, eyes carrying a hint of "you still haven't hit me" playfulness, occasionally a light eyebrow raise. Not lecherous, not sinister, no exaggerated villainous grin.
First time: slightly pleased with himself.
Second time: more obviously amused.
Third time: starting to deliberately provoke.
Fourth time: clearly knows the female lead is getting angrier, but still can't help showing a cheeky grin.
Fifth time: after the kiss, clearly shows a guilty "I went too far" smile, then immediately runs away.
【Dialogue Rules】
The film has very minimal dialogue — only the following 4 Korean lines are allowed. Do not add any other dialogue, narration, or subtitles.
After the first cheek kiss, the female lead says briefly, startled and annoyed: "야!"
After the second cheek kiss, the male lead says while casually backing away with a playful tone: "또 실패."
After the fourth cheek kiss, the female lead grits her teeth and suppresses her anger, saying: "너 진짜..."
She doesn't finish the sentence before resuming her attack.
After the fifth cheek kiss succeeds, the male lead immediately turns and runs away, shouting in a not-very-sincere casual tone: "미안!"
The third cheek kiss has no dialogue at all — only a clear kiss sound, the female lead's expression, and the immediate chase create the comedic beat.
All lines must strictly remain in the original Korean text above. No Chinese dialogue, no English dialogue, no auto-translation.
【Timeline】
0–4 seconds
Action starts within the first 1 second. The female lead is already attacking relentlessly in the high-rise living room area, fast and sharp: jabs, spinning elbow strikes, mid-level kicks in continuous succession. The male lead barely throws any offense — he only sidesteps, ducks, glides, and leans back to dodge, moving very calmly. Camera follows closely, establishing the basic dynamic of "female lead keeps attacking, male lead keeps dodging."
4–6 seconds | First cheek kiss
The female lead continues pressing forward, throwing a punch then a kick to close in on the male lead. After dodging, the male lead briefly steps in from roughly a 45-degree angle behind the female lead's side — not directly behind her. He then quickly delivers the first brief kiss on her right cheek. Must rapidly push in to a close-up of the face within the same shot, clearly showing the kiss. The female lead freezes for a moment, then says irritably: "야!" The male lead immediately backs off, showing his first smug, cheeky expression.
6–10 seconds
The female lead is visibly angrier, raising her attack tempo. She uses low sweeps, high kicks, hooks, and spinning pursuits. The male lead keeps dodging, not counter-attacking, only retreating and evading, occasionally using very brief redirections to change the direction of her attacks, but never launching a real counter.
10–12 seconds | Second cheek kiss
The female lead throws a high kick that the male lead barely dodges by ducking low. Seizing the half-beat as her leg comes down, he cuts in from the other side and delivers another kiss on her left cheek. Must again rapidly push in to a close-up of the face within the same shot, clearly showing the moment. After the kiss, he casually backs away saying: "또 실패." His expression is more obviously amused. The female lead gets angrier and immediately charges.
12–16 seconds
The female lead begins chasing more fiercely, with bigger, faster, more aggressive movements. She continuously presses forward, spins, throws high kicks, chasing the male lead from the living room to the glass window and then turning toward the open walkway. The male lead almost only performs near-miss dodges: leaning back, slipping around close, ducking low, sidestepping away. The action must feel like "one more moment of hesitation and he would get hit."
16–18 seconds | Third cheek kiss
After the female lead's two consecutive attacks miss, the male lead suddenly steps in at the instant she turns her head, delivering another shorter, more sudden, crisper cheek kiss. Still must rapidly push in to a close-up of the face within the same continuous shot, clearly showing the action. No dialogue here — only a brief natural "smack" sound and the female lead's expression. The male lead's expression now carries deliberate provocation and cheekiness. The female lead's eyes clearly turn fiercer, and she resumes the chase.
18–22 seconds
The female lead's anger continues to build, and the pursuit becomes more vicious. She uses high kicks, spinning crescent kicks, and knee strikes to close in, almost without pausing. The male lead dodges throughout, retreating into the hallway, weaving around furniture edges, never counter-attacking — only reading her movements and narrowly avoiding them. The male lead's composure begins to carry a hint of comedic panic from being chased, but he still dodges very skillfully.
22–24 seconds | Fourth cheek kiss
One of the female lead's attacks grazes past the male lead's face. Using the momentum of her attack, he suddenly steps in close and delivers the fourth cheek kiss. The camera again rapidly pushes in to a close-up within the same shot, clearly showing the action. The male lead still can't help showing a cheeky grin. The female lead grits her teeth, suppressing her anger, and says: "너 진짜..." Before she finishes the sentence, she immediately resumes the chase, not giving him a moment to breathe.
24–27 seconds
The female lead enters her most intense pursuit state, with maximum speed and pressure. She throws high kicks and spinning attacks almost without pause, driving the male lead deeper into the hallway. The male lead still does not counter — only continuously retreating at the last moment, sidestepping, ducking, leaning back, like he's dodging for his life.
27–29 seconds | Fifth cheek kiss
In the fastest chase and exchange of the entire film, the male lead looks like he's about to get kicked, but dodges at the very last instant, then suddenly and very quickly returns to the female lead's close range, completing the fifth — and most exaggerated, most cheeky — cheek kiss. Must use the clearest continuous push-in facial close-up, clearly showing the action. After the kiss, the male lead shows a guilty "I went too far" smile, then immediately turns and runs away, shouting with a hint of guilt and feigned casualness: "미안!"
29–30 seconds
The male lead immediately flees deeper into the hallway. The female lead chases after him without hesitation, still ready to continue attacking. The camera follows both of them rushing out at high speed, and CUT TO BLACK on the clear sense that "this chaos isn't over yet."
【Performance Focus】
Female lead: On the offensive the entire time. Brief shock at the first kiss; increasingly irritated afterward; by the latter half she's essentially chasing him down. Her movements cannot be soft — they must carry real aggression.
Male lead: Never truly counters — only dodges, steps in close, kisses once, then backs off. His expression goes from composed and playful to carrying a funny "I might have overdone it" look later, while still maintaining an expert-level ease of movement.
Both must move like people who actually know how to fight — realistic movement, realistic momentum, realistic breathing, realistic fabric movement.
【Environment & Physics】
The space must be realistically usable: floor-to-ceiling glass, living room, open floor, long hallway, city nightscape. Light collisions with furniture edges, sudden stops and turns, gliding steps, and weaving around objects are allowed, but no large-scale destruction. No weapons, no guns, no bloody injuries. Action sound effects must be realistic: footsteps, air whooshes, fabric friction, breathing, spatial reverb.
【Sound】
No weird BGM, no piano-style scoring. An extremely subtle low rhythmic ambient bed is acceptable, but action and performance take priority. Key sound effects: footsteps, breathing, punch/kick air whooshes, fabric friction during close dodges, spatial reverb. Each cheek kiss must have a natural brief "smack" sound. The beat of stillness after each kiss must be clear, but no slow motion.
【Strict Prohibitions】
Do not portray the male lead as actively counter-attacking;
Do not have the male lead throw consecutive punches or consecutive kicks;
Do not shoot the action as mutual fighting;
No prolonged standoffs;
No slow motion;
No animated feel;
No game CG feel;
No weapons;
No blood;
No extra characters;
No subtitles;
No watermarks;
No daytime scenes;
No old, worn-down, dilapidated scenes;
No vulgar intimate scenes;
No romantic long kisses;
No mouth-to-mouth kissing;
No kissing hair;
No kissing ears;
No kissing air;
No facial clipping/merging;
No character identity drift;
No face distortion;
No deformed limbs;
No unclear action;
Do not shoot the whole piece as a pure romance scene;
Do not shoot the whole piece as a slapstick farce — the action itself must remain fast, professional, and cinematic.
Cinematic AI Video Prompt
A frustrated young male creator sits alone in a dark creative studio, staring at a completely blank computer monitor. He rests his hand on his forehead, looking stuck and out of ideas. The room is moody and cinematic, with soft monitor glow and deep shadows.
The blank screen suddenly begins transforming into a vivid cinematic world. The camera smoothly pushes toward the monitor and transitions seamlessly inside it, revealing an enchanting fantasy forest filled with enormous ancient trees, glowing purple-pink foliage, exotic plants, soft mist, and a crystal-clear stream reflecting warm rays of sunlight. Magical particles float gently through the air as the camera slowly travels forward through the forest.
The scene then dramatically transitions into a sleek futuristic black supercar speeding through a winding mountain road at dusk. The car accelerates aggressively around the curves, tires producing subtle smoke and sparks, glowing red taillights reflecting across the wet asphalt. Massive mountains surround the road with a bright full moon in the background.
Ultra-cinematic commercial look, photorealistic details, dynamic camera movement, smooth transitions, volumetric lighting, atmospheric fog, realistic reflections, shallow depth of field, dramatic contrast, premium VFX, realistic motion blur, 4K quality.
End with a powerful tracking shot behind the supercar as it disappears into the mountain road.
Duration: 15 seconds.
Aspect ratio: 16:9.
No text, no logos, no watermark.
HANOI TRAVEL VLOG — 30 SEC
No reference image. The woman below is described in words and must be built from the text alone.
SUBJECT: one Korean woman, early 20s, the only person the film follows. Small oval face with a soft jawline, large round eyes with double lids, a small straight nose, full coral-red lips, fair skin. Black wavy hair tied low at the nape and pulled forward over one shoulder, with loose strands left at the temples. A small gold hoop in each ear. Real Korean skin texture with visible pores — no doll-like AI smoothness, no beauty filter. The face established in the first shot is repeated exactly in every following shot: same features, same hair, same colour. She never morphs into a different person and no lookalike appears.
OUTFIT (identical throughout): burgundy ribbed knit halter top, cream high-waist flared trousers, tan leather belt with a gold buckle, burgundy platform heeled sandals, a thin gold bracelet on one wrist. There is no changing shot.
CAMERA: she is filming herself on her own phone. Most shots are selfies with her arm extended; objects and food are shot by turning the phone toward them. 16:9 landscape throughout, no black bars. 26mm phone lens, deep focus, neutral iPhone colour, handheld jolt on every step, autofocus landing a beat late, exposure lagging when she moves from bright to dark, a fingerprint smudge bleeding beside strong lights. No cinematic grading, no studio lighting, no glow fx. Background people stay out of focus and nobody looks at the camera for long.
DIALOGUE: lines are short and land in one breath. An interjection plus one short sentence is fine, but never two complete sentences strung together. There is breath between them and they are never rushed into each other. Punctuation is used once — no doubled question marks or exclamation marks. The two Vietnamese greetings are not fluent — just phrases she has picked up.
(0:00–0:02) — Hotel window, Hanoi morning reveal. Selfie, low angle. She shoves the curtain sideways and narrow tube houses under bundled cables open up behind her. The window blows out white, then the exposure catches up. Bright grin. Dialogue: "Xin chào~ 하노이 왔다!" Transition: Hard cut on beat
(0:02–0:04) — Old Quarter junction, a river of motorbikes. Selfie, stepping out between the bikes. She swings the phone to her face while a stream of helmets and conical hats flows behind her, horns overlapping. Overcast daylight, deep focus. Dialogue: "이걸 어떻게 건너?" Transition: Fast cut on beat
(0:04–0:06) — Street café. Selfie, arm short. She pushes a glass with condensed milk settled at the bottom toward the lens until it fills half the frame, eyebrows raised. Low backless stools, a stained yellow French wall behind. Transition: Snap cut
(0:06–0:08) — Alley bánh mì cart. Insert, phone pointed down at her hands. A baguette split over the charcoal is handed to her in newspaper and she peels the paper down with her fingertips. Steam rises and a few sprigs of coriander stick out of the split. White tube light on the cart, a parked motorbike beside it. Transition: Cut on natural movement
(0:08–0:09.75) — Huc Bridge reveal. She raises the phone from her face and tilts it up until the red wooden bridge crossing the lake fills the frame. Natural daylight, sky blown out behind a banyan tree. Her mouth opens slightly. Transition: Soft cut mid-upward motion
(0:09.75–0:12) — Train Street café. Selfie. On a low stool beside the track she is holding a coffee glass in both hands; hearing it come she sets the glass down and pulls her stool back against the wall. She looks at the lens, then sideways, and her face collapses into a laugh as a train comes out of the far end of the lane and passes an arm's length away. Dialogue: "헐 진짜 온닷ㅋㅋㅋㅋ" Transition: Cut on the end of the laugh
(0:12–0:13.75) — Old Quarter, the 36 streets. Follow from behind, phone low. She shoulders through the packed lane with the bánh mì, finishes the last piece and folds the newspaper away. Unlit lanterns and signs stacked tight overhead, people pushing past. Overcast afternoon light. Transition: Cut on forward movement
(0:13.75–0:15.5) — Back seat of a cyclo. Her hands are empty now — the phone low at knee height in one hand, the rail in the other. The street slides past under the folded canopy, the driver's back swaying in front. She looks out, not at the lens. The music drops a beat, leaving only wheels and horns. Transition: Snap cut
(0:15.5–0:18) — Phở shop. From here the image is camcorder: blooming tape, light noise, late focus. The camera is fixed at eye level across the counter, no phone or camera visible in frame. Medium. She sits down on a low blue stool under a hand-painted Phở sign. The owner sets a steaming bowl of beef phở down — thin-sliced beef, rice noodles, spring onion, a side plate of sprouts and basil. Dialogue: "쌀국수! 이거 먹으러 왔어~" Transition: Cut on the bowl landing
(0:18–0:20) — Lifting the noodles. Macro, phone lowered to the bowl, shallow focus. She draws rice noodles up out of the clear broth with her chopsticks, steam climbing. The steam crosses the lens and one side of the frame hazes over and clears. Only the broth and the slurping at the next stool. Transition: Cut with the noodles still raised
(0:20–0:22) — First mouthful. Selfie, handheld. She pulls in a big mouthful, closes her eyes and opens them. Her shoulders drop and she nods once at the lens. Just after sundown, the white fluorescent light of the shop. Dialogue: "미쳤다 진짜." Transition: Cut where her face softens
(0:22–0:23.5) — The broth. Selfie, close. She lifts the bowl in both hands, tips it back, sets it down and lets out a breath. The lens fogs with steam. Sound only. Transition: Cut back to phone image quality
(0:23.5–0:27) — Tạ Hiện beer street ★signature. She ste
Create a 15-second ultra-realistic cinematic YouTube-style fashion vlog featuring the girl from the uploaded reference image as the exact character reference. Preserve her facial identity, eye color, skin tone, hairstyle, body proportions, and overall appearance throughout the entire video.
The video begins with a handheld selfie-style shot inside her cozy bedroom. She looks directly into the camera and says with playful indecision, "I'm going to a party tonight, but I seriously have no idea what to wear." She looks toward her wardrobe, walks behind the curtain, and quickly returns wearing a stylish cropped top with a mini skirt.
She flicks her fingers toward the camera — snap! Her outfit instantly transforms into a fitted crop top with high-waisted jeans. She looks down at herself, poses, then gestures toward the camera as if asking, "This one?"
Another finger flick — snap! The jeans transform into a trendy short skirt with a different top. She spins once, lets the skirt move naturally, then gives a playful uncertain expression.
She flicks again — snap! Her outfit changes into a stylish party dress with heels and accessories. She strikes a confident pose, checks herself in the mirror, then looks back at the camera with an excited smile.
For the final beat, she rapidly flicks her fingers once more and changes into a completely different glamorous party outfit. She looks at the camera, points at herself, then gestures "Which one?" with a playful expression as the camera pushes in.
Fast seamless outfit transformations, no jump cuts during the actual changes, realistic fabric morphing, natural hair movement, consistent bedroom environment, handheld creator-vlog energy, expressive reactions, quick camera movements, smooth whip pans, natural daylight mixed with warm bedroom lighting, realistic skin texture, cinematic depth of field, photorealistic, ultra-detailed, 4K HDR, 24fps, 16:9.
Audio: Natural room ambience, footsteps, wardrobe movement, finger snaps with satisfying transformation sound effects, subtle upbeat vlog-style background music, natural spoken dialogue. No subtitles, no logos, no watermark, no distorted hands, no duplicate people, no face changes.
Duration: 14 Seconds | Aspect Ratio: 16:9
STYLE: Ultra-photorealistic REAL LIVE-ACTION, AAA Hollywood supernatural assassin action, premium realistic Semi-CGI VFX, ARRI ALEXA 65, IMAX, Panavision anamorphic. Fast controlled choreography, brief micro slow-motion ONLY for phase-dodge.
CHARACTER: @Image1 = KAIA. Preserve exact face, hairstyle, body proportions and ORIGINAL CLOTHING. TWIN CURVED DAGGERS. Personality: COLD, FOCUSED, EFFICIENT, UNHURRIED—never frantic.
LOCATION: Night atop a snowbound mountain fortress: stone battlements, frost-slick walkways, watchtowers, torch braziers under a full moon. Sentries patrol separately. Fortress stays quiet—NO alarm, crowd, or large battle.
00:00–00:03 — SILENT HUNT
Camera already moving, low handheld pursuit behind Kaia as she crosses a frosted walkway. Two sentries patrol ahead along the battlement. She waits, motionless in shadow, until one turns his back.
WHOOM— short burst of SWIRLING SNOW-ASH. Kaia's REAL BODY physically accelerates past him, trailing translucent frost-grey afterimages. She stops at blade range behind him.
SHK-SHK— TWO precise dagger strikes. She catches and quietly lowers him. The second sentry begins turning—Kaia is already gone.
00:03–00:06 — BLIND SPOT
Camera slides around a watchtower as the second sentry scans the dark, breath visible in the cold. Kaia silently emerges from his blind spot, pins his weapon arm—
SHK-SHK. TWO precise strikes. She quietly lowers him and looks toward three sentries near the gatehouse, expression unreadable.
00:06–00:08 — PHASE-DODGE
Kaia crosses silently behind the gatehouse sentries. One unexpectedly notices and swings at her FROM BEHIND.
Camera rushes toward the incoming blade—MICRO SLOW-MOTION. Just before impact, Kaia's physical body visibly DISSOLVES into DRIFTING SNOW-ASH and pale moonlit vapor. The blade passes harmlessly THROUGH her swirling form.
TIME SNAP—WHOOSH! The ash sweeps around the attacker as camera performs a curved whip-pan. Kaia visibly REFORMS directly behind him—feet → torso → arms → face → TWIN DAGGERS.
He turns too late.
CROSS-SLASH. ONE precise counter. Tiny camera impact bump. Kaia catches and silently lowers him.
00:08–00:11 — ASSASSIN CHAIN
Two sentries remain. Kaia disappears behind a brazier instead of charging.
One passes—snow-ash curls behind him. Kaia emerges: TWO STRIKES, catches him, then slips back into shadow.
Final sentry sees a fading translucent afterimage and follows it, sword raised. Camera rotates around him—the REAL Kaia is already in his blind spot.
SHK-SHK. ONE controlled exchange. She catches and quietly lowers the FINAL SENTRY.
ALL ENEMIES ARE DEFEATED. NONE REMAIN.
00:11–00:14 — SILENT AFTERMATH
Absolute quiet. Camera tracks backward along the battlement, slower than before. All defeated sentries lie silently along Kaia's infiltration path. Braziers still burn; gates remain intact; NO alarm.
Kaia calmly wipes her TWIN DAGGERS and sheathes them.
CLICK.
Her eyes shift toward distant torchlight. Snow-ash curls around her feet.
WHOOM— Kaia silently bursts into the storm. Camera rushes after her but catches only translucent frost-grey afterimages fading into the blizzard, then holds on empty, drifting snow.
END.
CORE ASSASSIN LOGIC
OBSERVE → BLIND SPOT → SILENT APPROACH → PRECISE STRIKE → CONTROL FALL → DISENGAGE → NEXT TARGET.
Never frontal brawling or prolonged blade exchanges.
PHASE-DODGE: Incoming attack → micro slow-motion → blade almost connects → Kaia visibly dissolves into snow-ash → attack passes through → ash travels around attacker → Kaia visibly reforms at blind spot → precise counter. Use ONLY when directly attacked; do not spam.
VFX / CAMERA / AUDIO
Speed VFX: Drifting snow-ash + pale moonlit vapor + translucent frost-grey afterimages. During normal bursts, REAL Kaia always physically leads; VFX trails behind. Phase-dodge requires visible dissolve → ash travel → physical reformation. NO lightning, electrical arcs, or teleportation.
Camera: NEVER static. Low pursuit, over-shoulder stalking, watchtower reveals, close reactions, reactive whip-pans, curved tracking. FAST during eliminations, immediately CALM afterward: QUIET → FAST → QUIET → FAST → QUIET.
Audio: Wind, distant howling, torch crackle, soft footsteps on snow, fabric movement, subtle ash WHOOSH, dagger draw/impact, controlled body movement. No loud battle music or alarms.
NEGATIVE: No frontal mass battle, prolonged blade exchange, running fight, lightning/electricity, teleportation, excessive phase-dodge, explosions/destruction, static camera, robotic movement, air-gap dagger hits, surviving enemies, face/body/clothing/dagger drift, broken anatomy, full-3D/game look, subtitles, logo, watermark.
Create a 30-second cinematic dance sequence as if directed for a high-budget international music film.
DIRECTOR'S VISION
The film takes place inside a massive minimalist studio at night. The space is almost completely dark, with carefully controlled #FF4900 orange practical lighting creating a striking visual identity. The atmosphere should feel sophisticated, dramatic, and expensive.
There is one adult female dancer. She remains the exact same person throughout the entire sequence. Her face, hairstyle, wardrobe, proportions, and styling never change.
OPENING — 0:00–0:05
Start on an extreme close-up of the dancer's face.
She stands completely still.
Only a thin orange light crosses her face. Hold the shot for a moment before slowly pulling the camera backward on a dolly.
The music begins quietly.
She makes the first controlled movement.
BUILD — 0:05–0:12
Cut to a 50mm medium shot.
The dancer begins a precise contemporary choreography.
The camera moves sideways with her rather than simply pointing at her. Let the movement of the camera and performer feel connected.
Orange practical lights gradually illuminate behind her.
Use shallow depth of field and a subtle focus pull from the background to her eyes.
MOMENTUM — 0:12–0:20
The music becomes more energetic.
Transition into a 35mm tracking shot.
The camera slowly circles around the dancer while she performs a sequence of turns, controlled footwork, coordinated arm movements, and a brief jump.
Do not over-edit.
Let the choreography breathe.
Use natural motion blur and realistic physical movement.
HERO MOMENT — 0:20–0:26
Move into a wide 24mm shot.
The dancer moves toward the center of the enormous studio.
As she reaches the beat, hundreds of small orange lights activate across the architecture behind her.
The camera performs a slow crane movement upward, revealing the scale of the environment.
The dancer remains the visual focus.
ENDING — 0:26–0:30
Everything suddenly becomes quiet.
Return to a 50mm shot.
The dancer stops and looks directly toward the camera.
Hold the composition for two seconds.
The orange lights behind her slowly fade except for one strong backlight.
Camera gently pushes in.
Cut to black.
CINEMATOGRAPHY
High-end feature-film cinematography, motivated camera movement, deliberate framing, realistic lens characteristics, controlled depth of field, natural motion blur, subtle film grain, realistic exposure, sophisticated contrast, volumetric atmosphere, practical lighting, physically accurate reflections.
PERFORMANCE DIRECTION
The dancer should perform like a professionally trained performer. Movements are precise, confident, rhythmic, and natural. No exaggerated body motion. No unnatural poses.
CONTINUITY
One performer throughout. Perfect facial consistency. Identical hairstyle, wardrobe, accessories, proportions, and appearance in every shot. Stable anatomy and hands. No duplicates. No morphing. No wardrobe changes.
FINAL LOOK
Photorealistic 4K, premium theatrical cinematography, sophisticated production design, realistic skin and fabric, cinematic color science, restrained visual effects, professional music-film aesthetic, emotionally controlled pacing, seamless continuity.
Create a cinematic, photorealistic lifestyle video of a young East Asian woman in a cozy modern apartment by the sea. The video has a warm, natural morning atmosphere with soft daylight coming through large windows, realistic skin texture, subtle facial expressions, natural body movements, shallow depth of field, and smooth cinematic camera motion.
Scene 1 — Bedroom:
A young woman with long straight dark hair, wearing a light beige/pink satin pajama set, sits on the edge of her bed looking sleepy. She gently yawns and rubs her eyes. The bedroom is minimal and modern, with a neatly made bed, wooden furniture, soft curtains, and large windows letting in diffused natural light.
Scene 2 — Bed:
She pulls and adjusts the duvet, then sits and stretches slightly on the bed. Capture her natural sleepy morning routine with realistic movements and a calm atmosphere.
Scene 3 — Leaving Bedroom:
She stands up and slowly walks toward the bedroom doorway. The camera remains cinematic and slightly distant, showing the warm wooden interior and softly illuminated bedroom in the background.
Scene 4 — Cooking:
Cut to the kitchen. Close-up of a black electric sandwich/waffle-style press on the kitchen counter. The woman pours smooth light-brown batter into the heated mold. Use detailed macro shots of the batter flowing into the appliance.
Scene 5 — Preparing Food:
She operates the sandwich maker on the kitchen counter and carefully checks the food while cooking. Show realistic hand movements, steam/heat details, kitchen reflections, and natural daylight coming through the nearby window.
Scene 6 — Eating:
She opens the appliance and removes a freshly cooked golden-brown waffle/pastry. She holds it with both hands, takes a bite, then smiles naturally with a satisfied expression.
Visual style: photorealistic, cinematic lifestyle commercial, natural morning lighting, warm neutral color palette, realistic Asian facial features, authentic skin texture, detailed hair strands, realistic fabric physics, soft shadows, subtle film grain, shallow depth of field, professional cinematography, smooth transitions, realistic handheld camera movement, 4K quality.
Camera: combination of medium shots, close-ups, macro food shots, slow push-ins, gentle tracking shots, and shallow-depth-of-field portrait shots.
Mood: cozy, peaceful, warm, relaxing morning routine, premium lifestyle advertisement.
Aspect ratio: 16:9
Duration: approximately 20 seconds
No text, no subtitles, no watermark, no logo, no distorted hands, no extra fingers, no unnatural facial movements, no cartoon/anime appearance.
Translate the following AI generation prompt into English.
Output ONLY the translated text — no quotes, no markdown, no explanations.
Preserve line breaks and formatting.
Use the uploaded reference image as the exact character identity reference. Keep the same woman throughout the entire video — identical face, hairstyle, skin tone, body proportions, makeup style, and recognizable appearance. She is a professional adult fashion model preparing for a stylish outdoor photoshoot.
Create a fast-paced, cinematic fashion behind-the-scenes sequence where the photographer is never visible in frame. The camera is always positioned from the photographer's perspective, making it feel like we are seeing the photoshoot through their camera.
The video opens inside a stylish bedroom or makeup studio. She finishes her makeup, fixes her hair, checks herself in the mirror, puts on her accessories, grabs a small fashion bag, and confidently walks outside for the shoot.
She arrives at the first location: a beautiful street-side café. Wearing a stylish fitted crop top with a mini skirt, she stands beside the café window, adjusts her hair and looks directly toward the unseen photographer.
She says naturally: “Okay, here… now click.”
A realistic camera shutter sound plays — CLICK.
Instant cinematic transition to the next location.
Now she is sitting casually on a vintage scooter parked along a charming street, wearing a different fashionable outfit. She crosses one leg over the other, turns her face toward the camera, plays with her hair and gives several confident poses.
She looks toward the unseen photographer and says: “This angle… now click.”
CLICK.
Quick flash-style transition.
Next, she appears in a glamorous resort-style outfit with a tasteful two-piece swimwear look, posing beside a luxurious poolside setting. She walks toward the camera, turns over her shoulder, then gives a confident editorial pose.
She smiles and says: “Wait… this one. Click.”
CLICK.
Cut to a busy urban street. She is now wearing fitted jeans with a stylish crop top and heels. She leans casually against a street pole, changes poses with every camera movement, looks away, then suddenly looks directly into the lens.
“Okay, hold it… now click.”
CLICK.
Rapid transition to another location: a colorful outdoor café terrace. She wears a completely different elegant outfit, sits at a table with a coffee, crosses her legs, takes a small sip, then poses naturally while looking toward the unseen photographer.
“One more… click.”
CLICK.
Final location: golden-hour city street. She wears her strongest fashion look, walks toward the camera, stops under warm sunlight, turns slightly, lets her hair move naturally in the breeze and gives a confident final model pose.
She looks directly into the camera with a playful smile:
“That’s the one.”
One final realistic camera shutter — CLICK.
Use seamless match cuts, camera-flash transitions, natural body movement, realistic hair and clothing physics, shallow depth of field, cinematic lens changes, handheld photographer perspective, natural street ambience, realistic lighting, premium fashion-commercial aesthetics, photorealistic skin texture, elegant color grading, strong temporal consistency, and smooth transitions between locations.
No photographer visible, no camera visible, no logos, no subtitles, no watermarks, no distorted hands or face, no identity changes. 16:9 cinematic composition, high-end fashion editorial photography, realistic camera shutter sounds and natural location ambience.
A cozy Studio Ghibli-inspired cinematic cooking animation showing a delicious grilled chicken shawarma wrap and refreshing mint-lime soda being prepared step by step. Begin with slicing juicy shawarma meat on a vertical rotisserie, cutting fresh limes, adding ice, mint, and sparkling soda to a chilled glass, then assembling a warm tortilla with grilled chicken, lettuce, tomatoes, onions, purple cabbage, and creamy garlic sauce before rolling it into a perfect wrap. Finish with garnishing the drink with mint and lime, showcasing fizzy bubbles and the final plated meal on a rustic wooden counter in warm golden lighting, with highly detailed food, smooth camera movements, soft depth of field, and a magical hand-painted Ghibli aesthetic.
Montage, multi-shot handheld home-video vlog, avoid a single camera angle or single cut — 7 shots. Shot one-handed on a phone, snapshot-like realism, slightly tilted framing, visible handheld shake, warm late-afternoon indoor light, fine film grain, photorealistic.
A woman (Image) folds laundry alone in her living room while a kettle heats on the stove nearby. (image) provides only her face and hairstyle her clothing follows this description entirely: she wears a soft heather-grey cotton short-sleeve shirt, sleeves fully covering her shoulders and upper arms, tucked loosely into cream-colored loungewear pants. She is the only person who appears in the video throughout. The setting is a cozy living room corner: a low wooden coffee table stacked with warm towels fresh from the dryer, a wicker laundry basket half-emptied on the floor, a folded blanket draped over a nearby armchair, soft evening light slanting through sheer curtains; in the background, just visible through a kitchen doorway, a stove and kettle sit, steam beginning to curl from its spout. Her reactions are quiet, natural, small domestic movements. The whole sequence moves from a stack of warm, slightly rumpled laundry to neatly folded piles, timed loosely with the kettle building toward its whistle. The dialogue is casual, everyday spoken Korean, an immediate reaction to the action in the moment.
**Shot 1 (0-2s):** She pulls a warm towel from the basket, presses it briefly to her cheek to feel the heat, then starts folding it in practiced motions. She murmurs, content: "아 따뜻해~" ("Ahh, so warm~"). The camera is close and slightly low, drifting with her hands.
**Shot 2 (2-4s):** Close insert shot from above — she stacks the folded towel onto a growing pile, corners lining up neatly, her fingers smoothing the top edge flat. The camera follows her hands from above, with slight shake.
**Shot 3 (4-6s):** She picks up a wrinkled shirt, folds it once, frowns slightly at the crooked sleeves, and unfolds it to redo it more carefully. She mutters: "아니 이게 아니지" ("No wait, that's not right"). The camera is straight-on to her face and hands, only subtle handheld tremor.
**Shot 4 (6-8s):** She refolds the shirt properly this time, smoothing it flat with the side of her hand, satisfied with the neat rectangle. She says lightly in English: "Much better." The camera angles slightly upward, pushing in slowly.
**Shot 5 (8-10s):** In the background through the doorway, the kettle begins to rattle faintly, steam thickening and catching the low light; she glances up toward the kitchen, only half-focused, still folding. The camera is straight-on to her, swaying gently with her breathing.
**Shot 6 (10-13s):** The kettle's whistle rises sharply; she sets down the towel mid-fold, wipes her hands on her pants, and starts to rise from the floor, glancing toward the kitchen with a small huff. She says: "어, 나온다 나와" ("Oh, it's ready, it's ready"). The camera tilts and follows her shifting weight as she starts to stand.
**Shot 7 (13-15s):** She pauses halfway up, looking back at the neat stacks of folded laundry on the table, a small satisfied nod, the kettle still whistling faintly behind her. The camera drifts back and slightly up, lingering on her and the folded piles before she moves off.
**Sound (SFX):** No music, only live ambient sound — soft fabric rustling and folding, the towel pile settling, a small frustrated sigh, fabric smoothed flat, the kettle building from a low rattle to a sharp whistle in the background, her footsteps shifting on the floor, a quiet murmur.
No subtitles, no on-screen text, no logos, no watermarks. Do not depict the reference image itself; do not duplicate/copy the subject.
A young snowboarder drops into an alpine terrain park, carves smoothly down the slope, performs a controlled 180° jump and freestyle butter, then finishes with a clean stop. Ultra-realistic winter sports documentary, physically accurate snowboard physics, authentic human movement, natural English lip-sync, bright winter daylight, immersive mountain ambience, and seamless story continuity throughout.
30-SECOND CINEMATIC SHORT
Create a 30-second cinematic sequence with the visual discipline of a major feature film. The world is entirely made from paper, cardboard, ink, and delicate handcrafted materials.
00:00–00:05 — THE FIRST FOLD
Begin with an extreme macro shot of a blank sheet of textured paper.
A single fold slowly forms across its surface.
Lens: 100mm macro
Camera: Completely static
Soft daylight moves across the paper as the fold continues.
Cut precisely as the paper reaches its final shape.
00:05–00:11 — THE CITY EMERGES
Transition into a 50mm shot.
The folded paper begins forming miniature streets and architectural structures.
Buildings rise naturally from the surface through carefully constructed paper folds.
The camera slowly tracks sideways.
Tiny windows catch the light.
Nothing feels computer-generated; every surface should have believable paper fibers, folds, shadows, and imperfections.
00:11–00:17 — THE REVEAL
Move into a 35mm shot.
The camera pulls backward.
The small arrangement is revealed as an enormous handcrafted paper city stretching across a large table.
Roads connect different districts.
Bridges cross miniature rivers.
Paper trees move gently from an unseen breeze.
00:17–00:23 — MORNING
Transition into a slow overhead crane shot.
Warm sunlight gradually spreads across the paper city.
Shadows from the buildings become longer and more defined.
Tiny paper windows begin reflecting the sunlight.
The camera continues rising, revealing the geometric relationship between the streets and buildings.
00:23–00:27 — THE DETAIL
Cut to an 85mm close-up.
Focus on a tiny paper clock mounted on one building.
The clock moves forward by one minute.
Rack focus from the clock to the miniature skyline behind it.
00:27–00:30 — FINAL IMAGE
Return to an extreme wide shot.
The entire paper city sits beneath a large studio window.
The morning light completely fills the miniature world.
Camera slowly pulls backward until the city becomes a small object within the larger room.
Fade to black.
CINEMATOGRAPHY
Large-format feature-film aesthetic.
100mm macro for texture.
85mm for detail.
50mm for natural perspective.
35mm for environmental shots.
Controlled dolly and crane movements.
Realistic depth of field.
Natural focus transitions.
Soft motion blur.
Subtle lens imperfections.
No artificial camera shake.
MATERIAL DIRECTION
Every surface must visibly behave like physical paper.
Visible paper fibers.
Natural folds.
Tiny imperfections.
Soft cardboard edges.
Realistic contact shadows.
Believable paper thickness.
Natural material deformation.
LIGHTING
Soft morning daylight.
Large window as the primary source.
Natural bounce light.
Gentle shadows.
Subtle warm highlights.
No artificial neon effects.
No excessive visual effects.
CONTINUITY
The same paper city throughout the entire sequence.
Buildings maintain identical shapes and positions.
Roads remain connected.
Paper materials remain consistent.
No random transformations.
No flickering.
No duplicated structures.
No sudden environmental changes.
DIRECTOR’S NOTE
Treat this as a miniature feature film, not a visual-effects demonstration.
The camera should discover the world gradually.
Start intimate.
Reveal scale.
Return to detail.
Finish wide.
The audience should feel that they are watching a real handcrafted world photographed with a cinema camera.
Photorealistic 4K • tactile materials • feature-film cinematography • sophisticated lighting • realistic miniature photography • cinematic depth • precise camera movement • seamless visual continuity.
Ultra-realistic F1-style race car speeding through abandoned city streets at night. Missiles strike buildings nearby, raining glass and fire across the road. Driver never slows. Camera mounted inches above asphalt as the car drifts through debris-filled intersections. Tunnel collapses ahead. Final frame: car launching through flames into open daylight.
Ultra-realistic behind-the-scenes phone footage inside a giant VFX water-tank studio. One single continuous unbroken wide shot, no cuts, no close-ups, no coverage. A miniature coastal city and long airport runway sit inside the tank, surrounded by a huge green-screen wall with tracking markers, overhead rigging, studio lights, a Jimmy Jib, side camera trolleys on tracks, monitors, cables, and crew members in FX-studio shirts filming the setup.
The handheld phone camera starts from an elevated side angle, clearly showing the runway, miniature skyline, water tank, and crew. Someone off-camera shouts, “Action!” A fighter jet on the runway begins accelerating fast. At the exact same time, an enormous tsunami wave rises behind the city and runway. The phone camera stays wide and smoothly pans to follow the action as the jet races forward, lifts off at the last second, and climbs just as the wall of water smashes onto the runway behind it. The plane barely escapes, then banks hard over the miniature city while the tsunami engulfs buildings below, flooding streets and smashing through the skyline. Crew members step back, camera operators keep rolling, and the wave fills the set with spray and chaos. End while the jet is turning above the destruction and the tsunami is still tearing through the city.
Sound: off-camera “Action!”, deep wave-machine rumble, jet engine roar, crew chatter, rolling trolley sounds, crashing water, miniature buildings collapsing, spray, startled shouting.
Ultra-realistic futuristic colosseum packed with roaring crowds. Two massive combat robots circle in the sand arena, armor torn open and sparking. One drives the other into the wall, collapsing stone columns. Camera spins low around their feet as debris flies. Final frame: victorious robot raises burning fist under stadium lights.