Video & audio

Build an AI-generated vlog

Prepare the source assets, follow the supplied generation workflow, and inspect the result frame by frame before exporting.

Adapted from TUT-017, dated 2026-04-13.

What you'll make

A checked final video that follows the source workflow and is ready for export.

Build an AI-generated vlog

Who this is for

Creators, marketers, founders, and small teams producing short AI-assisted videos from source media.

Step by step

Run the workflow

  1. 01

    The full workflow behind how I made this exact video

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Follow the shown setup, keep the source order unchanged, and run one controlled test.
    Source example for build an ai-generated vlog.
    Use this visual to confirm the expected setup or result.
    Source example for build an ai-generated vlog.
    Use this visual to confirm the expected setup or result.
    Source example for build an ai-generated vlog.
    Use this visual to confirm the expected setup or result.
    Source example for build an ai-generated vlog.
    Use this visual to confirm the expected setup or result.
    Source example for build an ai-generated vlog.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  2. 02

    Prepare the source material

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Capture as many reference images as you need. You do not need professional camera gear for your character.
    Source example for the full workflow behind how i made this exact video.
    Use this visual to confirm the expected setup or result.
    Source example for the full workflow behind how i made this exact video.
    Use this visual to confirm the expected setup or result.
    Source example for the full workflow behind how i made this exact video.
    Use this visual to confirm the expected setup or result.
    Source example for the full workflow behind how i made this exact video.
    Use this visual to confirm the expected setup or result.
    Source example for the full workflow behind how i made this exact video.
    Use this visual to confirm the expected setup or result.
    Source example for the full workflow behind how i made this exact video.
    Use this visual to confirm the expected setup or result.
    Source example for the full workflow behind how i made this exact video.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  3. 03

    Open the image generator

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Create your scenes or environment with Gemini Nano Banana Pro
    2. Go to Google Flow and select "Create with Flow."
    3. Select + New project
    Source example for prepare the source material.
    Use this visual to confirm the expected setup or result.
    Source example for prepare the source material.
    Use this visual to confirm the expected setup or result.
    Source example for prepare the source material.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  4. 04

    Enter your prompts to generate the images

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. You may receive free credits (varies by account).
    Prompt setup used for open the image generator.
    Check the prompt field, reference order, and visible settings.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  5. 05

    Create Character Cheatsheet

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Image Model: NanoBanana Pro
    2. create a character cheatsheet for this person, front, back ,full body, mid shot, wide shot. all on white background
    Expected result for enter your prompts to generate the images.
    Use this source result as the visual comparison point for the step.
    Expected result for enter your prompts to generate the images.
    Use this source result as the visual comparison point for the step.
    Expected result for enter your prompts to generate the images.
    Use this source result as the visual comparison point for the step.
    Expected result for enter your prompts to generate the images.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  6. 06

    Create The Scenes

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Create The Scenes prompt
    A large multi-level traditional Japanese interior atrium inside a dark historic wooden building, photographed as a real- life cinematic live-action scene. The space is empty with no people anywhere. Wide symmetrical establishing shot from an elevated wooden platform looking across a deep open central void. Multiple staircases rise and descend along both sides of the structure. Layered corridors, balconies, platforms, beams, railings, sliding shoji- style wall panels, and stacked architectural sections create a dense labyrinth-like interior. Warm amber light glows softly from large paper wall panels and hidden interior practical lights, while deep shadows fill the recesses of the building. Dark polished timber floors in the foreground, aged wood grain throughout, subtle atmospheric haze in the air, quiet mysterious mood, dramatic depth, realistic scale, physically accurate materials, authentic Japanese architectural detailing, natural light falloff, cinematic contrast, live-action film still, ultra realistic, real photography, 35mm cinema lens, no people, no furniture focus, no text.

    Expected result for create character cheatsheet.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  7. 07

    Generate a reference image 1

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    A vast traditional Japanese interior hall photographed as a real-life cinematic live- action space. Low-angle wide shot from floor level, looking upward into a towering multi-story wooden atrium. The architecture is dense and layered, with stacked rooms, elevated platforms, balconies, staircases, timber beams, railings, lattice windows, shoji-style wall panels, and interconnected walkways rising high into the structure. The foreground shows smooth dark wooden flooring, while the rest of the space is built from warm aged timber, paper screens, and wooden frames. Soft golden amber light glows from the lower left and from hidden interior openings, casting warm illumination across the walls and floor. Deep shadows fill the upper recesses and corners, creating strong contrast and depth. The overall mood is quiet, grand, mysterious, and atmospheric, like an old Japanese bathhouse, inn, or historical interior maze. Realistic wood grain, authentic paper textures, natural light falloff, subtle dust in the air, physically accurate materials, believable architectural proportions, ultra realistic, real photography, live-action film still, cinematic, no people, no characters, no text.

    Expected result for create the scenes.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  8. 08

    Generate a reference image 2

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    A surreal distorted reality cityscape built from traditional Japanese architecture, photographed as a real-life cinematic live- action environment. Impossible buildings rise from every direction at once, surrounding the viewer in a dense vertical maze. Large wooden structures, multi-story inns, balconies, staircases, bridges, corridors, rooftops, and illuminated shoji- style windows extend upward, downward, sideways, and diagonally as if gravity has broken. Some buildings are stacked on walls, some hang overhead, some intersect at impossible angles, and some appear to grow out of other buildings. The viewpoint is high above the city, looking across an endless labyrinth of architecture folding into itself from all sides. Warm amber and golden interior light glows from countless windows and paper panels, contrasting against deep black voids between the structures. Rich aged wood, dark timber beams, layered platforms, railings, tiled roofs, and narrow walkways create a haunting, dreamlike sense of scale. The environment feels like a warped pocket dimension, an impossible city, a collapsing architectural dream, realistic but physically surreal. Ultra realistic, live- action film still, real photography, cinematic lighting, atmospheric haze, strong depth, physically detailed wood and paper textures, no people, no text.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  9. 09

    Generate a reference image 3

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    An impossible architectural void filled with endless stairways coming from every direction, captured as ultra realistic real-life cinematic photography. The viewer looks into a vast dark interior space where countless staircases rise, descend, cross, twist, and disappear into an enormous maze of traditional Japanese wooden buildings. Stairways cut diagonally through the frame from above, below, left, right, and deep into the background, creating a gravity- defying spatial paradox. Around them, dense layers of wooden corridors, balconies, shoji-style walls, illuminated windows, platforms, beams, and structural frames form an endless vertical city inside a single enclosed space. Warm amber and orange light glows softly from windows, lantern-like fixtures, and hidden interior rooms, while deep shadows swallow the gaps between structures. The atmosphere is mysterious, dreamlike, oppressive, and beautiful, like an infinite labyrinth folded into itself. Authentic aged timber, realistic paper screens, dark wooden railings, cinematic haze, physically believable textures, natural lens depth, real architectural detail, live-action film still, ultra realistic, no people, no text.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  10. 10

    Generate a reference image 4

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    A colossal impossible city-like interior suspended inside a vast dimensional void, photographed as ultra realistic real-life cinematic architecture. There is no single upright direction anywhere in the scene. Entire building masses, corridors, platforms, bridges, staircases, and windowed facades are attached to different axes, as if each section of the city obeys its own gravity. Some structures stand vertically, some hang sideways, some are inverted overhead, and others extend diagonally into empty space. The viewer cannot tell what is floor, wall, or ceiling because the entire environment is built on conflicting orientations. The space is made of dense traditional Japanese- inspired architecture fused into an endless labyrinth of glowing wooden buildings. Dark timber facades, shoji-lit windows, narrow balconies, suspended corridors, ladder-like access paths, protruding platforms, support frames, and layered building blocks stretch infinitely into the background. Large architectural volumes jut out from the sides of other structures at impossible angles. Entire streets of buildings appear mounted onto vertical surfaces. Deep in the distance, countless glowing structures fade into heavy golden haze, making the space feel endless and bottomless. The viewpoint is from a dark ledge or platform near one cluster of buildings, looking outward across a gigantic atmospheric void filled with floating or gravity-broken architecture. Warm amber, orange, and gold light pours from thousands of windows and lantern-like openings, creating a sea of glowing points suspended in darkness and fog. Thick atmospheric haze softens the far distance, while foreground structures remain heavy, dark, and sharply detailed. The mood is monumental, dreamlike, oppressive, and awe-inspiring, like standing inside an infinite city folded into itself where orientation has collapsed. This is not a normal city and not a normal building interior. It is a spatially broken architectural dimension where no single horizon line exists and no direction can be trusted. Every cluster of buildings should feel mounted on a different rotational axis. The viewer should feel disoriented, as if the world has been rotated and duplicated in every direction at once. Ultra realistic, live-action film still, real photography, physically believable materials, authentic aged wood, dark timber beams, glowing paper and glass window panels, cinematic haze, natural lens depth, large-scale atmospheric perspective, no people, no text, no stylization.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  11. 11

    Generate a reference image 5

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    A colossal impossible city-like interior suspended inside a vast dimensional void, photographed as ultra realistic real-life cinematic architecture. There is no single upright direction anywhere in the scene. Entire building masses, corridors, platforms, bridges, staircases, and windowed facades are attached to different axes, as if each section of the city obeys its own gravity. Some structures stand vertically, some hang sideways, some are inverted overhead, and others extend diagonally into empty space. The viewer cannot tell what is floor, wall, or ceiling because the entire environment is built on conflicting orientations. The space is made of dense traditional Japanese- inspired architecture fused into an endless labyrinth of glowing wooden buildings. Dark timber facades, shoji-lit windows, narrow balconies, suspended corridors, ladder-like access paths, protruding platforms, support frames, and layered building blocks stretch infinitely into the background. Large architectural volumes jut out from the sides of other structures at impossible angles. Entire streets of buildings appear mounted onto vertical surfaces. Deep in the distance, countless glowing structures fade into heavy golden haze, making the space feel endless and bottomless. The viewpoint is from a dark ledge or platform near one cluster of buildings, looking outward across a gigantic atmospheric void filled with floating or gravity-broken architecture. Warm amber, orange, and gold light pours from thousands of windows and lantern-like openings, creating a sea of glowing points suspended in darkness and fog. Thick atmospheric haze softens the far distance, while foreground structures remain heavy, dark, and sharply detailed. The mood is monumental, dreamlike, oppressive, and awe-inspiring, like standing inside an infinite city folded into itself where orientation has collapsed. This is not a normal city and not a normal building interior. It is a spatially broken architectural dimension where no single horizon line exists and no direction can be trusted. Every cluster of buildings should feel mounted on a different rotational axis. The viewer should feel disoriented, as if the world has been rotated and duplicated in every direction at once. Ultra realistic, live-action film still, real photography, physically believable materials, authentic aged wood, dark timber beams, glowing paper and glass window panels, cinematic haze, natural lens depth, large-scale atmospheric perspective, no people, no text, no stylization.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  12. 12

    Generate a reference image 6

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    A vast high-ceilinged interior sanctuary built around enormous indoor lotus ponds and elevated wooden bridges, photographed as ultra realistic real-life cinematic architecture. The space is monumental in scale, with an immense vertical volume rising far above the walkways and water. The ceiling should feel extremely high and distant, disappearing into darkness, haze, or dim architectural layers overhead. The viewer should immediately feel the size of the place, like standing inside a giant sacred hall, abandoned palace interior, or ritual chamber built around water. Large rectangular lotus pools dominate the floor plan, filled with dark blue-green water, dense lily pads, and scattered blooming lotus flowers in pale pink and white. Pale wooden walkways and bridge-like platforms stretch across the ponds in long clean geometric lines, forming a maze of elevated paths over the water. The bridges should feel small compared to the scale of the hall, emphasizing how huge and empty the surrounding space is. Dark support pillars descend into the water and rise upward into the towering void above. The architecture should feel expansive, sparse, and oppressive rather than cozy. Surrounding the ponds are broad open deck areas, tall structural columns, shadowy recesses, and distant upper architectural layers barely visible through haze. The hall should not feel like a small courtyard. It should feel like an enormous enclosed chamber with vast negative space above and around the ponds. The high ceiling may include dark beams, hidden rafters, suspended walkways, openings, or layered structural frames fading into the gloom, but it should remain mostly obscured to preserve the sense of height. The mood is eerie, sacred, and monumental. Beautiful lotus flowers float across the still water, but the overall feeling is uncanny and slightly haunting. Cool blue-green light rises softly from the ponds, while dim pale light brushes across the wood bridges and decks. Subtle haze fills the air, making the height and distance feel even more dramatic. Reflections in the water should feel deep and mysterious, with soft ripples and believable surface texture. The whole place should resemble a forgotten sacred water hall, a ritual lotus sanctuary, or a hidden chamber inside a colossal temple complex. Ultra realistic, real photography, live-action film still, physically believable materials, pale wood grain, dark structural beams, still water, lotus flowers, lily pads, massive enclosed volume, towering ceiling, atmospheric haze, cinematic depth, no people, no text, no anime, no CGI, no illustration, no stylization.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  13. 13

    Generate a reference image 7

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    A vast elevated ritual platform suspended inside an enormous lantern-lit architectural void, photographed as ultra realistic real-life cinematic architecture. The location is a giant dark wooden chamber floating high above an endless labyrinth of layered East Asian-inspired buildings, bridges, balconies, staircases, glowing windows, and suspended corridors. At the center of the composition is a large square platform made of polished dark timber, isolated in open space and surrounded by immense vertical depth on all sides. The platform feels like a ceremonial stage or sacred chamber hidden inside a colossal impossible megastructure. At the far end of the platform stands a tall vertical wall panel of dark aged wood. Across this wall, black branch-like forms spread outward in all directions like dead roots, veins, or cracks, creating a haunting organic pattern that dominates the backdrop. These branching lines should feel unnatural and ominous, stretching wide and high across the wall surface as if the architecture itself has been infected or scarred. The rest of the platform remains mostly empty, emphasizing the stillness, scale, and tension of the space. The viewpoint is a high-angle wide shot looking slightly downward, showing the full expanse of the platform and the surrounding void of architecture beyond it. The polished timber floor should feel broad, empty, and slightly reflective under low warm light. The surrounding world is a monumental maze of dark wooden facades, lantern-lit window grids, narrow bridges, layered rooms, and shadowy vertical shafts fading into darkness and amber haze. The location should feel endless, as if this one platform is only a tiny part of a much larger impossible architectural world. Lighting is warm, low, and cinematic. Soft amber light spills from surrounding lanterns and interior windows, reflecting faintly across the platform floor and the nearby wall. Deep shadows dominate most of the chamber, with subtle glow catching edges of wood, railings, and distant structures. Atmospheric haze should soften the far architecture and enhance the sense of vast vertical depth. The overall mood is sacred, eerie, oppressive, and still, like a hidden ritual chamber inside a giant haunted bathhouse-like city. Materials must feel completely real: polished dark timber floorboards, aged wall panels, realistic wood grain, subtle wear, natural reflections, physically believable warm lighting, deep cinematic shadow falloff, and large-scale atmospheric perspective. The final image should look like a still frame from a high-budget live-action dark fantasy thriller, not concept art. Ultra realistic, real photography, live-action film still, monumental scale, ominous ritual atmosphere, dark wood, glowing lanterns, black root-like branches across the wall, high-angle composition, no people, no text, no anime, no CGI, no illustration, no stylization.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  14. 14

    Generate a reference image 8

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    A vast multi-level interior maze of traditional East Asian wooden architecture, photographed as ultra realistic real-life cinematic architecture. The viewpoint is from a narrow staircase and landing in the foreground, looking across a deep open vertical void filled with stacked rooms, platforms, walkways, balconies, and layered structural volumes. The staircase is made of aged dark wood with worn steps, smooth handrails, vertical balusters, and a slightly glossy finish catching warm light. The viewer should feel as if they are standing midway up the stairs inside a giant impossible bathhouse-like megastructure. Across the open shaft, dense architectural layers rise and descend in all directions. Box-like rooms, shoji- style wall panels, wood-framed facades, paper windows, balcony edges, and narrow corridors are stacked tightly together, forming a towering vertical labyrinth. The buildings feel fused into one giant interior city, with no obvious exterior, just endless connected architecture. Some floors jut out into empty space, some small landings disappear behind walls, and some rooms open onto narrow walkways that overlook the dark central void. Warm amber lantern light glows from windows, doorways, alcoves, and hidden interior spaces, casting soft orange-gold illumination across the wooden surfaces. The overall atmosphere is moody, enclosed, and slightly eerie, with deep shadows filling the lower recesses and distant corners. The open void between the architectural blocks should feel deep and dangerous, emphasizing the enormous scale of the structure. The lighting should be low and cinematic, with warm highlights on railings, stair edges, and wall panels, and softer shadow falloff into the depth below. At the upper right area, a large wall panel or room facade can feature dark branch-like root patterns or ink-like organic cracks spreading across the surface, adding a haunting visual accent without becoming the main focus. This should feel like part of the architecture itself, as if the building carries an old scar or strange ritual residue. The materials must look completely real: aged timber beams, dark polished stair treads, paper screen textures, realistic joinery, subtle wear, physically believable warm lighting, slight haze in the depth, and natural cinematic lens rendering. The whole location should feel like a forgotten sacred inn, a towering haunted bathhouse interior, or an impossible vertical temple-city hidden inside darkness. Ultra realistic, real photography, live-action film still, warm lantern glow, dark wood, deep architectural shaft, stairway foreground, layered rooms and balconies, ominous but beautiful atmosphere, no people, no text, no anime, no CGI, no illustration, no stylization.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  15. 15

    Generate the video

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Now bring your frame to life in Seedance 2.0.
    2. I am using Higgsfield for Seedance 2.0, but you can use any other platform that offers Seedance 2.0 to achieve the same result.
    3. Open Higgsfield → click Video → Select Seedance 2.0 1
    4. Upload all the assets needed for your scenes, then write a prompt describing them.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  16. 16

    Generate a reference image 9

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    15-second vertical reel, realistic iPhone front-facing selfie camera vlog quality, not cinematic, not polished, not refined. The entire video is from a handheld front-facing selfie camera perspective, as if an Asian man is filming himself at arm's length while walking through the real-life Demon Slayer Infinity Castle. The recording device is never visible in frame. No phone, no camera, no selfie stick visible. The footage should look like real iPhone selfie camera footage, with slightly amateur framing, natural hand shake, mild front-camera distortion, soft autofocus breathing, low-light grain, slight motion blur, uneven exposure, casual framing, and imperfect stabilization. He is whispering the whole time while carefully walking and exploring. He is always talking to the camera in a nervous vlog style while occasionally angling the front-facing selfie view to show parts of the creepy space around him without losing the handheld selfie feel. The Infinity Castle feels dark, eerie, abandoned, and haunted, with endless wooden stairs, hanging corridors, weak lantern light, drifting fog, deep shadows, dust in the air, and impossible architecture stretching in all directions. Scene 1 0.0 to 3.0 He walks slowly through the eerie wooden corridor in a continuous handheld selfie-vlog moment. His face stays close in frame most of the time, slightly off-center in an amateur way, with dim lantern-lit hallways and impossible stairs behind him. He glances around nervously and whispers, "Okay... I'm inside the place where the Demon Slayer movie happened... and this place is way creepier in real life..." Scene 2 3.0 to 6.0 Still walking carefully in the same front-facing handheld selfie camera shot, he slowly angles the camera just enough to show endless stairs, wooden balconies, and impossible architecture above and around him while part of his face remains in frame. He whispers, "Look at this... it goes everywhere... it doesn't even make sense how this place is built..." Scene 3 6.0 to 9.0 He keeps walking, then slows when he notices something beside him. Still in the same selfie-vlog style, he lowers and turns the camera slightly to reveal scattered broken film gear on a nearby platform while keeping part of his face in frame. Visible on the floor are a fallen camera rig, broken monitor, tipped tripod, loose cables, cracked light stand, damaged production cases, shattered props, and faint blood traces smeared across the floorboards and gear. He whispers, "Wait... there's film gear everywhere... no crew... no people... nothing..." Scene 4 9.0 to 12.0 He continues walking past the wreckage in the same handheld front-facing selfie camera perspective. As he rounds a corner, he reveals dropped swords and more broken props scattered near a lantern-lit landing, as if everything was abandoned during something violent. He looks back into the camera, then toward the scene, whispering, "And there's swords here too... it really looks like everyone disappeared after something happened..." Scene 5 12.0 to 15.0 He keeps moving carefully in selfie mode, backing or sidestepping slowly while dark staircases and lantern-lit voids loom behind him. His breathing is quiet but tense. He glances over his shoulder, then back into the lens and whispers, "I'm just trying to show you what happened here... this place feels wrong..."

    Expected result for generate the video.
    Use this source result as the visual comparison point for the step.
    Expected result for generate the video.
    Use this source result as the visual comparison point for the step.
    Expected result for generate the video.
    Use this source result as the visual comparison point for the step.
    Expected result for generate the video.
    Use this source result as the visual comparison point for the step.
    Expected result for generate the video.
    Use this source result as the visual comparison point for the step.
    Expected result for generate the video.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  17. 17

    Generate a reference image 10

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    The black hair man from the reference image and his cat 12-second vertical video, realistic iPhone front-facing selfie camera vlog quality, handheld at arm's length, slightly amateur framing, recording device never visible in frame, natural hand shake, soft autofocus breathing, low-light phone grain, casual off-center selfie composition, not cinematic, not polished, real-life look. A realistic Asian man is whispering the entire time while slowly exploring a creepy abandoned wooden hall from the real-life Demon Slayer fight location. The space is built over still dark water filled with lotus leaves and pale lotus flowers. Tall wooden beams rise into shadow, faint bluish light filters through mist, dust drifts in the air, and the whole place feels silent, eerie, abandoned, and haunted. Scene 1, 1.0 to 4.0. Front-facing selfie-vlog perspective. He walks slowly and carefully through the wooden walkway while keeping the front camera on himself. His face stays close in frame in a natural slightly off- center way, with the creepy hall, water, and lotus flowers behind him. He whispers to the camera, "This is one of the places where the fights happened... look at this place..." As he walks, he slightly turns his wrist so the viewer can glimpse more of the eerie hall behind him. His ragdoll cat are walking around the area, Scene 2, 5.0 to 8.0. Still in the same handheld front-facing selfie shot, he keeps walking slowly and whispering while lowering the framing slightly so more of the narrow wooden walkway, still water, lotus leaves, drifting dust, and cat near his legs can be seen while part of his face remains in frame. The place feels completely silent except for soft footsteps He whispers, "It's so quiet here... there's dust everywhere..." The shot should stay natural and unrefined, like real iPhone selfie footage, with slight shake and imperfect framing. Scene 3, 9.0 to 12.0. He continues moving in the same selfie-vlog perspective and slowly turns the camera just enough to reveal faint blood stains on the dusty wooden floorboards near the walkway junction. One ragdoll cat pauses near the stain walks deeper into the haunted hall. He brings the frame slightly back toward his face, visibly uneasy, and whispers, "There's even blood stain on the floor..." The background shows empty wooden structures, still water, lotus flowers, floating dust, and dark silent space all around him. He keeps walking carefully as the clip ends. Ragdoll cat is a fluffy long-haired cat with cream-white fur, soft gray points on the ears and face, bright blue eyes, pink noses, and large plume-like tails. They move naturally around the walkway and meow softly. The entire video must remain realistic front-facing selfie camera footage with no phone visible in frame. Only one cat in the video

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  18. 18

    Generate a reference image 11

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    The black hair man from the reference image and his cat 15-second vertical video, realistic iPhone front-facing selfie camera vlog quality, handheld at arm's length, slightly amateur framing, recording device never visible in frame, no phone visible, no selfie stick visible, no camera visible, natural hand shake, soft autofocus breathing, low-light phone grain, slight motion blur, uneven exposure, casual off-center selfie composition, not cinematic, not polished, real-life look. A realistic Asian man is whispering the entire time while carefully exploring the real-life Demon Slayer Infinity Castle. The whole place is eerie, unstable, and terrifying. Endless wooden buildings, bridges, corridors, and stairways are constantly sliding and shifting like they are mounted on a giant rotary belt machine. The buildings never stop moving. The bridges slowly slide past each other. The floor beneath him is unstable, causing him to wobble and struggle to keep his balance while filming himself. The environment is filled with warm lantern light, drifting dust, mist, deep shadows, and impossible architecture moving in every direction. He hears a distant low growl echoing from somewhere in the castle. He is frightened and whispers to the camera that he does not understand why the place is moving and that he needs to get out soon. No cats visible anywhere in this scene. No animals in frame. No extra characters. Scene 1, 1.0 to 5.0. Front-facing selfie-vlog perspective. He is already walking carefully on an unstable wooden platform while filming himself from the off-screen front camera. His face is close in frame, slightly off-center, and the moving Infinity Castle fills the background behind him. The floor shifts beneath his feet, making him stumble slightly and shake as he tries to stay balanced. He nervously turns his wrist just enough to show buildings and bridges slowly sliding past in the background, as if the entire place is running on a giant mechanical belt. He whispers, "I don't understand why this place is moving..." His breathing is tense and quiet. Scene 2, 6.0 to 10.0. Still in the same front-facing handheld selfie shot, he keeps walking while trying to steady himself. The wooden floor keeps sliding under him, forcing him to adjust his footing and grab his balance. He slightly pans the front-facing view to reveal more of the haunted space around him, with lantern- lit buildings drifting sideways, bridges shifting out of place, and stairways slowly rotating away into the fog. The movement of the castle feels unnatural and continuous. He hears a distant growl somewhere far away and immediately freezes for a moment, eyes widening. Then he whispers, "Did you hear that...?" The whole place feels alive and hostile. Scene 3, 11.0 to 15.0. He starts moving again, shakier now, still in the same front-facing selfie-vlog perspective. The unstable floor makes his handheld framing even more nervous and imperfect as he tries to walk faster without falling. He glances over his shoulder, then slightly angles the camera so the viewer can see more buildings and bridges still sliding endlessly through the lantern-lit darkness behind him. Dust drifts through the air and the distant growl echoes again. He looks back into the camera, visibly scared, and whispers, "I need to get out of this place soon... it's scaring me..." He keeps trying to balance himself as the moving castle continues shifting all around him. The entire video must stay realistic, grounded, and unsettling, like true iPhone front-facing selfie footage recorded by one person trapped inside a moving haunted structure. The recording device must never be visible in frame. No cats visible anywhere in this video.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  19. 19

    Generate a reference image 12

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    The black hair man from the reference image and his cat 6-second vertical video, realistic iPhone front-facing selfie camera vlog quality, handheld at arm's length, slightly amateur framing, recording device never visible in frame, no phone visible, no selfie stick visible, no camera visible, natural hand shake, strong motion shake from running, soft autofocus breathing, low-light phone grain, slight motion blur, uneven exposure, casual off-center selfie composition, not cinematic, not polished, real-life look. A realistic Asian man is whispering in panic while running and trying to escape the real-life Demon Slayer Infinity Castle. The entire environment is moving and sliding around him the whole time. The buildings are moving, the bridges are moving, the corridors are moving, the stairways are moving, and the floor beneath him is also shifting and sliding like a giant mechanical maze that never stops. The whole place feels alive, unstable, and impossible to escape. He is struggling to run because the moving floor keeps throwing off his balance. He looks genuinely terrified and keeps glancing around while the entire castle shifts behind him. He is carrying exactly one single ragdoll cat in one arm while running. Only one cat visible in the entire video. No extra cats, no duplicate cats, no other animals. The cat is a fluffy long-haired ragdoll cat with a cream-white coat, soft gray points on the ears and face, bright blue eyes, a pink nose, and a large plume-like tail. The cat is clearly visible, held tightly against his chest, and it meows anxiously while he runs. Front-facing selfie-vlog perspective the whole time. His scared face fills most of the frame while the moving castle slides behind him. The entire environment is visibly shifting and sliding in the background, with buildings drifting past, bridges sliding out of place, corridors moving, and the floor shaking under his feet. The camera shakes heavily from his footsteps and from the unstable moving ground. He is breathing hard, whispering in fear, "I'm trying to get out... this place won't stop moving..." He nearly loses balance as a bridge slides past behind him and the floor shifts again under his feet. He clutches the cat tighter, keeps running, and the cat lets out a frightened meow. He looks back into the camera, visibly panicked, and whispers, "I'm scared... I need to get out now..." The entire scene must feel urgent, chaotic, realistic, and trapped inside a haunted moving structure where the whole environment, all buildings, and all bridges are constantly moving and sliding.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  20. 20

    Generate a reference image 13

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Generate a reference image prompt
    The black hair man from the reference image and his cat 6-second vertical video, realistic iPhone front-facing selfie camera vlog quality, handheld at arm's length, slightly amateur framing, recording device never visible in frame, no phone visible, no selfie stick visible, no camera visible, natural hand shake, strong motion shake from running, heavy shake from collapsing ground, soft autofocus breathing, low-light phone grain, slight motion blur, uneven exposure, casual off-center selfie composition, not cinematic, not polished, real-life look. A realistic Asian man continues running in panic through the real-life Demon Slayer Infinity Castle while carrying exactly one single ragdoll cat tightly against his chest. Only one cat visible in the entire video, no extra cats, no duplicate cats, no other animals. The cat is a fluffy long-haired ragdoll cat with a cream-white coat, soft gray points on the ears and face, bright blue eyes, a pink nose, and a large plume-like tail, and the cat meows anxiously while he runs. The entire environment is moving and sliding around him the whole time. The buildings are moving, the bridges are moving, the corridors are moving, the stairways are moving, and now the whole place also begins to crumble and shake violently. Wooden beams crack, platforms break apart, dust bursts into the air, lanterns swing wildly, debris falls, and sections of bridges collapse or split while the floor beneath him jolts and shifts under every step. Front-facing selfie-vlog perspective the whole time, his terrified face filling most of the frame while the collapsing castle slides and crumbles behind him. He struggles to keep his balance, almost falling as the ground bucks beneath him, clutching the cat tighter while breathing hard and whispering in panic, "It's collapsing... the whole place is falling apart..." He stumbles forward as the space shakes harder, the cat lets out a frightened meow, and he looks into the camera in pure fear and whispers, "I need to get out now..." The whole scene must feel urgent, chaotic, realistic, and trapped inside a violently collapsing haunted structure where the entire environment, all buildings, and all bridges are moving, sliding, shaking, and breaking apart.

    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Expected result for generate a reference image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  21. 21

    Set the edit rhythm

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Raw AI footage is your base. Post-production is where it becomes cinematic.
    2. Temporal Remapping (Speed Ramps):
    3. Use speed curves to create rhythm and emphasize key moments.
    Source example for generate a reference image.
    Use this visual to confirm the expected setup or result.
    Source example for generate a reference image.
    Use this visual to confirm the expected setup or result.
    Source example for generate a reference image.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  22. 22

    Grade the colour

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Visual Enhancement: Glow
    2. Apply a "Soft Glow" or "Edge Glow" effect to simulate high-end lens diffusion and luxury aesthetics.
    3. Color Grading: The Final Look
    4. Final color adjustments allow you to refine the mood beyond the AI's base
    Expected result for set the edit rhythm.
    Use this source result as the visual comparison point for the step.
    Expected result for set the edit rhythm.
    Use this source result as the visual comparison point for the step.
    Expected result for set the edit rhythm.
    Use this source result as the visual comparison point for the step.
    Expected result for set the edit rhythm.
    Use this source result as the visual comparison point for the step.
    Expected result for set the edit rhythm.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  23. 23

    Polish and export the final edit

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Polish is what makes it feel expensive. Fine-tuning is what maximizes impact and gives your commercial a premium finish.
    Source example for grade the colour.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.

Before you use it

Final checks

  • Check the full result for changed people, products, text, timing, or layout.
  • Confirm the tool, model, controls, and cost before generating.
  • Keep private or unreleased material out of third-party tools.

Share your work

Made it? Show me.

Post your finished result on Instagram and tag @terencesia. I'll share a few standout results with the community.

Tag @terencesia on Instagram (opens in a new tab)

Next step

Run one low-cost test, record the first visible failure, and revise only the instruction or source asset tied to that problem.

Back to the tutorial library

Sources

What this guide relies on