Video & audio

Change a talking-head scene with Gemini Omni

Prepare the source assets, follow the supplied generation workflow, and inspect the result frame by frame before exporting.

Adapted from TUT-032, dated 2026-06-25.

What you'll make

A checked final video that follows the source workflow and is ready for export.

Change a talking-head scene with Gemini Omni

Who this is for

Creators, marketers, founders, and small teams producing short AI-assisted videos from source media.

Step by step

Run the workflow

  1. 01

    Prepare the source material

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Record a talking video of yourself.
    Source example for change a talking-head scene with gemini omni.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  2. 02

    Copy the generated prompt for Phase 3

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Personal Creative Director
    2. Drop my prompt (Next Page) into ChatGPT and modify the timing and the changes to your desired scenes. Copy the generated prompt for Phase 3.
    Prompt setup used for prepare the source material.
    Check the prompt field, reference order, and visible settings.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  3. 03

    Start the video with the original clip with no changes

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Copy the prompt below and replace any project-specific subject, product, location, or format details.
    2. Attach the source assets in the same order described by the prompt, then run one controlled test.
    Working asset
    Start the video with the original clip with no changes prompt
    Start the video with the original clip with no changes. The face and hair must look original.
    
    At [x] seconds, after he snap his finger, transform the hairstyle into a nice looking korean parting hair.
    
    At [x] seconds, transform the clothing into a luxury tailored black suit.
    
    At [x] seconds, after he snap his finger, replace the background into a nice looking bar that matches the video lighting.
    
    * Preserve the exact same person. * Keep facial features unchanged. Face must look exactly like the same person * Maintain facial proportions, skin texture, age, ethnicity, and expressions. * Preserve eye shape, nose, mouth, jawline, and hairstyle transition realism. * Keep lip sync perfectly aligned with the original speech. * Do not modify the voice, dialogue, pacing, or performance.
    
    * Match lighting between subject and environment. * Maintain realistic shadows and reflections. * Preserve camera angle, framing, movement, and depth of field. * Keep all non-specified elements unchanged. * Photorealistic result with seamless transitions and no identity drift. * Do not change the framing throughout the video
    
    Generate a Gemini optimized video to video edit prompt for me based on these details.

    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.
  4. 04

    Assemble the video

    Use the source inputs to complete this stage and produce a result that can be checked before continuing.

    1. Now Edit your video with your newly generated image.
    2. With your assets ready, it's time to edit your footages!
    Prompt setup used for start the video with the original clip with no changes.
    Check the prompt field, reference order, and visible settings.
    Prompt setup used for start the video with the original clip with no changes.
    Check the prompt field, reference order, and visible settings.
    Prompt setup used for start the video with the original clip with no changes.
    Check the prompt field, reference order, and visible settings.
    Prompt setup used for start the video with the original clip with no changes.
    Check the prompt field, reference order, and visible settings.
    Common failure modes
    • The generated scene changes the person, camera, room, product, or timing beyond the requested effect.
    • Fast motion hides face, hand, clothing, text, or continuity errors that become visible when the clip is paused.

Before you use it

Final checks

  • Check the full result for changed people, products, text, timing, or layout.
  • Confirm the tool, model, controls, and cost before generating.
  • Keep private or unreleased material out of third-party tools.

Share your work

Made it? Show me.

Post your finished result on Instagram and tag @terencesia. I'll share a few standout results with the community.

Tag @terencesia on Instagram (opens in a new tab)

Next step

Run one low-cost test, record the first visible failure, and revise only the instruction or source asset tied to that problem.

Back to the tutorial library

Sources

What this guide relies on