Video & audio

Create a five-look AI wardrobe-transition reel

Build five consistent outfit frames, then use a motion reference to turn them into one fast vertical transition sequence.

Adapted from TUT-058, dated 8 Aug 2026. This web edition uses Seedream 5.0 Pro and Seedance 2.5.

What you'll make

A reviewed vertical reel with one consistent identity, five distinct outfits, and a 7.4-second transition sequence.

Recreate a timed wardrobe-transition effect with one consistent identity and five controlled looks.

Who this is for

Creators and small marketing teams who want a repeatable multi-reference fashion or outfit-transition workflow.

Step by step

Run the workflow

  1. 01

    Map the finished sequence

    Understand what each reference controls before generating any new media.

    1. Watch the supplied motion-reference video from beginning to end before opening a generation tool.
    2. Write down the five looks in order: cream pajamas, green outfit with grey vest, orange pullover, grey-and-white stripes, then white star T-shirt and blue shorts.
    3. Separate identity from performance: your identity photo controls the person's face, each outfit image controls styling and pose, and the video controls movement, camera, timing, overlays, and transitions.
    4. Plan for a 9:16 vertical output. The choreography ends at 7.40 seconds inside a 9-second generation range.
    Source example for review the workflow.
    Use this visual to confirm the expected setup or result.
    Source example for review the workflow.
    Use this visual to confirm the expected setup or result.
    Source example for review the workflow.
    Use this visual to confirm the expected setup or result.
    Source example for review the workflow.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • The reference video is treated as an identity source, so the output copies the wrong person.
    • The outfit order is not recorded, causing image numbers and timeline instructions to drift apart.
  2. 02

    Prepare the identity photo and image workspace

    Create one stable identity source that will be reused across all five looks.

    1. Choose a sharp, front-facing identity photo with visible facial features, neutral light, and no beauty filter.
    2. Open Higgsfield, choose Create image from the top navigation, and select Seedream 5.0 Pro.
    3. Upload the identity photo first. Keep it in the same Image 1 position for every look so the prompts remain consistent.
    4. Check the current model label and displayed generation cost before continuing.
    Source example for generate the source images.
    Use this visual to confirm the expected setup or result.
    Source example for generate the source images.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • A filtered, angled, cropped, or low-resolution portrait gives the model an unstable identity reference.
    • The identity photo changes between looks, creating five similar but visibly different people.
  3. 03

    Generate the cream-pajama opening frame

    Create the neutral opening look while preserving the same identity.

    1. Download the cream-pajama reference and add it as Image 2 beside the identity photo.
    2. Paste the Image 1 prompt below. Keep the tagged references in the correct order: identity is Image 1 and outfit reference is Image 2.
    3. Generate the smallest useful test set, then compare the face, hairstyle, framing, pajamas, room, and neutral pose with both inputs.
    4. Save the best result as 01-cream-opening so it remains first in the video stage.
    Working asset
    Image 1 - cream-pajama opening frame
    Replace the man in @[Image 2](image_2) with the man from @[Image 1](image_1).
    
    Copy the expression and posture from @[Image 2](image_2).
    
    Do not change the background or outfit from @[Image 2](image_2).

    Expected result for first image.
    Use this source result as the visual comparison point for the step.
    Expected result for first image.
    Use this source result as the visual comparison point for the step.
    Expected result for first image.
    Use this source result as the visual comparison point for the step.
    Expected result for first image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • Image 1 and Image 2 are reversed, so the tool preserves the wrong identity.
    • The result changes the room, adds text, or blends the source person's face with the reference person.
  4. 04

    Generate the green-vest frame

    Create the first transition look with both palms raised beside the face.

    1. Keep the same identity in Image 1 and replace only Image 2 with the green-vest reference.
    2. Paste the Image 2 prompt below and confirm that it asks for the face replacement while preserving the source framing, posture, expression, clothing, and room.
    3. Generate a test and inspect both open hands, all fingers, facial proportions, earrings, vest edges, and trouser shape.
    4. Save the best result as 02-green-vest.
    Working asset
    Image 2 - green outfit with grey vest
    Replace only the man's face in @[Image 2](image_2) with the man from @[Image 1](image_1).
    
    Copy the framing, posture, and expression from @[Image 2](image_2).
    
    Do not change anything else from @[Image 2](image_2).

    Expected result for second image.
    Use this source result as the visual comparison point for the step.
    Expected result for second image.
    Use this source result as the visual comparison point for the step.
    Expected result for second image.
    Use this source result as the visual comparison point for the step.
    Expected result for second image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • Raised hands produce missing, fused, or duplicated fingers that become more obvious during the close-up.
    • The vest, sleeves, or trousers blend with the opening outfit instead of forming one clean second look.
  5. 05

    Generate the orange-pullover frame

    Create the cheek-pointing pose used after the second spin transition.

    1. Keep the same identity as Image 1 and load the orange-pullover reference as Image 2.
    2. Paste the Image 3 prompt below, then generate one controlled test set.
    3. Check that one index finger touches the cheek, the opposite hand rests on the hip, and the subject stays front-facing.
    4. Inspect the pullover collar, sleeves, dark trousers, face, and room, then save the best frame as 03-orange-pullover.
    Working asset
    Image 3 - orange-pullover look
    Replace the man in @[Image 2](image_2) with the man from @[Image 1](image_1).
    
    Copy the framing, posture, and expression from @[Image 2](image_2).
    
    Do not change anything else from @[Image 2](image_2).

    Expected result for third image.
    Use this source result as the visual comparison point for the step.
    Expected result for third image.
    Use this source result as the visual comparison point for the step.
    Expected result for third image.
    Use this source result as the visual comparison point for the step.
    Expected result for third image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The pointing finger merges with the cheek or the hand-on-hip arm bends unnaturally.
    • The expression or camera distance differs too much from the motion reference, producing a jump in the final reel.
  6. 06

    Generate the striped-outfit frame

    Create the inward-pointing hand pose used near the end of the sequence.

    1. Keep the identity photo in Image 1 and load the striped-outfit reference as Image 2.
    2. Paste the Image 4 prompt below and generate a controlled test.
    3. Check both hands at chest height: the index fingers should point inward and nearly meet without merging.
    4. Inspect stripe continuity across the torso and sleeves, then save the best result as 04-striped-outfit.
    Working asset
    Image 4 - striped-outfit look
    Replace the man in @[Image 2](image_2) with the man from @[Image 1](image_1).
    
    Copy the framing, posture, and expression from @[Image 2](image_2).
    
    Do not change anything else from @[Image 2](image_2).

    Expected result for fourth image.
    Use this source result as the visual comparison point for the step.
    Expected result for fourth image.
    Use this source result as the visual comparison point for the step.
    Expected result for fourth image.
    Use this source result as the visual comparison point for the step.
    Expected result for fourth image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The fingertips fuse, cross, or multiply at the centre of the pose.
    • Horizontal stripes warp around the hands or change direction across the body.
  7. 07

    Generate the white-star T-shirt frame

    Create the final peace-sign close-up that ends the transition sequence.

    1. Keep the same identity in Image 1 and load the white-star T-shirt reference as Image 2.
    2. Paste the Image 5 prompt below and generate one controlled test set.
    3. Check the raised V sign, slight smile, direct eye contact, star graphic, white shirt, and blue shorts.
    4. Save the best result as 05-white-star-final and place all five selected frames in one folder in numerical order.
    Working asset
    Image 5 - white-star T-shirt look
    Replace the man in @[Image 2](image_2) with the man from @[Image 1](image_1).
    
    Copy the framing, posture, and expression from @[Image 2](image_2).
    
    Do not change anything else from @[Image 2](image_2).

    Expected result for fifth image.
    Use this source result as the visual comparison point for the step.
    Expected result for fifth image.
    Use this source result as the visual comparison point for the step.
    Expected result for fifth image.
    Use this source result as the visual comparison point for the step.
    Expected result for fifth image.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • The final V sign contains extra fingers or an unnatural wrist angle.
    • The fifth image looks correct alone but reveals a different face, hairline, or body build when compared with the first four.
  8. 08

    Download and label the motion references

    Prepare the video and audio that control the transition's movement and rhythm.

    1. Download the motion-reference video and watch it with sound on and sound off.
    2. Download the supplied audio if you want to use it as a rhythm reference.
    3. Name the files Reference-Video-1 and Reference-Audio-1 so their role stays obvious after upload.
    4. Confirm that the reference video contains the swipe overlay, spin direction, pullbacks, push-ins, and final poses described in the prompt.
    Source example for download the reference video + audio too.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • The motion reference is treated as a final visual source, causing the original person or outfit to leak into the result.
    • The audio and motion reference have different timing, so the transitions fall out of sync.
  9. 09

    Open Seedance 2.5

    Move from image preparation into the current multi-reference video workflow.

    1. Return to Higgsfield and choose Video from the top navigation.
    2. Select Seedance 2.5 for the final multi-reference video stage.
    3. Check the current model availability, supported inputs, duration options, resolution, and displayed credit cost before uploading anything.
    Source example for generate the video.
    Use this visual to confirm the expected setup or result.
    Common failure modes
    • A different model is selected, but the supplied reference map and limits are assumed to work unchanged.
    • Generation begins before the account's credit cost, input limits, and output duration are checked.
  10. 10

    Upload the seven reference assets

    Give Seedance a complete, correctly ordered reference map.

    1. Click Upload Media and add the five selected images, the motion-reference video, and the optional audio if you are using it.
    2. Keep the image order exact: cream pajamas, green vest, orange pullover, striped outfit, then white star T-shirt.
    3. Assign the motion reference as Reference Video 1. Assign the audio separately rather than treating it as a visual reference.
    4. Open every thumbnail once and confirm that no draft, duplicate, wrong person, or unintended file entered the media library.
    Upload interface used for upload the reference assets.
    Check the selected files and their order before continuing.
    Upload interface used for upload the reference assets.
    Check the selected files and their order before continuing.
    Upload interface used for upload the reference assets.
    Check the selected files and their order before continuing.
    Common failure modes
    • The green and orange looks are swapped, so the prompt calls for the wrong outfit at each timestamp.
    • A rejected draft remains in the library and is selected instead of the final frame.
    • The audio is mistaken for the motion reference and assigned to the wrong input.
  11. 11

    Paste and verify the video prompt

    Map every selected reference to an exact pose, transition, and timestamp.

    1. Open SIA's supplied video prompt and paste it into the Seedance prompt field.
    2. Confirm that Reference Video 1 and Reference Images 1-5 point to the files described in the prompt's reference map.
    3. Read every timestamp once. The motion reference controls the swipe line, spins, zooms, timing, and transition rhythm; the five images control identity, clothing, and close-up poses.
    4. Remove or rewrite any text, music, motion, or styling instruction that does not belong in your version.
    5. Keep the strict negative rules. They are the acceptance boundary for face drift, extra limbs, outfit blending, background movement, and invented transitions.
    Working asset
    Five-look wardrobe-transition video prompt
    Use Reference Video 1 as the master for timing, camera movement, spin direction, zoom speed, overlay animation, framing, and transition rhythm. Recreate the 7.4-second vertical sequence as closely as possible, but replace the woman with the same man from Reference Images 1-5. The video controls motion and effects only. The images control identity, outfit, and close-up pose.
    
    REFERENCE MAP
    Reference Video 1 = exact 7.4-second motion and effect reference.
    Reference Image 1 = cream pajama opening look, neutral front-facing pose.
    Reference Image 2 = green outfit with grey knit vest, both open palms raised beside face.
    Reference Image 3 = orange pullover with dark trousers, one index finger touching cheek, opposite hand on hip.
    Reference Image 4 = grey-and-white striped outfit, both index fingers pointing inward with fingertips nearly touching.
    Reference Image 5 = white star T-shirt with blue shorts, final V or peace-sign pose beside face.
    
    Keep the studio identical to Reference Video 1: warm room, wooden floor, shelves, artwork, plants, red ladder, furniture, and white circular platform. Keep the man centered and photorealistic. Preserve his exact face, hairstyle, earrings, skin tone, and masculine proportions across every transformation.
    
    0.00-3.25s
    Start immediately in Reference Image 1's cream pajama outfit, medium framing, front-facing, and almost completely still. Reproduce the large "Wave your Finger" title and downward arrow in the same position and style. Copy the thick white hand-drawn finger or swipe line exactly from the video: it repeatedly enters from the right or lower edge, sweeps horizontally across the lower torso or waist area, loops back, exits, then repeats with the same direction changes, speed, curve, and timing. The man does not perform this gesture. Only the overlay moves; camera and subject stay locked.
    
    3.25-3.55s
    Begin the first transition exactly with the video. The man turns rapidly toward profile while the camera pulls sharply backward from medium framing to distant full body. Hide the clothing change inside the fast rotational blur, switching from Reference Image 1 to Reference Image 2. The title, arrow, and white swipe line disappear during this transition. Reveal the full white circular platform.
    
    3.55-3.90s
    Green outfit, full-body distant framing. Continue the same rotation on the platform, moving through front and side angles with the exact speed and direction from the reference.
    
    3.90-4.35s
    Rapid camera push-in while the man finishes turning toward camera. End in the Reference Image 2 close pose: both hands raised beside his face, palms forward, fingers spread. Hold the pose briefly with direct eye contact.
    
    4.35-4.60s
    Immediately rotate again and pull the camera back. Hide the second outfit swap inside the side or back-facing motion blur, changing to Reference Image 3.
    
    4.60-4.90s
    Orange outfit, distant full-body framing on the platform, briefly facing front before continuing the turn.
    
    4.90-5.45s
    Camera rapidly pushes in while he rotates back toward camera. End in Reference Image 3's close pose: one index finger touching his cheek, opposite hand on hip, confident front-facing expression. Hold briefly.
    
    5.45-5.70s
    Fast spin and camera pullback. Change to Reference Image 4 only while his body is turned away or in profile and blurred.
    
    5.70-6.05s
    Grey striped outfit, distant full-body framing, centered on the platform, continuing the same rotation.
    
    6.05-6.55s
    Fast push-in to medium-close framing. End in Reference Image 4's pose: both hands at chest height, index fingers pointing inward with fingertips almost touching. Hold the pose briefly.
    
    6.55-6.80s
    Final spin and pullback. Hide the outfit change in rotational blur, switching to Reference Image 5.
    
    6.80-7.10s
    White star T-shirt and blue shorts, distant full-body framing, centered on the platform. Continue the same turn through front to side.
    
    7.10-7.40s
    Final rapid push-in while he turns back toward camera. End close to camera in Reference Image 5's pose with one hand raised beside his face in a V or peace sign, slight natural smile, and direct eye contact. End exactly at 7.40 seconds.
    
    STRICT RULES
    Reference Video 1 controls all movement and timing. Do not invent extra spins, poses, zooms, or transitions. Clothing changes happen only during fast side or back rotational blur. Keep one man only. No face morphing, feminisation, mannequin features, extra limbs, distorted fingers, outfit blending, background movement, flashes, particles, hard cuts, or slow motion. Preserve realistic fabric movement, natural motion blur, and the original effect pacing.

    A detailed multi-reference wardrobe-transition prompt entered in Higgsfield.
    Paste the full prompt only after the media numbers are correct. A precise timeline cannot fix a mislabelled reference.
    Common failure modes
    • The prompt is pasted before the reference order is checked, producing the right motion with the wrong outfits.
    • A timeline section is shortened or removed, causing the spin, swap, close-up pose, and next transition to overlap.
    • The copied prompt retains text or audio that does not belong in the new project.
  12. 12

    Configure, generate, and inspect

    Run one controlled test, then compare the identity and motion frame by frame.

    1. Set the output to 9:16 vertical and select the 9-second range. Choose a practical test resolution before spending on a higher-quality rerun.
    2. Check the displayed model, duration, resolution, audio option, and credit cost, then generate one result.
    3. Watch once at normal speed for rhythm. Watch again at reduced speed and pause before, during, and after every outfit swap.
    4. Compare the output with all five selected images and the motion reference. Reject identity drift, incorrect outfit order, broken hands, outfit blending, camera changes, background changes, or invented transitions.
    5. The choreography ends at 7.40 seconds. Inspect any trailing frames inside the 9-second output and trim them in your video editor after the generation passes review.
    Expected result for choose the output settings.
    Use this source result as the visual comparison point for the step.
    Expected result for choose the output settings.
    Use this source result as the visual comparison point for the step.
    Common failure modes
    • Fast rotational blur hides face, limb, or clothing defects that become visible when the clip is paused.
    • The model invents extra spins, zooms, flashes, particles, or slow motion instead of following the motion reference.
    • A polished result still contains the wrong person, room, text overlay, audio, or music.

Before you use it

Final checks

  • Reference-based generation can still change facial features, body proportions, clothing details, and room geometry between frames.
  • Review every frame for face drift, broken hands, outfit blending, altered text, background changes, and misleading edits.
  • Tool access, model names, input limits, credits, and retention terms can change. Check Higgsfield before uploading or generating.

Share your work

Built the wardrobe transition? Show me.

Post your finished reel on Instagram and tag @terencesia. I'll share a few standout results with the community.

Tag @terencesia on Instagram (opens in a new tab)

Next step

Run one test with your identity photo. Record every visible failure by timestamp, then revise only the reference or prompt line tied to that failure.

Back to the tutorial library

Sources

What this guide relies on