Skip to content

AI content · September 2026

The hard part of AI video is the second shot

One clip is a solved problem. Five clips that read as one person in one room is not. Here is a thirty-second skincare ad, shot by shot, with every prompt and the two kinds of joins that hold it together.


Generating one good AI clip stopped being interesting a while ago. Give a model a still and a careful paragraph and you get fifteen usable seconds. That is a solved problem, and we wrote up how we do it in the last post.

The second shot is where it falls apart. Generate two clips independently and you get two slightly different women in two slightly different rooms, and the cut between them reads as a mistake. Everything below is about that cut. The ad is a thirty-second skincare piece: five shots, one face, one bathroom, one afternoon.

The finished cut

The finished cut
5 clips, 30 seconds generated, cut to 29. Subtitles and the trims are the only things added afterwards. Tap the speaker for the voice.

Nothing in it was filmed. There is no creator, no product on a shelf and no bathroom. The only things added after generation are the subtitles and the trims between shots.

The five shots

Each of these was a separate generation, made in a separate request, minutes apart. Watch them run straight through:

Shot 1 / 5

10 seconds · on camera

Introduction

She holds the capped bottle beside her face and says the opening line straight to the lens.

Do you like seeing the texture before trying a serum? Today, let’s take a look at Typology’s L thirty-two.

Started from a new stillThe only shot with nothing before it. Its first frame is a still built from the character reference and the product reference.

010s of 30s generated · 5 separate clips

The sequence works because of one rule: a shot never invents its own opening frame. Every shot is handed the exact image it must begin on, and the whole craft is in deciding where that image comes from.

Four joins, two kinds

Five shots means four joins. Two of them cost nothing and two of them are the actual work.

  1. The last frame of shot 1.
    Out of shot 1
    The first frame given to shot 2.
    Into shot 2

    Carried frame — free

    Nothing changes between the two shots except time, so the last frame of the first shot is the first frame of this one.

    Prompt file: “Start frame: a clean final frame from Scene 1.

  2. The last frame of shot 2.
    Out of shot 2
    The first frame given to shot 3.
    Into shot 3

    Bridge still — generated

    Between the two shots the cap comes off, the pipette goes on and the hand rises. No model animates that reliably, so the finished state was generated as a still instead.

    Prompt file: “Start frame: the approved close-up with the open bottle and pipette above her cheek.

  3. The last frame of shot 3.
    Out of shot 3
    The first frame given to shot 4.
    Into shot 4

    Bridge still — generated

    Same problem again: the pipette has to leave her hand and two fingers have to arrive on the same spot. The still does the swap; the shot only has to move fingers.

    Prompt file: “Start frame: the prepared massage image, with the pipette out of view and two fingertips already touching the serum trace. Do not use the pipette image for this shot.

  4. The last frame of shot 4.
    Out of shot 4
    The first frame given to shot 5.
    Into shot 5

    Carried frame — free

    Her hand is already on her cheek at the end of the previous shot, which is exactly where this one needs to begin.

    Prompt file: “Vertical 9:16 image-to-video continuation.

A carried frame is the easy kind. Nothing changes between the two shots except time, so you export the last frame of one and hand it to the next. Shot 2 picks up from shot 1 with the bottle still at her cheek. Shot 5 picks up from shot 4 with her fingers still on it. Both are free, and both are seamless because the two frames are literally the same picture.

A bridge still is the interesting kind. Between shot 2 and shot 3, a cap has to come off, a pipette has to go on, a hand has to rise to her face and a drop has to form on the tip. Ask a video model to perform that and you get melted fingers and two bottles. So we do not ask. We generate the result of that action as a still, and the next shot starts there.

  • The bridge frame: a pipette held above her cheek with one droplet at the tip.

    Shot 3 — the bridge to the drop

    Cap off, pipette on, droplet already hanging. The state the third shot has to start in.

  • The bridge frame: two fingertips resting on the serum trace, the pipette out of shot.

    Shot 4 — the bridge to the massage

    Pipette gone, two fingertips already touching the serum. The state the fourth shot has to start in.

A bridge still is a piece of continuity work that happens to look like an image prompt. It inherits everything from the frame before it — the same face, the same angle, the same window light, the same serum trace on the same cheek — and changes exactly one thing. Read the two prompts below and you will see most of their length is spent on what must not change.

The payoff is that each shot then has one job. Shot 3 releases a drop that is already hanging. Shot 4 moves two fingers that are already touching skin. Neither has to invent an object or a grip, which is where these models fail.

The voice ladder

Only the first shot has her speaking on camera. Shots 2 through 4 are voiceover over the same take, and shot 5 is silent.

That is deliberate. Lip-sync is the least stable thing a video model does, and it gets worse the longer it runs and the more the face moves. Spending it all in one ten-second shot, while the face is big, still and doing nothing else, is the safest place to put it. After that the voice keeps going and the picture is free to be hands and product, which is the half these models are good at. The last shot drops the voice entirely so the ad ends on a look rather than a line.

Where it breaks: small text

Here is the failure worth knowing about. This is the same bottle in shot 1 and in shot 2, five seconds apart in the finished cut.

The label in shot 1, small in frame, reading SÉRUM ÉCLAT, Complexe vitamine C.
Shot 1 · label small in frameSÉRUM ÉCLAT — Complexe vitamine C 15% + Vitamine E 1% — (Radiance Serum)
The same label in shot 2, filling the frame, now reading SÉRM RUM ÉCLAIRCHISSEUR.
Shot 2 · same bottle, five seconds laterSÉRM RUM ÉCLAIRCHISSEUR — RIE + VITAMINE C — + ACIDE FERULIQUE

The brand name and the product code survive. Everything else is rewritten, and rewritten into words that are not quite French. The model is not reading the label and reproducing it; it is drawing something label-shaped, and the closer the bottle gets to the lens the more of that shape it has to invent.

There is no prompt that fixes this. Both prompts say to preserve the reference wording and neither one is obeyed at close range. What you do instead is treat it as a shot decision: keep product close-ups short, keep the label small in frame, or composite the real label back over the clip in the edit. We left it in here because the post is more useful with the failure visible than without it.

Every prompt

The five video prompts, in order. Each one is the file as it was written, minus its header.

SCENE 1 — INTRODUCTION. Start frame: the approved bathroom portrait holding the capped serum bottle beside her face.

1. IntroductionImage-to-video · 9:16 · 10s
Animate the supplied image as a photorealistic vertical 9:16 iPhone UGC video. One continuous shot, fixed phone camera.

The woman looks into the lens and says in relaxed, conversational English:
“Do you like seeing the texture before trying a serum? Today, let’s take a look at Typology’s L thirty-two.”

Synchronize her lips naturally with the dialogue. Give the opening question a curious, friendly tone, followed by a small pause. Avoid an announcer voice or exaggerated enthusiasm.

She keeps the bottle beside her face, label facing the camera. Include subtle breathing, one natural blink and a small conversational head movement. Her fingers maintain their existing grip. End with her mouth relaxed and closed, holding the bottle steady.

Preserve the starting frame’s exact identity, natural facial asymmetry, skin texture, curls, charcoal T-shirt and bathroom. Preserve the bottle’s geometry, amber glass, black cap and label artwork.

Match the existing window daylight, facial shadows and white balance. Retain ordinary iPhone video detail and a readable background. No beauty filter, skin smoothing, artificial glow, exposure pulsing, camera movement, extra fingers, subtitles, graphics or music. Only her voice and faint indoor room tone.

And the four image prompts behind them. The first one is worth reading even if you skip the rest: it is an edit pass over a single portrait, run before anything else, and it is the reason the same face survives seven images and five clips.

Locks the face. An edit pass over one portrait, run before anything else, so every later image has the same person to copy.

Character revisionImage · Source ratio
Edit the attached portrait into a photorealistic beauty character reference. Keep the changes subtle and confined to the face.

Preserve her long dark-brown curls, hairline, hazel-green eyes, skin tone, apparent adult age and body proportions. Retain the original framing, raised-arm pose, white ribbed zip top, necklace and room background.

Gently soften the angular taper of her chin, give her cheeks slightly fuller contours, and make her eyebrows a little straighter while retaining their natural thickness and individual hairs. Keep her nose and lip proportions unchanged. Reduce the appearance of heavy mascara and lip liner for a lightly made-up skincare look. Maintain direct eye contact and a relaxed, closed-mouth expression.

Preserve natural facial asymmetry, fine skin texture, subtle under-eye detail and believable lip texture. Avoid smoothing the skin into a uniform surface or exaggerating pores.

Match the reference’s soft daylight, warm-neutral color and gentle facial shadows. Keep both eyes and the facial features clearly in focus, with the room softly readable behind her. Preserve individual curls and flyaway hairs.

Produce one edited portrait in the original aspect ratio. No product, added lettering, watermark or collage.

The same two references, a whole campaign

Everything so far came from two reference images: one portrait and one product shot. Once those exist, the vertical ad is only one of the things they can produce. Same references, three more prompts:

Studio packshot of the amber serum bottle on an olive-grey sweep.
Packshot · 4:5Studio product shot, no person.
The bottle on a stone ledge in a shaft of late sunlight, the upper half of the frame left empty.
Story ad · 9:16Vertical, with the top half left empty for copy.
Wide studio portrait: she holds the bottle beside her face on the right, with dark empty space on the left.
Campaign hero · 16:9Wide, with the left 40% reserved for a headline.

Studio product shot, no person.

PackshotImage · 4:5
Create a photorealistic vertical 4:5 advertising packshot of the exact Typology L32 serum bottle shown in the reference.

Preserve its rectangular amber-glass body, rounded shoulders, thick glass base, black cylindrical collar and rounded black dropper bulb. Reproduce the cream label’s layout and wording from the reference, including “L32” and “Typology.” Do not change the packaging, formula text or proportions.

Place the closed bottle upright on a matte, desaturated olive-grey seamless studio surface. Center it horizontally, occupying approximately two-thirds of the image height, with comfortable space above the dropper and below the base. The entire bottle is visible. Turn the body only slightly to reveal its right-hand depth while keeping the front label clearly readable.

The foreground is an uncluttered stretch of the same surface. The bottle is the sharp focal subject. Behind it, the seamless sweep falls gradually into deep olive-charcoal, with no visible horizon or props. Clear air.

Photograph at label height with a restrained 85mm-equivalent product perspective and straight verticals. Keep the complete bottle, from dropper bulb to glass base, sharply resolved.

One small studio light above and to camera-left creates a defined shadow extending toward the lower right. Let warm light reveal the amber glass edges and thick base. A passive white reflector on camera-right provides just enough detail in the dark collar without creating a second cast shadow. Keep bright reflections away from the label.

Render the glass with believable thickness and refraction, the collar with subtle surface texture, and the label with a matte paper finish. Preserve the label’s pale cream color and neutral black typography. Deep, detailed shadows and restrained saturation.

The bottle rests firmly on the surface with a clear contact shadow. No levitation, liquid splash, droplets on the packaging, decorative ingredients, additional products, headline text, border or watermark.

That is the part clients tend to under-estimate. The expensive step is locking a face and a product you can shoot repeatedly. After that, a packshot, a story frame and a hero banner are three more paragraphs, not three more shoot days.

What we would tell you before you start

  • Lock the character first, in its own pass, before any scene exists.
  • Never let a shot choose its own first frame.
  • When the physical state has to change between shots, generate the change as a still. Do not animate it.
  • Give each shot one job: one movement, one line, one focus change.
  • Spend your lip-sync early, while the face is still and close.
  • Keep small text away from the lens, or plan to composite it back in.

Closing

30 seconds of generated footage, five clips, seven images, one afternoon. The prompts above are the whole method, and they are yours. If you want a run of these for your own products — variations across shots, languages and formats — talk to us.

Talk to us

The NOQTA Team