AI video tools reward the same thing image tools do: description. The difference is that video adds time — motion, pacing, and what changes between the first second and the last.

Here is how to write a prompt a video model can actually follow.

The blocks of a video prompt

  1. Subject and action — who or what, doing exactly what.
  2. Camera — static, slow push in, tracking shot, handheld, overhead drone.
  3. Setting and light — location, time of day, weather, light source.
  4. Style — cinematic film look, documentary handheld, clean product render, animation style.
  5. Pacing — what happens at the start, middle and end of the clip.
  6. Constraints — no text, no camera shake, single continuous shot, aspect ratio.

Before and after

Before: "A person working on a laptop."

After: "Slow push-in on a woman in her late thirties working at a kitchen table, morning light from a window on the left, she pauses, smiles slightly, and keeps typing. Warm cinematic color, shallow depth of field, single continuous shot, no camera shake, no text, 16:9."

Keep one clip to one idea

The most common mistake is asking for a story in a five-second clip. Video models handle one action well and three actions badly. Write several short prompts and edit them together rather than trying to get a scene out of a single generation.

Motion words that work

  • Camera: "slow dolly in," "static locked-off shot," "tracking alongside," "gentle handheld drift"
  • Subject: "turns toward the camera," "sets the cup down," "walks out of frame to the right"
  • Environment: "steam rising," "curtains moving slightly," "rain on the window"

Practical prompts

Product b-roll: "Static macro shot of [product] on a matte surface, slow rotation, soft studio lighting, subtle reflection, clean background, no text, no hands."

Social hook clip: "Handheld shot following a [person] walking into [place], natural light, documentary style, subject glances at camera once, 9:16 vertical, no text overlay."

Explainer background: "Abstract slow-moving gradient in black and cyan, subtle particle motion, seamless loop, no text, 16:9."

Fix one variable at a time

If the motion is wrong, change only the camera line. If the mood is wrong, change only the lighting. Changing four things at once makes the result feel random and teaches you nothing about what actually worked.

Plan the edit, not the miracle

Treat generated clips as raw footage. Write three to five short, well-described prompts, generate a few versions of each, and assemble the ones that landed. That workflow produces far better results than chasing one perfect generation.