AI Video Prompts, Motion Budgets, and Clip Costs
AI video prompts fail on motion, not description. Vendor duration limits side by side, the cost of a kept clip, and the one-action rule that fixes most of it.

The front page for ai video prompts is libraries again. One promises new prompts daily, one has eighty of them for content creators, a marketplace advertises eight thousand, and the only page in the set that explains anything is Adobe's own documentation, which is a vendor teaching you to use their tool. There's a Reddit thread in there too, from someone describing how they engineered one good prompt, and it's the most useful result on the page.
So, the direct answer. An ai video prompt is an image prompt plus a motion budget, and the motion budget is tiny. You get one action, one camera move, and somewhere between four and twenty seconds depending on whose model you're using. Everything that goes wrong past that point goes wrong because you asked for a second thing to happen. Describe the frame the way you'd describe a photograph, then add exactly one change over time, then stop writing.
That's the whole rule and I still break it most weeks.
The Part That Fails Is Never The Description
Watch what happens when a clip disappoints you. The subject is usually right. The setting is right, the lighting is more or less what you asked for, the colour treatment came through. Then somewhere around the middle a hand becomes something that isn't a hand, or a person walks and their legs swap, or an object in the background quietly changes shape while you weren't looking at it.
None of that is a description failure. The still-image half of the prompt worked.
What broke is the part where the model had to keep every one of those things consistent across dozens of frames while also making something move, and the more things you gave it to track, the more likely one of them drifts. Which means the lever you have isn't better adjectives. It's fewer moving parts.
I'd put it this way. Every element in the frame that can plausibly move is a liability, and you're paying for each one whether you asked it to move or not.
The Two Vendors Publish Different Boxes
Both of the big documented options tell you exactly what shape a clip can be, and the shapes are not the same, which matters more than any wording advice because it determines what you can even ask for.
| Constraint | Google Veo 3.1 | OpenAI Sora 2 and Sora 2 Pro |
|---|---|---|
| Clip durations | 4, 6, or 8 seconds | 16 and 20 second generations |
| Longer durations required for | 8 seconds needed for 1080p, 4k, or reference images | 1080p exports need sora-2-pro |
| Resolutions | 720p default, 1080p, 4k | 1280x720, 1920x1080, 1080x1920 |
| Aspect ratios | 16:9 default and 9:16 | 16:9 and 9:16 sizes listed |
| Extending a clip | Video extension listed as a feature | Add up to 20 seconds, six times, 120 seconds maximum total |
| Named prompt fields | Subject, action, style, camera positioning, composition, focus and lens effects, ambiance | Shot type, subject, action, setting, lighting |
| Content limits published | Standard policy | No real people including public figures, no copyrighted characters or music, content suitable for under 18 |
Both of those come off the vendors' own documentation pages, loaded today.
The row I'd stare at is the first one. Veo tops out at eight seconds per generation and Sora starts at sixteen, and that single difference changes how you write, because an eight-second box forces you into shot-thinking whether you like it or not while a twenty-second box tempts you into writing a scene with a beginning and a middle. The temptation is the trap.
Second row worth noticing is the prompt-field list. Google names seven fields and OpenAI names five, and the overlap is subject, action, and lighting or ambiance. The two fields Google adds that OpenAI's short guidance leaves out are camera positioning and focus, which happen to be the two most reliable controls in my own use, so I write them into Sora prompts anyway and they land fine.
One Action, And Then Stop
Here is the same intent written two ways. The brief was a shot for a book trailer, a hallway, somebody's shadow arriving before they do.
The version I sent first:
A dark hallway at night. A shadow stretches across the floor as someone approaches from around the corner, then a door at the end of the hall slowly opens and light spills out.
Two actions. A shadow moving and a door opening. What came back had the shadow doing something roughly correct for about two seconds and then the door situation resolving itself in a way that involved the wall bending. I ran it three times to be sure it wasn't a seed problem. Three for three.
The rewrite:
Locked-off shot down a dark, narrow hallway at night. A single warm practical light at the far end. A long human shadow slides slowly across the floor from the left. Nothing else in the frame moves. Handheld drift, almost imperceptible. Four seconds.
One action, one very small camera behaviour, and an explicit statement about what stays still. That last sentence does more than it looks like it does. Saying "nothing else in the frame moves" seems redundant, since you didn't ask for anything else to move, but in practice the model treats an unmentioned background as a place where it's allowed to invent, and naming the stillness closes that door.
The door opening became its own clip. That's the trade, and it isn't a compromise, it's how shots have always worked.
Camera Moves Land, Subject Moves Gamble
This is my own ranking rather than anybody's benchmark, built from something like forty clips across a couple of models while I was making trailer material. Small sample. I'd expect the ordering to hold and the details to wobble.
| Motion instruction | How often it does what I asked | Notes |
|---|---|---|
| Locked off, static camera | Nearly always | The most reliable thing in video prompting, and it's free |
| Slow push in | Usually | Speed is a guess, so say "very slow" if you mean it |
| Handheld, slight drift | Usually | Adds life without asking the model to track anything new |
| Slow pan left or right | Often | New scenery enters frame, which is where invention creeps in |
| Orbit or arc around subject | Sometimes | Asks for a consistent subject from angles it never saw |
| Subject walks across frame | Sometimes | Legs and gait are still the weak point |
| Subject picks something up | Rarely clean | Hands plus contact plus object permanence, all at once |
| Two subjects interacting | Rarely | I've mostly stopped trying |
| Anything with visible hands doing a task | Rarely | Cut away before the hands do the work, it's an old film trick anyway |
The pattern is that instructions about the camera are cheap and instructions about the subject are expensive, because the camera is one coherent thing and a subject is a hundred parts that all have to stay themselves. Google's own field list points the same direction, since camera positioning and composition get their own named slots while action is a single line.
If you want the equivalent breakdown for stills, where none of this applies and lighting carries the weight instead, that's in ai image prompt.
What A Clip Actually Costs You
Nobody in that SERP does this arithmetic, which is odd, because the arithmetic decides your whole workflow.
Published per-second rates, then multiplied out. My keep rate on video sits at roughly one in four across the clips I've run, which is worse than my keep rate on stills, so the last column assumes four generations per usable clip.
| Route | Per second | One clip | Cost per kept clip at 1 in 4 |
|---|---|---|---|
| Veo 3.1 Lite, 720p, 8s | $0.05 | $0.40 | $1.60 |
| Veo 3.1 Fast, 720p, 8s | $0.10 | $0.80 | $3.20 |
| Veo 3.1 Fast, 1080p, 8s | $0.12 | $0.96 | $3.84 |
| Veo 3.1 standard, 1080p, 8s | $0.40 | $3.20 | $12.80 |
| Veo 3.1 standard, 4k, 8s | $0.60 | $4.80 | $19.20 |
| Sora 2, 720p, 16s | $0.10 | $1.60 | $6.40 |
| Sora 2 Pro, 720p, 16s | $0.30 | $4.80 | $19.20 |
| Sora 2 Pro, 1080p, 16s | $0.70 | $11.20 | $44.80 |
Now scale it. A one-minute cut made of eight-second shots needs eight kept clips, so at one in four that's thirty-two generations. On Veo 3.1 Lite at 720p you're looking at about $12.80 for the whole minute. On Veo 3.1 standard at 1080p, the same minute is $102.40. Same prompts, same number of attempts, eight times the bill.
And the extension ceiling is its own line item. Sora documents a maximum total of 120 seconds through six extensions, and at the Pro 1080p rate that's $84 of billed seconds for a single two-minute video before you've rerolled anything at all.
Which produces the workflow that almost nobody follows, though it's the same one that works for stills. Draft at the cheapest tier that exists, settle the wording there, and only spend the top tier on a prompt that has stopped changing. The composition and the timing are readable at 720p. What the expensive tier buys is fidelity, not a different clip.
I generate a fair amount locally on an M4 Pro, which makes my marginal cost basically electricity and makes me an unreliable narrator about restraint. When the meter is running, people stop testing and start reasoning from the last thing that happened to work. That's how prompt folklore gets made.
Write The Still First
The habit that improved my clips most had nothing to do with video wording.
Generate the frame as a still image. Get it right there, where an attempt costs cents instead of dollars and comes back in seconds instead of minutes, then use that image as the first frame and write a motion instruction on top of it. Both vendors support starting from an image, Sora documents input references as first frames explicitly, and Veo lists image-based direction among its features.
Doing it that way splits one hard problem into two easy ones. The look is settled before any motion exists, so a bad clip is now unambiguously a motion problem, which is a thing you can fix in one line. Compare that to rerolling a text-to-video prompt where the light changed and the subject changed and the movement changed all at once and you have no idea which of your words did it.
The other benefit is that stills are where you can afford to be systematic, and video is where you can't. Everything I know about lighting phrases I learned from image generations I could throw away by the dozen. None of it came from video.
There's a fuller treatment of how the three prompt types diverge in ai prompts, and the subtraction side, which behaves differently again in video, is in negative prompts.
Questions People Ask
How long should an ai video prompt be? Shorter than an image prompt, which surprises people. I aim for the visual description at maybe five or six clauses, then one motion sentence, then one stillness sentence. Past that the descriptors start competing for a model that's already juggling time.
Does saying the duration in the prompt do anything? I think so, mildly, on pacing rather than on actual length, since length is set by the parameter and not by your text. That's a hunch from a handful of pairs, not a finding, and I'd want a lot more runs before I'd defend it.
Can I get a consistent character across clips? Partly, and not through wording. Sora documents reusable characters and says they work best with short two to four second clips at 720p to 1080p, which tells you the feature has a comfortable range and it's narrow. Text alone will not hold a face.
Why do the vendors publish different prompt fields? Because they're describing their own models and the fields reflect what each one responds to. The honest read is that the union of the two lists is a decent checklist, and the intersection, which is subject plus action plus light, is the part you can't skip.
Should I write negatives for video? Some systems take them, some don't, and the ones that do want them short. "Nothing else in the frame moves" has done more for me than any list of things to avoid.
Is Sora or Veo better for prompting? Wrong question for me to answer, honestly. They have different clip shapes, and the shape decides what kind of thing you can make. Eight seconds is a shot. Twenty seconds is a scene, and scenes are where models still fall over. There's a Sora-specific breakdown in sora prompts.
If You Change One Habit
Add the sentence that says what does not move.
It's the cheapest edit in this whole article, it costs six words, and it removes the class of failure where the background quietly reorganises itself while you were watching the subject. Everything else here is refinement. The wider framework for how prompts differ across text, images, and video sits in chatgpt prompts.


