Sora Prompts, Clip Length, and the Extend Trap
Sora prompts are shaped by a 16-second floor and a 120-second ceiling. The seam math, the batch discount nobody mentions, and what the policy makes you rewrite.

Page one for sora prompts opens with a Reddit post from somebody who ran 32 experiments and wrote up what worked, which is the best result on the page and also the one that will be stale in six months. Under it, a Medium piece about viral prompts, a site that exists to collect them, a gist, and two YouTube videos, one of which promises fifty prompts you're probably using wrong.
Here's what I'd want told to me first. Sora prompts are shaped less by wording than by two published numbers, a floor and a ceiling. The floor is that generations come in 16 and 20 second lengths, so the shortest thing you can make is already long enough to need internal structure. The ceiling is that a single video can be extended up to six times for a maximum total length of 120 seconds. Everything about how you write follows from those two facts, and almost nothing on that first page mentions either.
The wording guidance from OpenAI's own docs is short enough to quote in full, near enough. Describe shot type, subject, action, setting, and lighting.
Their Own Examples Are Doing Something Specific
Two example prompts appear in the video generation guide. I'd read them closely, because they're the model's makers showing you the shape they expect.
The first is "Wide shot of a child flying a red kite in a grassy park, golden hour sunlight". The second is "Close-up of a steaming coffee cup on a wooden table, morning light through blinds".
Count what's in there. One shot type, one subject, one implied ongoing action, one setting, one lighting phrase. No camera move at all in either. No second subject. No sequence of events, no "and then", no cut, and crucially no description of anything changing over the length of the clip other than the thing that was already happening when the clip started.
That last part is the design lesson. Both examples describe a continuous state rather than an event. A kite is being flown. Steam is rising. Neither prompt asks for something to begin or finish, which is exactly why they're the examples, and it's a much narrower template than the "cinematic scene" language you'll find in the prompt collections.
I'd start there and add exactly one thing at a time. The general framing for how motion prompts differ from still prompts is in ai video prompts.
Sixteen Seconds Is Longer Than It Sounds
Try a small experiment before you generate anything. Watch a film you like and count how long the shots are. Most of them will be under six seconds, plenty under three.
So a 16-second generation is not one shot in any normal sense. It's three or four shots' worth of screen time delivered as a single unbroken take, and you can't cut inside it, because the model produced it as one continuous piece. Which leaves you two options and it's worth deciding on purpose which one you're taking.
Option one is to write a genuine long take. Something with almost no cutting logic in it, a held frame with slow internal movement, the kind of shot that earns its length by being still. These work.
Option two is to generate long and trim in an editor afterwards, which means you're paying for 16 seconds and keeping maybe five. That's a normal thing to do and it's how the number in my cost table below gets ugly.
| Target output | If you use full generations | If you trim to the good 5 seconds |
|---|---|---|
| One 5-second shot | Not possible, 16s minimum | One 16s generation |
| A 30-second cut, six shots | Two generations, no cutting | Six generations, 96 billed seconds for 30 kept |
| A 60-second cut, twelve shots | Three or four generations | Twelve generations, 192 billed seconds for 60 kept |
The trim column is the honest one for anything with pacing. Three quarters of what you pay for goes in the bin, which is also true of live-action shooting, so it isn't outrageous. It just needs to be in the budget rather than a surprise.
The Extension Ceiling And What A Seam Costs
Extensions are the feature people get excited about and they carry a cost that isn't money.
The documented behaviour is that you can add up to 20 seconds at a time, a single video can be extended up to six times, and the maximum total length is 120 seconds. Six extensions means up to six joins, and every join is a place where the model had to pick up from a frame and continue, which is precisely the operation it finds hardest.
| Finished length | Extensions used | Seams in the result | Billed seconds |
|---|---|---|---|
| 20 seconds | 0 | 0 | 20 |
| 40 seconds | 1 | 1 | 40 |
| 60 seconds | 2 | 2 | 60 |
| 100 seconds | 4 | 4 | 100 |
| 120 seconds | 6 | 6 | 120 |
At the Sora 2 Pro 1080p rate of $0.70 per second, that bottom row is $84.00 of billed seconds for one two-minute video, assuming every single extension came back usable on the first attempt. Mine don't. Multiply by however many rerolls you think is realistic and the number stops being a hobby number.
My own view, offered as judgment rather than as a finding, is that extensions are for when continuity genuinely cannot be faked. A held shot that needs to run longer than a generation allows. Anything where you'd otherwise cut, cut instead, because a cut is free and a seam is not.
The Batch Discount Is Exactly Half
This one took ten seconds to spot on the pricing page and I haven't seen anyone mention it.
| Model and resolution | Standard, per second | Batch, per second |
|---|---|---|
| sora-2, 720p | $0.10 | $0.05 |
| sora-2-pro, 720p | $0.30 | $0.15 |
| sora-2-pro, 1024p | $0.50 | $0.25 |
| sora-2-pro, 1080p | $0.70 | $0.35 |
Every row is exactly half. Not roughly half, not half on some tiers, half everywhere, and the docs describe queueing renders through the Batch API for offline workflows.
What that means practically is that the entire exploration phase of a project belongs in a queue. You are not sitting there watching a render bar for a prompt you're about to throw away. Write eight variants, queue them, go do something else, come back and pick. The only thing you're giving up is immediacy, and immediacy is worth very little when your keep rate is what mine is.
The docs also position the two models this way. sora-2 is described as being for speed and flexibility and ideal for the exploration phase, sora-2-pro as the better choice when you need production-quality output. Combine that with the batch pricing and the cheap path is obvious. Explore on sora-2 at 720p in batch, at five cents a second, and spend the pro rate exactly once, on the prompt that stopped changing.
The Policy Rewrites Your Prompt Before You Do
Three published restrictions matter for how you write, not just for what you're allowed to make. Content is restricted to what's suitable for audiences under 18, copyrighted characters and copyrighted music are blocked, and generating real people including public figures is prohibited.
The third one reshapes more prompts than people expect, because "a famous scientist explaining" or "a well-known athlete running" is a specification most of us reach for without noticing, and it's out. So is the lazy shortcut of naming a celebrity to convey a physical type.
| What you wanted | Why it's blocked | What to write instead |
|---|---|---|
| A named public figure | Real people are prohibited | Describe age, build, clothing, bearing, and what they're doing |
| A recognisable film character | Copyrighted characters | Describe the silhouette and the world, not the property |
| A known song under the clip | Copyrighted music | Generate or license audio separately, then mix |
| A celebrity lookalike as shorthand for a face | Same restriction | Reference image of a person you have rights to |
| Anything adult-adjacent | Content suitable for under 18 | Reframe the scene entirely, tone workarounds don't work |
The second column is why the substitutions in the third column are better prompts anyway. "A man in his sixties, heavy build, cardigan, chalk on his sleeve, talking with his hands" specifies more visual information than any name ever did, and it doesn't rely on the model having a strong association you can't inspect.
Reference images are the other route, and Sora supports images as the first frame. Which means the fastest path to a specific look is often to solve it as a still, cheaply, and hand the finished frame over. There's a whole method for that in ai image prompt.
Characters Have A Documented Sweet Spot
Reusable characters get a note in the docs that reads like an admission and I mean that kindly. They work best with short 2 to 4 second clips in 16:9 or 9:16, at 720p to 1080p.
Two to four seconds. That's the window where a consistent non-human subject holds together, which lines up with everything I've seen in my own attempts, where the drift always arrives late in a clip rather than early. Faces and hands go first, backgrounds go second, and a shot that's fine at second two is often unusable by second seven.
So if a recurring character is central to what you're making, build around short beats. Lots of small clips assembled, rather than fewer long ones. That's more editing work and it's the difference between something you'd publish and something you'd apologise for.
I'd add the obvious caveat that I've run maybe thirty clips total on this kind of thing, which is enough to notice a pattern and not enough to call it a rule.
Questions
How long should a Sora prompt be? Shorter than you want it to be. Their own examples are one sentence. I'd write two, with the second one being about the camera or about what stays still.
Do negative statements work? Saying what doesn't move has been more useful to me than saying what shouldn't appear. Exclusion behaves differently across every system, and the general treatment is in negative prompts.
Can I control the timing of an action inside a clip? Badly, in my hands. Words like "then" and "after" are asking for a sequence, and a sequence is two events, and two events is where clips fall apart. I'd rather make two clips.
Is 1080p worth it? Only at the end. The difference between 720p and 1080p is fidelity and the difference between a good clip and a bad one is almost always composition or motion, both of which are perfectly legible at 720p.
Does the model understand film vocabulary? Shot type and lighting, clearly yes, since both are in their own guidance. Focal lengths and specific camera bodies, less reliably in video than in stills. I'd stick to shot size and let the rest go.
Why do my results vary so much on the same prompt? Because generation is random. Reroll twice before you conclude a word did anything, or you'll credit an edit for a coin flip. I've done it. More than once.
What I'd Do With A First Prompt
Write it the way their examples are written. One shot type, one subject, one continuous state, one setting, one lighting phrase, and nothing else. Queue four of them on the cheap model in batch.
Then look at what came back and add a single sentence about the camera. That's round two. The wider framework for how prompts differ across text, images, and video is in chatgpt prompts, and the one-variable-at-a-time habit that makes any of this learnable is in ai prompts.


