/ AI Prompts / AI Image Generation Prompts Are a Distribution

AI Image Generation Prompts Are a Distribution

Ai image generation prompts don't have an output, they have a spread. The odds table, how many renders before you judge, and writing for variance or lock-in.

AI Image Generation Prompts Are a Distribution

Here's the mistake underneath most complaints about ai image generation prompts, and almost nothing on page one mentions it. People write a prompt, render it once or twice, look at what comes back, and decide the prompt is bad.

A prompt doesn't produce an image. It produces a distribution of images, and a single render is one draw from that distribution. Which means judging a prompt on one or two attempts is like judging a coin on one flip, and the arithmetic on that is worse than most people's intuition suggests. If your prompt has a one in four keep rate, which is roughly where mine sits, there's about a thirty-two percent chance that four renders in a row give you nothing worth keeping. A third of the time, a perfectly good prompt looks broken.

So the loop matters as much as the wording. Here's the loop.

How Many Renders Before You're Allowed An Opinion

I worked this out because I kept abandoning prompts too early and then rediscovering them by accident weeks later.

The maths is a straightforward geometric distribution. Given a keep rate, the odds of getting nothing across n renders is one minus the rate, raised to n.

Your keep rate Renders for even odds of one keeper Renders for 90% odds Chance 4 renders give you nothing
5% 14 45 81%
10% 7 22 66%
20% 4 11 41%
25% 3 9 32%
33% 2 6 20%
50% 1 4 6%

The row I'd point at is ten percent, because that's roughly where an ambitious prompt sits, meaning something with an unusual composition or a hard subject. Two thirds of the time, four renders give you nothing. If your habit is to render four and rewrite, you will spend your whole life rewriting prompts that were fine.

Nine renders before judging is the number I settled on for anything I care about. It's not a magic figure, it's just where the odds get tolerable at the keep rates I actually see.

There's an obvious objection, which is that nine renders costs real money on a hosted service. That's true and it's why the drafting tier exists. I generate locally on an M4 Pro so the cost of another attempt rounds to electricity, and I'd say plainly that this has made me more patient than someone watching a meter can afford to be.

Four Inputs, And Words Are Only One

The prompt gets all the attention because it's the part you type. It isn't the only thing steering the result.

Input What it controls Can words substitute for it
The prompt text Subject, light, composition, style This is the words
Parameters Resolution, quality tier, ratio, sampler and steps where exposed No, and people waste hours trying
Seed Which draw from the distribution you get No, entirely outside language
Reference image Identity, style, or starting frame Partially, and badly

The seed row is the one that changes how you work. If your tool exposes a seed, locking it while you edit the prompt turns a noisy comparison into a nearly clean one, because you're holding the random draw constant and changing only the words. Release it once the wording has settled and you get variations within the look rather than variations of the look.

Half the frustration people describe as "the prompt isn't working" is actually seed variance, and the fix isn't better English.

Writing For Variance Versus Writing For Lock-In

Two different jobs, and they want opposite prompts. I don't think I've seen this distinction made anywhere and it changed how I write.

Written for variance Written for lock-in
Purpose Exploring, you don't know what you want Producing, you know exactly what you want
Subject clause Loose, "a market stall" Tight, "a wooden market stall, six crates, awning half rolled"
Lighting Named loosely or omitted Named precisely, direction and quality
Composition Left open Framing and subject placement stated
Style block Absent or one word Fixed, pasted byte for byte every time
Renders per attempt Many, judged as a grid Few, judged one at a time
What success looks like One surprising direction to chase Ten images that belong together

Most people write the lock-in version while doing the exploration job, then get frustrated that everything looks the same, or write the loose version while producing and wonder why the set doesn't match. The wrong prompt for the phase you're in feels like the model being uncooperative.

I flip between them on the same project, usually loose for an evening and tight for the next one. And the tight version isn't better writing, it's just a different instrument.

The Contact Sheet Habit

Judge images in a grid, never one at a time.

This sounds procedural and it isn't. Looking at renders individually makes you evaluate each one against your imagination, which is unwinnable. Looking at nine at once makes you evaluate them against each other, and comparison is the thing human eyes are actually good at. The winner is obvious in a grid and ambiguous in isolation.

My version is a folder and a file browser set to large icons, which is free and beats anything I've paid for. Sort by name so variants stay adjacent, generate four to nine per prompt version, and look at the whole sheet before you form a view.

The related habit is keeping the ones you rejected. I've gone back to a discarded render more than once because the project changed underneath me, and deleting aggressively cost me work I'd already paid for in time.

There's a whole software category built around logging this properly, and it's aimed at text rather than images, which I went through in prompt engineering tools.

The Best Render Is An Outlier, And That Matters

Something that took me a while to accept about the shape of the spread.

Across nine renders you get a cluster of similar, competent, slightly boring results, and then one or two that sit outside it. Sometimes outside in a good direction and sometimes in a strange one. The keeper is nearly always one of the outliers, not the middle of the cluster, because the middle of the cluster is by definition the most predictable thing the prompt could produce.

Which has an unpleasant consequence. If you found a great image and you want another one like it, running the same prompt again will not reproduce it, and it won't even reproduce the near miss. You'll get the cluster back. The great one was a tail event, and tail events don't repeat on demand.

So save it. Save the seed if you have one, save the exact prompt string, and save the file somewhere that isn't the tool's history, because chat interfaces lose things in a scroll and I've lost work that way more than once. I now keep a plain text file of the prompts that produced anything I was pleased with, which is the least sophisticated system imaginable and has never failed me.

The related trap is trying to improve an outlier by tweaking its prompt. You had a lucky draw. Changing the words changes the distribution, so the thing you're tweaking toward may not exist in the new one at all, and you can spend an hour chasing an image you already had.

When To Stop Rewriting And Start Editing

A stopping rule, because the temptation is to keep rewriting forever.

If nine renders produced nothing in the right neighbourhood, the prompt is wrong and rewriting is correct. If nine renders produced one image with the right composition and a wrong detail, stop rewriting. You've got the picture, and the remaining problem is a local one that a fresh generation will throw away along with everything that worked.

That distinction took me an embarrassing amount of wasted generation to internalise, and I still catch myself rewriting a whole prompt to move a background object. The editing side, meaning supplying the image and describing a change, is covered in ai image prompt.

The other stopping condition is fatigue, which sounds soft and is real. After about forty renders in a sitting my judgment goes, everything starts looking equally acceptable, and the keepers I picked at the end of a long session are the ones I most often discard the next morning.

Video Changes The Economics Of All Of This

Everything above assumes attempts are cheap. On video they aren't, and the whole loop reorganises around that.

Google's Veo runs clips at 4, 6, or 8 seconds, with 8 required if you want 1080p or 4K or want to hand it a reference image as the first frame. That last part is the important one for workflow, because it means you can settle the look as a still image, where attempts are cheap and fast, and only then spend a video render animating the frame you already approved.

Their documentation puts it plainly enough, advising you to "select an image closest to what you envision as the first scene of your video to animate everyday objects, bring drawings and paintings to life, and add movement and sound to nature scenes."

So the sensible video loop is two loops. Iterate the still until the composition, palette, and subject are right. Then iterate the motion prompt against a fixed first frame, which collapses the search space enormously because you're no longer gambling on the look and the movement at the same time. More on the motion half in ai video prompts.

Things I Get Asked

How many images should I generate per prompt? Nine if it matters, four if you're browsing, one if you're just checking the prompt parses. Fewer than four and you're reading noise.

What keep rate should I expect? Mine is around one in four overall, worse on faces, better on simple objects and minimalist compositions. Yours depends mostly on how demanding your brief is, and knowing your own number is more useful than knowing mine.

Does a longer prompt improve the distribution? Up to a point, then it flattens it in a bad way. Past roughly nine strong clauses the descriptors start averaging rather than each doing its job, and you get consistent mush instead of varied attempts.

Should I use the same prompt across generators? As a starting point, yes, since the vocabulary is mostly shared photographic and art language. Expect different strengths on the same words and expect the parameter layer to differ completely.

Is there a way to make results less random? Lock the seed. That's the only real lever, and it's a parameter rather than a phrase, which is why no amount of prompt rewriting achieves it.

Why do my results get worse when I add more detail? Because you're spending clause budget. Every additional descriptor competes with the ones already there, and the ones that lose are usually the ones you cared about, which is covered in ai photo prompt.

The Habit Worth Building

Render nine, look at them together, then decide.

That single change fixed more of my output than any wording trick, because it stopped me discarding prompts that were working and stopped me keeping images that only looked good against nothing. The vocabulary layer, meaning which words steer which kind of image, is in ai image prompt, and the wider framework is in chatgpt prompts.