How to Write an AI Photo Prompt That Renders Right
An AI photo prompt is a caption, not an instruction. Here's the slot order, the lighting vocabulary that does the work, and which word causes each failure.

Every result on page one for ai photo prompt is a bank. Two hundred and sixty free prompts here, thirty thousand there, a Pinterest board, an Instagram account, an app on the Play Store. I went through several of them properly. Not one explains what any of the words are doing, which means you can copy a prompt that produces a beautiful picture and still have no idea how to produce a different one.
So, the answer up front. A photo prompt is a caption, not an instruction. Describe the picture as though it already exists and you're labelling it for an archive. Front-load the subject, then the lighting, then the shot and lens, then the style, then whatever you're subtracting. Lighting is doing more work than any other slot and most beginners leave it out entirely, which is why their results look like stock photography nobody chose.
I'll go through each slot, with what changes when you swap the words.
Why "Caption, Not Instruction" Is The Whole Trick
Consider what these models learned from. Enormous piles of images with text attached, and that text is captions, alt text, photo credits, gallery descriptions. Human beings describing what is in a picture.
Nobody has ever written "make this look amazing and professional" underneath a photograph. So when you type that, you're pointing the model at the region of its training data where those words appear, which is advertising copy and stock library metadata, and you get the visual equivalent back. Flat lighting, centred subject, empty smiling faces.
Whereas "low afternoon sun through a dusty window, film grain, shallow focus" is exactly the sort of thing a person writes underneath an actual photograph, so that's the neighbourhood you land in.
Google says roughly this in their own documentation for Gemini image generation. Describe the scene rather than listing keywords, and the more specific you are, the more control you have over the result. Their photorealistic template names its fields as shot type, subject, setting, lighting, and camera angle or lens, which is close to what I'd arrived at independently, and I take that as reassurance rather than coincidence.
The Order I Put Them In
Subject. Lighting. Shot and lens. Setting. Style and medium. Aspect ratio. Exclusions.
That order isn't arbitrary. Words earlier in an image prompt get weighted more heavily, so whatever you'd be most annoyed about getting wrong goes first. For me that's the subject and then the light, every time.
A worked one, for a product shot I needed for a listing:
A chipped white ceramic mug, half full of black coffee, steam just visible. Low warm side light from the left, long soft shadow across the surface. Close up, 85mm, shallow depth of field. On a scratched oak desk near a window, papers blurred behind. 35mm film photograph, muted colour, fine grain. 4:5. No text, no watermark, no hands.
Seven slots, one sentence each, about sixty words. That's near the top of where I'd go. Past roughly nine strong visual clauses I find the model starts averaging descriptors rather than honouring them, and the picture goes soft in a way that's hard to diagnose.
Lighting Is The Entire Game
I tested this the boring way. Same subject, same lens, same everything, one phrase swapped. Rerolled each version twice so I wasn't crediting a word for what was just a different seed.
| Lighting phrase | What comes back |
|---|---|
| Low warm side light, long shadows | Late afternoon feel, strong shape, most flattering to texture |
| Overcast, flat, diffused | Even and dull, everything visible, nothing emphasised, reads as documentation |
| Hard direct sun, high contrast | Sharp black shadows, blown highlights, summer, slightly aggressive |
| Soft window light, north facing | Gentle gradient, the default of most good interior photography |
| Low key, single practical light source | Most of the frame dark, subject picked out, mood arrives immediately |
| Backlit, sun behind subject, lens flare | Rim light, silhouette risk, atmosphere at the cost of detail |
What surprised me the first time is that swapping the lighting doesn't give you the same picture lit differently. It gives you a different picture. Composition shifts, the framing changes, sometimes the subject's implied age or the room's implied decade shifts with it, because the model learned that hard midday sun and single-source low key belong to different kinds of photographs taken by different people for different reasons.
Which means lighting isn't a finishing touch you add at the end. It's a genre selector.
Distance And Lens Are Not The Same Control
People use these interchangeably and they do different jobs.
Distance is where the camera stands. Lens is how compressed the space looks and how much falls out of focus.
| Term | What it sets | Watch out for |
|---|---|---|
| Extreme close up | Fills the frame with part of the subject | Loses all context, easy to make unreadable |
| Close up | Head and shoulders, or one object | The safest default for portraits |
| Medium shot | Waist up, some environment | Where most portraits should actually live |
| Wide shot | Full figure plus setting | Faces degrade, don't use for likeness |
| 24mm | Wide, exaggerated depth, edges stretch | Distorts faces badly, great for rooms |
| 50mm | Roughly what your eye sees | Boring in the good way, very reliable |
| 85mm | Compressed, flattering, background falls away | The portrait default for a reason |
| 200mm | Heavy compression, background looks close | Reads as sports or surveillance |
| Shallow depth of field | Blurred background | Overused, will hide detail you wanted |
| Deep focus, f/16 | Everything sharp | Good for landscapes and product flat lays |
I went looking for the official parameter documentation on one of the big commercial generators to check a couple of value ranges against a primary source and their docs returned a 403 to me twice today, so I'm not going to quote numbers I couldn't personally load. The vocabulary above is model-agnostic anyway. It's photography language, and every one of these systems learned it from the same captions.
Grain, Stock, And Why Clean Renders Look Fake
Here's a counterintuitive one. Making a photo prompt more technically perfect usually makes the output look more obviously generated.
Real photographs have flaws. Grain, slight motion blur, a highlight that clipped, chromatic fringing at the edges, dust. When you ask for a pristine 8K ultra detailed masterpiece you get something that reads as a render, because renders are the pristine things in the training data and photographs are not.
So I add imperfection on purpose. "35mm film photograph, fine grain" does a lot on its own. "Slight motion blur", "shot handheld", "minor lens flare", "slightly underexposed" all push toward believability. For portraits, texture words on skin matter more than any beauty term, and I'd take "visible pores, uneven skin tone, small blemish" over "flawless skin" every single time if the goal is a picture somebody believes. More on that in ai portrait prompts.
The same logic runs through ai selfie prompts, where the whole genre depends on looking casual and unplanned, and the moment a selfie prompt gets too polished the illusion goes.
Getting The Same Thing Twice
The first time I needed a set rather than a single image, I found out how little a good prompt is worth on its own.
It was covers for a run of books I was publishing, and they had to look like they belonged together. Same treatment, same palette, same implied camera, different scene each time. I had a prompt that produced a picture I liked. Running it again with the subject swapped gave me something that could have come from a different photographer in a different decade, and doing that six times gave me six unrelated books.
What fixed it, mostly, was separating the prompt into a part that never changes and a part that does. The fixed part carries lighting, lens, film stock, colour treatment, aspect ratio, and exclusions, and I paste it identically every time, word for word, in the same position. The variable part is the subject and the setting. Keeping the fixed block byte-identical matters more than it sounds, because rewording it even slightly, "warm side light" one day and "warm light from the side" the next, is enough to drift the look.
The other lever is seed. If your tool exposes one, locking it while you change the subject holds a surprising amount of the composition steady, and unlocking it once you've settled the wording gives you variations within the look rather than variations of the look.
Neither of those gets you all the way. For a genuinely consistent face across many images you need reference-based methods rather than prompting alone, since text is a lossy way to specify a person and always will be. But for a set of objects, places, or scenes, a fixed style block plus a swapped subject line will carry you further than any single perfect prompt.
I'd add one small thing I got wrong for a while. Write the fixed block down somewhere outside the tool. I kept mine in chat history, lost it in a scroll, and rebuilt it from memory badly.
Which Word Caused That Failure
The diagnostic table I wish somebody had handed me. These are my own mappings from working through this, not anybody's research.
| What went wrong | Usual cause | What to change |
|---|---|---|
| Looks like stock photography | Adjectives like professional, high quality, stunning | Delete them, add specific lighting instead |
| Everything mushy, nothing sharp | Too many descriptors, model averaging | Cut to under nine clauses |
| Subject correct, mood wrong | No lighting phrase at all | Add one, put it second |
| Composition keeps drifting | Aspect ratio unset | Set it, the model recomposes rather than crops |
| Random text appears in frame | Model filling signage habits | Add "no text, no watermark, no signage" |
| Faces off in wide shots | Distance, not prompt quality | Move closer or accept it |
| Weird extra limbs or objects | Complex pose or crowded scene | Simplify the action, one thing happening |
| Style ignored | Style term buried at the end | Move it earlier, or drop competing terms |
That last row is the one I'd bet most people are hitting without knowing. If you've got both "oil painting" and "photorealistic" in a prompt, one of them is going to lose, and which one loses depends partly on where it sits in the sentence.
The subtraction side deserves its own treatment, since it works differently across models and it's routinely misunderstood as a quality dial. It isn't one. It's covered in negative prompts.
What The Banks Are Actually Good For
I'm being hard on the prompt libraries, so let me be fair for a paragraph.
As vocabulary sources they're genuinely useful. Scroll two hundred image prompts with their outputs attached and you'll absorb terms you didn't have, which is real learning even if it isn't structured. That's more than I'd say for text prompt lists, where the words are ordinary English and there's nothing to absorb.
What they can't do is tell you why a prompt worked, so you can copy one and be stuck the moment your subject differs. The fix is small. When a bank prompt gives you something good, delete one clause and regenerate. Then delete a different one. Ten minutes of that teaches you more than the next two hundred prompts in the collection.
Things People Ask
How long should an AI photo prompt be? Shorter than most banks suggest. I aim for six to nine clauses, and I'd say the failure mode of a long prompt is worse than the failure mode of a short one, since a short prompt gives you something clean you can build on and a long one gives you mush you can't diagnose.
Do I need to name a specific camera? No. "Shot on a Canon 5D" is doing far less than "85mm, shallow depth of field", because the focal length is a real visual fact and the body name is mostly metadata. Film stock names are an exception and they do carry a look.
Does negative prompting work everywhere? No. Some systems have a dedicated field for it, some want exclusions written into the prompt, and some largely ignore them. Test it on whichever one you're using rather than assuming.
Why do my results change when I change nothing? Because generation is random by design. Which is also why you should reroll twice before you conclude a word did anything. I've fooled myself into believing a tweak was a breakthrough when it was just a different seed, more than once.
Is it worth generating locally? It changed how I work, so yes for me, with a caveat. The value isn't saving money exactly, it's that experiments become free, and free experiments are the only way anyone learns this. Mine run on an M4 Pro and the marginal cost of another attempt is basically nothing, so I test instead of guessing. If every generation costs credits you'll guess, and guessing is how people end up with prompt superstitions they defend for years.
If You Only Change One Thing
Add a lighting phrase and move it to second position.
That's it. That's the highest-return edit available to almost anyone writing photo prompts today, and it's the slot the banks leave out most often because their prompts were tuned on a subject you're not using. Everything else in this piece is refinement on top.
If you want the wider framing for how image prompts differ from chat prompts, it's in chatgpt prompts, and the general image case is in ai image prompt.


