AI Image Prompt Rules for Photos, Art, and Text
An ai image prompt changes shape with the image type. The class vocabulary table, the size limits both vendors publish, and why editing beats regenerating.

Search ai image prompt and you get banks. Two hundred and sixty free ones on one site, four thousand on another, sixty-nine copy-paste examples somewhere else, a Pinterest board, and a Reddit thread of somebody asking where to find good ones because none of the above worked for them. That last result is the honest one.
The thing all of those collections hide is that an image prompt isn't a single format. It's at least six, and the words that steer one of them are inert in the others. Lighting, lens, and film stock control a photograph and do nothing for a flat vector mark. Line weight, fill style, and palette control an illustration and do nothing for a photo. If your results feel generically wrong rather than specifically wrong, the usual cause is that you're writing photography vocabulary at a target that isn't a photograph.
So decide the class first. Then write with that class's controls. Here's the map.
Six Classes And What Actually Steers Them
I built this out of my own work rather than from a source, mostly by noticing which words kept doing nothing. The last column is the failure I'd expect if you reached for camera language in that row.
| Image class | What actually controls it | Words that do nothing here | If you use photo vocabulary |
|---|---|---|---|
| Photograph | Lighting direction, lens, distance, film stock, grain | Line weight, cel shading, flat colour | Correct, this is the one it's for |
| Illustration or comic | Line weight, ink style, cel shading, palette, panel framing | Aperture, focal length, depth of field | Half-rendered look, painterly mush, no clean line |
| Flat vector or icon | Geometry, stroke width, colour count, negative space, symmetry | Anything about light or lens | Soft shadows and gradients creeping into a mark that needs none |
| 3D render | Material names, roughness, subsurface, studio lighting rig, camera height | Film stock, grain, handheld | Half plastic half photo, an uncanny middle |
| Text-forward design | The exact string, font character, placement in frame, contrast | Bokeh, atmosphere, most mood words | Beautiful unreadable lettering |
| Diagram or infographic | Labels, arrows, layout, hierarchy, restraint | Every atmospheric word you know | Decorative nonsense with confident fake labels |
The row I'd stare at is flat vector. People trying to make a logo write "professional minimalist logo, clean, modern, 4k" and get a soft gradient blob with a drop shadow, then conclude the model can't do logos. What it can't do is guess that you wanted two colours, one closed shape, and no light source, because you never said any of that.
The diagram row is the other one worth a warning. Models will produce a beautiful chart with authoritative labels that mean nothing, and it looks correct at thumbnail size, which is exactly how it ends up in a deck.
The Photographic Case Has Its Own Piece
Photography is the deepest of the six and the one most people want, so it lives separately.
The short version is that a photo prompt is a caption rather than an instruction, lighting carries more weight than any other slot, and distance and lens are two different controls that people use interchangeably. The full treatment, with the lighting phrase comparisons and the lens table, is in how to write an ai photo prompt. I'm not going to repeat it.
Everything below is the other five rows.
Illustration Runs On Line, Fill, And Palette
Nobody writes illustration prompts well on the first go, myself included, because the vocabulary is less familiar than camera vocabulary. Most people know what a close-up is. Fewer people can name the difference between a heavy uniform ink line and a tapered brush line, and that difference is the whole look.
Google publishes a template for this in their image generation docs, and it's worth reading as a checklist even if you never use their model. Their skeleton reads "A [style] of a [subject, with details about accessories or actions] doing [activity]. The design features [visual qualities, e.g., bold outlines, cel-shading, etc.] and [color/background preference]."
Four slots. Style, subject with detail, visual qualities, colour and background. The one people leave empty is visual qualities, which is the one carrying the look.
Terms that have earned a place in my own illustration prompts, roughly in order of how reliably they do something:
Bold uniform outlines. Tapered ink lines. Cel shading with two tones. Flat colour, no gradient. Limited palette of four colours. Halftone texture. Cross-hatching for shadow. Visible pencil underdrawing. Screen-print misregistration. Thick black keyline.
The last one is a cheat I use for stickers and small marks, because a heavy keyline forces the model to close its shapes and you get something that survives being shrunk.
What consistently fails is naming an artist. Partly it's blocked or filtered in various places, partly the results are a vague averaged impression rather than the actual technique, and describing the technique gets you closer than naming the person who used it. The style side of this has more depth to it than I'm giving it here, and it's covered in ai art prompt.
Text Inside The Image Is Its Own Skill
Two vendors, two different tones about the same problem, which tells you roughly where the technology sits.
Google's documentation for their current image model lists advanced text rendering as a capability, saying it's capable of generating legible, stylized text for infographics, menus, diagrams, and marketing assets, and they ship a dedicated template for it that reads "Create a [image type] for [brand/concept] with the text "[text to render]" in a [font style]. The design should be [style description], with a [color scheme]." OpenAI's image generation guide, covering their newer models, notes that the model can still struggle with precise text placement and clarity.
Both statements are the vendor describing their own product, so the truthful read is somewhere between them. Short strings are fine now. Long strings are a gamble. Exact typography is not a thing you can specify.
What I do, in descending order of confidence:
Put the exact string in quotes inside the prompt. Keep it under about four or five words. Say where in the frame it sits, because unplaced text drifts. Describe the font by character rather than by name, so heavy condensed sans rather than a typeface you love. Generate at a higher resolution than you need, since small text degrades first. And if the wording has to be exactly right, generate the image clean and set the type yourself afterwards.
That last one sounds like giving up and it's what anyone shipping real work does. I make covers and promotional frames for books I self-publish, and title text generated inside the image landed maybe one attempt in six, with a letterform subtly off even in the keepers. Adding type in an editor took two minutes and worked every time. I wasted an embarrassing number of evenings before I accepted that.
Size Is A Prompt Control, Not An Export Setting
This is the part I've never seen written down in one place, so I put both vendors' published limits side by side.
| Control | Google, Gemini image models | OpenAI, gpt-image-2 |
|---|---|---|
| Aspect ratios | 1:1, 3:2, 2:3, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9 | Any size meeting the constraints below |
| Named resolutions | 512px, 1K, 2K, 4K, gated by model | 1024x1024, 1536x1024, 1024x1536, 2048x2048, 2048x1152, 3840x2160, 2160x3840, auto |
| Hard limits | Flash Lite is 1K only, 512px on Flash Image only, 4K on Flash Image and 3 Pro Image | Max edge 3840px, both edges multiples of 16px, long to short ratio no more than 3:1, total pixels between 655,360 and 8,294,400 |
| Quality tiers | Set by model choice | low, medium, high, auto |
A few things fall out of that once it's in one table.
The widest ratio Google offers is 21:9, which works out to about 2.33 to 1, comfortably inside OpenAI's 3:1 cap, so anything genuinely cinematic is available on both. The 16-pixel multiple rule is the sort of constraint that bites you once and then you remember it forever. And the pixel floor of 655,360 means there's no tiny-image escape hatch on that API, which matters if you were hoping to draft cheaply by generating small.
The part that actually changes your writing is that aspect ratio recomposes rather than crops. Ask for a scene at 9:16 and you don't get the 16:9 version with the sides cut off, you get a different photograph with more headroom and a tighter horizontal read. So set it before you write the rest, not after, and if you need both orientations of one idea, write two prompts rather than one prompt and a crop.
What An Attempt Costs, And Why That Changes How You Write
Reroll budget is the invisible variable in every prompting article, including the good ones.
| Route | Published cost per image |
|---|---|
| Gemini 3.1 Flash Lite Image, 1K, batch | $0.0168 |
| Gemini 3.1 Flash Image, 1K, batch | $0.034 |
| Gemini 3.1 Flash Image, 1K, standard | $0.067 |
| Gemini 3.1 Flash Image, 4K, standard | $0.151 |
| Gemini 3 Pro Image, 1K or 2K | $0.134 |
| gpt-image-2, 1024 square, low | $0.006 |
| gpt-image-2, 1024 square, medium | $0.053 |
| gpt-image-2, 1024 square, high | $0.211 |
Those are the published sheets as of today, and the spread between the cheapest and dearest row is more than thirty-fold. My keep rate on serious images sits somewhere around one in four, so at high quality that's roughly eighty-odd cents a usable frame, and at the low tier it's under three cents and you'd re-render the winner anyway.
Which suggests an obvious workflow that almost nobody follows. Draft at the cheapest tier available, iterate the wording there, and only re-render at quality once the prompt has stopped changing. The composition and the content are mostly settled at low resolution. What you're buying at the top tier is fidelity, not a different picture.
I run generation locally on an M4 Pro, so my own marginal cost is basically electricity, and I want to be upfront that this makes me a bad model for how careful you should be. Free experiments are how I learned any of this. When the meter's running people stop testing and start reasoning from the last thing that happened to work, which is where prompt folklore comes from.
Editing Is A Prompt Too
The thing that took me longest to internalise is that regenerating is often the wrong move.
Google's image documentation names a set of operations that aren't text-to-image at all. Adding and removing elements while preserving the original style. Inpainting through semantic masking, where you describe the region rather than painting a mask. Style transfer onto an existing photograph. Composing multiple input images together. Preserving specific details across an edit.
Every one of those is a prompt where you supply a picture and describe a change. And in each case the thing you're protecting is the ninety percent of the image that already worked, which a fresh generation would throw away and re-roll from scratch.
My own rule now is that if I like the composition and dislike one element, I edit. If I dislike the composition, I regenerate. That sounds trivial written down. It took me a lot of wasted generations to stop rewriting whole prompts to fix a background object.
The subtraction side, where you specify what shouldn't be there at all, works differently across systems and gets misread as a quality dial constantly. It's covered properly in negative prompts.
Questions
How long should an image prompt be? Six to nine strong clauses is where my own hit rate peaks. Beyond that the descriptors start blending into each other instead of each doing its own job, and the picture goes soft in a way you can't trace back to any single word. Personal number, modest testing, not a rule.
Does word order matter? Much more than in a chat prompt, yes. Whatever you'd be angriest about getting wrong belongs in the opening clause, and anything sitting at the tail end is the first thing to lose an argument with an earlier term. I ran the front-loading test properly and wrote it up in ai prompts.
Can I get the same style across a whole set? Partly. Keep a fixed style block that you paste byte for byte and swap only the subject line. Rewording the fixed block slightly, even "warm side light" becoming "warm light from the side", is enough to drift the look.
Why does my logo prompt keep producing a gradient blob? Because logo isn't a class the model understands as constrained geometry, it's a word attached to thousands of soft branding mockups. Specify colour count, stroke behaviour, and negative space instead of the word professional.
Are the prompt banks worth anything? As vocabulary, yes, genuinely. Scroll two hundred prompts with their outputs and you'll absorb terms you didn't have. As things to copy, less so, since they carry someone else's subject. More worked ones with the reasoning attached are in ai image prompt examples.
Do these rules apply to video? The visual half carries over almost entirely. Motion is a separate axis and models are visibly weaker at it.
If You Change One Habit
Name the class before you write the first word. Photograph, illustration, vector, render, text-forward, diagram.
Half the frustration in this whole subject comes from writing a photograph prompt at a target that was never going to be a photograph, and the fix costs you nothing but the two seconds of deciding. The overall framework for how image prompts differ from chat prompts is in chatgpt prompts, if you want the wider map.


