AI Prompts Across Text, Image, and Video
AI prompts don't transfer between models. Here's one brief rewritten three ways, what carries across, and the word-order rule that only applies to images.

Almost everything written about ai prompts assumes you're typing into a chat window. MIT Sloan's guide does. The prompt marketplaces mostly do. Meanwhile the same person who wrote that chat prompt this morning is going to ask for a picture this afternoon, and they'll carry over habits that actively hurt them.
Direct answer, since that's what you probably came for. Prompts don't transfer cleanly across text, image, and video, because you're not doing the same thing in each case. Text models read your prompt as instructions. Image models read it as a description of something that already exists. Video models read it as a description that also has to survive being unrolled over time. Roles and politeness help the first one and do nothing for the other two. Lighting and lens vocabulary do nothing for the first and carry most of the weight in the second.
Below is the same brief written three ways, plus what actually transfers.
The Three Machines Behind The Same Text Box
They look identical from the outside. One box, one Generate button, and increasingly the same app, since ChatGPT will happily produce a picture in the middle of a conversation about your taxes.
Underneath they want different things. A text model has been trained to continue and to follow instructions, so an imperative like "rewrite this in 80 words, no exclamation marks" is meaningful to it. An image model was trained on pictures paired with captions, so what it wants is a caption. Nobody ever captioned a photograph "please make this professional and eye-catching". Write that into an image prompt and you've spent tokens on words that appeared in marketing copy, not under photographs, which is the mismatch in one line.
Video sits on top of the image case with an extra axis. You're describing a frame and also describing what happens to that frame, and the second part is where models are still visibly weaker.
One Brief, Three Prompts
I'll use a real one. I needed a cover-adjacent image for a book I was publishing, plus a short clip for the trailer, plus the copy for the listing. Same brief, three completely different pieces of writing.
The brief in plain English. A lone figure at the edge of a small town at dusk, uneasy, quiet, not fantasy, no monsters, meant to feel like something is about to go wrong.
| Target | The prompt I'd actually send | What lands |
|---|---|---|
| Text model | "You're a book blurb writer for literary suspense. Write 90 words of back-cover copy for a novel about a lone figure returning to a small town at dusk. Reader is browsing on a phone. No rhetorical questions, no 'but nothing is as it seems'. Match the restraint of this paragraph: [pasted]" | Usable draft in one or two rounds |
| Image model | "A lone figure standing at the edge of a small town at dusk, seen from behind, low warm side light, long shadows across empty asphalt, telephone poles receding, wide shot, 35mm film photograph, muted desaturated colour, slight grain, 16:9. No text, no people in the background." | A frame I'd keep about one time in four |
| Video model | "Locked-off wide shot. A lone figure stands at the edge of a small town at dusk, seen from behind, facing away down an empty road. Almost no movement. Wind moves the grass at the roadside. Slow, very slight push in over four seconds. Warm low side light, long shadows, muted film look." | A short clip, usually needing two or three rerolls |
Look at what happened to the role line. It's load-bearing in row one and it evaporated in rows two and three, because there's no photographer in the picture. Look at what happened to the constraints. "No rhetorical questions" is a real instruction to a text model, and its equivalent for the image is "no text, no people in the background", which is a subtraction from the scene rather than a rule about behaviour.
And notice the video row is the image row with the motion sentences added and the descriptive load slightly reduced. That's on purpose. Complexity that a still image handles fine tends to fall apart once things have to move.
What Carries Over And What Doesn't
I put this together after enough rounds of catching myself using the wrong habit in the wrong box.
| Technique | Text | Image | Video |
|---|---|---|---|
| Assigning a role | Works well | No effect worth the tokens | No effect |
| Specific context about audience | Central | Irrelevant | Irrelevant |
| Examples to fix tone | Strong lever | Use a reference image instead | Reference image or first frame |
| Naming the format explicitly | Essential | Aspect ratio does this job | Aspect ratio plus shot length |
| Saying what to avoid | Works, put it last | Works, keep it short | Works, keep it shorter |
| Politeness and encouragement | Harmless, mostly useless | Actively wasteful | Actively wasteful |
| Lighting and lens vocabulary | Meaningless | Most of the result | Most of the result |
| Sequencing and timing words | Meaningless | Confusing | Essential |
| Long, detailed prompts | Usually helps | Helps to a point, then mush | Hurts sooner than you'd think |
Two rows there are the ones I'd argue about with somebody. The examples row, because plenty of people do paste style descriptions into image prompts and get somewhere, and I'd say a reference image does the same job with less ambiguity. And the long-prompts row, since the fashion right now is toward very long image prompts, and my own experience is that past a certain density the model starts averaging your descriptors instead of honouring them. Where that point sits probably depends on the model. I haven't tested it carefully enough to give you a number and I'd rather say so.
The Words That Change Meaning Between Boxes
"Detailed."
Type that into a text prompt and you get more sentences. Type it into an image prompt and you get more texture, more clutter in the frame, more small objects the model felt obliged to add, and often a worse picture, because detail and composition pull against each other and the model has no idea which one you cared about.
There's a small vocabulary of words like this and they cause more confusion than any technique I know. "Clean" reads as clear and uncluttered writing in one box and as a bright, minimal, high-key photograph in the other. "Professional" is nearly meaningless in both, but it does something in image models, which is worse than nothing, because it drags you toward stock photography with the flat corporate lighting that comes with it. "Realistic" is the one I'd warn hardest about, since it can mean photographic, or it can mean plausible, and image models tend to hear the first while you meant the second. If what you want is a scene that makes sense, describe why the objects are where they are, and leave "realistic" out of it entirely.
"Cinematic" I use anyway, guiltily, because it does reliably shift lighting and aspect toward something moodier even though it's exactly the kind of vague adjective I'd tell somebody else to replace with specifics. It's shorthand that happens to work. I'd still rather write "low key, warm practicals, shallow depth of field" when I know that's what I mean, and I don't always know.
The general habit that helps is to ask whether a word describes the writing or describes the picture. Words about the output as a piece of communication belong in text prompts. Words about light, distance, material, and surface belong in image prompts. Adjectives that could go either way, and that's most of the flattering ones, are usually doing less than you think in both places.
Word Order Only Matters In One Of Them
Here's the thing that took me longest to notice, and I only noticed it because I was rerolling the same prompt cheaply on my own machine and could afford to be systematic.
In a text prompt, moving a clause from the end to the beginning changes almost nothing. The model reads the whole thing. In an image prompt, it changes the picture. Terms near the front get more weight, and terms near the back get treated as decoration that can be dropped when they conflict with something earlier.
I ran the same subject with the same eight descriptors in two orders. Front-loaded lighting gave me consistent lighting and inconsistent styling. Front-loaded style gave me a consistent look with the lighting drifting all over the place. Same words. Nothing else changed except where they sat in the sentence.
That's a practical rule rather than a theory. Put the thing you'll be angriest about getting wrong in the first clause. For me that's usually the subject and then the lighting, which is why my image prompts read like a police description followed by a weather report.
One caveat worth stating. This is my own observation across a few dozen generations on the models I run, not a benchmark, and I've seen enough randomness in image generation to know that a few dozen isn't many. Treat it as a starting hypothesis you can test in an afternoon, which is roughly what it cost me.
The Library Problem
Prompt libraries are enormous now. One of the sites ranking for this exact phrase advertises over thirty thousand prompts, updated daily, organised by model and by use case. There's a paid marketplace on the same results page. Free community collections everywhere.
They're not useless. As a source of visual vocabulary, browsing a few hundred image prompts genuinely teaches you words you didn't have, which is more than I can say for browsing text prompt lists. But the library model has a built-in ceiling, because the thing that makes a prompt work is the part specific to your situation, and a library entry by definition contains none of that.
What I do instead, and this is barely a system, is keep a short file of my own prompts that worked, with a note on what I changed to get there. Maybe fifteen of them. The note matters more than the prompt. Six months later the prompt might not work on the current model, but "front-load the lighting" still will.
If you want worked ones rather than a bank, ai prompt examples walks through several with the reasoning attached.
Where I Still Waste Time
Video. Every time.
I'll write something ambitious with two actions in it, get back a clip where the first action happens and the second turns into soup, and rewrite it three more times before admitting I should have made two clips. This is a solved problem and I keep not solving it.
Questions That Come Up
Is there one prompt framework that covers everything? At the level of "say what you want specifically", sure, and that's true enough to be useless. At the level of what you actually type, no. The six-slot structure in the chatgpt prompts guide covers text well and image prompts need a different set of slots entirely.
Do different text models need different prompts? Somewhat. The structure is portable, the formatting isn't. Anthropic's documentation recommends XML tags for marking sections in Claude prompts, and OpenAI's developer docs mention Markdown and XML for showing logical boundaries. So the same prompt is worth reformatting when you move it. I go through the specifics in claude prompts.
Are negative prompts a real thing or superstition? Real, but model-dependent, and much less powerful than people hope. They remove things, they don't add quality. Piling on twenty negatives to make an image "better" is a known way to waste a generation.
What's the fastest way to get better at image prompts? Change one word at a time and reroll each version at least twice. It's tedious and it works, and it's why cheap generation matters more than it sounds. When each attempt costs real money you start guessing instead of testing, and guessing is how people end up with prompt superstitions they defend for years. Mine run locally, so the only thing an experiment costs me is the afternoon.
Should I learn video prompting yet? If you need it, yes, and lower your expectations about clip length. If you don't need it, the vocabulary overlaps enough with images that you'll pick it up later without much pain. ai video prompts covers the motion side properly.
The Short Version
Text prompts fail because you didn't say enough about your situation. Image prompts fail because you said too much, or said the wrong kind of thing, or buried the important word at the end of the sentence. Video prompts fail because you asked for two things to happen.
Different failure modes, so different fixes, and carrying one set of habits into all three boxes is the mistake almost everyone makes including me. The ai image prompt piece is where I'd go next if pictures are what you're here for.


