/ AI Prompts / Midjourney Prompts, Draft Mode, and Style Refs

Midjourney Prompts, Draft Mode, and Style Refs

Midjourney prompts split control between prose and flags. The draft-mode arithmetic, a four-round iteration ladder, and what breaks when you paste elsewhere.

Midjourney Prompts, Draft Mode, and Style Refs

Three of the eight results on page one for midjourney prompts are Midjourney's own documentation. The rest are a gallery of other people's prompts, a Reddit collection, a listicle of a hundred examples, and a tips post. Which tells you the search is split between people who want the manual and people who want somebody else's words to paste.

The thing neither group tends to say out loud is that Midjourney prompts aren't one piece of writing. They're two, and the two obey different rules. There's prose, which the model reads the way any image model reads a caption, and there are parameters, which are flags that set the job rather than describe the picture. Most of the frustration I see comes from people trying to solve a parameter problem with adjectives, or writing a beautiful description and then wondering why the shape of the output never changes.

Get that split straight and the rest is iteration. Which, as of this summer, has an official cheap path.

Two Halves, One Text Box

Prose describes what's in the frame. Parameters describe the job.

That distinction sounds pedantic until you notice which category your current complaint lives in. If the picture is the wrong shape, no adjective will fix it. If the picture is too stylised, adding "realistic" to the prose is fighting a setting rather than changing one. If you want four variations of the same idea rather than four different ideas, that's a job property too.

There's a third input people forget, which is that you can hand it images as part of the prompt rather than only words. A reference image is a much denser specification than a paragraph, and it sidesteps the whole problem of not having the vocabulary for a look you can see in your head but can't name. I'd reach for that before I'd reach for another sentence.

I Could Not Load The Parameter Documentation

Being upfront here. Midjourney's documentation site returned a 403 to my tooling every time I tried it today, both the parameter list and the prompt basics page, and the main site did the same. So I'm not going to quote value ranges or defaults for flags I couldn't personally read, because the whole point of this blog is that the numbers are real.

What did load is their update log, which is a public changelog and a primary source, and it happens to contain the most useful recent thing they've shipped for anyone learning this. So that's what the next section is built on, and everything I say about specific flags comes from posts I actually opened.

If you're following along, open the docs yourself in a browser. They open fine for humans. It's automated fetching they don't like.

Draft Mode Changed What The Right Workflow Is

Their post from June 16 describes draft mode as generating 24 images at a lower resolution and quality, and then says something I had to reread. Their words, quoted, are that draft jobs "use half as many fast hours as V8.1 SD jobs (even though they generate 24 images!)". The workflow they describe from there is to click Vary on the ones you like to render those at full quality and full resolution.

Run the arithmetic on that and it's fairly stark.

What you run Fast hours, relative Images out Fast hours per image
One standard V8.1 job 1.0 Standard job output 1.0 divided by that count
One draft job 0.5 24 About 0.021
Ten draft jobs 5.0 240 About 0.021
Draft, then one full render of the winner 1.5 24 previews plus 1 final Mixed

I've left the standard row's image count blank on purpose rather than filling it in from memory, since that's exactly the kind of number I couldn't verify today. The relative figure that matters is in row two. Half the compute cost of one job, twenty-four pictures out of it.

What that does to your prompting is more interesting than what it does to your bill. When exploration is nearly free, the correct move stops being "write the best possible prompt" and becomes "write a deliberately underspecified prompt and see what the space looks like". You're no longer trying to land a shot. You're sampling a region and then walking toward the part you liked.

There's a --preview flag mentioned in the same post for trying unreleased models, with a caveat straight from them that preview images "might be a little unpolished and the jobs are not guaranteed to run consistently over time". Fun to poke at. Not something to build a look on.

A Four Round Ladder

Here's the iteration order I'd use now, which is different from the order I used before drafts were cheap. It moves from widest to narrowest, and each round changes exactly one category.

Round What you change What you're looking for When to move on
1 Nothing, deliberately. Subject only, no style words Does the model already have a strong default for this subject You can name the default you're fighting
2 Style and medium terms only, subject frozen Which visual world this idea lives in Two or three drafts you'd defend
3 Composition, distance, and framing Whether the idea survives being framed properly The shape stops bothering you
4 Parameters and one full-quality render Fidelity, nothing else It's done, or you go back to round 2

Round one is the one people skip and it's the cheapest information available. Every subject has a default, and the default is what the model reaches for when your prompt doesn't argue. Cars default to hero angles. Interiors default to a real estate listing. Portraits default to a flattering three-quarter with soft light. If you don't know what your subject's default is, half your later prompt is going to be spent unknowingly reinforcing it.

Round four is last for a reason. I spent a long time adjusting settings early, which meant I was tuning fidelity on compositions I'd later throw away.

Random Style References As A Vocabulary Machine

The June 25 post adds something I like more than I expected. They document including --sref random in a draft-mode prompt to create 24 images with different styles, and draft mode itself is triggered with --draft or the lightning icon in the prompt bar.

Twenty-four style samples of one subject, at half the compute of a normal job. That's not a gimmick, that's a teaching tool, and it fixes the single most common problem beginners have with image prompting, which is not having words for looks.

The way I'd use it, and this is a method rather than a result since I've only run it on a couple of subjects, is to keep the subject line boring and unchanging and let the styles vary. Then pick three you can't stop looking at and try to write down, in ordinary words, what's actually different about them. Line weight. Colour count. Whether there's texture. Where the light is. Whether the edges are hard.

That exercise builds vocabulary faster than reading any list of style terms, because you're describing something in front of you rather than trying to imagine what a word means. The style side of this, written out properly, is in ai art prompt.

One honest caveat. Style references are their own system with their own behaviour, and I'm describing the discovery loop rather than claiming I've mapped how the feature works internally. I haven't.

What Breaks When You Paste It Somewhere Else

People carry Midjourney prompts to other generators constantly and get confused results, so here's the translation table. This is my own mapping from moving prompts around, not a vendor's.

Midjourney habit What happens elsewhere What to do instead
Trailing parameter flags Read as literal text, or ignored Move the intent into prose or into the API field
Aspect ratio as a flag Usually a separate setting or field Set it before you write, since it recomposes rather than crops
Very short, comma-stacked prose Thinner results on models trained toward description Write full descriptive sentences
Style reference images No equivalent on some systems Describe the technique attributes instead
Relying on the four-image grid to sample You get one image Ask for several generations, or accept fewer samples
Assuming a house aesthetic underneath Other models are flatter by default Add the aesthetic explicitly, since it was never yours

That last row is the one that catches people. A Midjourney prompt often looks short and magical because the model brings a lot of taste to the table on its own. Move the same words to a system with less of an opinion and the output goes plain, and the reasonable conclusion is that your prompt was bad, when actually your prompt was always thin and something else was covering for it.

Their own July 24 note on version 8.2 says the update focuses on aesthetics, image quality, and personalization, and that low-quality outputs should be "dramatically reduced". That's a vendor describing their own product, so read it accordingly, but the direction is clear enough. More opinion, not less.

Questions I Get

How long should a Midjourney prompt be? Shorter than a Gemini or gpt-image prompt, in my experience, because more of the work is done by the model's defaults and by parameters. Six to nine strong clauses is where my hit rate sits across image models generally, and Midjourney sits at the lower end of that.

Do the giant prompt galleries help? As vocabulary, genuinely yes, and I say that as someone who complains about prompt banks constantly. Scroll a few hundred with the images attached and you'll absorb terms you didn't have. As things to copy, they carry somebody else's subject.

Should I use image references or describe the style? References, when you have one. Text is a lossy way to specify a look and always will be. Describe it when you're inventing something that doesn't exist yet as a picture.

Does personalization change how I should prompt? Probably, and I'd want to see it over months rather than a week before saying how. The version 8.2 note says profiles can now be built from a larger and improved pool of images, and anything that learns your taste is by definition making your short prompts mean more, which is convenient and slightly dangerous if you ever need to hand a prompt to someone else.

Is negative prompting a thing here? Exclusion behaves differently across every system I've used, and it's routinely misread as a quality dial. That's covered properly in negative prompts.

What transfers from Stable Diffusion habits? Less than you'd think, since the weighting syntax and the token budget are different animals. I went into that in stable diffusion prompts.

The One Change

Spend a draft round on the subject alone, with no style words, before you write a real prompt.

You'll find out what the model already thinks your subject looks like, and every later decision gets easier because you're now steering away from something specific instead of describing into a void. The general structure that sits underneath all of this, across models, is in chatgpt prompts and the image-specific version is in ai image prompt.