/ AI Prompts / AI Portrait Prompts and the Default Face Problem

AI Portrait Prompts and the Default Face Problem

Most ai portrait prompts fail on the slots you left blank. The default-drift table, the skin texture terms that held up in testing, and the gaze controls.

AI Portrait Prompts and the Default Face Problem

Page one for ai portrait prompts is prompt banks and a Reddit thread. Somebody's ten go-to portrait prompts, a site with two hundred and sixty free ones, a Pinterest board, twenty cinematic Gemini styles. I read through most of them. They all specify the same handful of things and stay silent on the same handful, which is why they all produce the same face.

Here's the answer up front, and it's slightly annoying. The parts of ai portrait prompts that decide whether the picture looks like a person are usually the parts you didn't write. Age, skin condition, expression intensity, gaze direction, and what's reflected in the eyes all get filled in for you if you leave them out, and the fill is always the same narrow default. So the useful move isn't adding more adjectives. It's finding out which slots are silently defaulting and putting a value in them.

I'll show you the slots, then the vocabulary that moves each one.

What The Model Does With The Slots You Left Empty

Write "portrait of a woman, professional lighting, high quality" and you'll get a specific picture. Not a random one. The same one, more or less, forever.

I went through this properly because it was annoying me. Six unfilled slots, one prompt each with that slot left blank, four renders per prompt so I wasn't reading a single seed as a pattern. Twenty-four images, which is a small sample and I'd want someone to redo it at ten times the scale before treating the right-hand column as law.

Slot you left blank What arrives by default Term that actually overrides it
Age Mid twenties, every time A specific decade, "in her fifties", not "older"
Skin Retouched, poreless, even tone "visible pores", "uneven skin tone", "faint freckling"
Expression Slight closed-mouth smile Name the muscle behaviour, "brow slightly drawn", "lips parted"
Gaze Straight down the lens "looking off frame left", "eyes lowered"
Body language Squared to camera, shoulders level "turned three quarters away, head back toward camera"
Wardrobe Dark neutral knit or blazer Fabric and colour, "mustard corduroy shirt, collar open"

The age row is the one I'd stare at. If you want anyone who isn't in their twenties you have to say a number or a decade, and vague words like older or mature barely register, they land you at maybe thirty-two with slightly different lighting. Naming the decade works. "A man in his sixties" reliably produces a man in his sixties.

Wardrobe surprised me less but costs more, because the default blazer is the single loudest signal that a portrait came out of a generator rather than off a camera. Nobody in real life is wearing that.

Skin Is The Tell, And Beauty Words Make It Worse

Every beauty adjective you add makes the face less believable. That's not a paradox, it's just what the words are attached to in the training data.

Flawless, smooth, radiant, glowing, perfect complexion. Those phrases live under retouched commercial photography, so asking for them puts you in the region of images that were already smoothed by a person with a wacom tablet, and the output arrives pre-smoothed on top of a model that already leans that way.

The terms that pushed the other direction in my testing, roughly in order of how much they moved the image:

Visible pores. Uneven skin tone. Fine lines around the eyes. Faint stubble shadow. Small blemish on the cheek. Slightly oily t-zone. Sun damage across the nose. Chapped lips.

Pores did the most work by a wide margin. Adding that one phrase changed the render more than the other seven combined in the versions I ran, which I did not expect, and I've since started putting it in as a default rather than a fix.

One caution. Overloading the imperfection words swings you into a different failure, where the face reads as deliberately weathered, an actor cast as a fisherman. Two of these terms is usually enough. Three is a character study. The broader argument for why technically perfect prompts produce less believable pictures sits in realistic ai prompts.

Where The Eyes Point, And What Is Reflected In Them

This is the section I'd have wanted three years ago.

Two controls, both nearly absent from every prompt bank I checked. Gaze direction and catchlight.

Gaze is where the subject is looking. Unspecified, it's down the barrel, which is the single most common composition in the training data and also the one that reads most like a headshot on a company about page. Sending the eyes somewhere else changes the whole register of the picture. "Looking past the camera to the left" turns a headshot into a candid. "Eyes lowered, reading" removes the viewer from the scene entirely. "Looking directly at the lens, chin slightly down" is the confrontational version, and it's a different photograph from the default even though both are technically eye contact.

Catchlight is the reflection of the light source in the eye. Real portraits have one. Its shape tells you what lit the picture, a soft rectangle for a window, a ring for a beauty dish, a small hard dot for a bare bulb. Generated portraits often have them too, since the model learned from portraits, but they're inconsistent and sometimes there's a catchlight in one eye and not the other, which is the sort of thing you don't consciously notice and still feel.

Specifying it helps. "Single soft rectangular catchlight in each eye, from a window to camera left" does two jobs at once, it fixes the eye detail and it constrains the lighting setup so the rest of the face is lit consistently with it.

Head position is the third one, and it's cheap. Straight-on, three quarter, profile, and the small tilts between them. Three quarter with the far eye slightly hidden is the default of most good portrait photography and almost nobody types it.

The Vendor Template, Bent Toward Faces

Google publishes a photorealistic skeleton in their image generation documentation, and it reads "A photorealistic [type of shot] of a [subject description] in a [setting description]. [Description of the light]. Shot from a [camera angle] with a [lens type]."

Five slots. It's a decent skeleton and it's aimed at scenes rather than faces, so here's how I fill it when the subject is a person, with the extra fields the template doesn't have.

The shot type slot takes distance, and for a face that's close-up or medium rather than anything wider. Subject description is where the age, skin, expression, and wardrobe from the table above all go, and it ends up being the longest clause by far, maybe half the prompt. Setting can be almost nothing for a studio look, or it can carry the story. The light description does more than any other sentence in the prompt and it's covered properly in ai photo prompt so I won't repeat it here. Camera angle and lens control face shape more than people expect, since a low angle plus a wide lens will give you a jaw that doesn't exist.

Then I add two fields the template omits. Gaze, and catchlight. Both go immediately after the light description, because they're part of the lighting logic rather than part of the subject.

Expression Words Are Mostly Broken

Smiling. Serious. Confident. Thoughtful. Mysterious.

Those five do very little, and confident and mysterious do close to nothing at all in what I've run. They're abstractions, and the model has no reliable mapping from an abstraction to a set of facial muscles, so it falls back on the average face associated with the word in captions, which is the stock expression again.

What works is describing the face mechanically. Not the emotion, the geometry. "Brow slightly drawn together, mouth relaxed" gets you thoughtful without asking for thoughtful. "One corner of the mouth higher than the other" gets you wry. "Lips parted, eyes wide, caught mid-sentence" gets you the candid thing that "natural expression" never produces.

Teeth deserve a warning of their own. Asking for a broad open smile is the single fastest way to break a generated face, because teeth are small repeating structures and small repeating structures are exactly where these models fall apart. Closed mouth or slightly parted is much safer. If the picture needs a real laugh, expect to burn a lot more attempts on it, and check the canines specifically.

Hands, Glasses, And The Rest Of The Accident List

Ranked by how often they wrecked a frame I otherwise liked, from my own keeper pile rather than any study.

Element How often it fails Why What I do
Hands near the face Constantly Fingers plus occlusion plus foreshortening Keep hands out of frame, or one hand, relaxed, below the chin
Teeth in a wide smile Very often Small repeating structures Closed or slightly parted mouth
Glasses Often Frame geometry plus reflection plus the eye behind it Say "thin wire frames, no reflection"
Earrings Often Asymmetry between left and right Ask for one visible ear, or none
Necklaces and chains Sometimes Chains merge into skin Keep the neckline simple
Hair against a busy background Sometimes Edge separation Simple background, or rim light behind the head
Text on clothing Nearly always It's text Plain fabric, no logos, no print

The hands row is the one people already know about and still get caught by, because the failure isn't dramatic anymore. You don't get six fingers now. You get a hand that's subtly the wrong size for the face, or a thumb joint that doesn't bend where thumbs bend, and at thumbnail scale it passes and at full size it doesn't.

Likeness Is Not A Prompting Problem

Worth being blunt about this, because it's where a lot of people waste an evening.

You cannot describe a specific person into existence with words. Text is a lossy channel and a face is a very high-dimensional thing, so "a woman with green eyes, dark hair, high cheekbones and a small scar above her left eyebrow" narrows the space and does not land on one person, and running it twice gives you two different people who both fit.

OpenAI say something similar about their own image model in their documentation, noting that it "may occasionally struggle to maintain visual consistency for recurring characters or brand elements across multiple generations." That's the vendor describing the ceiling.

The consistency methods that do work are reference-based rather than text-based, meaning you supply an image of the face and the system carries it forward. For a set of portraits of the same invented person that's the route, and no amount of prompt precision substitutes for it. What prompting can hold steady across a set is everything except identity, so the light, the lens, the palette, the grain, all of that stays put if you paste a fixed style block byte for byte and change only the subject line.

The Check I Run Before I Keep One

Zoom to a hundred percent and look at these in this order, because that's roughly the order a viewer's eye finds them.

Eyes first. Are both catchlights present, the same shape, in the same relative position? Then the mouth corners, which is where asymmetry that reads as wrongness usually lives. Then the ears, which are frequently two different ears. Then hands if there are any. Then hair against the background edge. Then anything with text on it, which will be wrong.

Six checks, maybe twenty seconds. My keep rate on portraits is worse than on any other subject I generate, somewhere around one in five or six, and I'd rather find the problem before I've built a layout around the image.

What People Ask About Faces

What's the single highest-return addition to a portrait prompt? A skin texture phrase. "Visible pores, uneven skin tone" cost me nothing and changed more than any lighting swap I tried in the same session. Small sample, and I'd still bet on it.

Should I name a photographer or a portrait style? Naming a person is unreliable and increasingly filtered. Describing the technique they used gets you closer, so heavy side light and a long lens rather than a surname. The general style vocabulary question is worked through in ai art style prompts.

Why do all my portraits look like LinkedIn? Because the default wardrobe, the default gaze, and the default expression are all the LinkedIn versions, and leaving all three blank stacks them. Change any one and it breaks. Change all three and you're somewhere else entirely.

Do these apply to selfies? Partly. Selfies are a different genre with their own geometry and their own tells, and the polish that helps a portrait actively hurts a selfie. That's in ai selfie prompts.

How many attempts should a good portrait take? More than you want. I budget six to eight renders for anything I'm going to publish, and I've had faces that took twenty. Generation is random, so a bad first render is not evidence that the prompt is wrong.

If You Change One Line

Put age, skin, and gaze into the subject clause before you write anything else.

Three values, maybe eight words, and they're the three the banks leave out most consistently. The wider framework for how image prompts differ from chat prompts is in chatgpt prompts, and the general image case is in ai image prompt.