/ AI Prompts / LLM Prompt Engineering: What Changes Across Models

LLM Prompt Engineering: What Changes Across Models

LLM prompt engineering splits into the part that ports between models and the part that doesn't. One task written four ways from the vendors' own docs.

LLM Prompt Engineering: What Changes Across Models

Every llm prompt engineering guide on the front page treats the model as one thing. IBM runs eighteen named techniques against a single climate-change question. Google Cloud sorts prompts into four types and gives you a table per use case. Neither mentions that the companies making the models have published guidance that contradicts each other on the basics.

Here's the direct answer. The specification ports and the wrapper doesn't. What you want done, what the inputs are, what the constraints are, and what shape the output takes, all of that transfers between OpenAI, Anthropic, Google, and an open-weight Llama with almost no edits. What doesn't transfer is the container you put it in, how many examples you attach, whether you ask the model to reason, and which sampling knobs you're allowed to touch. Those four things are where the vendors disagree, in writing, and reading the disagreements is faster than discovering them one broken prompt at a time.

I pulled the current guidance from each vendor's own documentation on the same day. The rest of this is what it says, one task written four ways, and a table of what breaks in transit.

Where The Vendors Disagree, In Their Own Docs

OpenAI's guide says it outright. "Some prompt engineering techniques work with every model, like using message roles. But different model types (like reasoning versus GPT models) might need to be prompted differently to produce the best results." They go further and warn that even different snapshots within one family can behave differently, which is why they tell you to pin a snapshot and build an evaluation suite before you touch a production prompt.

Here's the comparison as I'd lay it out. Every cell is from the vendor's documentation as of the fetch date, and a blank means the page I read didn't cover it, which isn't the same as the model not caring.

Question OpenAI Anthropic Claude Google Gemini 3 Meta Llama 4
Where do standing instructions go A developer message, or the instructions parameter, which outranks the user turn A top-level system parameter beside the messages The System Instruction, or "at the very beginning" of the prompt A system role inside the chat template
How many examples Few-shot in the developer message for GPT models; reasoning models should "try zero shot first" Three to five examples "for best results", each in <example> tags "We recommend to always include few-shot examples", but too many will overfit Not covered on the format page
Should you ask it to think Helps GPT models; on reasoning models "think step by step" "can sometimes hinder" Adaptive thinking, where the model decides how much to think; the effort parameter sets depth Thinking is automatic; "Think very hard before answering" can help at a token cost Not covered on the format page
Where long inputs sit Context "near the end" of the developer message, so the cacheable part comes first Data at the top, query at the bottom, for 20k+ tokens "supply all the context first" and put the question "at the very end" Not covered on the format page
Structure markers Markdown headers and lists, XML tags for boundaries XML tags, consistent names, nested for documents XML or Markdown; the template is <role>, <constraints>, <context>, <task> Special tokens around each role header and turn
Sampling Not on the prompting page budget_tokens returns a 400 on Claude 4.7 and later; use effort Keep defaults on 3.x; temperature below 1.0 "can cause unexpected behavior, such as looping" Whatever your inference server sets

The examples row is the one I'd stare at. Google says always. OpenAI's reasoning guide says try without. Anthropic says three to five. My read is that they're describing three different model behaviours rather than three opinions about the same one, and that's the whole reason a cross-model guide needs to exist. A prompt tuned on Gemini arrives at an OpenAI reasoning model carrying examples the model didn't need, and OpenAI's guide warns that examples which don't "align very closely with your prompt instructions" produce poor results. So the examples aren't just wasted, they're a risk.

The long-input row looks like a disagreement and mostly isn't. OpenAI's ordering is about caching, since the reusable identity and instructions go first so the prefix matches on the next call, and the context that changes per request goes last. Anthropic and Google are talking about a different problem, which is quality when the input is long, and both say to put the document first and the question after it. You can honour both at once. Fixed instructions in the developer or system slot, the long document and then the question in the user turn.

One Task, Written Four Ways

The task is extracting action items from a meeting transcript into JSON. I picked it because the output shape matters, the input can be long, and the model has two chances to invent things, an owner nobody named and a date nobody said. The specification is identical in all four. Only the wrapper changes.

OpenAI, Developer Message Plus User Turn

OpenAI's guide lays a developer message out as Identity, Instructions, Examples, Context, in that order, with Markdown headers to mark the sections. The transcript is context, so it goes in the user turn.

## Identity
You extract action items from meeting transcripts for a small software team.

## Instructions
* Return only a JSON array. No prose before or after it.
* One object per action item with the keys owner, task, due, evidence.
* owner is the person named as responsible. If nobody is named, use "unassigned".
* due is the date phrase exactly as spoken, or null if none was spoken. Never guess a date.
* evidence is the shortest transcript quote that supports the item.
* If the transcript contains no action items, return [].

## Examples
<transcript_example>
Dana: I'll get the invoice template to Priya by Thursday.
Priya: And I'll look at the pricing page, no rush.
</transcript_example>
<output_example>
[{"owner": "Dana", "task": "Send the invoice template to Priya", "due": "Thursday", "evidence": "I'll get the invoice template to Priya by Thursday."},
{"owner": "Priya", "task": "Review the pricing page", "due": null, "evidence": "I'll look at the pricing page, no rush."}]
</output_example>

The user turn is then the transcript inside a <transcript> tag and nothing else. The second example line exists to show the model what null looks like, because a single example with a date in it teaches that every item has a date.

Two edits if the target is one of OpenAI's reasoning models. Drop the Examples block first and see whether the instructions alone are enough, which is what their reasoning guide tells you to do. And don't add anything about thinking step by step, since the model does that internally and the guide says asking "can sometimes hinder it". There's a third edit that only matters when you want formatted output rather than JSON. Reasoning models in the API avoid Markdown unless the first line of the developer message is Formatting re-enabled, which is the kind of detail that costs an afternoon if you don't know it. For a deeper pass on the reasoning versus GPT split, the chatgpt prompt engineering post goes through which classic techniques survived it.

Anthropic, System Parameter Plus XML

Claude takes the standing instructions in a system parameter that sits beside the messages rather than inside them. Everything else goes in the user turn, wrapped.

system:
You extract action items from meeting transcripts for a small software team. Respond directly without preamble.

user:
<instructions>
Return only a JSON array with one object per action item and the keys owner, task, due, evidence.
owner is the person named as responsible, or "unassigned" if nobody is named.
due is the date phrase exactly as spoken, or null. Never guess a date.
evidence is the shortest transcript quote that supports the item. This field exists so a reader can check every item against the transcript without rereading it, so quote precisely.
If there are no action items, return [].
</instructions>

<examples>
<example>
<transcript>Dana: I'll get the invoice template to Priya by Thursday.
Priya: And I'll look at the pricing page, no rush.</transcript>
<output>[{"owner": "Dana", "task": "Send the invoice template to Priya", "due": "Thursday", "evidence": "I'll get the invoice template to Priya by Thursday."}, {"owner": "Priya", "task": "Review the pricing page", "due": null, "evidence": "I'll look at the pricing page, no rush."}]</output>
</example>
</examples>

<transcript>
[PASTE]
</transcript>

Extract the action items from the transcript above.

Three Anthropic-specific choices in there. The sentence explaining why the evidence field exists is their "add context" advice, where the documented example is telling the model a response will be read aloud instead of just banning ellipses, and they claim it generalises from the explanation. The question sits after the transcript, which is their long-context rule for inputs past roughly 20k tokens, with a claim in the docs of up to 30 percent better quality that I can't verify and don't need to, since the ordering is free. And "Respond directly without preamble" in the system prompt is doing a job that used to be done by prefilling the assistant turn with an opening bracket. That trick returns a 400 error on Claude 4.6 and later, and the docs' own migration advice is exactly that sentence, or structured outputs if you need the schema enforced. The tag style has a whole post of its own in claude prompts.

Gemini, System Instruction Plus Role Tags

Google's Gemini API guide publishes a four-tag skeleton, <role>, <constraints>, <context>, <task>, and tells you to put behavioural constraints and format requirements in the System Instruction or at the very top. It also says to always include examples, so they go in.

System Instruction:
You are an assistant that extracts action items from meeting transcripts for a small software team. Output only a JSON array.

Prompt:
<role>
You extract action items from meeting transcripts.
</role>

<constraints>
1. Return only a JSON array, no prose.
2. Keys: owner, task, due, evidence.
3. owner defaults to "unassigned". due is the date phrase as spoken, or null. Never invent a date.
4. evidence is the shortest quote that supports the item.
5. If there are no action items, return [].
</constraints>

<examples>
Transcript: Dana: I'll get the invoice template to Priya by Thursday. Priya: And I'll look at the pricing page, no rush.
Output: [{"owner": "Dana", "task": "Send the invoice template to Priya", "due": "Thursday", "evidence": "I'll get the invoice template to Priya by Thursday."}, {"owner": "Priya", "task": "Review the pricing page", "due": null, "evidence": "I'll look at the pricing page, no rush."}]
</examples>

<context>
[PASTE TRANSCRIPT]
</context>

<task>
Based on the information above, list every action item as JSON.
</task>

"Based on the information above" is lifted from their guide, which calls it anchoring, a transition phrase after a large block of data so the model knows the data ended and the request began. Two things not to do here. Don't turn the temperature down to make the JSON more reliable, because Google's page says changing sampling on Gemini 3.x, "for example, setting the temperature below 1.0", can cause looping or degraded output, and they recommend leaving every parameter at its default. And don't expect a chatty answer, since the guide says Gemini 3 models "provide direct and efficient answers" by default and you have to ask for more. If the schema gets more complicated than four flat keys, their advice is the structured output feature rather than a longer constraints block.

Llama 4, The Raw Template

An open-weight model is the case where the wrapper is literally tokens. Meta's Llama 4 format page documents <|begin_of_text|> at the start, <|header_start|> and <|header_end|> around each role name, and <|eot|> at the end of each turn, with system, user, assistant, and tool as the roles.

<|begin_of_text|><|header_start|>system<|header_end|>

You extract action items from meeting transcripts for a small software team. Return only a JSON array with the keys owner, task, due, evidence. owner defaults to "unassigned". due is the date phrase as spoken, or null. Never invent a date. evidence is the shortest quote that supports the item. If there are no action items, return [].<|eot|><|header_start|>user<|header_end|>

Transcript:
[PASTE]

Extract the action items as JSON.<|eot|><|header_start|>assistant<|header_end|>

If you're calling Llama through a hosted API or a local runner that applies the chat template for you, you never type those tokens, you send roles like anywhere else. The raw form matters when you're driving the model directly, and it's where ports go wrong quietly, because a prompt missing its headers still gets an answer, just from a model that has no signal about where your turn ended. The format page also says something about system prompts that the other vendors don't put quite so plainly. "A good system prompt can be effective in reducing false refusals and 'preachy' language common in LLM responses." I'd take that as a hint that the system slot on an open-weight model is doing more of the work than it does elsewhere.

What Breaks When You Port A Prompt

The four blocks above look similar enough that you'd expect to swap them freely. This is what I'd check before assuming that. Every row comes from the documentation, none from a war story.

The move What goes wrong The documented fix
A GPT-tuned prompt to an OpenAI reasoning model "Think step by step" is redundant and may hinder; Markdown output disappears Delete the reasoning cue; put Formatting re-enabled on line one if you want Markdown
A few-shot prompt to an OpenAI reasoning model Examples that don't match the instructions degrade output Try zero-shot first; if examples stay, align them exactly with the instructions
A prefilled assistant turn to Claude 4.6 or later 400 error Move the constraint into the system prompt or use structured outputs
A budget_tokens thinking cap to Claude 4.7 or later 400 error Lower effort or cap with max_tokens
Any prompt to Gemini 3.x with a lowered temperature Looping or degraded output, per Google Leave sampling at the defaults
A prompt that assumes a long answer, to Gemini 3 Terse output Ask for the detail explicitly
A prompt that assumes a short answer, to Claude Opus 5 Longer output than earlier models, and effort doesn't reliably shorten it Prompt for conciseness in words
A hosted-API prompt to a raw open-weight model No chat template, no turn boundaries Supply the header and end-of-turn tokens, or use a runner that does
OpenAI instructions parameter across a multi-turn thread Instructions from earlier turns are absent when continuing with previous_response_id Resend the instructions on every request

The pattern across the rows is that the recent breakages are mostly vendors removing knobs, not adding them. Prefill gone, thinking budgets gone, temperature discouraged. The model's own judgment is taking over jobs that used to be yours, and the prompt is being asked to say what you want rather than to mechanically force it.

The Four Prompt Types Question

The People Also Ask box wants to know the four types of LLM prompt, and the answer it's looking for is Google Cloud's list. Direct or zero-shot, where you give the instruction and nothing else. One-, few-, and multi-shot, where you attach examples of the input and output. Chain of thought, where you ask for the intermediate reasoning. And zero-shot chain of thought, which is the instruction plus a line asking the model to reason step by step.

It's a fine list for a model that doesn't reason on its own. I'd describe the same four as two dials, the number of examples you attach and whether you request visible reasoning, and on the current reasoning models the second dial is already turned up by the vendor. OpenAI says so about their reasoning line, Google says Gemini 2.5 and 3 "automatically generate internal 'thinking' text", and Anthropic says thinking on the current Claude models is adaptive and always on for the newest ones. Which leaves you with one dial that still does something everywhere, and the vendors can't agree on where to set it.

Is ChatGPT An LLM Or Generative AI

Both, at different levels. Generative AI is the category, covering anything that produces text, images, audio, or video from a prompt. A large language model is one kind of generative model, the kind that produces text. ChatGPT is a product built on OpenAI's language models, with a chat interface, tools, and memory on top. So the model underneath is an LLM, the product is an application, and the whole thing sits inside generative AI. The question exists because people say "ChatGPT" when they mean the model and "the model" when they mean the product. The guidance in this article is about the models.

The Career Questions

Three of the related questions are about the job, and I'd sooner answer them honestly than pad them.

On salary, I'm not going to quote a number. The figures floating around online mix job titles that share the word "prompt" with roles that are really ML engineering or product work, and they span a wide enough range that any single number I put here would be a guess wearing a citation. If you're evaluating an offer, look at the posting's actual duties and compare against those.

On how to become one, and whether it's hard. The twelve moves every guide teaches take an afternoon. The part that takes longer is the evaluation habit, meaning being able to say whether an edit made the output better with something more than a feeling, and no technique list teaches it. I went through the courses one by one in prompt engineer course, and the syllabus I'd write instead is in there. If you can write a clear specification and you can set up a small fixed test set, you're most of the way to the skill, whatever the job title turns out to be.

On demand, my honest read is that the skill is in demand and the title is rarer than the search volume suggests. The prompting work in most teams is done by whoever owns the feature. That's not a reason to skip learning it. It's a reason to learn it alongside something.

What I'd Do First

Write the specification once, in plain text, before you open any vendor's playground. Task, inputs, constraints, output shape. Keep that file.

Then wrap it. Developer message for OpenAI, system parameter and tags for Claude, the four-tag skeleton for Gemini, the chat template for anything open-weight. Keep each wrapper as thin as you can. Everything you add to a wrapper is something you'll re-derive when the next model ships and the docs change under you, which they probably have between my reading them and your reading this.

And port it with a test, not a glance. Five fixed inputs, run on both sides, scored before you look. The general framework for prompts that sits underneath all four vendors is in ai prompts, and it hasn't changed much. What changed is the box you put it in.