Prompt engineering is the practice of designing the instructions, examples and context a language model receives so it performs a task reliably, then proving that with tests. In 2026 the wording matters less than it did; what you put in the context and how you measure results matter more.
This guide covers the principles that still hold, the patterns worth knowing, how to run prompt work as engineering, and what to expect if you bring in outside help.
Principles that hold across models
- Be explicit about the task and the audience. State what the output is for, who reads it, and what a good result looks like. Models are capable but cannot read your mind.
- Give context, not just commands. Explain why a rule exists ("answers are read aloud, so avoid lists") and the model generalizes better than with a bare instruction.
- Show examples. A few input-output pairs (few-shot examples) communicate format and judgment faster than paragraphs of description. Vary them so the model does not copy one example's quirks.
- Specify the output format. For anything parsed by code, use the provider's schema-constrained structured output rather than asking nicely for JSON.
- Define the escape hatch. Tell the model what to do when it cannot answer: say so, ask a clarifying question, or hand off.
- Separate instructions from data. Mark user-supplied or retrieved text clearly (for example, inside tagged sections) and tell the model it is material to analyze, not instructions to follow. This reduces, but does not eliminate, prompt injection.
Role prompts ("You are an experienced tax accountant") can set tone and vocabulary, but they do not add knowledge the model lacks. Retrieval does that.
Prompt patterns by task
| Task | What works | Common failure |
|---|---|---|
| Information extraction | Schema-constrained output, field definitions, examples of missing values | Inventing values for empty fields |
| Summarization | State length, audience and what must be preserved (numbers, decisions, owners) | Dropping the one detail that mattered |
| Question answering over documents | Retrieved passages plus "answer only from these; cite sources; say if not found" | Answering from general knowledge instead of sources |
| Classification | Clear label definitions with borderline examples; a fixed label set | Inconsistent handling of edge cases |
| Code generation | Repository context, conventions, tests to satisfy | Plausible code that calls APIs that do not exist |
| Comparison and evaluation | Explicit criteria and a rubric; ask for reasoning before the verdict | Position bias (favoring whichever option came first) |
Reasoning models, which think before answering, changed one old habit: you rarely need to say "think step by step" anymore, and over-prescribing the reasoning steps can make results worse. Give them the goal, the constraints and the success criteria, and let them work.
Context engineering
For applications, the prompt is assembled at runtime from parts: system instructions, tool definitions, retrieved documents, conversation history, user profile data. Deciding what goes in, in what order, and what gets left out is where much of the quality comes from. Practical rules:
- Put stable content first so providers' prompt caching can reuse it; it lowers cost and latency.
- Retrieve fewer, better passages rather than everything that might be relevant.
- Summarize long conversation history instead of resending it in full.
- Write tool descriptions as carefully as instructions; models choose tools based on them.
The LLM application development guide covers the surrounding architecture, including retrieval and tool use.
Treat prompts as code
- Build an evaluation set of realistic inputs with expected behavior, including hard and adversarial cases.
- Define checks: exact checks where possible (valid schema, correct label), model-graded rubrics for fuzzy qualities, validated against human judgment.
- Version prompts in source control, with the model name and parameters alongside.
- Change one thing at a time and run the full suite. A fix for one case often breaks three others.
- Gate releases on eval results, and rerun everything when the model version changes.
- Sample production traffic regularly and add new failure cases to the set.
Without this loop, prompt engineering degrades into anecdote: someone tries a few examples, it looks better, it ships, and something else quietly breaks.
Models: write for the family, test for the version
The landscape the old "prompt engineering" pages described has moved on. Google's Bard became Gemini in 2024 and PaLM 2 was superseded; OpenAI's GPT models, Anthropic's Claude models and Google's Gemini models are the main hosted families, while open-weight families such as Meta's Llama, Mistral's models and others are widely self-hosted. Image models like Stable Diffusion have their own prompting conventions. Each family responds differently to formatting, examples and instruction style, and behavior shifts between versions. Do not assume a prompt transfers; rerun your evals. Check the provider's own prompting guide, for example OpenAI's prompt engineering guide, for the current recommendations of the model you use.
When prompting is not the answer
- The model lacks knowledge: add retrieval, not more instructions.
- You need a consistent style at high volume or a narrow classifier: consider fine-tuning once you have data and evals.
- The task is deterministic: write code. Do not ask a model to add up invoice totals.
- Errors are unacceptable: add verification, human review or a different design.
Hiring prompt engineering help
A standalone "prompt engineer" role has largely merged into AI engineering, but specialist help is still useful for audits and for standing up an evaluation process. Ask for deliverables you can keep: versioned prompts, an eval set, a test harness and a report showing before-and-after results on your data. Be skeptical of anyone selling prompt "libraries" or "secret prompts" as a product; value comes from fit to your task and from measurement. For broader AI projects, the guide to choosing an AI development partner lists evaluation questions that apply here too, and if you are adding a model to existing tools, see ChatGPT integration.
Frequently asked questions
Is prompt engineering still relevant with smarter models?
Yes, but it has shifted. Better models need less coaxing on wording, while context selection, output structure, tool descriptions and evaluation have become the main work.
How long should a system prompt be?
As long as it needs to be to state the task, constraints, format and edge-case handling, and no longer. Very long prompts with conflicting rules cause inconsistent behavior. Let your eval results decide.
Can prompts stop prompt injection?
They help but cannot guarantee it. Defend in depth: limit tool permissions, require confirmation for risky actions, validate outputs, and isolate untrusted content.
Do the same prompts work across different models?
Often partially. Core instructions transfer, but formatting preferences, example handling and verbosity differ. Always rerun your evaluation set when switching models or versions.