Prompt Engineering Best Practices for Developers
Prompt engineering best practices for developers: clear instructions, examples, XML structure, reliable output formats, tool descriptions and eval-driven iteration.

The prompt engineering best practices that matter most for developers are not clever tricks. They are clear instructions with context, a few good examples, a clearly separated prompt structure, enforced output formats, well-written tool descriptions, and a test set that shows whether a change actually helped. Modern models follow instructions closely, so prompt quality now depends mostly on how precisely you specify what you want.
1. Be explicit and give the reason
Treat the model like a capable new colleague with no background on your project. Say what you want, who it is for and what “done” looks like.
Weak: Summarise this ticket.
Better: Summarise this support ticket for the on-call engineer.
In 3 bullet points: the customer's problem, what they already tried,
and the error message verbatim. The engineer will read this on a phone
at 3 a.m., so keep each bullet under 20 words.
Anthropic’s prompting best practices note that explaining the motivation behind an instruction helps the model generalise. “Keep it short because it is read on a phone” gives better results than “keep it short”, because the model can apply the reason to cases you did not anticipate.
Checklist for instructions:
- State the audience, the goal and the constraints.
- Say what to do rather than only what not to do.
- Number the steps when order matters.
- Define ambiguous terms (“urgent” means the customer cannot log in or pay).
- Say what to do when information is missing: ask, or answer with a stated assumption.
2. Use a few well-chosen examples
Examples (few-shot prompting) are one of the most reliable ways to control format, tone and edge-case handling. Anthropic’s guidance recommends 3–5 examples, and they work best when they are:
- Relevant: drawn from real inputs, not toy cases.
- Diverse: covering different categories and at least one tricky case, so the model does not copy one pattern.
- Clearly marked: wrapped in
<example>tags so they are not confused with instructions.
<examples>
<example>
<input>I was charged twice for order 4412</input>
<output>{"category": "billing", "priority": "high"}</output>
</example>
<example>
<input>How do I change my profile photo?</input>
<output>{"category": "account", "priority": "low"}</output>
</example>
</examples>
If the model copies surface details from your examples too closely, add more variety instead of more instructions.
3. Structure prompts with tags or sections
When a prompt mixes instructions, reference data, examples and user input, give each part its own labelled section. XML-style tags work well with Claude and are clear to other models too. Markdown headings are a reasonable alternative.
<role>You are a support assistant for Acme's billing team.</role>
<documents>
<document index="1">
<source>refund-policy.md</source>
<document_content>{policy_text}</document_content>
</document>
</documents>
<instructions>
Answer using only the documents above. If the answer is not there,
say so and offer to connect the user with a human.
</instructions>
<user_question>{question}</user_question>
Two long-context tips from the same guidance:
- Put long documents near the top, above the instructions and the question.
- Ask for relevant quotes first on long-document tasks, then the answer. This keeps the model grounded in the text.
Structure also protects you. Keeping untrusted user input inside its own tag makes it clearer to the model that it is data, not instructions. This is not a complete defence against prompt injection, but it helps. See our web app security checklist for the application-side controls.
4. Enforce output formats with the API, not with pleading
If your code parses the output, do not rely on “please return valid JSON”. Use the provider’s structured output features:
- The Claude API supports structured outputs with a JSON schema, plus
strict: trueon tools, using constrained decoding. - OpenAI’s API offers Structured Outputs with
strict: trueand a JSON schema.
A comparison of approaches:
| Approach | Reliability | When to use |
|---|---|---|
| Instructions only (“return JSON”) | Usually works, occasionally breaks | Prototypes, human-read output |
Output inside tags (<answer>…</answer>) |
Easy to extract, flexible content | Mixed prose + a clear final answer |
| Tool / function calling | High; arguments follow a schema | Actions and structured extraction |
| Structured outputs with strict schema | Highest; schema enforced at decode time | Anything your code parses |
Two related notes:
- Prefill is going away on newer models. Anthropic’s docs state that, starting with Claude 4.6-generation models, prefilling the final assistant turn returns an error. Use structured outputs, tags or direct instructions instead.
- Match the prompt style to the output style. A prompt written in heavy Markdown tends to produce Markdown. If you want plain prose, write the prompt in plain prose and say so.
5. Write tool descriptions like documentation
For agents, tool definitions are prompts, and often the most important ones. Anthropic’s engineering post on writing tools for agents matches what we see in practice:
- Fewer, workflow-shaped tools beat one tool per API endpoint.
schedule_meetingis better thanlist_users+list_events+create_event. - Namespace related tools:
crm_search_contacts,crm_update_contact. - Describe when to use the tool and when not to, and explain each parameter with its format and an example value.
- Return meaningful, compact results: names and summaries rather than raw IDs and huge JSON blobs, with pagination and sensible defaults.
- Make errors actionable: “date must be YYYY-MM-DD; received 03/10” lets the model fix its own call.
{
"name": "orders_lookup",
"description": "Look up a customer's order by order number. Use when the user mentions a specific order. Do not use for general shipping-policy questions; answer those from the policy documents.",
"input_schema": {
"type": "object",
"properties": {
"order_number": {
"type": "string",
"description": "Order number as shown on the receipt, e.g. 'A-10442'."
}
},
"required": ["order_number"]
}
}
The same principles apply to MCP servers. See our MCP servers guide.
6. Break complex work into steps
One large prompt that retrieves, reasons, formats and checks everything at once is hard to debug. Split it:
- Prompt chaining: extract, then analyse, then write. Each step has its own prompt and test cases.
- Let the model think before answering on hard tasks. Use the provider’s reasoning or “thinking” controls where available. Otherwise ask for reasoning in a separate tag that you strip before showing the user.
- Add a self-check step for high-stakes output: “Verify each figure against the source table before finalising.”
These patterns are the building blocks of agents. Our agentic AI design patterns article covers them in detail.
7. Iterate with evals, not intuition
A prompt change that looks better on three hand-picked examples can be worse on the other 97. The loop we use:
- Collect 30–100 realistic test inputs with expected behaviour.
- Run the current prompt and record pass/fail per case.
- Change one thing (an instruction, an example, a tool description).
- Re-run the full set and compare per case, not just the average.
- Keep the change only if it improves the target without regressions elsewhere.
- Version prompts in git, like code.
Our LLM evaluation guide explains how to automate this in CI.
8. Common mistakes we fix
- Contradictory instructions added over months (“be concise”… “always explain in detail”).
- ALL-CAPS warnings everywhere. Modern models tend to over-apply shouted rules. Calm, specific wording works better.
- Hidden requirements that exist only in someone’s head. If a reviewer rejects an output for a reason, that reason belongs in the prompt or the eval.
- Prompt-only fixes for non-prompt problems. If the right document is not retrieved, no prompt will fix the answer. If latency is the issue, a different model may be the answer.
- Model-specific prompts without re-testing. When you upgrade models, re-run your evals. Prompts tuned for an older model may need less instruction, not more.
Prompt engineering best practices: key takeaways
- Be explicit about audience, goal, constraints and the reason behind rules.
- Use 3–5 diverse, clearly tagged examples.
- Separate instructions, documents, examples and user input into labelled sections, with long documents first.
- Enforce machine-readable output with structured outputs or tool schemas.
- Treat tool descriptions as documentation for the model.
- Split complex tasks, and measure every change against a test set.
Prompts are production code and deserve the same review, versioning and testing. If you are integrating LLMs into a product and want prompts, tools and evals that hold up in production, see our LLM integration service or contact us.