AI Solutions

How to Evaluate LLM Applications: A Practical Guide

How to evaluate LLM applications: choosing metrics, building test sets, using LLM-as-judge safely, picking eval tools, and running regression tests in CI.

GPTLabAI team 7 min read

To evaluate LLM applications well, you need three things: a test set of realistic inputs, clear pass/fail criteria for each one, and an automated way to run them every time you change a prompt, model or retrieval setting. Public benchmarks tell you little about your chatbot or agent. A few hundred well-chosen cases from your own domain tell you much more. This guide shows how we set that up.

Why LLM evaluation is different from normal testing

Traditional software tests check deterministic outputs: add(2, 2) must return 4. LLM outputs are variable, open-ended and often have several acceptable answers. That leads to three problems:

  • You cannot string-match most answers. “Your order ships Monday” and “It will be dispatched on Monday” are both correct.
  • Small changes have wide effects. A prompt tweak that fixes one case can quietly break ten others.
  • Quality has several dimensions. An answer can be accurate but rude, or friendly but wrong.

A good evaluation setup mixes cheap deterministic checks with model-based grading and a small amount of human review.

Step 1: Define what “good” means

Before choosing tools, write down success criteria that someone could actually check. Anthropic’s guide on defining success criteria makes the same point: criteria should be specific and measurable.

Vague goal Measurable criterion
“Answers should be accurate” Answer is supported by the retrieved documents; no unsupported claims
“Be helpful” Resolves the user’s question without asking for info already given
“Stay on brand” No competitor recommendations; no pricing promises outside the price list
“Return valid data” Output parses against the JSON schema; all required fields present
“Be fast” 95th percentile latency under a target you set
“Be safe” Refuses out-of-scope requests (legal advice, medical dosing) politely

Step 2: Build a test set from real traffic

Your eval set is the most valuable asset in the whole process. Sources, in order of value:

  1. Real user queries from logs or support tickets, with personal data removed.
  2. Known failures: every bug report becomes a test case.
  3. Edge cases written by domain experts: ambiguous questions, missing info, adversarial phrasing, other languages.
  4. Synthetic variations generated by an LLM to widen coverage. Review these; do not trust them blindly.

Practical tips:

  • Start with 50–100 cases. A small set you actually run beats a large one you never finish.
  • Tag each case by category (billing, returns, technical) so you can see where quality moves.
  • Store expected behaviour, not just expected text: “must mention the 30-day return window”, “must not promise a refund”.
  • Version the test set alongside your code.

For retrieval systems, also record which documents should be retrieved. That lets you measure retrieval separately from generation. See our RAG best practices for more.

Step 3: Choose metrics that match the task

Deterministic checks (cheap, run on everything)

  • Schema validity, required fields, allowed enum values
  • Exact or fuzzy match for classification and extraction
  • Regex or keyword checks (“contains a link to the returns page”)
  • Length limits, forbidden phrases, language detection
  • Latency, token usage and cost per request

Retrieval metrics (for RAG)

  • Recall@k: did the right document appear in the top k results?
  • Precision / MRR: how high did it rank?
  • Context relevance: how much of the retrieved text is actually useful?

Model-graded metrics (for open-ended output)

  • Faithfulness / groundedness: are claims supported by the provided context?
  • Answer relevance: does it answer the question asked?
  • Rubric scores: tone, completeness, following policy
  • Pairwise preference: is version B better than version A?

Agent metrics

For tool-using agents, grade the trajectory as well as the final answer. Did it call the right tools with valid arguments, avoid unnecessary steps and stop when done? See agentic AI design patterns for the architectures these tests cover.

Step 4: Use LLM-as-judge carefully

Using a strong model to grade outputs scales far better than human review. The research is encouraging but also comes with caveats. The paper Judging LLM-as-a-Judge found that strong judges agree well with human preferences, and it also documented position bias (favouring the first answer), verbosity bias (favouring longer answers) and self-enhancement bias (favouring outputs from the same model).

How we reduce those problems:

  • Grade against a rubric, not “is this good?” Give the judge explicit criteria and ask for a short reason before the score.
  • Prefer binary or small scales. “Pass/fail: does the answer mention the return window?” is more reliable than a 1–10 score.
  • Swap positions in pairwise comparisons and count only consistent wins.
  • Use a different model family from the one being tested, where practical.
  • Provide the reference answer or source context so the judge checks facts instead of guessing.
  • Calibrate against humans. Have a person label 30–50 cases and check that the judge agrees before you trust it at scale.
  • Return structured output (JSON with score and reason) so results are machine-readable.

Example judge prompt:

You are grading a customer-support answer.

<context>{retrieved_documents}</context>
<question>{user_question}</question>
<answer>{model_answer}</answer>

Criteria:
1. Every factual claim is supported by <context>.
2. The answer addresses the question directly.
3. No promises about refunds or delivery dates not stated in <context>.

For each criterion, give a one-sentence reason, then PASS or FAIL.
Return JSON: {"criteria": [{"id": 1, "reason": "...", "result": "PASS"}], "overall": "PASS"}

Step 5: Pick an evaluation tool

You can start with a spreadsheet and a script. Once you have more than a few dozen cases, a framework saves time. As of September 2026, these are the options we see most often:

Tool Type Licence / hosting Strong at
promptfoo CLI + YAML configs Open source (MIT); now part of OpenAI Prompt/model comparisons, red-teaming, CI
DeepEval Python, pytest-style Apache 2.0 Unit-test style evals, G-Eval, RAG and agent metrics
Ragas Python toolkit Apache 2.0 RAG metrics, test data generation
Inspect Python framework (UK AI Security Institute) MIT Rigorous, reproducible evals; agent and tool tasks
Langfuse Tracing + evals platform Open source core (MIT), self-hostable Production traces, datasets, LLM-as-judge on live data
Arize Phoenix Observability + evals Elastic License 2.0, self-hostable OpenTelemetry tracing, experiments

How to choose:

  • Developers who want tests next to code: DeepEval or promptfoo.
  • RAG-heavy systems: Ragas metrics, or the RAG metrics in DeepEval.
  • Research-grade reproducibility: Inspect. See also benchmarking LLMs for research.
  • Production monitoring plus evals in one place: Langfuse or Phoenix, especially if data must stay on your own servers.

Step 6: Run regression tests in CI

Evaluation pays off when it runs automatically. The pattern we use:

  1. On every pull request that touches prompts, model config, retrieval or tools, run a fast subset of 50–100 cases, mostly deterministic checks plus a few judged cases.
  2. Compare against the main branch baseline, not an absolute number. Fail the build if the pass rate drops by more than a threshold you choose, or if any “critical” tagged case fails.
  3. Nightly or before release, run the full suite, including slower judged metrics and multiple samples per case to measure variance.
  4. Post a summary on the PR: pass rate by category, newly failing cases, cost and latency change.
  5. Pin model versions in config so that a provider update does not silently change results. Re-baseline deliberately when you upgrade.

Keep costs under control by caching model responses for unchanged inputs and using a smaller model as judge for simple criteria. Our LLM cost optimization guide has more.

Step 7: Close the loop with production data

Offline evals catch regressions. Production tells you what you missed.

  • Log inputs, outputs, retrieved context and tool calls, with appropriate privacy controls.
  • Collect lightweight feedback (thumbs up/down, “did this solve your problem?”).
  • Sample a small percentage of live traffic for judged scoring each week.
  • Turn every confirmed failure into a new test case.

Key takeaways checklist

  • Written, measurable success criteria for each quality dimension
  • A versioned test set built from real queries and known failures, tagged by category
  • Deterministic checks first; LLM-as-judge only where needed, with rubrics and human calibration
  • Retrieval measured separately from generation (for RAG)
  • Evals running in CI on every prompt/model change, compared against a baseline
  • Pinned model versions and deliberate re-baselining
  • Production traces and feedback feeding new test cases

An evaluation suite turns LLM development from trial and error into ordinary engineering. If you want help building evals for a chatbot, RAG system or agent, or integrating them into your CI pipeline, see our AI solutions or talk to us.

6 min

Prompt Engineering Best Practices for Developers

Prompt engineering best practices for developers: clear instructions, examples, XML structure, reliable output formats, tool descriptions and eval-driven iteration.

Read article

7 min

How Much Does an AI Chatbot for Business Cost?

What drives AI chatbot cost for a business: build, model API usage, hosting and maintenance, build vs buy, plus a simple formula to estimate your token costs.

Read article

Have a project in mind? Let’s talk.

Whether you run a business or a research group, tell us what you need built, fixed or evaluated. You get a free consultation and a clear written estimate — no obligation.

  • Free consultation
  • Written scope and estimate
  • We reply within one working day
Contact us