AI Solutions

LLM Cost Optimization: How to Cut API Costs Without Losing Quality

Practical LLM cost optimization: model routing, prompt caching, batch APIs, shorter context, retrieval, output limits and monitoring that cut spend safely.

GPTLabAI team 7 min read

LLM cost optimization is mostly about sending fewer, cheaper tokens to the right model, not about haggling over price per million tokens. In our experience the biggest savings come from routing easy requests to smaller models, caching repeated prompt prefixes, batching work that is not urgent and trimming the context you send. Done carefully, none of these reduce answer quality, and you can prove that with a small evaluation set.

Below is the order we usually work through when a client says “our AI bill is growing faster than our usage”.

Start by measuring where the money goes

You cannot optimise what you do not measure. Before changing anything, log for every request:

  • the feature or endpoint that made the call
  • model name
  • input tokens, cached input tokens and output tokens
  • latency and whether the call succeeded
  • a request ID you can join to user feedback or evaluation scores

Both OpenAI and Anthropic return token usage in every response, including how many input tokens were served from cache. Store those numbers next to your application logs, then build a simple table: cost per feature, per day. It is common to discover that one feature, one oversized system prompt or one retry loop accounts for most of the spend.

At the same time, build a small evaluation set: 50–200 real inputs with known good outputs or grading criteria. Every optimisation below should be checked against it. Our LLM evaluation guide covers how to set one up.

1. Route requests to the smallest model that works

The single largest lever is not using your most capable model for everything. Typical traffic is a mix of:

  • Easy tasks — classification, extraction, short rewrites, routing, simple FAQ answers.
  • Medium tasks — summarisation, drafting, most RAG answers.
  • Hard tasks — multi-step reasoning, complex code, long agentic runs.

Small and mid-tier models handle the first two groups well for a fraction of the price of flagship models. A router can be as simple as a rule (“ticket classification always uses the small model”) or a cheap classifier that picks a tier. A common pattern is cascading: try the cheaper model first, validate the output (schema check, confidence, a rule), and escalate only when it fails.

def answer(question: str) -> str:
    draft = call_llm(model=SMALL_MODEL, prompt=question, max_tokens=400)
    if passes_checks(draft):
        return draft
    return call_llm(model=LARGE_MODEL, prompt=question, max_tokens=800)

Check each route against your evaluation set. If the small model’s score is close enough for that task, keep it.

2. Use prompt caching for repeated prefixes

Most production prompts share a long, stable prefix: system instructions, tool definitions, few-shot examples, a policy document. Prompt caching lets the provider reuse the processed prefix instead of charging full price every time.

As of September 2026:

  • OpenAI prompt caching is on by default for supported models; cached input tokens are billed at a fraction of the normal input rate, and a stable prefix is required for a cache hit.
  • Anthropic prompt caching is controlled with explicit cache breakpoints. Cache reads cost a small fraction of the base input price (10% on many models), while cache writes cost somewhat more than normal input. The default cache lifetime is five minutes, with a longer option available.

Rules of thumb that make caching work:

  1. Put static content first (instructions, tools, examples, reference documents) and dynamic content last (user message, retrieved passages, timestamps).
  2. Do not inject the current date, user name or request ID at the top of the system prompt — one changed character at the start breaks the cache.
  3. Keep tool definitions in a stable order.
  4. Check the cached-token fields in responses to confirm hits are actually happening.

3. Batch everything that is not interactive

If nobody is waiting for the answer — nightly document tagging, bulk summaries, evaluation runs, data enrichment — use a batch API. As of September 2026, both the OpenAI Batch API and Anthropic’s Message Batches API advertise a 50% discount in exchange for asynchronous processing (OpenAI’s completion window is up to 24 hours). Caching can also apply to batched requests on supported models, so the two savings can stack.

Moving back-office jobs off the real-time API is often the easiest win because it does not touch the user-facing product at all.

4. Send less context

Input tokens usually dominate cost, and context tends to grow silently over time. Common culprits:

  • Whole documents pasted into the prompt “just in case”.
  • Full chat history sent on every turn of a long conversation.
  • Verbose tool results, such as raw JSON API responses with hundreds of unused fields.
  • Huge system prompts that accumulated rules nobody remembers adding.

Fixes:

  • Summarise or truncate older conversation turns; keep the last few verbatim.
  • Strip tool outputs down to the fields the model needs.
  • Review the system prompt quarterly and delete rules that no longer change behaviour (test with your evaluation set).
  • Use structured, compact formats instead of prose where possible.

5. Retrieve instead of stuffing

Long context windows are tempting: put the whole knowledge base in and let the model find the answer. It works, but you pay for every token on every request, and answer quality can drop when the relevant passage is buried in noise.

Retrieval-augmented generation sends only the few passages that matter. A well-tuned RAG pipeline often uses a small fraction of the tokens of a “stuff everything” approach and gives more grounded, citable answers. See our RAG best practices for chunking, hybrid search and re-ranking tips.

6. Limit and shape the output

Output tokens are typically priced several times higher than input tokens, so verbose answers are expensive.

  • Set max_tokens (or the equivalent) per feature, based on what the UI actually shows.
  • Ask for the format you need: “Return JSON with fields X, Y, Z” instead of an explanation plus JSON.
  • Use structured outputs / JSON schema so you do not need a second call to fix formatting.
  • For reasoning models, use the lowest reasoning effort setting that passes your evaluation set; extra thinking tokens are billed as output.

7. Cache whole answers at the application level

Provider caching reuses prompt prefixes; your own cache can reuse complete answers. If many users ask the same questions (support bots, product FAQs, documentation search), store responses keyed by a normalised question plus the relevant document version. Semantic caching — matching similar rather than identical questions with embeddings — can extend this, but set a strict similarity threshold and invalidate entries when source documents change.

8. Remove waste in the plumbing

Not every cost is about prompts:

  • Retries: a bug that retries failed calls five times multiplies spend. Use exponential backoff and cap retries.
  • Duplicate calls: front-end double submits or re-renders that trigger the same request twice.
  • Agent loops: an agent without a step limit can burn thousands of calls on one task. Always set a maximum number of steps and a per-task budget.
  • Embedding re-computation: only re-embed documents that changed.

9. Consider self-hosting only when the numbers work

Running an open-weight model on your own GPUs can make sense at high, steady volume or when data must not leave your infrastructure. It rarely saves money at low or spiky volume once you count GPU rental, engineering time and monitoring. Our guide to open-weight models you can self-host covers the trade-offs.

Cost levers at a glance

Technique Typical effort Risk to quality Best for
Model routing / cascading Medium Low if evaluated Mixed-difficulty traffic
Prompt caching Low None Long, stable system prompts and tools
Batch API Low None Non-interactive jobs
Trimming context Low–medium Low Chat history, tool outputs
RAG instead of stuffing Medium–high Often improves quality Large knowledge bases
Output limits / structured output Low Low Every feature
Application-level answer cache Medium Low with good invalidation Repeated questions
Retry and loop limits Low None Agents, flaky integrations

Key takeaways

  • Measure cost per feature before optimising anything, and keep an evaluation set to catch quality regressions.
  • Route easy work to small models; escalate only when checks fail.
  • Structure prompts static-first so provider caching actually hits.
  • Move non-urgent work to batch APIs for a large, low-risk discount.
  • Send less: trim history, strip tool output, retrieve instead of stuffing whole documents.
  • Cap output length, reasoning effort, retries and agent steps.

Getting help

Most of these changes are a few days of engineering, not a rebuild, and they compound: routing plus caching plus batching together can transform a bill. If you want a second pair of eyes on your prompts, token logs or architecture, our LLM integration service covers exactly this kind of audit and implementation. Get in touch and tell us what you are running today.

6 min

Prompt Engineering Best Practices for Developers

Prompt engineering best practices for developers: clear instructions, examples, XML structure, reliable output formats, tool descriptions and eval-driven iteration.

Read article

7 min

How Much Does an AI Chatbot for Business Cost?

What drives AI chatbot cost for a business: build, model API usage, hosting and maintenance, build vs buy, plus a simple formula to estimate your token costs.

Read article

Have a project in mind? Let’s talk.

Whether you run a business or a research group, tell us what you need built, fixed or evaluated. You get a free consultation and a clear written estimate — no obligation.

  • Free consultation
  • Written scope and estimate
  • We reply within one working day
Contact us