AI Solutions

How Much Does an AI Chatbot for Business Cost?

What drives AI chatbot cost for a business: build, model API usage, hosting and maintenance, build vs buy, plus a simple formula to estimate your token costs.

GPTLabAI team 7 min read

The cost of an AI chatbot for business has four parts: the one-off build (or setup of an off-the-shelf tool), ongoing model API usage, hosting and data storage, and maintenance. Build cost depends mostly on scope: how many systems the bot connects to and how much it must be trusted to do. Running cost depends mostly on conversation volume and how much text goes into each model call. Below we break down each driver and give a formula you can use with your own numbers.

We deliberately do not quote “average chatbot prices”. The range is too wide to be useful, and model prices change several times a year. Once you know the drivers, you can get a realistic estimate for your case.

The four components of AI chatbot cost

Component One-off or recurring Main drivers
Build / setup One-off (plus improvements) Scope, integrations, data preparation, security and compliance needs
Model / API usage Recurring, usage-based Conversations per month, tokens per conversation, model choice
Hosting & infrastructure Recurring App server, vector database, logging, backups; self-hosted model GPUs if used
Maintenance & improvement Recurring Content updates, monitoring, evals, model upgrades, bug fixes

1. Build cost: what makes a chatbot more expensive

Roughly from simplest to most complex:

  1. FAQ bot on a fixed set of pages. Answers from your website or help centre. Little integration work.
  2. Knowledge-base bot (RAG). Answers from internal documents, manuals or tickets, with citations. This needs document ingestion, chunking, a vector database and retrieval tuning. See our RAG best practices.
  3. Transactional bot. Looks up orders, books appointments or creates tickets through your APIs. Each integration adds design, security and testing work.
  4. Agent-style assistant. Chains multiple tools, takes actions and needs approval flows, audit logs and careful guardrails. See agentic AI design patterns.

Other factors that increase build effort:

  • Messy or scattered source data. Cleaning PDFs, wikis and spreadsheets often takes more time than the chatbot itself.
  • Access control. Answers that respect user permissions (HR documents only for HR) need extra design.
  • Channels. Website widget, Slack or Teams, WhatsApp, mobile app: each one is additional work.
  • Languages and tone requirements.
  • Compliance. Personal data, logging policies, data residency and AI transparency obligations. See our GDPR and AI Act checklist.
  • Human handover to live agents with conversation context.
  • Evaluation. A test set and quality checks before launch. This is not optional if the bot talks to customers. See our LLM evaluation guide.

2. Model / API usage: how to estimate token costs

LLM APIs charge per token, a chunk of text (roughly three-quarters of an English word on average, though this varies by language and model). Input tokens (what you send) and output tokens (what the model writes) are usually priced differently, with output costing more per token.

In a chatbot, input is usually much larger than output. Each call sends:

  • the system prompt (instructions, tone, rules),
  • retrieved document chunks (for RAG),
  • tool definitions (for bots that call APIs),
  • the conversation history so far,
  • the user’s new message.

The formula

For one model call:

cost per call = (input tokens  × input price per token)
              + (output tokens × output price per token)

For a month:

monthly model cost = conversations per month
                   × calls per conversation
                   × cost per call

Remember that input grows during a conversation, because the history is re-sent on every turn. Use the average input size across all turns, not the first turn.

A worked example (illustrative prices)

The prices below are made-up round numbers to show the arithmetic, not any provider’s real pricing. Replace them with current prices from your chosen provider’s pricing page.

Assumptions:

  • 3,000 conversations per month
  • 4 user messages per conversation, so 4 model calls
  • Average input per call: system prompt 800 + retrieved chunks 2,000 + history 700 + user message 50 = 3,550 input tokens
  • Average output per call: 250 tokens
  • Illustrative prices: $2.00 per million input tokens, $10.00 per million output tokens

Per call:

input  = 3,550 × ($2.00 / 1,000,000)  = $0.0071
output =   250 × ($10.00 / 1,000,000) = $0.0025
total per call                        = $0.0096

Per month:

3,000 conversations × 4 calls × $0.0096 = $115.20

Two things stand out. First, input tokens make up about three-quarters of the cost here, mainly because of retrieved context. Second, the result scales linearly: ten times the conversations means ten times the bill, and moving to a model priced five times higher multiplies it by five. Run the formula for your realistic and peak volumes, and for two or three candidate models.

Add a margin for things the simple formula misses: retries, longer-than-average conversations, internal testing and evaluation runs, and embedding costs for indexing documents (usually small compared with chat, but not zero).

Ways to reduce model cost

  • Right-size the model. Many support questions do not need the most capable model. Route simple queries to a smaller, cheaper one.
  • Retrieve less, retrieve better. Five relevant chunks beat fifteen loosely relevant ones, for both cost and answer quality.
  • Use prompt caching where your provider offers it. A long, stable system prompt and tool list can often be cached at a reduced input rate.
  • Summarise long histories instead of re-sending every turn.
  • Cap output length for chat answers.
  • Use batch APIs for non-interactive work, such as nightly ticket summaries, where providers offer discounted batch pricing.

Our LLM cost optimization guide goes deeper.

3. Hosting and infrastructure

For API-based chatbots, hosting is usually modest:

  • a small web or API server for the chat backend and widget,
  • a database for conversations and a vector store for documents (often PostgreSQL with pgvector; see vector databases compared),
  • logging, monitoring and backups.

It gets more expensive if you self-host the language model for privacy or data-residency reasons. GPU servers are a significant fixed monthly cost whether the bot is busy or idle. That only pays off at high, steady volume or when policy requires it. See top open-weight LLMs to self-host for the trade-offs.

4. Maintenance: the cost people forget

A chatbot is not “done” at launch. Budget time every month for:

  • Content updates: new products, changed policies, re-indexing documents.
  • Reviewing conversations to find wrong or unhelpful answers.
  • Expanding the test set and re-running evals after every change.
  • Model upgrades. Providers retire older models, and each migration needs testing.
  • Security updates and dependency upgrades for the web app.
  • Integration changes when connected systems change their APIs.

A small bot might need a few hours a month. A customer-facing bot with several integrations needs more. Either way, plan for it rather than treating it as a surprise.

Build vs buy

Factor Off-the-shelf chatbot platform Custom-built chatbot
Time to launch Fast Slower
Upfront cost Low Higher
Recurring cost Per-seat, per-conversation or per-resolution fees Model API + hosting + maintenance
Integrations Whatever connectors the vendor offers Anything with an API
Data control Stored with the vendor You choose: your servers, your region, your model
Customisation Within platform limits Full control of flows, UI and guardrails
Vendor lock-in Higher Lower, if well architected

Buy when your needs are standard (FAQ plus handover to your existing helpdesk), volume is modest, and the vendor’s data handling meets your requirements.

Build when you need deep integration with your own systems, strict data control (see private RAG for business), unusual workflows, or when per-conversation platform fees at your volume would exceed the cost of running your own.

A hybrid is common too: a custom retrieval and integration backend behind a standard chat UI or helpdesk.

Checklist: what to prepare for an accurate quote

  • The top 20–50 questions the bot must answer, and who asks them
  • Where the answers live (website, PDFs, wiki, database) and how often they change
  • Systems the bot must read from or act on (CRM, orders, booking, ticketing)
  • Channels (website, Slack/Teams, WhatsApp, app)
  • Expected conversations per month, now and in a year
  • Languages
  • Data protection and hosting requirements
  • Handover process to humans
  • Who will own content updates after launch

Key takeaways

  • Chatbot cost = build + model usage + hosting + maintenance. Do not budget only for the first.
  • Build cost is driven by integrations, data quality, access control and compliance, not by “AI” itself.
  • Model cost = conversations × calls × (input tokens × input price + output tokens × output price). Input, especially retrieved context and history, usually dominates.
  • Self-hosting a model changes usage-based cost into a fixed GPU cost. It is worth it only at scale or when policy requires it.
  • Buy for standard needs, build for deep integration and data control.

If you would like a concrete estimate for your use case, or want us to build a chatbot on your own documents and systems, see our AI chatbot development service or send us your requirements.

Have a project in mind? Let’s talk.

Whether you run a business or a research group, tell us what you need built, fixed or evaluated. You get a free consultation and a clear written estimate — no obligation.

  • Free consultation
  • Written scope and estimate
  • We reply within one working day
Contact us