AI customer support: cost per conversation, RAG and human handoff

Work out cost per conversation without dropping below your quality bar

Data reviewed on Sep 8, 2026 · LLM Stats indexes and official prices

The data, today

The context

A support assistant is judged conversation by conversation, not reply by reply. A typical chat runs several turns, the history grows with each one, and the model also receives passages from your knowledge base. So the metric that matters is cost per resolved conversation, paired with a quality floor you set yourself: answers that are unacceptable however cheap they are, like inventing a refund policy or promising a delivery date that does not exist.

The setup that tends to work has three parts. A knowledge base indexed with embeddings, so the model answers from your policies rather than from guesswork. A model chosen for its quality on your own conversations and for its official price. And a clear path to a human when the case calls for one. This page covers the decisions behind each part and what to measure before putting it in front of customers.

What to decide

  1. 1

    How do I calculate cost per conversation?

    Take real conversations and add up input tokens across every turn, including system prompt, history and retrieved passages, plus output tokens. Price them in the calculator at official rates. Divide by conversations resolved without escalation; escalated ones also carry the cost of a human agent.

  2. 2

    How do I define the quality floor?

    Write down the unacceptable failures: invented policies, promised refunds, requests for sensitive data, going off topic. Build a test set of conversations designed to trigger them and require the model to avoid every one. Compare price and latency only among the models that pass.

  3. 3

    What does the knowledge base need?

    Current articles, each with a review date and an owner. Split them by question or section, embed them and keep the source link so answers can cite it. Measure how often a real customer question retrieves a relevant passage; if retrieval misses, no model will answer well.

  4. 4

    When should the bot hand off to a person?

    Set explicit rules: refunds above a threshold, repeated complaints, angry tone, legal topics, or the customer asking for it. Pass the human agent a summary and the history so they do not start from zero. Track the escalation rate and how many escalated cases the bot could have handled.

  5. 5

    How much context does the assistant need?

    Add up system prompt, retrieved passages and the latest turns of the conversation. In long chats, summarize older turns rather than sending them verbatim. Typical support does not need a huge window, and paying for context you never use makes every reply more expensive.

Common mistakes

  • Measuring cost per reply instead of per full conversation with all its history.
  • Launching the assistant without a visible way to reach a human when the customer needs one.
  • Indexing outdated articles and blaming the model for answers that reflect stale information.

Tools and comparators

Guides to go deeper

Want a recommendation for your case?

Tell us your volume and the options you are weighing. We reply in writing with the numbers of your real usage; no commitment.

Request advice

No provider pays for its position. Indexes come from LLM Stats; prices from each provider's standard API. How we measure

The Codifly brief

A clearer perspective.
In your inbox.

Analysis, guides and new technology comparisons.

You can unsubscribe whenever you like.