Top 3 · Customer support
- 01Gemini 3.8 FlashGoogle$1.50 USD / M tokens
- 02GPT-5.6 TerraOpenAI$4.50 USD / M tokens
- 03GPT-5.6 SolOpenAI$8.00 USD / M tokens
Work out cost per conversation without dropping below your quality bar
A support assistant is judged conversation by conversation, not reply by reply. A typical chat runs several turns, the history grows with each one, and the model also receives passages from your knowledge base. So the metric that matters is cost per resolved conversation, paired with a quality floor you set yourself: answers that are unacceptable however cheap they are, like inventing a refund policy or promising a delivery date that does not exist.
The setup that tends to work has three parts. A knowledge base indexed with embeddings, so the model answers from your policies rather than from guesswork. A model chosen for its quality on your own conversations and for its official price. And a clear path to a human when the case calls for one. This page covers the decisions behind each part and what to measure before putting it in front of customers.
Take real conversations and add up input tokens across every turn, including system prompt, history and retrieved passages, plus output tokens. Price them in the calculator at official rates. Divide by conversations resolved without escalation; escalated ones also carry the cost of a human agent.
Write down the unacceptable failures: invented policies, promised refunds, requests for sensitive data, going off topic. Build a test set of conversations designed to trigger them and require the model to avoid every one. Compare price and latency only among the models that pass.
Current articles, each with a review date and an owner. Split them by question or section, embed them and keep the source link so answers can cite it. Measure how often a real customer question retrieves a relevant passage; if retrieval misses, no model will answer well.
Set explicit rules: refunds above a threshold, repeated complaints, angry tone, legal topics, or the customer asking for it. Pass the human agent a summary and the history so they do not start from zero. Track the escalation rate and how many escalated cases the bot could have handled.
Add up system prompt, retrieved passages and the latest turns of the conversation. In long chats, summarize older turns rather than sending them verbatim. Typical support does not need a huge window, and paying for context you never use makes every reply more expensive.
Tell us your volume and the options you are weighing. We reply in writing with the numbers of your real usage; no commitment.
No provider pays for its position. Indexes come from LLM Stats; prices from each provider's standard API. How we measure
Analysis, guides and new technology comparisons.