How to reduce AI costs: blended pricing, routing and measurement

Measure tokens per request and pay for big models only when needed

Data reviewed on Sep 8, 2026 · LLM Stats indexes and official prices

The data, today

The context

An AI bill grows for three reasons: the model you picked, the tokens per request and the number of requests. Teams usually attack only the first by switching to a cheaper model, when the other two often have as much room or more. A bloated system prompt, a conversation history resent in full, or output longer than anyone reads all get paid for on every call, multiplied across your entire traffic.

To compare models with a single number we use a blended price, mixing input and output prices at a typical ratio. Our prices are each provider's published standard rates, without caching or batch discounts. Those discounts exist and can lower your real cost, but they depend on how your traffic behaves. That is why the first step is always to measure your own requests before you optimize anything at all.

What to decide

  1. 1

    What should I measure first?

    Log input tokens, output tokens, the model used and the product feature behind every request. A week of data shows which features drive spend and what your real input-to-output ratio looks like. Without that baseline, any optimization is a guess you cannot verify afterwards.

  2. 2

    How does model routing work?

    Sort tasks by difficulty and send the easy ones, such as extraction, classification or short answers, to a low-cost model, keeping the expensive one for complex reasoning. Use a test set to confirm the cheap model holds your quality floor for its group. Track how many requests escalate and the resulting average cost.

  3. 3

    When do caching and batch processing help?

    Prompt caching pays off when many requests share a long prefix, like fixed instructions or reference documents. Batch processing suits work that can wait hours. Our prices exclude these discounts, so read each provider's terms and model your case with the share of tokens that genuinely repeats.

  4. 4

    How do I cut tokens without hurting quality?

    Trim the system prompt, summarize history instead of resending it, retrieve only relevant passages and cap output length. Change one thing at a time, and compare quality on your test set and average tokens per request before and after each tweak.

  5. 5

    How do I estimate cost before switching models?

    Enter your average input and output tokens and monthly volume in the calculator and compare candidates at official prices. Then run the new model on a real sample: the same task may produce longer outputs or need more retries, and that shifts the final number.

Common mistakes

  • Switching to a cheaper model without checking whether quality drops below your acceptable floor.
  • Resending the full conversation history on every call without summarizing or trimming it.
  • Counting on a cache or batch discount without confirming your traffic meets its conditions.

Tools and comparators

Guides to go deeper

Want a recommendation for your case?

Tell us your volume and the options you are weighing. We reply in writing with the numbers of your real usage; no commitment.

Request advice

No provider pays for its position. Indexes come from LLM Stats; prices from each provider's standard API. How we measure

The Codifly brief

A clearer perspective.
In your inbox.

Analysis, guides and new technology comparisons.

You can unsubscribe whenever you like.