Top 3 · Cost-benefit
- 01Gemini 3.8 Flash$1.50 · blended price34.0 pts/USD
- 02GPT-5.6 Terra$4.50 · blended price11.4 pts/USD
- 03GPT-5.6 Sol$8.00 · blended price6.9 pts/USD
Measure tokens per request and pay for big models only when needed
An AI bill grows for three reasons: the model you picked, the tokens per request and the number of requests. Teams usually attack only the first by switching to a cheaper model, when the other two often have as much room or more. A bloated system prompt, a conversation history resent in full, or output longer than anyone reads all get paid for on every call, multiplied across your entire traffic.
To compare models with a single number we use a blended price, mixing input and output prices at a typical ratio. Our prices are each provider's published standard rates, without caching or batch discounts. Those discounts exist and can lower your real cost, but they depend on how your traffic behaves. That is why the first step is always to measure your own requests before you optimize anything at all.
Log input tokens, output tokens, the model used and the product feature behind every request. A week of data shows which features drive spend and what your real input-to-output ratio looks like. Without that baseline, any optimization is a guess you cannot verify afterwards.
Sort tasks by difficulty and send the easy ones, such as extraction, classification or short answers, to a low-cost model, keeping the expensive one for complex reasoning. Use a test set to confirm the cheap model holds your quality floor for its group. Track how many requests escalate and the resulting average cost.
Prompt caching pays off when many requests share a long prefix, like fixed instructions or reference documents. Batch processing suits work that can wait hours. Our prices exclude these discounts, so read each provider's terms and model your case with the share of tokens that genuinely repeats.
Trim the system prompt, summarize history instead of resending it, retrieve only relevant passages and cap output length. Change one thing at a time, and compare quality on your test set and average tokens per request before and after each tweak.
Enter your average input and output tokens and monthly volume in the calculator and compare candidates at official prices. Then run the new model on a real sample: the same task may produce longer outputs or need more retries, and that shifts the final number.
Tell us your volume and the options you are weighing. We reply in writing with the numbers of your real usage; no commitment.
No provider pays for its position. Indexes come from LLM Stats; prices from each provider's standard API. How we measure
Analysis, guides and new technology comparisons.