Output prices differ 250-fold between the cheapest and the priciest model: estimate your bill with your own input-to-output ratio, not with the input price or a generic average.
1. September prices, model by model#
Language model APIs charge per million tokens, with one rate for the tokens you send (input) and another, almost always higher, for the tokens the model generates (output). The table lists standard prices for 17 models from Codifly's open dataset, reviewed on September 8, 2026 against each provider's official page. These are list prices before caching, batch processing or enterprise agreements.
| Model | Provider | Input (USD/M) | Output (USD/M) | Blended 3:1 | Context (tokens) |
|---|---|---|---|---|---|
| Ministral 3 · 14B | Mistral AI | $0.20 | $0.20 | $0.20 | 256,000 |
| Mistral Small 4 | Mistral AI | $0.15 | $0.60 | $0.2625 | 256,000 |
| Codestral · 25.08 | Mistral AI | $0.30 | $0.90 | $0.45 | 128,000 |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | $0.45 | 1,050,000 |
| DeepSeek V4 Flash | DeepSeek | $0.44 | $1.32 | $0.66 | 1,000,000 |
| Mistral Large 3 | Mistral AI | $0.50 | $1.50 | $0.75 | 256,000 |
| Gemini 3.8 Flash | $0.75 | $3.75 | $1.50 | 1,048,576 | |
| DeepSeek V4 Pro | DeepSeek | $1.32 | $3.96 | $1.98 | 1,000,000 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $2.00 | 200,000 |
| Grok 4.6 | xAI | $2.00 | $6.00 | $3.00 | 500,000 |
| Mistral Medium 3.5 | Mistral AI | $1.50 | $7.50 | $3.00 | 256,000 |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $4.00 | 1,000,000 |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | $4.50 | 1,050,000 |
| GPT-5.6 Sol | OpenAI | $4.00 | $20.00 | $8.00 | 1,050,000 |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | $10.00 | 1,000,000 |
| Claude Fable 5.1 | Anthropic | $10.00 | $50.00 | $20.00 | 1,000,000 |
| GPT-6 Astra | OpenAI | $10.00 | $50.00 | $20.00 | 1,050,000 |
The spread is huge. The cheapest output costs $0.20 per million (Ministral 3 · 14B) and the most expensive $50 (GPT-6 Astra and Claude Fable 5.1): 250 times more. Input ranges from $0.15 (Mistral Small 4) to $10, about 67 times. Nearly every model prices output 3 to 6 times higher than input; Ministral 3 · 14B is the exception, charging the same rate both ways.
Two changes landed after that review. DeepSeek's pricing page now lists deepseek-flash as DeepSeek-V4.1-Flash, at $0.30 input and $1.20 output per million at peak (the table shows $0.44 and $1.32), and Anthropic now recommends Claude Opus 5.5, at $4 and $20, for most workloads, listing Claude Opus 5 among its legacy models that remain available. Confirm those rows on the official page before you budget with them.
Documentation: OpenAI: API pricing ↗ · Anthropic: Claude models and pricing ↗ · Google: Gemini API pricing ↗ · DeepSeek: API pricing ↗ · xAI: Grok 4.6 ↗ · Mistral: Ministral 3 14B ↗ · Mistral: Mistral Small 4 ↗ · Mistral: Codestral 25.08 ↗ · Mistral: Mistral Large 3 ↗ · Mistral: Mistral Medium 3.5 ↗
2. Why output tokens drive the bill#
A typical application sends far more text than it gets back: system instructions, conversation history, chunks retrieved from your documents and the user's question. Output is still usually the largest line item, because each generated token costs several times more than each token read. With an output rate five times the input rate, a 400-token answer costs as much as 2,000 tokens of context.
That makes the input price, the figure most announcements lead with, a poor basis for comparison. To rank models on a single number, Codifly uses a blended price that assumes 3 input tokens per output token: (3 × input + output) / 4. Claude Sonnet 5 comes out at (3 × $2 + $10) / 4 = $4 per million; GPT-5.6 Luna at (3 × $0.20 + $1.20) / 4 = $0.45.
- The 3:1 ratio is a convention for comparing models, not a forecast of your usage.
- A classifier that returns a single label can run at 50:1, so the input price dominates.
- A drafting or code generation workload can approach 1:1 or even flip the ratio, so the output price dominates.
- If the model reasons before it answers, check the provider's documentation for how those tokens are billed, and measure actual output in the usage field of each response.
The rule of thumb: measure your real ratio for a week using the token counts the API returns, then re-rank the models with it. Two models with the same blended price can drift apart once your workload moves away from 3:1, as Grok 4.6 and Mistral Medium 3.5 do, both at $3.
3. Worked example: a support bot for one month#
Take a support assistant that handles 100,000 conversations a month. Each conversation sends 1,500 input tokens (instructions, history and help articles) and receives 400 output tokens. That adds up to 150 million input tokens and 40 million output tokens. The math is simple: input tokens times the input rate plus output tokens times the output rate, divided by one million.
PRICES = { # USD per million tokens: (input, output)
"Claude Fable 5.1": (10.00, 50.00),
"Claude Sonnet 5": (2.00, 10.00),
"Gemini 3.8 Flash": (0.75, 3.75),
"GPT-5.6 Luna": (0.20, 1.20),
"Mistral Small 4": (0.15, 0.60),
}
CONVERSATIONS = 100_000
INPUT, OUTPUT = 1_500, 400 # tokens per conversation
def monthly_cost(input_price: float, output_price: float) -> float:
input_tokens = CONVERSATIONS * INPUT # 150 M
output_tokens = CONVERSATIONS * OUTPUT # 40 M
return (input_tokens * input_price + output_tokens * output_price) / 1_000_000
for model, (inp, out) in PRICES.items():
print(f"{model:18} ${monthly_cost(inp, out):>9,.2f}")| Model | Input (150 M) | Output (40 M) | Monthly total | Per conversation |
|---|---|---|---|---|
| Claude Fable 5.1 | $1,500.00 | $2,000.00 | $3,500.00 | $0.035 |
| Claude Sonnet 5 | $300.00 | $400.00 | $700.00 | $0.007 |
| Gemini 3.8 Flash | $112.50 | $150.00 | $262.50 | $0.002625 |
| GPT-5.6 Luna | $30.00 | $48.00 | $78.00 | $0.00078 |
| Mistral Small 4 | $22.50 | $24.00 | $46.50 | $0.000465 |
Output is only 21% of the tokens, yet it makes up 57% of the bill on Claude Fable 5.1, Claude Sonnet 5 and Gemini 3.8 Flash, and 61.5% on GPT-5.6 Luna. If answers grow to 800 tokens, the Claude Sonnet 5 bill rises from $700 to $1,100 (+57%); if the context doubles to 3,000 tokens instead, it rises to $1,000 (+43%). Trimming answers saves more than trimming the prompt.
The blended price gives a rough estimate too: 190 million tokens at Claude Sonnet 5's $4 comes to $760, 8.6% above the exact figure. The gap exists because this workload has 3.75 input tokens per output token, more than the assumed 3:1. For a budget, always use both rates separately.
Documentation: Anthropic: Claude Fable 5.1 and Claude Sonnet 5 pricing ↗ · Google: Gemini 3.8 Flash pricing ↗ · OpenAI: GPT-5.6 Luna pricing ↗ · Mistral: Mistral Small 4 pricing ↗
4. Terms that change the list price#
The table shows standard prices and leaves out caching and batch discounts. Those discounts, and a few pricing conditions, can move your bill more than switching models, so read them on each provider's page before you decide.
- Batch: OpenAI and Anthropic document a 50% discount for asynchronous requests sent through their batch APIs. It suits classification, summaries or evaluations that can wait.
- Prompt caching: OpenAI bills cached input at 10% of the standard input price. Anthropic bills cache reads at 0.1 times the base input price (0.025 times on Claude Fable 5.1) and cache writes at 1.25 times (5-minute cache) or 2 times (1-hour cache).
- Time of day: DeepSeek charges half outside its peak hours, which run 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday, excluding Chinese public holidays. The table shows the peak rate.
- Prompt length: Grok 4.6 charges $4 input and $12 output per million once a prompt exceeds 200,000 tokens, double its base rate.
- Announced changes: Google lists Gemini 3.8 Flash at $0.75 and $3.75 through December 31, 2026; from January 1, 2027 the rates become $1.50 and $7.50.
Some good news does not show up in a table either: Anthropic confirmed that Claude Sonnet 5's $2 and $10, announced as introductory pricing, are now its standard price and that the increase planned for September 1 will not happen. Record the date you verified each condition alongside your budget.
Documentation: OpenAI: batch and cached input ↗ · Anthropic: prompt caching and Batch API ↗ · DeepSeek: peak and off-peak rates ↗ · xAI: Grok 4.6 rates by prompt length ↗ · Google: Gemini 3.8 Flash price change ↗
5. How to choose by task#
Price is a filter, not the final criterion. A cheap model that fails and forces retries can end up costing more per completed task. The sensible order is to describe the shape of your workload, shortlist on price and context, and decide with an evaluation on your own cases.
| Task | What drives the bill | What to check first |
|---|---|---|
| Classify, extract or route | Input: answers of a few dozen tokens | Input price and structured output; Mistral Small 4, Ministral 3 · 14B and GPT-5.6 Luna compete here |
| Chat and support | Output, plus context that grows every turn | Output price, caching of the history and a cap on answer length |
| Code generation | Long output and retries | Quality on your own tests; Codestral · 25.08 is cheap but has a 128,000-token context |
| Long documents and RAG | Very large input per request | Context window and length surcharges, such as Grok 4.6 above 200,000 tokens |
| Agents with tools | Input: context is re-sent at every step | Prompt caching and a step limit per task |
For the shortlist, Codifly's price leaderboard ranks models by cost and the value leaderboard weighs cost against performance. The AI cost calculator runs the math from section 3 with your volume and your input-to-output ratio, and the AI model comparison shows the limits and terms on each model card. Once you have two or three candidates, run the evaluation: the guide to choosing an LLM for production walks through the full process.
Go deeperHow to Choose an LLM for Production Without Relying on Leaderboards
Documentation: Anthropic: Claude Haiku 4.5 context window ↗ · Mistral: Codestral 25.08 context ↗
6. Keep your numbers current#
AI API prices change faster than almost any other cloud rate: new models ship, older ones move to legacy status and some providers announce dated price changes. A budget built in March can be out of date by September.
The figures in this guide come from Codifly's open dataset of AI API prices, published as JSON and CSV under a CC BY 4.0 license: use it in notebooks, articles or products as long as you credit the source. Each row carries input, output, blended price, context window, max output and links to the official pages, and the reviewed_at field gives the date of the last review. To show the prices on your own site, the open data page offers an embeddable widget that accepts a language, a theme and the number of models.
curl -s https://codifly.co/data/ai-models.json \
| jq -r '.reviewed_at, (.models | sort_by(.blended_usd_per_mtok) | .[:5][]
| "\(.name)\t\(.input_usd_per_mtok)\t\(.output_usd_per_mtok)")'- Automate a weekly query of the dataset and compare reviewed_at with the date of your last budget.
- Log billed tokens per model and per feature: without that data you will not know which price change affects you.
- Before signing a commitment, open the provider's official page; the dataset is a starting point, not a contract.
Documentation: Codifly: open dataset of AI API prices (JSON) ↗ · Creative Commons: CC BY 4.0 license ↗
Sources and scope
Documentation checked on September 27, 2026. Examples and decision criteria are editorial proposals; adapt them to your application's contract and validate them in an authorized test environment.
- OpenAI: batch and cached input ↗
- Anthropic: Claude Haiku 4.5 context window ↗
- Google: Gemini 3.8 Flash price change ↗
- DeepSeek: peak and off-peak rates ↗
- xAI: Grok 4.6 rates by prompt length ↗
- Mistral: Ministral 3 14B ↗
- Mistral: Mistral Small 4 pricing ↗
- Mistral: Codestral 25.08 context ↗
- Mistral: Mistral Large 3 ↗
- Mistral: Mistral Medium 3.5 ↗
- Anthropic: prompt caching and Batch API ↗
- Codifly: open dataset of AI API prices (JSON) ↗
- Creative Commons: CC BY 4.0 license ↗