Ranking · AI models

Best AI models for legal work

Contracts, due diligence and long case files: the model must read whole documents and reason over them.

Criteria models with a published reasoning index and at least 1,000,000 tokens of context, ranked by window; equal windows by reasoning.

Data reviewed on Sep 8, 2026 · 6 of 17 models publish this figure.

  1. 01
    1.05MContext window
  2. 02
    1.05MContext window
  3. 03
    1.05MContext window
#ModelContext windowBlended priceCost-benefit
01GPT-6 AstraOpenAI1.05M$20.003.0 pts/USD
02GPT-5.6 SolOpenAI1.05M$8.006.9 pts/USD
03GPT-5.6 TerraOpenAI1.05M$4.5011.4 pts/USD
04Gemini 3.8 FlashGoogle1.05M$1.5034.0 pts/USD
05Claude Opus 5Anthropic1.00M$10.005.5 pts/USD
06Claude Fable 5.1Anthropic1.00M$20.002.8 pts/USD

How it is calculated

  • The indexes (overall, reasoning, coding) are the ones LLM Stats publishes; Codifly does not run these tests.
  • Prices are each provider's official standard API prices, without caching or batching.
  • The blended price weighs 3 input tokens per output token; cost-benefit divides the overall index by that price.
Methodology and sources

Which one fits your case?

The top of a ranking is not always the most cost-effective at your volume. Estimate your cost or ask us for a written recommendation.

Frequently asked questions

Which is the best AI model for legal?

GPT-6 Astra (1.05M) leads this ranking, followed by GPT-5.6 Sol (1.05M) · GPT-5.6 Terra (1.05M). Data reviewed on Sep 8, 2026.

Which model in this ranking has the best price-performance?

Gemini 3.8 Flash: 34.0 LLM Stats index points per dollar, at a blended price of $1.50 per million tokens.

How is this ranking calculated?

models with a published reasoning index and at least 1,000,000 tokens of context, ranked by window; equal windows by reasoning. The indexes (overall, reasoning, coding) are the ones LLM Stats publishes; Codifly does not run these tests. Prices are each provider's official standard API prices, without caching or batching. The blended price weighs 3 input tokens per output token; cost-benefit divides the overall index by that price.