AI for legal teams: long contracts, context, RAG and confidentiality

Pick models that read long contracts without dropping clauses

Data reviewed on Sep 8, 2026 · LLM Stats indexes and official prices

The data, today

The context

Legal work with AI revolves around long documents: contracts with dozens of schedules, case files that pile up years of filings, due diligence rooms with hundreds of documents. The model has to find the clause that contradicts another, point to where it sits, and never paper over gaps with plausible text. That makes two factors decisive, both visible in LLM Stats indexes: how much context a model accepts and how well it reasons over it.

Once the volume outgrows any window, retrieval takes over: split documents, generate embeddings and bring back only the relevant passages. And before a single file is uploaded there is confidentiality, which is settled by the provider's contract rather than a toggle in the interface: data retention, training use, processing region and subprocessors. This page helps you work through those questions in order before you commit to a model.

What to decide

  1. 1

    Long context or RAG?

    If the matter fits comfortably in the window and you need to connect scattered clauses, long context keeps things simple. If you work across a large document archive, RAG is unavoidable. Test both against questions you already know the answers to and measure correct citations, omissions and input tokens per query.

  2. 2

    How do I check the model is not making things up?

    Require every statement to cite document and section, and verify automatically that the quoted text exists. Build a question set with known answers, including some the documents cannot answer, and measure how often the model says so instead of improvising. Rerun it whenever you switch models.

  3. 3

    How should contracts be chunked for RAG?

    Split along the document's structure, by clause or section, not by a fixed character count, and store the heading, numbering and contract reference with every chunk. That lets the model cite precisely. Check whether the right clause shows up in the top results before you touch the prompt.

  4. 4

    What should I require from the provider on confidentiality?

    Read the data processing agreement: whether inputs are used for training, how long logs are kept, which region processes them and which subprocessors are involved. Ask for zero retention if offered and get it in writing. Codifly compares prices and capabilities; contractual sign-off belongs to your legal team.

  5. 5

    What does it cost to review a full case file?

    Estimate the tokens in the file, the number of questions per matter and the length of answers. With long context you pay for the whole document on every query; with RAG you pay for embeddings once and only for retrieved passages per question. Run both scenarios in the calculator with official prices and compare cost per matter.

Common mistakes

  • Assuming a large context window means the model reads the whole document well.
  • Uploading client documents before reviewing retention and data use terms in the provider's contract.
  • Chunking contracts by fixed length and losing the numbering you need to cite clauses.

Tools and comparators

Guides to go deeper

Want a recommendation for your case?

Tell us your volume and the options you are weighing. We reply in writing with the numbers of your real usage; no commitment.

Request advice

No provider pays for its position. Indexes come from LLM Stats; prices from each provider's standard API. How we measure