Document search and RAG: embeddings, chunking and pgvector

Find the right passage before asking the model for an answer

Data reviewed on Sep 8, 2026 · LLM Stats indexes and official prices

The data, today

The context

Most RAG systems that give bad answers do not have a model problem; they have a retrieval problem. The passage containing the answer never reaches the prompt. That is why it pays to split the system in two and evaluate each half on its own. First, given a question, do you retrieve the right passages? Second, with those passages in front of it, does the model answer correctly and cite the source? Blending both questions makes it impossible to know what to fix.

The infrastructure choices are simpler than they look. You pick an embedding model for its quality on your documents and its price per million tokens. The vectors can live in PostgreSQL with pgvector up to fairly large volumes. And the original documents go to object storage, where you can reprocess them every time you change embedding model or chunking strategy without hunting for copies.

What to decide

  1. 1

    Which embedding model should I pick?

    Compare price per million tokens and vector dimensions, which affect storage and search speed. Then test two or three models on your documents with a set of real questions where the correct passage is labeled. Measure recall in the top five or ten results; that number makes the call.

  2. 2

    How should I chunk documents?

    Follow the structure: sections, paragraphs, table rows. Start with chunks of a few hundred tokens with some overlap, and prepend the document and section title to each. Change the size only when your retrieval evaluation improves, not on a hunch.

  3. 3

    pgvector or a dedicated vector database?

    If you already run PostgreSQL, pgvector saves you another system and lets you filter on metadata with SQL in the same query. Consider a dedicated store once you measure p95 latency or index memory that PostgreSQL cannot sustain at your volume. Compare monthly cost and operational effort as well.

  4. 4

    How do I evaluate retrieval?

    Build a set of fifty to a hundred real questions, each labeled with the passage that answers it. Measure recall at k and the average rank of the correct passage. Try hybrid search with keywords and a reranker. Only once retrieval is solid should you tune the generation prompt.

  5. 5

    Where should source documents be stored?

    In object storage, versioned, with a stable ID the index can reference. That lets you reindex when you change embedding models without relying on stray copies, and show users the original being cited. Estimate stored GB, reads per reindex and egress if you serve the files.

Common mistakes

  • Tuning the prompt for weeks when the right passage never makes it into the context.
  • Mixing vectors from two different embedding models in the same index.
  • Throwing away the source documents and keeping only vectors, which makes reindexing impossible.

Tools and comparators

Guides to go deeper

Want a recommendation for your case?

Tell us your volume and the options you are weighing. We reply in writing with the numbers of your real usage; no commitment.

Request advice

No provider pays for its position. Indexes come from LLM Stats; prices from each provider's standard API. How we measure