Most affordable · Embeddings
- 01Voyage 4 LiteVoyage AI$0.02 USD / M tokens
- 02Mistral EmbedMistral AI$0.10 USD / M tokens
- 03Codestral EmbedMistral AI$0.15 USD / M tokens
Find the right passage before asking the model for an answer
Most RAG systems that give bad answers do not have a model problem; they have a retrieval problem. The passage containing the answer never reaches the prompt. That is why it pays to split the system in two and evaluate each half on its own. First, given a question, do you retrieve the right passages? Second, with those passages in front of it, does the model answer correctly and cite the source? Blending both questions makes it impossible to know what to fix.
The infrastructure choices are simpler than they look. You pick an embedding model for its quality on your documents and its price per million tokens. The vectors can live in PostgreSQL with pgvector up to fairly large volumes. And the original documents go to object storage, where you can reprocess them every time you change embedding model or chunking strategy without hunting for copies.
Compare price per million tokens and vector dimensions, which affect storage and search speed. Then test two or three models on your documents with a set of real questions where the correct passage is labeled. Measure recall in the top five or ten results; that number makes the call.
Follow the structure: sections, paragraphs, table rows. Start with chunks of a few hundred tokens with some overlap, and prepend the document and section title to each. Change the size only when your retrieval evaluation improves, not on a hunch.
If you already run PostgreSQL, pgvector saves you another system and lets you filter on metadata with SQL in the same query. Consider a dedicated store once you measure p95 latency or index memory that PostgreSQL cannot sustain at your volume. Compare monthly cost and operational effort as well.
Build a set of fifty to a hundred real questions, each labeled with the passage that answers it. Measure recall at k and the average rank of the correct passage. Try hybrid search with keywords and a reranker. Only once retrieval is solid should you tune the generation prompt.
In object storage, versioned, with a stable ID the index can reference. That lets you reindex when you change embedding models without relying on stray copies, and show users the original being cited. Estimate stored GB, reads per reindex and egress if you serve the files.
Tell us your volume and the options you are weighing. We reply in writing with the numbers of your real usage; no commitment.
No provider pays for its position. Indexes come from LLM Stats; prices from each provider's standard API. How we measure
Analysis, guides and new technology comparisons.