RAG, retrieval-augmented generation, is a pattern that fetches relevant documents from an external store and passes them to a language model as context before it generates a response. The model answers from what it retrieves, not only from what it learned during training.

Why it exists

Language models have a fixed knowledge cutoff and no awareness of private data. A model trained on public text cannot answer questions about your internal documentation, recent events, or proprietary records unless that information is supplied at inference time. RAG is the standard way to supply it.

The alternative, fine-tuning, embeds knowledge into the model's weights. That is expensive, slow to update, and poorly suited to information that changes frequently. RAG keeps the knowledge store and the model separate, so you can update documents without touching the model.

How it works

At a high level, a RAG pipeline has three stages:

  1. Indexing. Source documents are split into chunks, converted to vector embeddings, and stored in a vector database. Metadata is often stored alongside the vectors to support filtering.
  2. Retrieval. When a query arrives, it is also embedded. The system searches the vector store for chunks whose embeddings are close to the query embedding, typically by cosine similarity or dot product. Some pipelines add a reranking step to improve precision.
  3. Generation. The retrieved chunks are inserted into the prompt as context. The language model reads both the query and the retrieved text, then generates a response grounded in that material.

This is sometimes called the context window as a working surface: you are populating the model's context window with retrieved facts rather than relying on parametric memory.

What can go wrong

RAG is not automatic accuracy. Several failure modes are common.

Retrieval misses. If the relevant chunk is not retrieved, the model either hallucinates or says it does not know. Chunk size, embedding model choice, and query formulation all affect recall.

Context poisoning. Retrieved text that is outdated, contradictory, or off-topic can mislead the model even when the retrieval step technically succeeds.

Prompt injection. Malicious content embedded in a retrieved document can attempt to redirect the model's behaviour. This is a real risk when the document corpus is not fully controlled. See what is prompt injection for more detail.

Attribution errors. Models sometimes blend retrieved content with training knowledge in ways that are difficult to trace. Grounding claims back to source chunks is good practice but requires deliberate pipeline design.

Where RAG fits in a broader system

RAG is often one component in a larger architecture. An MCP server can expose retrieval as a tool that an agent calls dynamically, rather than retrieving on every request. Agentic coding environments sometimes use RAG to pull in codebase context that exceeds what fits in a single prompt.

Evaluation matters here. A RAG pipeline needs testing across retrieval quality, answer faithfulness, and answer relevance, as three separate concerns. A single aggregate metric obscures which component is failing. See what is an eval for how evaluation design applies.

What we test for

Engineers working with RAG pipelines need to reason clearly about where a failure originates: the retrieval step, the context passed to the model, or the model's use of what it is given. Our vetting assesses this directly. In the AI-native assessment, candidates are scored on AI output verification and AI-accelerated debugging, both of which surface whether someone understands the full pipeline or is treating it as a black box. The full approach is described at how we vet.

Short answers

What is the difference between RAG and fine-tuning?

Fine-tuning embeds knowledge into model weights and requires retraining to update. RAG retrieves knowledge at inference time from an external store, making it easier to keep current and more suitable for private or frequently changing information.

Does RAG eliminate hallucination?

No. RAG reduces hallucination by grounding responses in retrieved text, but the model can still misread, blend, or ignore that text. Retrieval failures also leave the model without the context it needs, which can produce confident but wrong answers.

What kind of database does RAG use?

Most RAG pipelines use a vector database to store and search document embeddings by similarity. Some implementations add relational metadata stores or keyword search alongside vector search to improve retrieval precision.

Let's talk

Get a shortlist within five working days

You share the roles and the stack in a short form or a thirty-minute call. Within five working days you get named senior engineers to review, each with both scorecards.

Reviewed onClutch4.9 out of 5 from 36 reviews
ISO 27001
Certified

Book thirty minutes with Dale

The calendar is provided by HubSpot, which sets its own cookies. Load it here, or book on HubSpot's page.

Open booking page