What Is Retrieval-Augmented Generation in Practice
SkillVeris Team
AI Research Team

Retrieval-augmented generation (RAG) fetches relevant documents at query time and feeds them to an LLM as context, so answers are grounded in real sources instead of the model's memory alone.
In this guide, you'll learn:
- RAG solves two problems at once: outdated training data and hallucination, because the model reasons over text you supply rather than guessing.
- A practical RAG system has two phases — offline indexing of your documents into a vector store, and online retrieval plus generation for each user query.
- Chunking, embedding quality, and retrieval relevance matter more than the choice of LLM for most production RAG systems.
- RAG lets you update knowledge by re-indexing documents, with no expensive model fine-tuning or retraining required.
1What Is Retrieval-Augmented Generation?
Retrieval-augmented generation (RAG) is a technique that connects a large language model to an external knowledge source, retrieving relevant text at query time and passing it to the model as context before it answers. Instead of relying only on what the model learned during training, RAG lets it reason over documents you control — a knowledge base, product docs, or a set of PDFs.
In practice this means the model answers questions about information it was never trained on, and it can quote the exact passage it used. That grounding is what makes RAG the default architecture for question-answering assistants, internal search, and support bots in 2026.
2Why RAG Matters
RAG addresses the two most common complaints about raw LLMs: they hallucinate, and their knowledge has a cutoff date. By supplying fresh, relevant text at inference time, you sidestep both problems without touching the model's weights.
- Accuracy: the model answers from supplied evidence, reducing confident-but-wrong responses.
- Freshness: update the knowledge base and answers update immediately — no retraining.
- Traceability: you can show which document produced an answer, which builds user trust.
- Cost: re-indexing documents is far cheaper than fine-tuning or training a model.
🔑Key Idea
RAG changes the question from 'what does the model know?' to 'what did we give the model to read?' — and the second question is one you can control.
3How RAG Works: The Two Phases
Every RAG system splits into an offline phase and an online phase. Getting the offline phase right is what makes the online phase feel like magic.
- Offline indexing: split documents into chunks, convert each chunk into an embedding vector, and store the vectors in a vector database.
- Online retrieval: embed the user's question, find the most similar chunks by vector distance, and collect the top matches.
- Augmentation: paste the retrieved chunks into the prompt alongside the question and clear instructions.
- Generation: the LLM produces an answer grounded in the supplied chunks, ideally citing them.
Where Quality Is Won or Lost
Most RAG failures trace back to retrieval, not generation. If the right chunk never reaches the prompt, even the best model can only guess. Tuning chunk size, embeddings, and ranking pays off more than swapping the LLM.
4The Core Components
A minimal RAG stack has four moving parts, each with mature open-source and hosted options.
- Embedding model: turns text into vectors (e.g. OpenAI text-embedding-3, or open models like BGE and E5).
- Vector store: indexes and searches vectors (e.g. pgvector, Pinecone, Qdrant, Weaviate, Chroma).
- Retriever: runs the similarity search and optional re-ranking.
- Generator: the LLM that reads the context and writes the answer.
💡Pro Tip
Start with pgvector if you already run Postgres — it keeps your documents and vectors in one database and avoids a whole extra service in early prototypes.
5Chunking: The Underrated Step
Chunking is how you split source documents into retrievable pieces, and it quietly determines retrieval quality. Chunks that are too large dilute relevance and waste context; chunks that are too small lose the surrounding meaning a passage needs to make sense.
A common starting point is a few hundred tokens per chunk with a small overlap so ideas that straddle a boundary are not cut in half. Splitting on natural structure — headings, paragraphs, list items — usually beats a blind fixed-length cut.
- Aim for roughly 200-500 tokens per chunk as a starting range, then measure.
- Add a small overlap (for example 10-15 percent) to preserve cross-boundary context.
- Split on semantic boundaries — headings and paragraphs — before falling back to length.
- Store metadata (title, source URL, section) with each chunk for filtering and citations.
6Assembling the Prompt
Once you have retrieved chunks, you build a prompt that tells the model to answer only from the supplied context and to say when it does not know. Clear instructions here prevent the model from drifting back to its training data.
A typical template concatenates a system instruction, the retrieved passages labelled with their sources, and the user's question. Asking the model to cite the source of each claim turns a black-box answer into a verifiable one.
7Common Mistakes to Avoid
Teams new to RAG tend to trip over the same issues, and most are fixable without changing models.
- Ignoring retrieval quality: if the right chunk is not retrieved, no prompt trick will save the answer.
- Chunks too big or too small: both wreck relevance — test a few sizes on real questions.
- No source citations: without them users cannot verify answers and trust erodes.
- Stuffing too many chunks: more context is not always better; irrelevant passages distract the model.
- Forgetting to re-index: stale documents produce stale answers even though nothing looks broken.
⚠️Watch Out
RAG does not eliminate hallucination — it reduces it. A model can still misread supplied context, so keep citations visible and evaluate answers against a test set.
8Evaluating a RAG System
You cannot improve what you do not measure, and RAG has two things to measure: whether retrieval found the right passages, and whether the final answer is faithful to them.
- Retrieval quality: does the correct chunk appear in the top results for known questions?
- Answer faithfulness: does the answer stay within the retrieved evidence, or invent extra claims?
- Answer relevance: does the response actually address the user's question?
- Build a small labelled test set of question-and-source pairs and re-run it after every change.
9Key Takeaways
The essentials of RAG reduce to a few durable ideas you can carry into any implementation.
- RAG grounds an LLM in external documents fetched at query time, improving accuracy and freshness.
- The pipeline is index offline, then retrieve and generate online.
- Retrieval quality — driven by chunking and embeddings — usually matters more than the LLM choice.
- Cite sources so answers are verifiable, and evaluate with a labelled test set.
- Update knowledge by re-indexing, not retraining.
10Frequently Asked Questions
Q: How is RAG different from fine-tuning? A: Fine-tuning changes the model's weights to bake in new behaviour or knowledge, which is expensive and slow to update. RAG leaves the model untouched and supplies knowledge at query time, so you update it by re-indexing documents. Many teams use RAG for knowledge and fine-tuning only for style or format.
Q: Does RAG stop hallucination completely? A: No. It substantially reduces hallucination by grounding answers in real text, but a model can still misinterpret supplied context. Visible citations and evaluation against a test set are how you catch the remaining errors.
Q: What do I need to build a first RAG prototype? A: An embedding model, a vector store such as pgvector or Chroma, a retriever, and an LLM. You can wire these together with a framework like LangChain or LlamaIndex, or by hand in a few dozen lines of Python.
Q: How do I keep RAG answers up to date? A: Re-index documents when they change. A scheduled job that re-embeds new or edited content keeps the vector store fresh, and answers update immediately with no model retraining.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
AI Research Team
Our AI team covers the latest in machine learning, generative AI, and emerging tech — clearly and accurately.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.