Ayan Ali.
6 min readUpdated

What Is RAG and Why Do LLMs Hallucinate?

An LLM predicts likely text, it does not look facts up. RAG fetches the relevant documents first and makes the model answer from them. Here is how it works and where it breaks.

RAGLLMAI
A formless grey particle cloud on the left resolving into an ordered lattice of matte cubes on the right, one face lime, with the OpenAI logo on a plate in the foreground.

Short answer: RAG, short for Retrieval-Augmented Generation, searches your own documents for passages relevant to a question, adds them to the prompt, and asks a language model to answer from them. LLMs hallucinate because they generate plausible text rather than looking facts up, so giving them the right source material removes much of the guessing. It does not remove all of it, which the last sections cover. The examples come from a RAG recommender I built over 838 dramas.

What is RAG?

RAG is a way to answer questions with a language model using your data, without retraining the model. The system retrieves relevant text first, then generates an answer from that text.

The name describes the order: retrieval, then augmentation (adding the retrieved text to the prompt), then generation. The model itself is unchanged. What changes is what it reads before it answers, which is why a RAG system can use new information the day you add it to the index.

Why do LLMs hallucinate?

A language model predicts likely next words, and likely is not the same as true. It was trained on large amounts of text and learned patterns, and it has no step that checks a claim against a source before writing it.

Three situations make this worse. The model may not have seen the fact at all, such as your private documents or anything after its training cutoff. It may have seen conflicting versions of the fact. Or the question may ask for something specific, like a title, a price or a citation, where a plausible-sounding invention reads exactly like a real answer. The output sounds equally confident in every case, which is why a hallucination is hard to spot without a source to compare against.

How does RAG work, step by step?

A RAG system runs four steps for each question. Each step is a separate piece you can test.

  1. Embed the question. An embedding model turns the question into a list of numbers that represent its meaning.
  2. Retrieve passages. A vector database finds stored passages whose meaning is closest to the question.
  3. Build the prompt. The system puts those passages in the prompt with an instruction to answer only from them.
  4. Generate the answer. The language model writes a response using the passages.

In my drama recommender, step 2 returns about 30 candidates, a reranker narrows them to 6, and GPT-4o writes recommendations only about those 6. The user never asks the model to remember which dramas exist, because the system hands it the real ones.

What does RAG fix, and what does it not fix?

RAG fixes missing knowledge, and it does not fix everything else. It helps with facts the model never saw and with information that changes. It does not guarantee the model uses what it is given.

| Problem | Does RAG help? | |---|---| | Model lacks your private data | Yes, retrieval supplies it | | Information changes often | Yes, re-index instead of retraining | | Model invents a title or number | Reduces it, since real text is in the prompt | | Retrieval returns the wrong passage | No, the model answers from bad context | | Model ignores the supplied text | No, needs a stricter prompt and an output check | | Question has no answer in the data | Only if the prompt allows "I don't know" |

The honest summary is that RAG moves the failure point. Instead of the model guessing from memory, failures now come from retrieval quality, prompt wording, or the model drifting off the context, and each of those can be measured and improved.

When should I use RAG instead of fine-tuning or a long prompt?

Use RAG when you need the model to answer from facts that live in documents and that change over time. Use a long prompt when all the material fits in the context window. Use fine-tuning when you want to change style or behavior rather than add facts.

A RAG system is easy to update, because adding a document means indexing it, with no training run. Fine-tuning bakes information into the model and is slow to change. A long prompt is the simplest option for a small, fixed set of text. See RAG vs fine-tuning for the comparison.

What do I need to build a RAG system?

You need four things: documents cut into chunks, an embedding model, a vector database, and a language model. The pieces are standard, and quality comes from how you combine them.

In my project that was OpenAI's text-embedding-3-small for embeddings, Qdrant as the vector database, a cross-encoder reranker, and GPT-4o for generation, behind a FastAPI backend. The first decision that mattered most was chunking, since retrieval can only return what the chunks contain. See how to chunk documents for RAG.

How do I reduce hallucinations in a RAG system?

Layer several measures, since no single one is enough. The ones I used or would add are below.

  • Grounded prompt. Tell the model to answer only from the supplied passages and never invent titles or facts.
  • Good chunks. Cut documents so each chunk answers one kind of question.
  • Reranking. Rescore the top candidates so the best passages reach the prompt.
  • A no-answer path. Allow "none of these fit" when retrieval is weak.
  • An output check. Compare names, titles and numbers in the answer against the retrieved text.

The full build with the prompt and the check is in building a RAG chatbot that doesn't hallucinate.

How do I tell if my RAG system is working?

Test it on a fixed set of questions and count the failures. Include questions whose answer is not in your documents, since those show whether the system invents one.

Track how often retrieval returns the right passage, and how often the final answer contains something that was not in the retrieved text. I have not published a measured rate for my project, and I would not trust one that was not tested on a set like this. Re-run the set after every change to chunking, retrieval or the prompt.

If you want a RAG chatbot built on your own catalogue or documents, see hire a RAG developer for a custom chatbot, the K-Drama RAG case study, or get in touch.

Frequently asked questions

What is RAG in simple terms?

Retrieval-Augmented Generation. The system searches your own documents for passages relevant to the question, puts them in the prompt, and asks the language model to answer using only those passages.

Why do LLMs hallucinate?

A language model generates the most plausible next words from patterns in its training data. It has no built-in lookup that checks a fact, so when it lacks the right information it can still produce a fluent, wrong answer.

Does RAG eliminate hallucinations?

No. It reduces them by supplying real source text, but retrieval can miss the right passage, and the model can still ignore or misread what it is given.

Do I need a vector database for RAG?

Usually, for more than a few documents. A vector database finds passages by meaning. For a very small set you can put the documents directly in the prompt.

When is RAG the wrong choice?

When the whole knowledge base fits in the prompt, or when you want to change the model's style or behavior rather than give it facts. In those cases, direct prompting or fine-tuning may be simpler.

Need a retrieval system that doesn't make things up?

Vector search, reranking, grounded answers — I have shipped this in production, including one engine over 81,000 titles running at about a cent a query. Tell me what you're building and I'll tell you what it takes.

Read next

How to Add an AI Chatbot to Shopify Grounded in Your Product Catalogue
A frosted speech-bubble form above a grid of product boxes, connected by thin lime lines, with the Shopify logo set into the floor and the OpenAI logo on the bubble.
Shopify

How to Add an AI Chatbot to Shopify Grounded in Your Product Catalogue

A Shopify chatbot that does not invent products needs four parts: a synced catalogue index, retrieval, a grounded prompt, and a Liquid section that talks to your backend. Here is each part.

Read
How to Build an AI Product Finder Section for Shopify
A wide field of dark cubes receding into shadow with a narrow lime light picking out exactly three, and the Shopify logo on a plate in the foreground.
Shopify

How to Build an AI Product Finder Section for Shopify

A product finder takes 'a waterproof jacket for light hiking under $150' and returns matching products. It needs vector search and filters, and it does not need a chat model.

Read
How to Chunk Documents for RAG (With a 6,234-Chunk Example)
A tall dark slab sliced into many thin horizontal layers fanned apart, four of them lime-edged, with the Qdrant logo on a plate beneath.
RAG

How to Chunk Documents for RAG (With a 6,234-Chunk Example)

One vector per document blends everything into a blur. Splitting each record into typed chunks, one per kind of question, kept my retrieval sharp. Here are the real counts and rules.

Read

Got a store or a build in mind?

Tell me what you're trying to ship — a Shopify section, a full storefront, a chatbot, or a web app. I'll tell you what it takes, honestly, before you commit.

ayanrjpoot@gmail.com