Ayan Ali.
7 min readUpdated

Building a RAG Chatbot That Doesn't Hallucinate

A RAG chatbot hallucinates less when the model only explains what retrieval found. Here is the pipeline I built over 838 dramas, the prompt rules, and the check I would add.

RAGLLMPythonQdrant
A thin lime line passing through four dark gates of narrowing aperture and exiting as one clean line, with the Qdrant logo at the final gate and the OpenAI logo at the first.

Short answer: a RAG chatbot stays honest when retrieval does the knowing and the language model only does the explaining. In my drama recommender, every request retrieves real shows from a vector database, reranks them, and passes only those into a prompt that forbids inventing titles. This reduces hallucination a lot, but a prompt rule is not a guarantee, so the last section covers the check I would add. All numbers below come from my own project, a RAG engine over 838 Korean and Chinese dramas.

How does a RAG chatbot stay grounded in real data?

It splits the job in two. Retrieval finds the facts, and the language model writes sentences about facts it was handed. The model is never asked to remember which dramas exist.

In my build, a request is embedded with OpenAI's text-embedding-3-small (1,536 dimensions), searched against Qdrant for about 30 candidates, deduplicated by show, reranked by a cross-encoder down to the top 6, and then placed in the GPT-4o prompt. The answer streams back token by token, and the titles that were retrieved are sent along in a response header so the interface can show real sources. The catalogue can grow by re-indexing, with no retraining.

What does the data look like before retrieval?

Retrieval quality depends on how the source data was cut up, so the first step is chunking. Each show becomes up to four typed chunks, each written for a different kind of question.

  1. Chunk type : What it holds : Count
  2. synopsis : Title, year, genres, cast, plot : 838
  3. mood : Tone and vibe tags : 838
  4. review : Individual viewer reviews : 3,809
  5. similarity : "Fans also liked" with comments : 749

That is 6,234 chunks across 838 shows. A mood search queries the mood and synopsis types, and a similarity search queries similarity and synopsis. The full reasoning is in how to chunk documents for RAG.

Why use two-stage retrieval?

Vector search is fast but approximate, and a cross-encoder is accurate but slow. Using both gives accuracy at bounded cost. The embedding model compares pre-computed vectors, so it can scan thousands of chunks, but it can rank a plausible match above a better one.

A cross-encoder reads the query and one candidate together and scores how well they fit. It cannot be pre-computed, so it only runs on the short list from vector search. I used ms-marco-MiniLM-L-6-v2 through fastembed and ONNX Runtime, which runs the same model weights in roughly a quarter of the memory of the PyTorch version. See the reranker article for the details.

What does the grounded prompt say?

The prompt does three things: it restricts the model to the supplied context, it asks for specific evidence, and it requires an honest caveat. The core of my system prompt reads:

Always base recommendations strictly on the provided context — never invent
or hallucinate show titles.
For each recommendation:
- State the show title and year
- Explain in 2-3 sentences why it fits the request using specific details
  from the context (plot elements, mood tags, review phrases)
- Add one honest caveat if relevant (pacing, sad ending, slow start)

Asking for specific details from the context matters. It pushes the model to quote retrieved text rather than fall back on general knowledge, and it makes a wrong answer easier to spot because the cited detail either appears in the context or does not.

How do I keep a multi-turn chat grounded?

Retrieve again on every turn, and put that turn's results into the system prompt. A follow-up such as "more like number 2" or "anything shorter?" changes what should be retrieved, so reusing the first turn's context would let the model drift.

My chat prompt says to recommend only shows present in the "Available shows" section for this turn, and the context for that turn is rebuilt from fresh retrieval and appended to the system message. If the user only chats or asks a clarifying question, the prompt tells the model to answer naturally rather than force a list.

What happens when nothing relevant is retrieved?

The model should say so rather than guess. This is the case most grounded prompts forget: the retrieval returns the nearest chunks even when none is a good match, and a model told to recommend will recommend from weak material.

Add an explicit instruction for it, such as "If none of the available shows fit, say so and ask a clarifying question." Better still, apply a minimum reranker score and skip generation below it. I have not set a score threshold in this project, so a vague query still gets the closest available shows.

How do I verify the answer after generation?

Check that every title in the answer was retrieved for that request. A prompt rule asks the model to behave, and a check proves it did. This is the safeguard I would add, since my current system relies on the prompt and the sources header.

def find_invented_titles(answer_titles: list[str], retrieved_titles: set[str]) -> list[str]:
    known = {t.casefold() for t in retrieved_titles}
    return [t for t in answer_titles if t.casefold() not in known]

invented = find_invented_titles(extracted_titles, retrieved_titles)
if invented:
    # drop those items, or regenerate with a stricter instruction
    ...

Extracting titles from free text is the hard part. The simple route is asking the model to return a numbered list in a fixed format, or to return JSON with a title field, so the check compares exact strings instead of parsing prose.

How do I measure whether it hallucinates?

Build a small test set and count failures. I have not measured a hallucination rate for this project, so I do not quote one. A useful test set has 30 to 50 queries, including some with no good match in the catalogue, and a script that runs each one and applies the title check above.

Track two numbers: how often an answer contains a title that was not retrieved, and how often it recommends something that does not match the query. Re-run the set after changing chunking, the reranker or the prompt, so each change has evidence behind it.

What does a query cost?

About one cent per query in my project: one embedding call and one GPT-4o call. Indexing the whole catalogue once cost around five cents. A browse carousel on the same data skips the language model and the reranker entirely, because showing the user a row of results does not need an explanation.

Matching cost to value like this keeps the cheap path cheap. The model is worth paying for when it explains a recommendation, and wasted when the user is only scrolling.

What did I learn building it?

Three lessons carried over to other projects. First, quality starts in the chunks, since a vector store full of blurred averages retrieves blurred answers. Second, the same vector database can behave differently in local and cloud modes: embedded Qdrant filtered without indexes, and Qdrant Cloud returned an error until I created payload indexes. Third, a restart loop with no stack trace pointed to the host killing the process for using too much memory, which the ONNX reranker fixed.

See the K-Drama RAG case study for the full build. If you want a chatbot grounded in your own catalogue or documents, see hire a RAG developer for a custom chatbot, or get in touch.

Frequently asked questions

Does RAG stop an LLM from hallucinating?

It reduces it, but does not guarantee it. RAG gives the model real text to answer from, and a grounded prompt tells it to use only that text. The model can still ignore the context or retrieve the wrong passage, so check the output.

What is the main rule that keeps a RAG chatbot grounded?

Put the retrieved passages in the prompt and instruct the model to recommend only from them. Retrieval supplies the facts, and the model only explains them.

How do I check a RAG answer for invented items?

Compare every title or entity in the answer against the set retrieved for that request. Anything outside the set is flagged or removed before the answer is shown.

Do I need a reranker?

It helps when vector search returns plausible but weaker matches. A cross-encoder rereads the query and each candidate together, so the best passages reach the prompt and the weaker ones do not.

How much does a RAG query cost?

In my project, roughly one cent per query: one embedding call and one GPT-4o call. The one-time indexing cost for about 6,234 chunks was around five cents.

Need a retrieval system that doesn't make things up?

Vector search, reranking, grounded answers — I have shipped this in production, including one engine over 81,000 titles running at about a cent a query. Tell me what you're building and I'll tell you what it takes.

Read next

How to Add an AI Chatbot to Shopify Grounded in Your Product Catalogue
A frosted speech-bubble form above a grid of product boxes, connected by thin lime lines, with the Shopify logo set into the floor and the OpenAI logo on the bubble.
Shopify

How to Add an AI Chatbot to Shopify Grounded in Your Product Catalogue

A Shopify chatbot that does not invent products needs four parts: a synced catalogue index, retrieval, a grounded prompt, and a Liquid section that talks to your backend. Here is each part.

Read
How to Chunk Documents for RAG (With a 6,234-Chunk Example)
A tall dark slab sliced into many thin horizontal layers fanned apart, four of them lime-edged, with the Qdrant logo on a plate beneath.
RAG

How to Chunk Documents for RAG (With a 6,234-Chunk Example)

One vector per document blends everything into a blur. Splitting each record into typed chunks, one per kind of question, kept my retrieval sharp. Here are the real counts and rules.

Read
Shopify AI Chatbot App vs Building Your Own RAG Chatbot
A sealed matte black monolith beside an open lime-edged frame showing its interior lattice, with the Shopify logo on each, under even light.
Shopify

Shopify AI Chatbot App vs Building Your Own RAG Chatbot

An app is faster to start, and a custom build gives you control over answers, data and cost. Here is how to decide, what a build includes, and what to ask a developer if you hire.

Read

Got a store or a build in mind?

Tell me what you're trying to ship — a Shopify section, a full storefront, a chatbot, or a web app. I'll tell you what it takes, honestly, before you commit.

ayanrjpoot@gmail.com