Ayan Ali.
6 min readUpdated

How to Chunk Documents for RAG (With a 6,234-Chunk Example)

One vector per document blends everything into a blur. Splitting each record into typed chunks, one per kind of question, kept my retrieval sharp. Here are the real counts and rules.

RAGChunkingPythonQdrant
A tall dark slab sliced into many thin horizontal layers fanned apart, four of them lime-edged, with the Qdrant logo on a plate beneath.

Short answer: cut each document by the kind of question it should answer, not only by length. I turned 838 dramas into 6,234 chunks of four types, with one fixed rule set per type, and each search mode queries only the types relevant to it. A single vector per show averages plot, mood and reviews into one blurred point. Counts, rules, sizes and one limit to check are below, taken from my own pipeline.

What does a good chunk look like?

A good chunk covers one idea, carries enough context to make sense alone, and is written for the kind of query that should find it. A chunk that mixes plot, mood and reviews matches many queries weakly and none strongly.

Each chunk in my project starts with the show's title, so a passage like a review is never anonymous once retrieved. Chunks also carry metadata alongside the text. Text is what gets embedded for similarity, and metadata is what you filter and group on.

Should I chunk by size or by structure?

Chunk by structure when your data has it, and by size only when it does not. Fixed-size splitting is the default for plain text such as articles or transcripts. Structured records, such as product or show data, already have fields that map to different questions.

My data was structured, so splitting by fields made more sense than splitting by character count. Fixed-size splitting would have cut a show's synopsis mid-sentence and mixed unrelated fields in one chunk. Where you must split long plain text, split at paragraph or sentence boundaries, and consider overlap so a sentence is not separated from the context that explains it.

How did I chunk the drama dataset?

Each show becomes up to four chunk types. Each type targets a different search intent.

| Chunk type | Built from | Count | Typical length (median) | |---|---|---|---| | synopsis | Title, year, country, episodes, rating, genres, cast, plot | 838 | 822 characters | | mood | Title, mood and vibe tags, genres, rating | 838 | 292 characters | | review | One viewer review per chunk | 3,809 | 1,542 characters | | similarity | "Similar shows" list with comments | 749 | 2,489 characters |

The total is 6,234 chunks across 838 shows. Every show gets a synopsis and a mood chunk. Reviews and similarity chunks exist only when the source has enough material, which is why similarity has 749 rather than 838.

What rules decide which chunks get created?

Small rules keep low-quality chunks out of the index. A chunk that is nearly empty still gets retrieved, and it still takes a slot in the prompt.

  • Reviews: up to 5 per show, only if the review is at least 100 characters, cut at 1,500 characters.
  • Similarity: only if the show has at least 3 recommendations with comments, using up to 12, each comment cut at 250 characters.
  • Synopsis and mood: always created.

The review chunk text starts with the show title, year and country, then the review content. That prefix keeps the review tied to its show when a chunk is retrieved alone.

def make_review_chunks(show):
    chunks = []
    for i, review in enumerate(show.get("reviews", [])[:5]):
        content = (review.get("content") or "").strip()
        if len(content) < 100:
            continue
        text = (f"Review of {show['title']} ({show['year']}, {show['country']}):\n"
                f"{content[:1500]}")
        chunks.append({"id": f"{show['id']}__review_{i}", "show_id": show["id"],
                       "chunk_type": "review", "title": show["title"], "text": text})
    return chunks

Why not one vector per document?

One vector per document averages every aspect of it into a single point. A query about mood then competes with plot and cast details that have nothing to do with it, and the match gets diluted.

Typed chunks keep each aspect sharp. A mood search queries the mood and synopsis types, and a similarity search queries similarity and synopsis. The mood chunk is short and made only of tone tags, so a query like "cozy and emotional, not too heavy" lands on it directly instead of on a long plot summary.

What metadata should every chunk carry?

Carry what you will filter on, group on, or display. Every chunk in my project holds the show ID, chunk type, title, year, country, source, rating, genres and tags.

Three uses justify each field. The show ID lets retrieval deduplicate, keeping at most 2 chunks per show so one popular show does not fill the prompt. The chunk type lets each search mode pick the right chunks. Country, rating, year and genre power user filters. One practical detail: Qdrant Cloud requires a payload index on any field you filter by, while embedded Qdrant filtered without one. My filtered queries worked locally and failed in the cloud until I created indexes on chunk_type, source, genres and rating.

How do I embed and update the chunks?

Embed each chunk's text with the same model you use for queries, and store the vector with its metadata. I used OpenAI's text-embedding-3-small, which produces 1,536 dimensions, and stored vectors in Qdrant with cosine distance.

The cost is small. Indexing the whole catalogue once cost around five cents. For updates, an incremental run chunks and embeds only new shows, about four chunks each, instead of re-embedding the corpus. If you move between vector stores, copy the existing vectors rather than regenerating them, since embeddings are just data you already own.

What is one limit I would check?

Check that your prompt builder does not cut chunks shorter than the chunking step made them. In my generator, each retrieved chunk is truncated to its first 600 characters before it goes in the prompt.

That is shorter than the median review chunk (1,542 characters) and similarity chunk (2,489 characters), so the model sees only the start of those. It may be fine, since the opening lines include the title and the strongest content, but I have not tested whether it hurts answers. The test is simple: raise the limit, re-run a fixed set of queries, and compare.

How do I know my chunking is good?

Test retrieval directly, before looking at the model's answers. For a set of queries where you know which chunk should match, check whether it appears in the top results.

Look at failures by type. If mood queries return plot chunks, the mood chunks are too thin or the query targets the wrong types. If one source dominates, the deduplication cap is too high. Chunking problems show up as retrieval problems first, which is faster to debug than reading finished answers. The whole system is described in building a RAG chatbot that doesn't hallucinate.

For a RAG system built on your own documents, see hire a RAG developer for a custom chatbot, the K-Drama RAG case study, or get in touch.

Frequently asked questions

What is chunking in RAG?

Splitting source documents into smaller passages that are each embedded and stored. Retrieval finds and returns chunks, so how you cut the documents sets the ceiling on answer quality.

What chunk size should I use for RAG?

There is no universal number. Start from the content: one coherent idea per chunk, short enough to stay on one topic. In my project the chunks ranged from about 300 characters (mood tags) to about 2,500 (similar-show lists).

Do I need overlap between chunks?

For long unstructured text, overlap helps keep sentences from being cut off from their context. For structured records where each chunk is built from separate fields, as in my project, overlap is not needed.

Why add metadata to every chunk?

Metadata lets you filter before ranking, deduplicate by source, and show where an answer came from. Each chunk in my project carries the show ID, chunk type, title, year, country, rating, genres and tags.

Do I have to re-embed everything when I add a document?

No. Chunk and embed only the new document. In my pipeline a new show adds about four chunks, so an incremental update re-embeds only those.

Need a retrieval system that doesn't make things up?

Vector search, reranking, grounded answers — I have shipped this in production, including one engine over 81,000 titles running at about a cent a query. Tell me what you're building and I'll tell you what it takes.

Read next

How to Add an AI Chatbot to Shopify Grounded in Your Product Catalogue
A frosted speech-bubble form above a grid of product boxes, connected by thin lime lines, with the Shopify logo set into the floor and the OpenAI logo on the bubble.
Shopify

How to Add an AI Chatbot to Shopify Grounded in Your Product Catalogue

A Shopify chatbot that does not invent products needs four parts: a synced catalogue index, retrieval, a grounded prompt, and a Liquid section that talks to your backend. Here is each part.

Read
Building a RAG Chatbot That Doesn't Hallucinate
A thin lime line passing through four dark gates of narrowing aperture and exiting as one clean line, with the Qdrant logo at the final gate and the OpenAI logo at the first.
RAG

Building a RAG Chatbot That Doesn't Hallucinate

A RAG chatbot hallucinates less when the model only explains what retrieval found. Here is the pipeline I built over 838 dramas, the prompt rules, and the check I would add.

Read
Shopify AI Chatbot App vs Building Your Own RAG Chatbot
A sealed matte black monolith beside an open lime-edged frame showing its interior lattice, with the Shopify logo on each, under even light.
Shopify

Shopify AI Chatbot App vs Building Your Own RAG Chatbot

An app is faster to start, and a custom build gives you control over answers, data and cost. Here is how to decide, what a build includes, and what to ask a developer if you hire.

Read

Got a store or a build in mind?

Tell me what you're trying to ship — a Shopify section, a full storefront, a chatbot, or a web app. I'll tell you what it takes, honestly, before you commit.

ayanrjpoot@gmail.com